I believe that the model does two things during the chain of thoughts. One is sampling, a form of search. Another however is really some form of reasoning where next tokens update the state to converge to a solution, a bit like human reasoning does, for info synthesis.