After the pretraining the model can partially learn to scale thinking more and better just with some SFT on reasoning traces. Like if the pretraining already created the potential for it.
After the pretraining the model can partially learn to scale thinking more and better just with some SFT on reasoning traces. Like if the pretraining already created the potential for it.