This VibeThinker-3B must be tested absolutely!
I bet he thinks like crazy to do a small change, but...
Let's try!
Francesco Bertolotti (@f14bertolotti)
Stellar performance from a 3B model. These results were achieved primarily through post-training refinements on Qwen2.5-Coder. The paper doesn't provide many details, but it appears they distill from RL ckpts and then do a final RL-based instruct RL.
🔗arxiv.org/abs/2606.16140
— https://nitter.net/f14bertolotti/status/2066752828505288902#m