Twitter/X

Author @ivanfioravanti urges testing VibeThinker-3B (a 3 billion-parameter model)…

Brief

VibeThinker-3B, a 3B-parameter model, is being promoted for immediate testing (post dated 2026-06-16) with the author claiming it exhibits powerful reasoning after minor changes. Francesco Bertolotti highlights that the model’s strong results stem from post-training refinements on Qwen2.5-Coder, apparently distilling from RL checkpoints and applying a final RL-based instruct fine-tune (see arXiv:2606.16140).

Why it matters

Author @ivanfioravanti urges testing VibeThinker-3B (a 3 billion-parameter model) and suggests it “thinks like crazy” after small changes; post date referenced is 2026-06-16.

Key details

  • Francesco Bertolotti (@f14bertolotti) reports “stellar performance” from the 3B model achieved mainly via post-training refinements on Qwen2.5-Coder, claiming the paper indicates distillation from RL checkpoints followed by a final RL-based instruct fine-tuning (arXiv:2606.16140).
Source evidence

This VibeThinker-3B must be tested absolutely!
I bet he thinks like crazy to do a small change, but...

Let's try!

Francesco Bertolotti (@f14bertolotti)

Stellar performance from a 3B model. These results were achieved primarily through post-training refinements on Qwen2.5-Coder. The paper doesn't provide many details, but it appears they distill from RL ckpts and then do a final RL-based instruct RL.

🔗arxiv.org/abs/2606.16140

— https://nitter.net/f14bertolotti/status/2066752828505288902#m