Twitter/X

Ornith-1.0 is an open-source LLM family specialized for agentic coding, released…

Brief

Ornith-1.0 is a new open-source family of models specialized for agentic coding (9B, 31B dense; 35B and 397B MoE), released under the MIT license on 2026-06-25. Post-trained on gemma4 and qwen3.5, it uses a novel RL self-improvement method that jointly optimizes task-specific scaffolds and solution rollouts, and reports SOTA open-source coding scores (Terminal-Bench 77.5, SWE-Bench 82.4 verified, NL2Repo 48.2, ClawEval 77.1).

Why it matters

Ornith-1.0 is an open-source LLM family specialized for agentic coding, released under the MIT license (announced 2026-06-25) with model sizes: 9B Dense, 31B Dense, 35B MoE, and 397B MoE.

Key details

  • Ornith-1.0 reports state-of-the-art open-source coding benchmark results for its size: Terminal-Bench 2.1 = 77.5; SWE-Bench = 82.4 (verified), 62.2 (pro), 78.9 (Multilingual); NL2Repo = 48.2; SWE Atlas = 41.2 (QnA), 42.6 (RF), 39.1 (TW); ClawEval = 77.1.
  • Ornith-1.0 was post-trained on gemma4 and qwen3.5 and uses a novel self-improving reinforcement-learning strategy that generates both solution rollouts and task-specific scaffolds, jointly optimizing scaffold and solution to improve agentic coding performance.
Source evidence

Most people will compare benchmark scores.

I’m paying attention to something else.

Open models built specifically for agentic coding are improving at an incredible pace.

Huge release from the Ornith team.

Ornith (@ornith_)

Aloha! 🌺 Meet Ornith-1.0, a family of open-source LLMs specialized for agentic coding.

Ornith-1.0 spans the full parameter sizes including 9B Dense, 31B Dense, 35B MoE, and 397B MoE. It achieves state-of-the-art performance among open-source models of comparable size on coding benchmarks including:
✅Terminal-Bench 2.1(77.5)
✅SWE-Bench(82.4 on verified, 62.2 on pro, 78.9 on Multilingual)
✅NL2Repo(48.2)
✅SWE Atlas(41.2 on QnA, 42.6 RF, 39.1 TW)
✅ClawEval(77.1)

Post-trained on top of gemma4 and qwen3.5, Ornith-1.0 employs a novel self-improving training strategy in which reinforcement learning is used to generate not only solution rollouts, but also the task-specific scaffolds that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model generate higher-quality solutions in agentic coding.😎

All models are released under the MIT license, enabling full commercial and research use.

📖Tech Blog: deep-reinforce.com/ornith1
🤗Huggingface: huggingface.co/collections/d…

— https://nitter.net/ornith_/status/2070148887067963854#m