Twitter/X

Ornith-1.0 is a family of open-source agentic-coding LLMs (9B dense, 31B dense…

Brief

Ornith-1.0 is an MIT-licensed open-source LLM family specialized for autonomous coding, released 2026-06-26. Built on Gemma 4 and Qwen 3.5, it uses a self-scaffolding RL pipeline to improve agentic reasoning and reports strong benchmark results (e.g., 397B: 82.4 SWE-Bench Verified; 35B: 64.2 Terminal-Bench 2.1), with models available on Hugging Face.

Why it matters

Ornith-1.0 is a family of open-source agentic-coding LLMs (9B dense, 31B dense, 35B MoE, 397B MoE) post-trained on Gemma 4 and Qwen 3.5 using a novel self-scaffolding reinforcement-learning pipeline that jointly optimizes task scaffolds and solutions.

Key details

  • Benchmark claims: Ornith 397B hits 82.4 on SWE-Bench Verified (claimed to match Claude Opus 4.7); Ornith 35B scores 64.2 on Terminal-Bench 2.1 (claimed to beat larger models); thread also reports Terminal-Bench 2.1(77.5), NL2Repo 48.2, ClawEval 77.1 and other SWE Atlas/SWE-Bench breakdowns.
  • The models are 100% free and released under the MIT license with code and weights published (Hugging Face collection and tech blog linked); announcement posted 2026-06-26.
Source evidence

🚨 Open-source models are rapidly closing the gap on frontier agentic coding.

The launch of the Ornith-1.0 family proves you don't need a closed API to get top-tier coding agent performance!

Built on top of Gemma 4 and Qwen 3.5, Ornith uses a self-scaffolding RL pipeline to drastically boost coding reasoning.

The benchmarks are hard to ignore:
→ Ornith 397B hits 82.4 on SWE-Bench Verified, matching Claude Opus 4.7
→ Ornith 35B scores 64.2 on Terminal-Bench 2.1, beating larger models
→ Ornith 9B brings frontier-level logic to edge devices

100% Free and open-source.

This is exactly what the open-source ecosystem needed to push autonomous dev tools forward.

repo in 🧵↓

Ornith (@ornith_)

Aloha! 🌺 Meet Ornith-1.0, a family of open-source LLMs specialized for agentic coding.

Ornith-1.0 spans the full parameter sizes including 9B Dense, 31B Dense, 35B MoE, and 397B MoE. It achieves state-of-the-art performance among open-source models of comparable size on coding benchmarks including:
✅Terminal-Bench 2.1(77.5)
✅SWE-Bench(82.4 on verified, 62.2 on pro, 78.9 on Multilingual)
✅NL2Repo(48.2)
✅SWE Atlas(41.2 on QnA, 42.6 RF, 39.1 TW)
✅ClawEval(77.1)

Post-trained on top of gemma4 and qwen3.5, Ornith-1.0 employs a novel self-improving training strategy in which reinforcement learning is used to generate not only solution rollouts, but also the task-specific scaffolds that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model generate higher-quality solutions in agentic coding.😎

All models are released under the MIT license, enabling full commercial and research use.

📖Tech Blog: deep-reinforce.com/ornith1
🤗Huggingface: huggingface.co/collections/d…

— https://nitter.net/ornith_/status/2070148887067963854#m