Kog open-sourced on @huggingface the 2B model that they used to show a model running at 3,000+ tokens per second. Very cool work! huggingface.co/blog/kogai/ko…
Kog open-sourced a 2B-parameter model on Hugging Face (referenced via a Hugging…
Brief
Kog released a 2-billion-parameter model on Hugging Face (announced in a Hugging Face blog) and used that model to demonstrate inference speeds above 3,000 tokens/second. Clement Delangue shared the announcement on Twitter/X on 2026-06-24, praising the release as “very cool work.”
Why it matters
Kog open-sourced a 2B-parameter model on Hugging Face (referenced via a Hugging Face blog) on 2026-06-24.
Key details
- Clement Delangue highlighted that the same 2B model was used to demonstrate inference throughput exceeding 3,000 tokens per second.