Twitter/X

Kog open-sourced a 2B-parameter model on Hugging Face (referenced via a Hugging…

Brief

Kog released a 2-billion-parameter model on Hugging Face (announced in a Hugging Face blog) and used that model to demonstrate inference speeds above 3,000 tokens/second. Clement Delangue shared the announcement on Twitter/X on 2026-06-24, praising the release as “very cool work.”

Why it matters

Kog open-sourced a 2B-parameter model on Hugging Face (referenced via a Hugging Face blog) on 2026-06-24.

Key details

  • Clement Delangue highlighted that the same 2B model was used to demonstrate inference throughput exceeding 3,000 tokens per second.
Source evidence

Kog open-sourced on @huggingface the 2B model that they used to show a model running at 3,000+ tokens per second. Very cool work! huggingface.co/blog/kogai/ko…