been working with these guys as they quietly built this in stealth for the last year. Super impressive result.
it's turns out with ternary weight models, you can make an extremely size-efficient model that can run on device extremely fast, and be smart enough to solve IMO problems.
the GIFs showing the speed are pretty mind blowing.
so excited the whole world gets to check this out!
DeepGrove (@deepgrove_ai)
We introduce Maple-Preview, an open-source 20B-A1B ternary-weight reasoning LLM, SOTA in its weight class.
It solves IMO-level problems and runs at 200+ tokens/s on a Mac Mini M4, 5–16× faster than efficient models like Gemma 4, Qwen3.5, and gpt-oss.
deepgrove.ai/maple-preview
— https://nitter.net/deepgrove_ai/status/2084727154928189783#m