Twitter/X

DeepGrove released Maple-Preview, an open-source 20B-A1B ternary-weight reasoning…

Brief

Maple-Preview is an open-source 20B-A1B ternary-weight reasoning LLM from DeepGrove that the team claims is SOTA in its weight class. DeepGrove says it solves IMO-level problems and runs at 200+ tokens/s on a Mac Mini M4—5–16× faster than Gemma 4, Qwen3.5, and gpt-oss. @agupta reports the project was developed in stealth for a year.

Why it matters

DeepGrove released Maple-Preview, an open-source 20B-A1B ternary-weight reasoning LLM that DeepGrove claims is state-of-the-art in its weight class.

Key details

  • Maple-Preview reportedly solves IMO-level problems and runs at 200+ tokens/s on a Mac Mini M4, delivering 5–16× higher throughput than Gemma 4, Qwen3.5, and gpt-oss.
  • @agupta says the team built the model in stealth over the past year and shared GIFs demonstrating the speed; the announcement was posted on 2026-08-04.
Source evidence

been working with these guys as they quietly built this in stealth for the last year. Super impressive result.

it's turns out with ternary weight models, you can make an extremely size-efficient model that can run on device extremely fast, and be smart enough to solve IMO problems.

the GIFs showing the speed are pretty mind blowing.

so excited the whole world gets to check this out!

DeepGrove (@deepgrove_ai)

We introduce Maple-Preview, an open-source 20B-A1B ternary-weight reasoning LLM, SOTA in its weight class.

It solves IMO-level problems and runs at 200+ tokens/s on a Mac Mini M4, 5–16× faster than efficient models like Gemma 4, Qwen3.5, and gpt-oss.

deepgrove.ai/maple-preview

— https://nitter.net/deepgrove_ai/status/2084727154928189783#m