Twitter/X

@alexocheema corrects the estimate to about 50 tokens/sec using MTP…

Brief

@alexocheema corrects their estimate to about 50 tokens/sec with MTP (multi-token-prediction), arguing MTP and speculative decoding (e.g., DFlash) give substantial speedups with no quality loss. Speedups depend on task acceptance rates (coding benefits from filler tokens). They criticize MLX's slow adoption, note MTPLX V2 supports MTP on Mac, and cite Youssof's 72+ TPS on Qwen 3.6 27B while urging adoption around Qwen 3.8 27B.

Why it matters

@alexocheema corrects the estimate to about 50 tokens/sec using MTP (multi-token-prediction), noting throughput varies by task because draft-model acceptance rates differ (e.g., coding has higher acceptance due to structured filler tokens like brackets).

Key details

  • Author claims MTP and speculative decoding methods (e.g., DFlash) deliver significant speedups with no quality loss and criticizes MLX for being slow to adopt MTP; they hope Qwen 3.8 27B's launch will push MLX to add support.
  • MTPLX already supports MTP on macOS (MTPLX V2 released); Youssof reports 72+ TPS on Qwen 3.6 27B running on a MacBook Pro M5 Max.
Source evidence

Correction: Should actually be closer to 50 tok/sec with MTP!!

MTP (multi-token-prediction) and other speculative decoding methods like DFlash are a free lunch. You get a significant speedup with no loss in quality. The speedup depends on the task (since the draft model's acceptance rates vary by task, e.g. coding can consist of lots of structured "filler" tokens like brackets so has higher acceptance rates and therefore more of a speedup.

The problem is MLX has been lazy and slow to adopt MTP.

There are projects like MTPLX that are ahead of the curve and already have support for MTP on Mac, see nitter.net/Youssofal_/status/2074…

I hope with the launch of Qwen 3.8 27B we will see a push to add support for MTP in the MLX ecosystem.

Thanks @Youssofal_ for pointing this out!

Youssof Al Toukhi (@Youssofal_)

72+ TPS on Qwen 3.6 27B on a Macbook pro M5 max.

MTPLX V2 out now!

The fastest way to run models on MLX.

Video

— https://nitter.net/Youssofal_/status/2074404826373623922#m