Correction: Should actually be closer to 50 tok/sec with MTP!!
MTP (multi-token-prediction) and other speculative decoding methods like DFlash are a free lunch. You get a significant speedup with no loss in quality. The speedup depends on the task (since the draft model's acceptance rates vary by task, e.g. coding can consist of lots of structured "filler" tokens like brackets so has higher acceptance rates and therefore more of a speedup.
The problem is MLX has been lazy and slow to adopt MTP.
There are projects like MTPLX that are ahead of the curve and already have support for MTP on Mac, see nitter.net/Youssofal_/status/2074…
I hope with the launch of Qwen 3.8 27B we will see a push to add support for MTP in the MLX ecosystem.
Thanks @Youssofal_ for pointing this out!
Youssof Al Toukhi (@Youssofal_)
72+ TPS on Qwen 3.6 27B on a Macbook pro M5 max.
MTPLX V2 out now!
The fastest way to run models on MLX.
Video
— https://nitter.net/Youssofal_/status/2074404826373623922#m