Twitter/X

DwarfStar DeepSeek V4 Flash 0731 mxfp4 ran on an M3 Ultra in a single-request, no…

Brief

Ivan Fioravanti reports DwarfStar DeepSeek V4 Flash 0731 mxfp4 on an M3 Ultra achieved ~41 tokens/sec decode in a single-request, no MPT/DSpark setup—breaking the 40 tps barrier. He mandates "same logits, 100% same result" for kernel-improving agents, credits Opus 5 (20%) and @Kimi_Moonshot K3 (80%), and will apply lessons to the M5 kernel (branch: github.com/ivanfioravanti/ds…).

Why it matters

DwarfStar DeepSeek V4 Flash 0731 mxfp4 ran on an M3 Ultra in a single-request, no MPT/DSpark setup and decoded at ~41 tokens/sec, claiming to have broken the 40 tps barrier.

Key details

  • "Same logits, 100% same result" is presented as a mandatory rule for any agents improving kernels — speed without mathematical precision is useless.
  • Author credits 20% to Opus 5 and 80% to @Kimi_Moonshot K3, studies K3's thinking trace during experiments, and plans to apply lessons to the M5 kernel (branch: github.com/ivanfioravanti/ds…).
Source evidence

Where there's a will, there's a way 😎
DwarfStar DeepSeek V4 Flash 0731 mxfp4 on M3 Ultra single request no MPT/DSpark ~41 toks/s decode!!! 🚀

40 tps barrier broken!

Same logits, 100% same result: this is The Mandatory Rule for any agents improving kernels that I think everyone should follow. Speed without mathematical precision is useless.

Thanks 20% Opus 5 and 80% @Kimi_Moonshot K3 (what a model!)

And just looking at the thinking trace of K3 you can learn tons of stuff! Here usually I have a separate chat with K3 and when I see something I want to dig deeper or ask about from the main thinking trace I do and in some cases I stop the experiment and pivot immediately to something else.

Now let's see if I can apply learned lessons to the M5 kernel! 🚀

Branch for any braves soul willing to test it: github.com/ivanfioravanti/ds…

Video