Twitter/X

Alex Cheema says Kimi K3 will require 4 × 512GB M3 Ultra Mac Studios (noting 2TB…

Brief

Alex Cheema says Kimi K3 will need four 512GB M3 Ultra Mac Studios (2TB @ 3.2TB/s) and estimates 30+ tokens/sec using MTP plus tensor parallelism over RDMA on Thunderbolt 5. He expects prefill to be slow until a 512GB M5 Ultra (due in October) yields ~5× faster prefill, and promises local.ai benchmarks and early-access codes; he already runs Kimi K2.5 at 24 tok/sec on two 512GB M3 Ultras (exolabs/MLX).

Why it matters

Alex Cheema says Kimi K3 will require 4 × 512GB M3 Ultra Mac Studios (noting 2TB @ 3.2TB/s) and he expects it can run at 30+ tokens/sec using MTP + tensor parallelism over RDMA via Thunderbolt 5.

Key details

  • He warns model prefill will be slow on current hardware but expects a 512GB M5 Ultra (arriving in October) to make prefill ~5× faster.
  • Cheema already runs Kimi K2.5 at 24 tokens/sec on 2 × 512GB M3 Ultra Mac Studios connected with Thunderbolt 5 (RDMA) using the exolabs/MLX backend, can run 'clawdbot', and will publish local.ai benchmarks and early-access codes when weights drop.
Source evidence

Running Kimi K3 on my desk?

It will require 4 x 512GB M3 Ultra Mac Studios (2TB @ 3.2TB/s).

Waiting on active parameter count, but based on the rumours I expect we can run it at 30+ tok/sec with MTP + Tensor Parallelism using RDMA over Thunderbolt 5.

Prefill will be slow but 512GB M5 Ultra should make that ~5x faster (expecting in October).

It will be on local dot ai with full benchmarks as soon as the weights drop. Comment below for early access - sending access codes out throughout the day.

Alex Cheema (@alexocheema)

Running Kimi K2.5 on my desk.

Runs at 24 tok/sec with 2 x 512GB M3 Ultra Mac Studios connected with Thunderbolt 5 (RDMA) using @exolabs / MLX backend.

Yes, it can run clawdbot.

Video

— https://nitter.net/alexocheema/status/2016404573917683754#m