Now run Kimi K3 with a 2.78 trillion parameter model on a single CPU with 8.24 GB of RAM. 100% Opensource.
It's called kimi-k3-in-c.
Most inference stacks assume you need a GPU cluster and 5+ terabytes of memory to touch a frontier MoE.
The entire engine is 176 KB of portable C99.
Runs without BLAS, PyTorch, CUDA, any framework, a GPU, or AVX-512.
The 1.56 TB checkpoint sits on NVMe. Only 16 of 896 experts fire per token, so the sleeping 93% streams in on demand and never touches RAM.
The dense trunk runs in whatever memory you give it. 8 GB, 32 GB, 128 GB, 224 GB. Same weights. Byte-identical output at every budget.
What it does:
→ Runs Kimi K3 inference on a laptop with 8 GB free
→ Reads MXFP4 weights directly, never dequantizes to float32
→ Streams the trunk from disk with O_DIRECT, faster than the page cache
→ Keeps KDA attention state fixed regardless of context length
→ Passes bit-identical output between scalar, OpenMP, and AVX2 paths
→ Verifies against the PyTorch reference on every kernel
Ships with a 45 MB source tree, six C files and one Python packing script.
That's the whole thing.