Twitter/X

A 3-bit MLX build of Laguna M.1 was made to run locally on Apple Silicon (M3 Max)…

Brief

Laguna M.1 was compiled into a 3-bit MLX build that runs locally on Apple Silicon (tested on an M3 Max with 128 GB unified memory), delivering roughly 26 tokens per second and hitting about 100 GB peak memory. The build was produced by @eauchs and published on Hugging Face (ox-ox/Laguna-…), showing local inference is possible though memory-intensive.

Why it matters

A 3-bit MLX build of Laguna M.1 was made to run locally on Apple Silicon (M3 Max) achieving ~26 tokens/sec with ~100 GB peak memory usage on a machine with 128 GB unified memory.

Key details

  • @eauchs produced the build and the model artifacts are shared on Hugging Face (ox-ox/Laguna-…).
  • The result demonstrates feasible local inference of Laguna M.1 on consumer Apple Silicon, albeit with high memory requirements.
Source evidence

someone asked if they could try getting Laguna M.1 running on a Mac.

we said yes.

they came back with a 3-bit MLX build running locally on Apple Silicon: ~26 tok/s, with ~100 GB peak memory on an M3 Max with 128 GB unified.

absolute GOAT behavior from @eauchs

huggingface.co/ox-ox/Laguna-…