Twitter/X

Published mixed q2/q3 quantizations of Laguna S2.1 that run on MacBook systems…

Brief

Laguna S2.1 now has mixed q2/q3 quantized builds that the author says run on MacBooks with 64 GB of RAM and deliver quality nearly matching q4. They require the DwarfStar laguna-s2.1 branch (or llama.cpp) to function; DwarfStar benchmarks show about 65 tokens/second generation and roughly 570 tokens/second prefill.

Why it matters

Published mixed q2/q3 quantizations of Laguna S2.1 that run on MacBook systems with 64 GB RAM and claim quality close to q4.

Key details

  • Requires the DwarfStar laguna-s2.1 branch (or llama.cpp) to work; with DwarfStar measured speeds are ~65 tokens/s generation and ~570 tokens/s prefill.
Source evidence

I just published mixed q2/q3 quants of Laguna S2.1 that work in MacBook systems with 64GB of RAM with quality not far from q4. You need DwarfStar laguna-s2.1 brach for them to work, or llama.cpp. With DwarfStar the speed is 65 t/s generation, ~570 t/s prefill.