Read more here, I think this is how we get to good, fast 2-3 bit models. This method has the lowest kld compared to any other compression method.
github.com/turboderp-org/exl…
Link
exllamav3/doc/exl3.md at master · turboderp-org/exllamav3
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs - turboderp-org/exllamav3
github.com