Pit stop · 2026-10-01

Which quant should I pick?

Q4_K_M, IQ4_XS, Q6_K: what the names mean and where quality starts to drop.

A quant stores each weight with fewer bits. Fewer bits mean a smaller file and faster generation, at the price of some accuracy. The number in the name is roughly the bits per weight.

QuantBits per weightSize vs Q8_0When to use it
Q8_08.5100 %Near lossless. When you have room to spare.
Q6_K6.5677 %Hard to tell from Q8_0.
Q5_K_M5.6967 %Safe choice for small models.
Q4_K_M4.8557 %The usual default. This site’s VRAM figures use it.
IQ4_XS4.2550 %Close to Q4_K_M, a bit smaller.
Q3_K_M3.9146 %Visible loss on small models.
IQ2_XXS2.0624 %Last resort, only for very large models.

Rules of thumb

  • A bigger model at Q4 usually beats a smaller model at Q8 of the same file size.
  • Small models (under about 8B) suffer more from low quants than large ones.
  • For code and maths, stay at Q5 or above when you can.
  • Prefer quants made with an importance matrix (often marked “imatrix” or “i1”), and the dynamic quants from Unsloth (marked “UD”). They keep the sensitive layers at higher precision.

Check the real file size

The GGUF page on Hugging Face lists the size of each file. Add the context cache on top: Will it fit does it for you.