Choosing a GGUF quantization
Most listings here exist in several quantizations. The right one depends on your RAM/VRAM and your quality tolerance.
The short version
- Q4_K_M — the default choice. ~4.7 bits/weight, minor quality loss, runs 8B models in ~6 GB.
- Q5_K_M — noticeably closer to full quality, ~20% larger. Choose it when you have headroom.
- Q6_K / Q8_0 — near-lossless; Q8_0 is the safe archival pick if you'll re-quantize later.
- Q2/Q3 — only when nothing else fits. Degradation is real and visible.
- F16 / BF16 — full precision. For archiving, fine-tuning, or serving from serious hardware.
Sizing rule of thumb
file size ≈ params × bits ÷ 8, plus context memory on top. An 8B model at Q4_K_M is ~4.9 GB and wants ~2 GB more for a long context. Leave your system 2–4 GB of breathing room.
Which to archive?
If you're seeding for posterity rather than running locally, prefer the original safetensors (or Q8_0 where the original is impractical). Every other quant can be regenerated from it; the reverse is not true.