17 entries · Updated 2026-10-01
Models
Open-weight models worth knowing, sorted by how much VRAM they need at 4-bit. Pick your card size to see what fits.
How the VRAM figure is computed
One rule for every model on this site: weights only, at Q4_K_M, which averages 4.85 bits per weight in llama.cpp.
GB = parameters (billions) × 4.85 ÷ 8
Mixture-of-experts models count all their parameters, since every expert has to sit in memory. Rounded to 0.5 GB below 10 GB, to 1 GB above. The context cache and runtime buffers come on top: Will it fit adds them for your context length and cards. Image and video models also need their text encoder.