17 entries · Updated 2026-10-01

Models

Open-weight models worth knowing, sorted by how much VRAM they need at 4-bit. Pick your card size to see what fits.

Your VRAM
video
%!g(uint64=5)B dense · Apache 2.0
plus text encoder and VAE
3 GB
·
chat
%!g(uint64=8)B dense · Apache 2.0
4B effective, the rest are per-layer embeddings
5 GB
·
FLUX.1 [dev]
Black Forest Labs · GGUF ↗
image
%!g(uint64=12)B dense · Non-commercial
plus T5 text encoder
7.5 GB
·
chat
%!g(uint64=12)B dense · Apache 2.0
7.5 GB
·
chat
%!g(uint64=14)B dense · Apache 2.0
8.5 GB
·
chat
%!g(uint64=21)B MoE, 3.6B active · Apache 2.0
13 GB
·
chat
%!g(uint64=26)B MoE, %!g(uint64=4)B active · Apache 2.0
16 GB
·
code
%!g(uint64=27)B dense · Apache 2.0
hybrid attention, only 16 of 64 layers keep a KV cache
16 GB
·
chat
%!g(uint64=31)B dense · Apache 2.0
19 GB
·
chat
%!g(uint64=35)B MoE, %!g(uint64=3)B active · Apache 2.0
21 GB
·
chat
%!g(uint64=117)B MoE, 5.1B active · Apache 2.0
71 GB
·
chat
%!g(uint64=119)B MoE, %!g(uint64=6)B active · Apache 2.0
72 GB
·
code
%!g(uint64=744)B MoE, %!g(uint64=40)B active · MIT
451 GB
·
Kimi K2.6
Moonshot AI · GGUF ↗
code
1.04T MoE, %!g(uint64=32)B active · Modified MIT
631 GB
·
speech
1.55B dense · MIT
1 GB
·
Phi-4-mini
Microsoft · GGUF ↗
chat
3.84B dense · MIT
2.5 GB
·
chat
70.6B dense · Llama Community
43 GB
·

How the VRAM figure is computed

One rule for every model on this site: weights only, at Q4_K_M, which averages 4.85 bits per weight in llama.cpp.

GB = parameters (billions) × 4.85 ÷ 8

Mixture-of-experts models count all their parameters, since every expert has to sit in memory. Rounded to 0.5 GB below 10 GB, to 1 GB above. The context cache and runtime buffers come on top: Will it fit adds them for your context length and cards. Image and video models also need their text encoder.