Pit stop · 2026-10-01
Out of memory: what to try, in order
The model does not load, or crashes on a long prompt. Seven fixes, from the cheapest to the most costly.
A model needs room for three things: its weights, the context cache, and some working buffers. When the total is bigger than your VRAM, something has to give. Try these in order.
1. Shorten the context
The context cache grows with every token of context. Going from 32K to 8K can free several GB.
llama-server -m model.gguf -c 8192
2. Quantize the context cache
llama-server -m model.gguf -fa on --cache-type-k q8_0 --cache-type-v q8_0
The cache takes about half the room, with very little quality loss. Quantizing the V cache needs FlashAttention (-fa on).
3. Close what else uses the card
nvidia-smi
A desktop session and a browser can hold 0.5 to 1.5 GB. On a headless machine that memory is yours.
4. Drop one quant level
Q4_K_M to IQ4_XS saves about 12 % of the weights. Below Q3_K_M, quality drops quickly. See Which quant should I pick.
5. Mixture-of-experts: move experts to RAM
llama-server -m model.gguf -ngl 99 --n-cpu-moe 8
6. Offload fewer layers
llama-server -m model.gguf -ngl 30
The remaining layers run on the CPU. It works, but speed falls sharply.
7. With Ollama
OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 OLLAMA_CONTEXT_LENGTH=8192 ollama serve
Not sure where you stand? Will it fit shows the split between weights, cache and buffers for your cards.