Pit stop · 2026-10-01

Out of memory: what to try, in order

The model does not load, or crashes on a long prompt. Seven fixes, from the cheapest to the most costly.

A model needs room for three things: its weights, the context cache, and some working buffers. When the total is bigger than your VRAM, something has to give. Try these in order.

1. Shorten the context

The context cache grows with every token of context. Going from 32K to 8K can free several GB.

llama-server -m model.gguf -c 8192

2. Quantize the context cache

llama-server -m model.gguf -fa on --cache-type-k q8_0 --cache-type-v q8_0

The cache takes about half the room, with very little quality loss. Quantizing the V cache needs FlashAttention (-fa on).

3. Close what else uses the card

nvidia-smi

A desktop session and a browser can hold 0.5 to 1.5 GB. On a headless machine that memory is yours.

4. Drop one quant level

Q4_K_M to IQ4_XS saves about 12 % of the weights. Below Q3_K_M, quality drops quickly. See Which quant should I pick.

5. Mixture-of-experts: move experts to RAM

llama-server -m model.gguf -ngl 99 --n-cpu-moe 8

6. Offload fewer layers

llama-server -m model.gguf -ngl 30

The remaining layers run on the CPU. It works, but speed falls sharply.

7. With Ollama

OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 OLLAMA_CONTEXT_LENGTH=8192 ollama serve

Not sure where you stand? Will it fit shows the split between weights, cache and buffers for your cards.