Updated 2026-10-02

Glossary

The words that decide whether a model runs: video memory, bandwidth, quantisation, KV cache, MoE, prefill and decode, explained without hand-waving.

Every argument about local AI is really an argument about memory. This page defines the terms that keep coming back, in the sense the rest of this site uses them.

Memory and capacity

VRAM — the memory attached to the GPU, and the only memory the GPU can read at full speed. Measured on this site in decimal GB, as vendors sell it.

Unified memory — a machine where the CPU and GPU share one pool instead of having separate memory (Apple silicon, AMD’s Strix Halo, NVIDIA’s DGX Spark). Capacity is large and cheap, and the GPU can address only part of it.

Usable memory — on a unified-memory machine, the share the GPU driver will actually allocate to models, typically 75-90% of the total. A 128 GB box is not a 128 GB GPU, and the site shows both numbers.

Memory bandwidth — how fast the GPU can read its memory, in GB/s. It is the single best predictor of generation speed: two cards with the same model and the same quantisation run tokens at a ratio close to their bandwidth ratio. It is not the same as compute, and it is usually the bottleneck. Compare cards on GPUs.

VRAM overhead — the memory a workload uses beyond the weights and the context cache: compute buffers, attention scratch space, CUDA graphs, the KV cache of other concurrent requests. This is why a model that “fits” arithmetically can still fail to load, and why Will it fit? asks for context length instead of only model size.

KV cache — the key and value tensors a transformer stores for every token it has already processed, so it does not recompute the whole context for each new token. It grows linearly with context length, with batch size, and with the number of concurrent sequences, and it is why long context is expensive. Quantising the KV cache to 8 or 4 bits cuts it roughly in half or four.

Context window — the maximum number of tokens the model can attend to at once, from the training configuration. Exceeding it does not error out cleanly; the model degrades. Longer context is not free: it is paid in KV cache.

Sliding-window attention — a design where most layers attend only to a recent window of tokens and a few layers keep the whole context. It makes long context much cheaper on the KV side. The shape matters so much that the site stores, per model, how many layers are full and how many are windowed.

GQA / MQA — grouped-query and multi-query attention: several query heads share one key/value head. Fewer KV heads means a smaller cache. It is why two 8 B models can have very different memory use at the same context length.

CPU offload — keeping some layers in system memory and running them on the CPU. It makes a model load on a smaller GPU, and generation speed collapses to CPU memory bandwidth. Useful to fit, painful to use.

Memory-mapped weights — reading the model file from disk instead of copying it into RAM first. Fast on an NVMe drive, which is why Storage matters as much as the card.

Quantisation

Quantisation — storing weights at lower precision than the training format, to make the model smaller and faster at some cost in quality. The practical range for local inference is 8 bits down to about 2 bits, with quality falling off faster below 4.

Bits per weight (bpw) — the honest unit for comparing quantisations, because names lie. A “Q4” scheme from one project is not the same size as a “Q4” from another.

Q4_K_M — the k-quant mixture at roughly 4.85 bits per weight that most local setups use as a default. It is the reference used across this site for “what fits”, and the figure that turns a parameter count into GB.

GGUF — llama.cpp’s single-file format for weights plus metadata, readable by most local runtimes and Hugging Face’s own tooling. Practically the lingua franca of local inference.

imatrix / IQ quants — importance-matrix-aware quantisation, which measures how each weight actually affects the output and spends precision where it matters. Lower bits for the same quality, slower to produce.

AWQ, GPTQ, EXL2, bitsandbytes — GPU-side quantisation families with their own formats and runtimes. They avoid GGUF’s CPU-oriented layout and are the usual choice inside vLLM or ExLlamaV2, at the cost of format lock-in.

Perplexity — the standard quality metric for a quantisation: lower is better, and a small change is usually invisible in use. Beyond benchmarks, the only honest test is your own prompts.

Mixed-precision KV cache — quantising the context cache (to 8 or 4 bits) while weights stay high precision. It buys context length cheaply, and it is a measurement, not a claim: check it on your own outputs.

Running and serving

Prefill — the first phase of a request, where the whole prompt is processed in parallel. It is compute-bound and usually fast.

Decode — generating tokens one at a time, each requiring a full pass over the weights. It is memory-bandwidth-bound and slow. Almost every speed complaint about local AI is a decode complaint.

TTFT — time to first token, dominated by prefill and by queueing. TPOT — time per output token, dominated by bandwidth. Tokens/s — the reciprocal of the second, and meaningless without stating batch size and context length.

Continuous batching — adding new requests to a running batch instead of waiting for it to finish, which raises total throughput and raises the latency of each individual request. The classic throughput-versus-latency trade, and the reason two people quoting tokens/s can both be right.

Tensor parallelism — splitting each layer across GPUs, which needs fast interconnects because every layer requires an all-reduce. Pipeline parallelism — splitting layers into stages across GPUs, cheaper on interconnect, worse on latency. Expert parallelism — for mixture-of-experts models, putting different experts on different GPUs.

Paged attention, flash attention — implementation techniques for the KV cache and attention: less memory wasted, less bandwidth spent recomputing. They are the difference between vLLM’s throughput and a naive loop.

Speculative decoding — a small model proposes several tokens and the large model verifies them in one pass. It helps most when the large model is bandwidth-bound and the small one is fast.

Models and architectures

Parameters — the learned weights, in billions. Active parameters — how many are used per token in a mixture-of-experts model; for memory, total parameters is what matters.

Dense vs MoE — a dense model uses all weights for every token; a mixture of experts routes each token to a few experts. MoE gives you the quality of a large model with the compute of a small one, and the memory of the large one.

Attention, feed-forward, layers — the repeated blocks that make a transformer. Layer count and hidden size are what make a model “big”.

RoPE and context extension — rotary position encoding is how the model knows token order, and how a model trained at 8k is stretched to 32k or more. Stretching degrades quality; the vendor’s own numbers are not the ones to trust at the edges.

Multimodal / vision-language — a model with an image encoder wired to a language model, so a screenshot can go in the prompt. Vision models spend memory on image tokens, which are cheaper than they look and not free.

Reasoning model — a model trained or tuned to emit deliberation tokens before its answer. Slower and more verbose, sometimes much better on maths and code, and always billed in tokens.

Distillation — training a smaller model to imitate a larger one. This is why an 8 B model from 2026 is not comparable to an 8 B model from 2023.

LoRA / adapter — a small set of extra weights trained on top of a frozen base model, applied at load time. Fine-tuning for behaviour, without shipping a second copy of the model.

Hardware

Backend — the compute stack a runtime uses: CUDA on NVIDIA, ROCm on AMD, Metal on Apple, SYCL or Vulkan for the cross-vendor cases. A card is only as useful as its backend’s maturity for the tool you need.

Clamshell memory — memory packages mounted on both sides of the card’s PCB. The stock RTX 4090 does not support it, which is why a 48 GB 4090 needs a different board: see The 48 GB RTX 4090.

Blower vs axial cooler — a blower pushes air through the card and out of the case, which is what racks need; axial fans are quieter in a desktop and dump heat inside the case. It decides whether a multi-GPU build is buildable at all.

PCIe lanes — a x8 link halves the bandwidth available to the card. It matters more for loading models from host memory and for multi-GPU traffic than for generation.

ECC memory — memory with error correction, standard on datacenter cards. It costs a little bandwidth and buys you trustworthy long runs.

Runtimes — the software that actually loads the model and serves tokens: llama.cpp, vLLM, ExLlamaV2, MLX, TensorRT-LLM and the rest. Compare them on Runtimes, and see what each is good at rather than which is “best”.