Tuning · 2026-10-01

Two different GPUs, one model

How llama.cpp splits a model across mismatched cards, and what it means for speed.

A 3080 Ti next to a 3060, a 3090 next to a 4060 Ti: mixed pairs are common in home rigs. llama.cpp handles them well, as long as you know how it splits the work.

How the split works

By default llama.cpp uses --split-mode layer. Each card gets a block of consecutive layers, with the context cache for those layers. For every new token, the data goes through the layers on card 0, then through the layers on card 1. The cards take turns, they do not work in parallel.

So the time per token is the time spent on the first card plus the time spent on the second. A slow card holding half the layers costs you about half its slowness.

Control the proportions

llama-server -m model.gguf -ngl 99 -c 16384 -fa on --tensor-split 1,1
  • -ngl 99 puts all layers on the GPUs
  • --tensor-split 1,1 gives each card the same share. With a 24 GB card and a 12 GB card, 2,1 matches their sizes.
  • -fa on turns on FlashAttention, which also saves memory

Give the faster card as many layers as its VRAM allows. When both cards are full anyway, the split is decided by memory, not by speed.

Keep the card numbers stable

CUDA can number cards differently from nvidia-smi. Set this before launching:

export CUDA_DEVICE_ORDER=PCI_BUS_ID

At load time, llama.cpp prints one buffer size per card (CUDA0 model buffer size, CUDA1 ...). Check that the split matches what you asked for.

Models bigger than your VRAM

For mixture-of-experts models, keep the expert weights of some layers in system RAM and everything else on the GPUs:

llama-server -m model.gguf -ngl 99 --n-cpu-moe 10

Raise the number until the model loads. It is much faster than offloading whole layers.