Tuning · 2026-10-01
Power-limit your card, keep your speed
Generation speed depends on memory bandwidth far more than on GPU power. A lower power limit often costs very little.
When a model writes text, each new token means reading the whole set of active weights from VRAM. The GPU cores spend much of that time waiting for memory. That is why token generation tracks memory bandwidth (GB/s) much more closely than it tracks power or clock speed.
Prompt processing is different. Reading a long prompt is compute work, and it slows down more when you cut the power.
Check your current limits
nvidia-smi -q -d POWER
This shows the current, default, minimum and maximum power limit of each card.
Set a lower limit
sudo nvidia-smi -pm 1
sudo nvidia-smi -i 0 -pl 280
The first line turns on persistence mode, the second caps GPU 0 at 280 W. Use -i 1 for the second card. Start around 70 to 80 % of the default limit, then measure.
Measure before and after
llama-bench -m model.gguf -ngl 99 -p 512 -n 128
Compare tg128 (generation) and pp512 (prompt) at each limit. A wall power meter tells you what you save in W. Keep the limit that gives up little tg128 for a clear drop in watts.
Make it survive a reboot
The limit resets at boot. A small systemd unit sets it again (create it with sudo vi /etc/systemd/system/gpu-power-limit.service):
[Unit]
Description=GPU power limit
After=nvidia-persistenced.service
[Service]
Type=oneshot
ExecStart=/usr/bin/nvidia-smi -i 0 -pl 280
[Install]
WantedBy=multi-user.target
Then sudo systemctl enable --now gpu-power-limit.service.
AMD cards
rocm-smi --showpower
sudo rocm-smi -d 0 --setpoweroverdrive 250
Same idea: lower the cap, run the benchmark, keep what works.