Updated 2026-10-01
Leaderboard
Measured generation speeds on real home hardware. Same method for every run, so the numbers can be compared.
No runs published yet. The first ones come from the editor's own rig, then from yours: the method is below.
How runs are measured
Every result comes from llama-bench, which ships with llama.cpp:
llama-bench -m model.gguf -ngl 99 -fa 1 -p 512 -n 128
- pp512: prompt processing speed on 512 tokens, in tokens/s
- tg128: generation speed over 128 tokens, in tokens/s. This is the number people feel.
Each run lists the exact GGUF file, the llama.cpp build, the power limit and, when measured, the wall power in W.
Send your run
Paste the full llama-bench output and the command line into an email to contact@needforvram.com, or add it to your Garage page.
For large community datasets, see also LocalScore and LocalMaxxing.