44 questions · 13 pages · Updated 2026-10-04

Questions

Every question this site answers, gathered on one page: what fits in VRAM, which quant to pick, whether a drive or a modified card is genuine, how to read a power limit. Each answer links to the page it comes from.

Answers are written from the same measurements as the pages they come from: no advice, no vendor numbers without a label. If a question is missing, send it.

About

Does this site track me?
No. There is no analytics, no cookie, no third-party request: the pages are static and the CSS and JavaScript are inlined. The only external content is a YouTube player you have to click to load.
How are the entries chosen?
By whether the item helps someone run AI on their own hardware. A factual claim carries its source, and vendor marketing numbers are labelled as such.
I am listed here and want out. How?
Write to contact@needforvram.com and the entry comes out. The same address takes corrections.

The 48 GB RTX 4090

Does a 48 GB RTX 4090 exist from NVIDIA?
No. NVIDIA never sold one. Every 48 GB 4090 is a 24 GB card with a replacement PCB and extra memory modules, built by third-party shops.
What does the modification actually cost?
From about $142 for a bare clamshell PCB kit with GDDR6X modules around $24 each, up to roughly $3,320-3,500 for a finished card in China and $3,750-4,100 from a US reseller, against a $6,800 reference for a genuine 48 GB RTX 6000 Ada.
What goes wrong with a modified card?
Intermittent failures hours into a CUDA load, bare-board scams where the PCB arrives with no GPU or memory fitted, unknown provenance with no NVIDIA warranty, features lost to a custom BIOS, and repair refusal on an already-modified card.
How do I check one before paying?
Run a compute benchmark instead of trusting the label or the vBIOS string, confirm the memory type and module count, and ask for the burn-in report and diagnostic sheet. The shops listed above publish theirs.

Leaderboard

How are the speed numbers measured?
With llama-bench from llama.cpp, on the same command line for every run, so cards, quants and builds can be compared. Each entry lists the exact GGUF file, the build and the power limit.
What is the difference between pp512 and tg128?
pp512 is prompt processing on 512 tokens, in tokens per second; tg128 is generation over 128 tokens. Generation is the number people feel at the keyboard.
How do I submit a run?
Paste the full llama-bench output and the command line to contact@needforvram.com, or put it on your own Garage page. Runs without the command line are not comparable and are not listed.

Models

How much VRAM does a 27B model need at 4-bit?
Roughly 0.57 GB per billion parameters at Q4_K_M, so about 15 to 16 GB of weights for 27B, plus the context cache: a few more GB at 32K with a q8_0 cache. Set your own card, quant and context in /will-it-fit/.
Why is this list sorted by VRAM?
Because memory is the constraint that decides whether a model runs at all on consumer hardware: size is the first filter, capability the second.
Where do the parameter counts come from?
From the safetensors metadata in the model’s own Hugging Face repository, not from the prose on the model card.

Out of memory: what to try, in order

Why does a model that fits still fail on a long prompt?
VRAM holds three things at once: the weights, the context cache and some working buffers. The cache grows with every token of context, so a model that loads easily at 8K can fail at 32K.
What is the first fix to try when a model will not load?
Shorten the context. Going from 32K to 8K can free several GB and it is the cheapest change there is: llama-server -m model.gguf -c 8192.
Can the context cache be compressed without hurting quality?
Yes. --cache-type-k q8_0 --cache-type-v q8_0 with -fa on roughly halves the cache for very little quality loss. Quantising the V cache needs FlashAttention turned on.
How do I run a model that is bigger than my VRAM?
Offload fewer layers (-ngl 30), or for a mixture-of-experts model keep the experts in RAM (--n-cpu-moe 8) and let the GPU hold the rest. It runs, but generation speed falls sharply.

Which quant should I pick?

What does the number in a quant name actually mean?
Roughly the bits per weight. Q4_K_M stores each weight in about 4.85 bits and the file is about 57% of the Q8_0 size; the table above lists the measured ratios.
Which quant should I pick by default?
Q4_K_M. It is the usual default and the one this site’s VRAM figures assume. Step up to Q6_K when the card has room, and stay above Q3_K_M on small models, where the loss becomes visible.
Is a bigger model at a low quant better than a smaller model at a high quant?
Usually yes: a larger model at Q4 beats a smaller one at Q8 of the same file size. Models under about 8B are the ones that suffer most from aggressive quantisation.
How do I know the file I downloaded is really that quant?
Check the real file size against the table. The quant name in the file name is a label; the byte count is the fact.

Runtimes

Which engine should I start with?
One card up to 24 GB, or CPU only: GGUF with llama.cpp, or a UI built on it such as Ollama, LM Studio or KoboldCpp, with a quantised KV cache and partial offload.
Do I need vLLM?
Only when you serve several users with long contexts, or hold full-precision weights across two to four cards and up. That is where batching and tensor parallelism pay for themselves; for one user llama.cpp is simpler and often faster.
What makes a coding agent heavy on memory?
It re-sends a growing transcript every turn, so its cost is a long-context problem rather than a model-size problem. The KV cache, not the weights, is what you run out of.

Choosing a drive for models and datasets

Is the advertised speed the speed I will get?
No. The spec sheet is a peak number. What decides a drive for models and datasets is sustained write speed after the cache empties, random read on datasets, and how the controller behaves once the drive is full.
TLC or QLC?
TLC for anything written repeatedly. QLC only for read-mostly model storage, where you accept the sustained-write cliff in exchange for capacity.
Are used datacenter drives a good idea?
They are the honest sweet spot when the numbers check out: enterprise U.2 drives with real DRAM and documented wear, verified with the same tests you would run on a new drive.

How drives get faked

What are the kinds of faked drive?
Four: fake capacity (the drive lies about its size), counterfeit branded drives (a real-looking product that is not the real thing), used drives sold as new (the biggest category), and marketing fakes, where the drive is genuine and the numbers are not.
Where do fake drives actually arrive from?
Third-party marketplace listings at volume, eBay-style sales with hollow enclosures, cheap cross-border marketplaces, “Renewed” listings, and high-capacity enterprise drives bought off unofficial channels.
Which checks still work?
Buy from the brand’s authorised sellers, treat a price far below market as a scam rather than a bargain, check the serial with the manufacturer before writing data, and test the whole surface before trusting the drive with anything irreplaceable.
How large is the problem in practice?
Measured rather than anecdotal: a 20% complaint rate about faked storage on Amazon USB drives, and a “16 TB SSD” that turned out to be a board with a 60 GB microSD card and weights glued in to feel heavy.

Test a new drive before you trust it

What should I check on a new drive before writing anything to it?
The SMART baseline: power-on hours, power-cycle count, reallocated sectors, current pending sectors, UDMA CRC errors and total LBAs written. A genuinely new drive shows near zero on all of them.
Is a benchmark enough to prove a drive is genuine?
No. Counterfeits that pass a capacity check also pass short benchmarks, and the FARM log has been altered in the wild. Verify the capacity, write the whole surface and read it back, then keep the baseline for later comparison.
How can I tell a used drive sold as new?
Per-head operating hours inside FARM (they survive a global counter reset), a manufacturing date more than about six months before your purchase date, missing or mismatched serial stickers, scuffs on a drive sold as sealed, and a negotiated interface generation lower than the model specifies.
Does the warranty protect me if the drive turns out to be used?
Only partly. The serial lookup decides everything, OEM and bulk drives carry no retail warranty, and a “Renewed” listing is a seller guarantee of 90 to 365 days, not a manufacturer term.

Two different GPUs, one model

Can two different GPUs run one model together?
Yes. llama.cpp splits the layers across them, and you can set the proportions yourself: --tensor-split 2,1 matches a 24 GB card and a 12 GB card instead of splitting evenly.
Why does my speed change after a reboot?
The card order changed, so the same split landed differently. Keep the card numbering stable so the proportions you measured are the proportions you get.
Can I run a model larger than my total VRAM?
Yes: the remainder goes to RAM or the CPU and you pay for it in speed. For mixture-of-experts models the cheaper trick is keeping the experts on the CPU.

Power-limit your card, keep your speed

Does lowering the power limit slow generation down?
Far less than expected. Token generation tracks memory bandwidth more than power, so 70 to 80% of the default limit often costs very little generation speed. Prompt processing is compute work and does slow down more.
How do I set a power limit and make it survive a reboot?
sudo nvidia-smi -pm 1 then sudo nvidia-smi -i 0 -pl 280, plus a small systemd unit that reapplies the limit at boot. Measure tg128 and pp512 with llama-bench before and after.
Which number should I optimise?
tg128, the generation speed: it is the number people feel. Keep the limit that gives up little tg128 for a clear drop in watts.

Will it fit?

How does the calculator decide whether a model fits?
It adds the weights at the chosen quant, the context cache at the chosen length and cache quant, and a working buffer, then compares the total with the VRAM of the cards you pick. It is the arithmetic from /pit-stop/out-of-memory/, run in reverse.
Why does my card show less VRAM than the box says?
Because a graphical session, a browser and the driver reserve some of it: 0.5 to 1.5 GB. The calculator works with the memory you can actually give to the model.
Which quant does the site assume?
Q4_K_M, the usual default. Change it in the picker if you run a different quant, or a different context length.