Storage · 2026-10-02
Choosing a drive for models and datasets
TLC versus QLC, DRAM versus HMB, the sustained-write cliff, used datacenter U.2 drives: which numbers about an SSD survive contact with a local AI workload.
A drive for local AI has three jobs, and they pull in different directions.
- Load huge model files. A 96 GB of weights has to come off the disk before the GPU sees it. Zero-copy formats like safetensors memory-map the file, so the first pass is bound by sequential read bandwidth and free capacity.
- Feed datasets. This is random, small-block traffic, not sequential. MLPerf’s storage benchmark reads roughly 315 KiB samples one at a time in random order, about 500 files per second per simulated accelerator for one of its workloads, and needs around 5.9 GiB/s of read bandwidth to keep a single simulated B200 fed with large sequential streams (MLPerf Storage).
- Survive writes. Checkpoints, logs, model swaps. This is where cheap drives fall apart, and MLPerf runs explicit checkpoint workloads for exactly that reason.
The spec sheet is a peak number, not a speed
The headline “up to 7,450 MB/s” is a burst figure, valid while the pseudo-SLC cache holds. Tom’s Hardware measured a 2 TB Crucial P310 writing at 6.3 GB/s for about 64 seconds — call it a 400 GB cache — before collapsing well below SATA speeds, and the effective cache shrinks as the drive fills (Tom’s Hardware). Sustained write on that class of drive was measured at 139 MB/s once the cache emptied.
So: judge writes by the post-cache rate, judge reads by 4 KiB random IOPS and queue-depth-1 latency, and ignore any number that does not say whether it is a peak or a sustained figure.
Two more traps:
- GB versus GiB. Vendors sell decimal GB (10^9 bytes), the OS may show binary GiB (2^30), so a 500 GB drive shows about 465.7 GiB — roughly 7.4% less (Wikipedia). Size your model library accordingly, on the OS number.
- TBW versus DWPD. TBW is total bytes written under warranty and scales with capacity, so the useful form is TBW per TB (600 TBW on a 1 TB drive implies about 600 write cycles). DWPD is full-drive writes per day over a usually five-year enterprise warranty, and it is the only capacity-independent comparison (KIOXIA).
NAND and controller: where the money is decided
TLC versus QLC. Three bits per cell versus four. QLC is cheaper and denser, but its write ceiling and endurance are lower. Datacenter QLC is engineered for read-heavy duty with read response close to TLC; consumer QLC with a small pSLC cache is not that.
DRAM versus HMB. A DRAM-less controller keeps its flash-translation table in host memory via Host Memory Buffer, which partly rescues reads. It has no power-loss protection, and on consumer drives it is the budget tier. Do not treat it as equivalent for write-heavy work (Wikipedia).
Controller and firmware. The controller runs the firmware that governs garbage collection, wear levelling and cache policy. Captive controllers (Samsung, Kioxia, Micron, Solidigm/SK hynix) are tuned to their own NAND; merchant parts (Phison, Silicon Motion, Marvell) range from excellent to budget depending on firmware. The same NAND on a different controller behaves differently.
Thermals. A bare controller at full speed has been modelled at about 145 °C versus about 133 °C with an 8 mm heatsink (ATP). Long reads and writes trigger firmware throttling, so a heatsink and airflow are performance parts, and datacenter U.2 drives need real airflow even at idle load.
Over-provisioning. Every drive hides spare area: a 500 GB drive exposes about 465.7 GiB, roughly 7.37% overprovisioned, which the controller needs for wear levelling. A drive that is 95% full has far less effective spare and slows down. Leave room on the drive that holds your models.
Used datacenter drives: the honest sweet spot
Second-hand U.2/E1.S enterprise SSDs are frequently barely used. In ServeTheHome’s production sample, 88% of drives had under 150 TBW written, and a guide based on 78 eBay drives found an average under 3 TB written (ServeTheHome). You get power-loss protection, 1 DWPD endurance and stable sustained performance for a fraction of new price.
The costs are real: 15 mm 2.5-inch U.2 or E1.S needs a backplane or adapter, they throttle without airflow, and some OEM-branded units expose only a partial SMART attribute set — which matters if you want to monitor wear.
Tiers, with a real drive in each
| Tier | What it is | Example |
|---|---|---|
| Budget DRAM-less QLC M.2 | Cheap, dense, fine as read-only model storage; poor for checkpointing | Crucial P310 2 TB (QLC, Phison E27T, DRAM-less, ~400 GB pSLC) |
| Mainstream consumer TLC with DRAM | The sensible default: TLC, DRAM cache, ~600 TBW per TB, wants a heatsink under long I/O | Samsung 990 Pro 2 TB (176-layer TLC, LPDDR4 DRAM, 5-year warranty) |
| Datacenter read-intensive, bought used | PLP, 1 DWPD, stable sustained performance; needs a U.2/E1.S host and airflow | KIOXIA CD8-R (PCIe 4.0 x4, TLC, up to 15.36 TB, up to 1,250K 4 KiB random-read IOPS) |
| Write-intensive SLC | For constant checkpointing where endurance is the constraint; small and expensive per TB | Solidigm D7-P5810 (SLC, up to 50 DWPD, up to 1.6 TB) |
| Max-capacity QLC | Very large model and dataset libraries at low cost per TB, read-mostly | Solidigm D5-P5336 (QLC, up to 122.88 TB, U.2 15 mm / E3.S / E1.L) |
Capacities in the last two rows are the vendors’ own maxima, from the product pages linked above; treat any “up to” figure, including those, as the top of the range rather than the model you are buying.
What this means in practice
- Store the model library on the biggest read-mostly device you can afford; it is a read workload.
- Put checkpoints, active datasets and anything write-heavy on TLC with DRAM, or on a used enterprise drive with PLP.
- Buy the heatsink. It is cheaper than the throttling.
- Keep 10-20% of the drive free, and do not fill a QLC drive and then wonder about the speed.
- Verify the link width and generation before blaming the drive for being slow.
- And check the drive is real before you fill it: How drives get faked, Test a new drive.