Updated 2026-10-01
Will it fit?
Pick your cards, a model, a quant and a context length. See what fits in VRAM, how much room is left, and the speed ceiling of your setup.
Pick your cards and a model.
How it is computed
Weights: parameters × bits per weight of the quant ÷ 8. Mixture-of-experts models count every expert.
Context cache: read from each model's config. Layers that keep the whole context grow with every token. Sliding-window layers stop at their window. Hybrid models only cache their full-attention layers.
Buffers: 0.6 GB per card for the driver and working memory, plus 2 % of the weights.
Speed ceiling: memory bandwidth ÷ bytes read per token (active weights plus the context cache). With several cards the layers run one after the other, so their times add up. Real runs usually land below this ceiling.
Card sizes are taken at their nominal GB, which leaves a few percent of margin. Everything runs in your browser. Nothing is sent anywhere.