Updated 2026-10-02
Consulting
Sizing a machine for a workload, or paying too much for inference? Write in and get introduced to people who do this for a living.
What this is
This site is a directory, not a company. I do not sell hardware, I do not take a commission on parts, and I am not an integrator.
Almost every email that arrives here is one of two problems. Someone has to size a machine for a workload and needs an answer they can defend to a boss. Someone else is already running models and the bill, the power draw or the latency is wrong.
What I do with it
Read the email, and if it is that kind of problem, introduce you to a specialist who works on exactly that: AI solution architecture, and performance and cost optimisation. People who do this for a living, whose work I have looked at myself.
No invoice from me, no cut of the deal.
What to send
One email, short:
- the workload: models, sizes, concurrency, and whether you are constrained by latency or by throughput
- what you have today: GPUs, servers, cloud spend per month, what it actually costs you
- the hard constraint: budget, power, rack space, data that cannot leave the building, deadline
A paragraph is enough. A screenshot of the bill and an nvidia-smi output say more than a page of prose.
Typical asks
- which cards, how much VRAM, and which serving engine for a given model and traffic pattern
- cost per million tokens today versus on owned hardware, including power and depreciation
- moving a workload off a cloud account without breaking latency or compliance
- quantisation, context length and KV cache: what to trade against what
- storage and network sized for multi-hundred-gigabyte models and many parallel users
What this is not
- not a sales channel: nobody here will pitch you a pre-packaged “AI solution”
- not a hosting service: this site has no cloud account, no referral links and no affiliate codes
- not free engineering: the specialist bills you for their work, and you agree the scope and the price with them, not with me
contact@needforvram.com