The local AI paddock · Updated 2026-10-01
Run more
with less.
A short directory for people who run AI on their own hardware. Who to follow, which models fit your card, who builds the tools, and where the community meets next.
Next on the calendar
All events →Fits in 24 GB
All models →Ministral 3 14B
≈9.5 GB at 4-bitMistral AI · 14B dense · Apache 2.0
chatPhi-4-mini
≈3.5 GB at 4-bitMicrosoft · 3.8B dense · MIT
speechWhisper large-v3
≈1.5 GB at 4-bitOpenAI · 1.5B dense · MIT
chatQwen3.6-35B-A3B
≈22 GB at 4-bitAlibaba · 35B MoE, 3B active · Apache 2.0
chatGemma 4 31B
≈20 GB at 4-bitGoogle · 31B dense · Apache 2.0
codeQwen3.6-27B
≈18 GB at 4-bitAlibaba · 27B dense · Apache 2.0
Start your feed here
All people →Georgi Gerganov
@ggerganovCreated llama.cpp and the GGML/GGUF stack that most local inference still runs on.
X buildersAwni Hannun
@awnihannunLeads MLX at Apple. The one to follow if you run models on a Mac.
X buildersPrince Canuma
@Prince_CanumaMaintains MLX-VLM and ports new vision and audio models to Apple silicon within days.
X buildersVaibhav Srivastav
@reach_vbWorks at Hugging Face. Spots new open releases early and shows how to run them.
X buildersSimon Willison
@simonwTries almost every new model the day it lands and writes up what works. His llm CLI is handy.
X researchersTim Dettmers
@Tim_Dettmersbitsandbytes and QLoRA. A lot of what we know about 4-bit on consumer GPUs started with him.
XFrom the garage
All rigs →Follow the paddock
New models, rig tours and benchmarks. Posted on X, filmed for YouTube.
Why this site exists
Everyone running models at home hits the same wall sooner or later. The model you want is a few gigabytes too big for the card you own. So you drop to a smaller quant, trim the context, push a few layers to system RAM, and watch the tokens per second fall.
That wall is where the interesting work happens. People fit 70B models on two used 3090s. Others run image and video pipelines on a laptop. Quantizers publish GGUF files a few hours after a release, and the inference engines get faster every month.
This site keeps track of the people doing that work, the models worth downloading, the companies building for local inference and the events where everyone meets. The list is short on purpose. Each entry is here because it helps someone who runs models on their own machine.
Missing someone? Suggest an entry.