26 entries · Updated 2026-10-01
People
Builders, researchers, quantizers and hardware tinkerers worth following. Links go straight to their profiles, with nothing embedded.
Georgi Gerganov
@ggerganovCreated llama.cpp and the GGML/GGUF stack that most local inference still runs on.
X builderscomfyanonymous
comfyanonymousCreated ComfyUI, the node editor most people use to run image and video models locally.
GitHub hardwareAnush Elangovan
@AnushElangovanRuns software at AMD. The place to hear where ROCm support for Radeon cards is heading, and to ask for it.
X hardwareDigital Spaceport
@DigitalSpaceportBuilds quad and eight-GPU home servers on camera, with parts lists, power draw and long-term reviews.
YouTube quantizersDaniel Han
@danielhanchenCo-founder of Unsloth. Single-GPU fine-tuning, dynamic quants, and frequent fixes for day-one release bugs.
X researchersTim Dettmers
@Tim_Dettmersbitsandbytes and QLoRA. A lot of what we know about 4-bit on consumer GPUs started with him.
X buildersJohannes Gäßler
JohannesGaesslerWrites much of the CUDA backend of llama.cpp, including multi-GPU and FlashAttention kernels.
GitHub buildersKijai
kijaiComfyUI wrappers for new video models, Wan included, often days after release and tuned for consumer cards.
GitHub buildersAwni Hannun
@awnihannunLeads MLX at Apple. The one to follow if you run models on a Mac.
X buildersPrince Canuma
@Prince_CanumaMaintains MLX-VLM and ports new vision and audio models to Apple silicon within days.
X buildersVaibhav Srivastav
@reach_vbWorks at Hugging Face. Spots new open releases early and shows how to run them.
X buildersSimon Willison
@simonwTries almost every new model the day it lands and writes up what works. His llm CLI is handy.
X researchersMaxime Labonne
@maximelabonneModel merging, abliteration and a free LLM course that a lot of people learned from.
X researchersTeknium
@Teknium1Nous Research. Hermes fine-tunes and open training data.
X researchersThomas Wolf
@Thom_WolfHugging Face co-founder. Open science, small models and open robotics.
X quantizersbartowski
bartowskiPuts out GGUF quants of most new models within hours, in every size from Q2 to Q8.
Hugging Face quantizerscity96
city96Brought GGUF to diffusion models with ComfyUI-GGUF. The reason FLUX runs on 8 GB cards.
Hugging Face hardwareJeff Geerling
@geerlingguyTests GPUs, clusters and odd boards to see what local AI runs on, and publishes the numbers.
X hardwareAlex Cheema
@alexocheemaCo-founder of exo, which splits one model across several machines on your desk.
X hardwareAhmad Osman
@TheAhmadOsmanBuilds multi-GPU rigs at home and shares parts lists, power draw and results.
X hardwareGeorge Hotz
@realGeorgeHotztiny corp and tinygrad. Sells multi-GPU boxes in AMD and NVIDIA versions and pushes hard on non-CUDA stacks.
X hardwareLevel1Techs
@Level1TechsWendell covers workstation GPUs, PCIe lanes and multi-GPU builds in real depth.
YouTube educatorsAndrej Karpathy
@karpathyExplains how LLMs work from the ground up. nanoGPT and llm.c make good weekend projects.
X educatorsSebastian Raschka
@rasbtCareful long-form breakdowns of new architectures, and the book Build a Large Language Model (From Scratch).
X communitiesr/LocalLLaMA
r/LocalLLaMAThe main forum for local inference. Benchmarks, rig photos and release-day threads.
Reddit communitiesClem Delangue
@ClementDelangueCEO of Hugging Face, where nearly every open model gets published.
X