26 entries · Updated 2026-10-01

People

Builders, researchers, quantizers and hardware tinkerers worth following. Links go straight to their profiles, with nothing embedded.

builders

Georgi Gerganov

@ggerganov

Created llama.cpp and the GGML/GGUF stack that most local inference still runs on.

X
builders

comfyanonymous

comfyanonymous

Created ComfyUI, the node editor most people use to run image and video models locally.

GitHub
hardware

Anush Elangovan

@AnushElangovan

Runs software at AMD. The place to hear where ROCm support for Radeon cards is heading, and to ask for it.

X
hardware

Digital Spaceport

@DigitalSpaceport

Builds quad and eight-GPU home servers on camera, with parts lists, power draw and long-term reviews.

YouTube
quantizers

Daniel Han

@danielhanchen

Co-founder of Unsloth. Single-GPU fine-tuning, dynamic quants, and frequent fixes for day-one release bugs.

X
researchers

Tim Dettmers

@Tim_Dettmers

bitsandbytes and QLoRA. A lot of what we know about 4-bit on consumer GPUs started with him.

X
builders

Johannes Gäßler

JohannesGaessler

Writes much of the CUDA backend of llama.cpp, including multi-GPU and FlashAttention kernels.

GitHub
builders

Kijai

kijai

ComfyUI wrappers for new video models, Wan included, often days after release and tuned for consumer cards.

GitHub
builders

Awni Hannun

@awnihannun

Leads MLX at Apple. The one to follow if you run models on a Mac.

X
builders

Prince Canuma

@Prince_Canuma

Maintains MLX-VLM and ports new vision and audio models to Apple silicon within days.

X
builders

Vaibhav Srivastav

@reach_vb

Works at Hugging Face. Spots new open releases early and shows how to run them.

X
builders

Simon Willison

@simonw

Tries almost every new model the day it lands and writes up what works. His llm CLI is handy.

X
researchers

Maxime Labonne

@maximelabonne

Model merging, abliteration and a free LLM course that a lot of people learned from.

X
researchers

Teknium

@Teknium1

Nous Research. Hermes fine-tunes and open training data.

X
researchers

Thomas Wolf

@Thom_Wolf

Hugging Face co-founder. Open science, small models and open robotics.

X
quantizers

bartowski

bartowski

Puts out GGUF quants of most new models within hours, in every size from Q2 to Q8.

Hugging Face
quantizers

city96

city96

Brought GGUF to diffusion models with ComfyUI-GGUF. The reason FLUX runs on 8 GB cards.

Hugging Face
hardware

Jeff Geerling

@geerlingguy

Tests GPUs, clusters and odd boards to see what local AI runs on, and publishes the numbers.

X
hardware

Alex Cheema

@alexocheema

Co-founder of exo, which splits one model across several machines on your desk.

X
hardware

Ahmad Osman

@TheAhmadOsman

Builds multi-GPU rigs at home and shares parts lists, power draw and results.

X
hardware

George Hotz

@realGeorgeHotz

tiny corp and tinygrad. Sells multi-GPU boxes in AMD and NVIDIA versions and pushes hard on non-CUDA stacks.

X
hardware

Level1Techs

@Level1Techs

Wendell covers workstation GPUs, PCIe lanes and multi-GPU builds in real depth.

YouTube
educators

Andrej Karpathy

@karpathy

Explains how LLMs work from the ground up. nanoGPT and llm.c make good weekend projects.

X
educators

Sebastian Raschka

@rasbt

Careful long-form breakdowns of new architectures, and the book Build a Large Language Model (From Scratch).

X
communities

r/LocalLLaMA

r/LocalLLaMA

The main forum for local inference. Benchmarks, rig photos and release-day threads.

Reddit
communities

Clem Delangue

@ClementDelangue

CEO of Hugging Face, where nearly every open model gets published.

X