64 entries · Updated 2026-10-02
Runtimes
The inference engines that actually run models on your own hardware: local runtimes, serving engines, typed-decision models, diffusion, speech and the interfaces on top.
llama.cpp
The C/C++ core most local AI runs on: GGUF, llama-server, llama-bench. Runs on CPU, CUDA, ROCm, Vulkan, SYCL and Metal.
github.com/ggml-org/llama.cpp </a> <a class="card rt" href="https://github.com/ikawrakow/ik_llama.cpp" rel="noopener noreferrer" data-cat="runtime"> runtime3,273ik_llama.cpp
Fork of llama.cpp with extra quants (IQK) and tweaks that often win on CPU and mixed CPU/GPU rigs.
github.com/ikawrakow/ik_llama.cpp </a> <a class="card rt" href="https://github.com/ollama/ollama" rel="noopener noreferrer" data-cat="runtime"> runtime182,049Ollama
One command to pull and run a model, with an OpenAI-compatible API. Easiest way in, llama.cpp underneath.
github.com/ollama/ollama </a> <a class="card rt" href="https://github.com/lmstudio-ai/lms" rel="noopener noreferrer" data-cat="runtime"> runtime5,329LM Studio
Desktop app plus a local server: pick a GGUF, set GPU offload, get an OpenAI endpoint. No terminal needed.
github.com/lmstudio-ai/lms </a> <a class="card rt" href="https://github.com/LostRuins/koboldcpp" rel="noopener noreferrer" data-cat="runtime"> runtime11,924KoboldCpp
Single binary with its own web UI: GGUF chat, image generation and TTS, CUDA/ROCm/Vulkan builds.
github.com/LostRuins/koboldcpp </a> <a class="card rt" href="https://github.com/mozilla-ai/llamafile" rel="noopener noreferrer" data-cat="runtime"> runtime26,154llamafile
Ships a model and the runtime in one executable file that runs on Linux, macOS and Windows.
github.com/mozilla-ai/llamafile </a> <a class="card rt" href="https://github.com/nomic-ai/gpt4all" rel="noopener noreferrer" data-cat="runtime"> runtime77,389GPT4All
Desktop app for running GGUF models locally. Repository has been quiet since May 2025.
github.com/nomic-ai/gpt4all </a> <a class="card rt" href="https://github.com/mudler/LocalAI" rel="noopener noreferrer" data-cat="runtime"> runtime49,369LocalAI
Self-hosted multi-modal engine: text, vision, voice and image models behind OpenAI-compatible endpoints.
github.com/mudler/LocalAI </a> <a class="card rt" href="https://github.com/containers/ramalama" rel="noopener noreferrer" data-cat="runtime"> runtime3,066RamaLama
Serves models from OCI containers, so the runtime and its drivers travel with the model.
github.com/containers/ramalama </a> <a class="card rt" href="https://github.com/oobabooga/textgen" rel="noopener noreferrer" data-cat="runtime"> runtime47,721Text generation web UI
Desktop app (formerly oobabooga) for text, vision and tool calling, with an OpenAI-compatible API.
github.com/oobabooga/textgen </a> <a class="card rt" href="https://github.com/ml-explore/mlx-lm" rel="noopener noreferrer" data-cat="runtime"> runtime7,198MLX-LM
Run and quantize LLMs on Apple silicon through MLX. The default path on unified-memory Macs.
github.com/ml-explore/mlx-lm </a> <a class="card rt" href="https://github.com/janhq/jan" rel="noopener noreferrer" data-cat="runtime"> runtime44,759Jan
Offline desktop assistant that can also expose a local OpenAI-compatible server.
github.com/janhq/jan </a> <a class="card rt" href="https://github.com/AtomicBot-ai/Atomic-Chat" rel="noopener noreferrer" data-cat="runtime"> runtime1,658Atomic Chat
Local AI app with its own inference engine, built for agents and multi-model workflows.
github.com/AtomicBot-ai/Atomic-Chat </a> <a class="card rt" href="https://github.com/tetherto/qvac" rel="noopener noreferrer" data-cat="runtime"> runtime657qvac
On-device AI SDK: GGUF models, RAG, images and music with no cloud and no API keys.
github.com/tetherto/qvac </a> <a class="card rt" href="https://github.com/cactus-compute/cactus" rel="noopener noreferrer" data-cat="runtime"> runtime6,084Cactus
Quantization, kernels and runtime aimed at phones, wearables and robots rather than desktops.
github.com/cactus-compute/cactus </a> <a class="card rt" href="https://github.com/vllm-project/vllm" rel="noopener noreferrer" data-cat="serving"> serving93,059vLLM
The reference serving engine: paged KV cache, continuous batching, tensor parallelism. Strong on both CUDA and ROCm.
github.com/vllm-project/vllm </a> <a class="card rt" href="https://github.com/sgl-project/sglang" rel="noopener noreferrer" data-cat="serving"> serving36,716SGLang
Serving framework with a prefix cache (RadixAttention) built for multi-turn, agents and structured output.
github.com/sgl-project/sglang </a> <a class="card rt" href="https://github.com/NVIDIA/TensorRT-LLM" rel="noopener noreferrer" data-cat="serving"> serving14,756TensorRT-LLM
NVIDIA's engine: compiles the model into optimized kernels and graphs. Fastest on recent GeForce and datacenter GPUs.
github.com/NVIDIA/TensorRT-LLM </a> <a class="card rt" href="https://github.com/huggingface/text-generation-inference" rel="noopener noreferrer" data-cat="serving"> serving10,884Text Generation Inference
Hugging Face's serving stack. Still used in production, but development has slowed since early 2026.
github.com/huggingface/text-generation-inference </a> <a class="card rt" href="https://github.com/turboderp-org/exllamav3" rel="noopener noreferrer" data-cat="serving"> serving1,571ExLlamaV3
Quantization and inference library tuned for one or two consumer GPUs, successor to ExLlamaV2.
github.com/turboderp-org/exllamav3 </a> <a class="card rt" href="https://github.com/turboderp-org/exllamav2" rel="noopener noreferrer" data-cat="serving"> serving4,633ExLlamaV2
The exl2 engine that made big models fit on 24 GB cards. Superseded by ExLlamaV3, quiet since March 2026.
github.com/turboderp-org/exllamav2 </a> <a class="card rt" href="https://github.com/mlc-ai/mlc-llm" rel="noopener noreferrer" data-cat="serving"> serving23,202MLC-LLM
Compiles models for a target device, from phones to servers, using TVM. Same stack as in-browser inference.
github.com/mlc-ai/mlc-llm </a> <a class="card rt" href="https://github.com/mlc-ai/web-llm" rel="noopener noreferrer" data-cat="serving"> serving19,214WebLLM
Runs quantized models inside the browser on WebGPU. No server, no install, model cached locally.
github.com/mlc-ai/web-llm </a> <a class="card rt" href="https://github.com/bentoml/OpenLLM" rel="noopener noreferrer" data-cat="serving"> serving12,552OpenLLM
BentoML's packaging layer: turn any open model into an OpenAI-compatible endpoint with a few lines.
github.com/bentoml/OpenLLM </a> <a class="card rt" href="https://github.com/ai-dynamo/dynamo" rel="noopener noreferrer" data-cat="serving"> serving8,206Dynamo
Distributed serving framework that splits prefill and decode across machines. Datacenter scale, Rust core.
github.com/ai-dynamo/dynamo </a> <a class="card rt" href="https://github.com/lightseekorg/tokenspeed" rel="noopener noreferrer" data-cat="serving"> serving2,189TokenSpeed
Recent high-throughput serving engine, focused on squeezing latency out of the decode path.
github.com/lightseekorg/tokenspeed </a> <a class="card rt" href="https://github.com/dphnAI/sonar" rel="noopener noreferrer" data-cat="serving"> serving1,869Sonar
Large-scale LLM inference engine, aimed at clusters rather than single rigs.
github.com/dphnAI/sonar </a> <a class="card rt" href="https://github.com/trymirai/uzu" rel="noopener noreferrer" data-cat="serving"> serving1,820uzu
Rust inference engine for AI models, portable across CPU and GPU backends.
github.com/trymirai/uzu </a> <a class="card rt" href="https://github.com/lucasjinreal/Crane" rel="noopener noreferrer" data-cat="serving"> serving488Crane
Pure Rust engine for LLM, VLM, TTS and OCR, built on Candle. Pitched as a simpler alternative to llama.cpp.
github.com/lucasjinreal/Crane </a> <a class="card rt" href="https://github.com/microsoft/sarathi-serve" rel="noopener noreferrer" data-cat="serving"> serving529Sarathi-Serve
Research serving engine built around chunked prefill for lower tail latency. Dormant since January 2026.
github.com/microsoft/sarathi-serve </a> <a class="card rt" href="https://github.com/Tiiny-AI/PowerInfer" rel="noopener noreferrer" data-cat="serving"> serving9,813PowerInfer
Serves local models by keeping hot neurons on the GPU and offloading the rest. Quiet since May 2026.
github.com/Tiiny-AI/PowerInfer </a> <a class="card rt" href="https://github.com/openvinotoolkit/openvino" rel="noopener noreferrer" data-cat="serving"> serving10,943OpenVINO
Intel's toolkit for CPU, integrated GPU and NPU inference. The practical path on non-NVIDIA hardware.
github.com/openvinotoolkit/openvino </a> <a class="card rt" href="https://github.com/alibaba/MNN" rel="noopener noreferrer" data-cat="serving"> serving16,166MNN
Alibaba's lightweight engine for mobile and edge devices, with its own quantized format.
github.com/alibaba/MNN </a> <a class="card rt" href="https://github.com/kvcache-ai/ktransformers" rel="noopener noreferrer" data-cat="serving"> serving19,555KTransformers
Runs large MoE models by keeping attention on the GPU and experts in CPU memory. Built for big sparse models on small VRAM.
github.com/kvcache-ai/ktransformers </a> <div class="card rt" data-cat="decision"> decisionJEV (System One)
TypeSafe AI's decision models: a typed question in, a typed decision out (choice, score, boolean) with a confidence. No text generation.
The category is decision, not chat: score candidate branches off one shared prefix.
typesafe.ai </div> <a class="card rt" href="https://github.com/feder-cr/jev" rel="noopener noreferrer" data-cat="decision"> decision1,175jevos
Open-source alternative to Jev for yes/no decisions. C++, GGUF, runs on a laptop CPU behind a FastAPI server.
github.com/feder-cr/jev </a> <a class="card rt" href="https://github.com/nokia-applied-research/AnyJev" rel="noopener noreferrer" data-cat="decision"> decision1,009AnyJev
Turns any LLM into a Jev-style decision model - typed decisions with real probabilities, no training.
github.com/nokia-applied-research/AnyJev </a> <a class="card rt" href="https://github.com/mizorewww/laya-mlx" rel="noopener noreferrer" data-cat="decision"> decision6,702laya-mlx
Native MLX runtime for Laya typed decision models: 7-14 ms short decisions on an M3 Max, no generation.
github.com/mizorewww/laya-mlx </a> <a class="card rt" href="https://github.com/TypeLLM/TypeLLM" rel="noopener noreferrer" data-cat="decision"> decision910TypeLLM
LLMs with type-safe generation: constrain the output to a declared type instead of parsing free text.
github.com/TypeLLM/TypeLLM </a> <a class="card rt" href="https://github.com/zwliJay/jev-forge" rel="noopener noreferrer" data-cat="decision"> decision119jev-forge
Training and inference stack for Jev-style models: score dynamic candidate branches from a shared prefix, with calibration.
github.com/zwliJay/jev-forge </a> <a class="card rt" href="https://github.com/lyuyiqi/open-jev-fast" rel="noopener noreferrer" data-cat="decision"> decision88open-jev-fast
Faster CUDA backend for Open-Jev-27B: fused kernels, prefix tree and CUDA graphs on bf16.
github.com/lyuyiqi/open-jev-fast </a> <a class="card rt" href="https://github.com/kikoncuo/jevfire" rel="noopener noreferrer" data-cat="decision"> decision71jevfire
Parallel decisions on top of an existing vLLM server: one context, many decisions, vLLM API.
github.com/kikoncuo/jevfire </a> <a class="card rt" href="https://github.com/Jwuthri/SelfJev" rel="noopener noreferrer" data-cat="decision"> decision61SelfJev
Qwen3.5-4B plus a LoRA that answers with typed decisions and probabilities on a single GPU.
github.com/Jwuthri/SelfJev </a> <a class="card rt" href="https://github.com/SAGAR-TAMANG/sarvam-jev" rel="noopener noreferrer" data-cat="decision"> decision58sarvam-jev
Decision models for Indic languages: constrained logit readout instead of generated JSON, runs in the browser.
github.com/SAGAR-TAMANG/sarvam-jev </a> <a class="card rt" href="https://github.com/Comfy-Org/ComfyUI" rel="noopener noreferrer" data-cat="diffusion"> diffusion135,830ComfyUI
Node-graph backend for image and video models: FLUX, Qwen Image, Wan and others, scriptable through an API.
github.com/Comfy-Org/ComfyUI </a> <a class="card rt" href="https://github.com/leejet/stable-diffusion.cpp" rel="noopener noreferrer" data-cat="diffusion"> diffusion7,498stable-diffusion.cpp
Diffusion inference in plain C/C++ (SD, FLUX, Wan, Qwen Image), the GGUF-minded sibling of llama.cpp.
github.com/leejet/stable-diffusion.cpp </a> <a class="card rt" href="https://github.com/AUTOMATIC1111/stable-diffusion-webui" rel="noopener noreferrer" data-cat="diffusion"> diffusion165,182Stable Diffusion web UI
The original A1111 interface. Enormous extension ecosystem, development quiet since March 2026.
github.com/AUTOMATIC1111/stable-diffusion-webui </a> <a class="card rt" href="https://github.com/lllyasviel/stable-diffusion-webui-forge" rel="noopener noreferrer" data-cat="diffusion"> diffusion13,044Stable Diffusion webui Forge
Memory-optimized fork of A1111 for smaller cards. No commits since July 2025.
github.com/lllyasviel/stable-diffusion-webui-forge </a> <a class="card rt" href="https://github.com/invoke-ai/InvokeAI" rel="noopener noreferrer" data-cat="diffusion"> diffusion28,330InvokeAI
Creative studio around diffusion models: canvas, layers and a professional workflow on top of the same weights.
github.com/invoke-ai/InvokeAI </a> <a class="card rt" href="https://github.com/huggingface/diffusers" rel="noopener noreferrer" data-cat="diffusion"> diffusion34,641Diffusers
Hugging Face's Python library for diffusion pipelines. The base most other image tools build on.
github.com/huggingface/diffusers </a> <a class="card rt" href="https://github.com/mflux-community/mflux" rel="noopener noreferrer" data-cat="diffusion"> diffusion2,433mflux
Apple MLX implementations of current image and video models, optimized for unified-memory Macs.
github.com/mflux-community/mflux </a> <a class="card rt" href="https://github.com/drawthingsai/draw-things-community" rel="noopener noreferrer" data-cat="diffusion"> diffusion576Draw Things
Local image generation app for Apple devices and desktop, with its own models and LoRA support.
github.com/drawthingsai/draw-things-community </a> <a class="card rt" href="https://github.com/ggml-org/whisper.cpp" rel="noopener noreferrer" data-cat="speech"> speech54,095whisper.cpp
Whisper speech recognition in C/C++, GGML-quantized. The default local transcriber.
github.com/ggml-org/whisper.cpp </a> <a class="card rt" href="https://github.com/SYSTRAN/faster-whisper" rel="noopener noreferrer" data-cat="speech"> speech25,673faster-whisper
CTranslate2 implementation of Whisper, several times faster than the reference at the same accuracy.
github.com/SYSTRAN/faster-whisper </a> <a class="card rt" href="https://github.com/k2-fsa/sherpa-onnx" rel="noopener noreferrer" data-cat="speech"> speech15,077sherpa-onnx
One ONNX runtime for speech recognition, synthesis, diarization and source separation, embeddable everywhere.
github.com/k2-fsa/sherpa-onnx </a> <a class="card rt" href="https://github.com/speaches-ai/speaches" rel="noopener noreferrer" data-cat="speech"> speech3,695Speaches
Self-hosted speech-to-text and text-to-speech behind OpenAI-compatible endpoints.
github.com/speaches-ai/speaches </a> <a class="card rt" href="https://github.com/rhasspy/piper" rel="noopener noreferrer" data-cat="speech"> speech11,298Piper
Fast local neural text-to-speech with small ONNX voices. Quiet upstream since August 2025.
github.com/rhasspy/piper </a> <a class="card rt" href="https://github.com/hexgrad/kokoro" rel="noopener noreferrer" data-cat="speech"> speech9,116Kokoro
Small, high-quality text-to-speech model, usually run through ONNX or MLX wrappers.
github.com/hexgrad/kokoro </a> <a class="card rt" href="https://github.com/open-webui/open-webui" rel="noopener noreferrer" data-cat="ui"> ui153,791Open WebUI
Self-hosted chat front end for Ollama, vLLM or any OpenAI-compatible endpoint, with users and RAG.
github.com/open-webui/open-webui </a> <a class="card rt" href="https://github.com/Mintplex-Labs/anything-llm" rel="noopener noreferrer" data-cat="ui"> ui66,668AnythingLLM
Desktop and server app for chatting with your own documents, local models or APIs.
github.com/Mintplex-Labs/anything-llm </a> <a class="card rt" href="https://github.com/LibreChat-AI/LibreChat" rel="noopener noreferrer" data-cat="ui"> ui45,203LibreChat
Multi-provider chat interface with agents and MCP support, deployable next to your own models.
github.com/LibreChat-AI/LibreChat </a> <a class="card rt" href="https://github.com/ome-projects/ome" rel="noopener noreferrer" data-cat="orchestrator"> orchestrator514OME
Kubernetes operator for LLM serving: GPU scheduling and model lifecycle as cluster resources.
github.com/ome-projects/ome </a> <a class="card rt" href="https://github.com/kserve/kserve" rel="noopener noreferrer" data-cat="orchestrator"> orchestrator6,058KServe
Standard inference platform for Kubernetes, covering both predictive and generative models.
github.com/kserve/kserve </a> <a class="card rt" href="https://github.com/ray-project/ray" rel="noopener noreferrer" data-cat="orchestrator"> orchestrator43,963Ray Serve
Composes multi-model pipelines and scales them across machines; used as the layer under several engines.
github.com/ray-project/ray </a>Four layers, not one thing
People say “the engine” and mean four different things. Keeping them apart makes every benchmark and every bug report easier to read.
- Kernels. The math: cuBLAS, ROCm, Metal, Vulkan, SYCL, AVX-512. Everything above inherits their speed and their limits.
- Engine. What loads the weights, allocates the KV cache and decides how to split work: llama.cpp, vLLM, SGLang, ExLlamaV3, MLX.
- Server. What exposes the model: an OpenAI-compatible HTTP port, a task queue, a gRPC endpoint.
- App. What a person actually touches: Open WebUI, ComfyUI, LM Studio, KoboldCpp.
A “slow model” is usually a slow choice at one of those four layers, not a slow GPU.
Which one, by what you have
- One card up to 24 GB, or CPU only. GGUF and
llama.cpp(or a UI built on it: Ollama, LM Studio, KoboldCpp). Quantized weights, KV cache inq8_0, partial GPU offload.ik_llama.cppwhen you want the extra quant formats. - Two to four consumer cards. Still llama.cpp for GGUF, or
vLLM/ExLlamaV3if you can hold full-precision weights. This is the range where tensor parallel and layer split start to matter. - 96 GB and up on one machine.
vLLMorSGLangwith batching: several users, long contexts, an OpenAI endpoint.TensorRT-LLMif the stack is NVIDIA-only and you want the last 10 %. - Apple silicon, 128 GB and up.
MLX/MLX-LM, or GGUF through llama.cpp with Metal. - No NVIDIA at all. ROCm with llama.cpp or vLLM on recent Radeons,
OpenVINOon Intel, Vulkan as the fallback. - Image and video.
ComfyUI, orstable-diffusion.cppif you want one binary and GGUF quantization. - Speech.
whisper.cpp(orfaster-whisper) in,Piper/Kokoroout,sherpa-onnxif you need it embedded. - A whole cluster.
OMEorKServeon Kubernetes,DynamoorRay Serveunderneath.
The decision engines: JEV and friends
JEV (TypeSafe AI’s System One models) is not a chat model and not a smaller vLLM. It takes a typed question and returns a typed decision — a choice, a score, a boolean — with a confidence, by reading logits from one forward pass over a shared prefix. No tokens are generated, so the VRAM and the latency are a different order of magnitude: a few milliseconds and a few gigabytes, on a laptop.
That makes it the right tool for classification, routing, rubric scoring, verification and
agent guardrails — the places where people otherwise pay for a chat model and then try to
parse its answer. Open implementations exist (jevos, AnyJev, laya-mlx), and the same
idea shows up in TypeLLM for type-safe generation.
How to read this page
Cards are grouped by type and can be filtered. formats is what the engine can load,
backends is what it really runs on today, multi-GPU says whether one model can be split
across cards, and KV cache says whether the context cache can be quantized — the cheapest
way to buy context on a small card. Star counts are a snapshot taken on 2026-10-02, kept only
as a rough popularity signal; a quiet repository is not always a dead one, and the notes say
when a project has gone silent.