64 entries · Updated 2026-10-02

Runtimes

The inference engines that actually run models on your own hardware: local runtimes, serving engines, typed-decision models, diffusion, speech and the interfaces on top.

<a class="card rt" href="https://github.com/ggml-org/llama.cpp" rel="noopener noreferrer" data-cat="runtime"> runtime130,123

llama.cpp

The C/C++ core most local AI runs on: GGUF, llama-server, llama-bench. Runs on CPU, CUDA, ROCm, Vulkan, SYCL and Metal.

Loads
gguf
Runs on
CPU · CUDA · ROCm · Vulkan · SYCL · Metal
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/ggml-org/llama.cpp </a> <a class="card rt" href="https://github.com/ikawrakow/ik_llama.cpp" rel="noopener noreferrer" data-cat="runtime"> runtime3,273

ik_llama.cpp

Fork of llama.cpp with extra quants (IQK) and tweaks that often win on CPU and mixed CPU/GPU rigs.

Loads
gguf
Runs on
CPU · CUDA · ROCm
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/ikawrakow/ik_llama.cpp </a> <a class="card rt" href="https://github.com/ollama/ollama" rel="noopener noreferrer" data-cat="runtime"> runtime182,049

Ollama

One command to pull and run a model, with an OpenAI-compatible API. Easiest way in, llama.cpp underneath.

Loads
gguf
Runs on
CPU · CUDA · ROCm · Metal
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/ollama/ollama </a> <a class="card rt" href="https://github.com/lmstudio-ai/lms" rel="noopener noreferrer" data-cat="runtime"> runtime5,329

LM Studio

Desktop app plus a local server: pick a GGUF, set GPU offload, get an OpenAI endpoint. No terminal needed.

Loads
gguf
Runs on
CPU · CUDA · ROCm · Vulkan · Metal
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/lmstudio-ai/lms </a> <a class="card rt" href="https://github.com/LostRuins/koboldcpp" rel="noopener noreferrer" data-cat="runtime"> runtime11,924

KoboldCpp

Single binary with its own web UI: GGUF chat, image generation and TTS, CUDA/ROCm/Vulkan builds.

Loads
gguf
Runs on
CPU · CUDA · ROCm · Vulkan
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/LostRuins/koboldcpp </a> <a class="card rt" href="https://github.com/mozilla-ai/llamafile" rel="noopener noreferrer" data-cat="runtime"> runtime26,154

llamafile

Ships a model and the runtime in one executable file that runs on Linux, macOS and Windows.

Loads
gguf
Runs on
CPU · CUDA · Metal
Multi-GPU
no
KV cache
quantizable
API
OpenAI-compatible
github.com/mozilla-ai/llamafile </a> <a class="card rt" href="https://github.com/nomic-ai/gpt4all" rel="noopener noreferrer" data-cat="runtime"> runtime77,389

GPT4All

Desktop app for running GGUF models locally. Repository has been quiet since May 2025.

Loads
gguf
Runs on
CPU · CUDA · Vulkan · Metal
Multi-GPU
no
KV cache
quantizable
API
OpenAI-compatible
github.com/nomic-ai/gpt4all </a> <a class="card rt" href="https://github.com/mudler/LocalAI" rel="noopener noreferrer" data-cat="runtime"> runtime49,369

LocalAI

Self-hosted multi-modal engine: text, vision, voice and image models behind OpenAI-compatible endpoints.

Loads
gguf · safetensors
Runs on
CPU · CUDA · ROCm · Metal
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/mudler/LocalAI </a> <a class="card rt" href="https://github.com/containers/ramalama" rel="noopener noreferrer" data-cat="runtime"> runtime3,066

RamaLama

Serves models from OCI containers, so the runtime and its drivers travel with the model.

Loads
gguf · safetensors
Runs on
CPU · CUDA · ROCm
Multi-GPU
yes
KV cache
full precision
API
OpenAI-compatible
github.com/containers/ramalama </a> <a class="card rt" href="https://github.com/oobabooga/textgen" rel="noopener noreferrer" data-cat="runtime"> runtime47,721

Text generation web UI

Desktop app (formerly oobabooga) for text, vision and tool calling, with an OpenAI-compatible API.

Loads
gguf · exl2 · safetensors
Runs on
CPU · CUDA · ROCm · Metal
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/oobabooga/textgen </a> <a class="card rt" href="https://github.com/ml-explore/mlx-lm" rel="noopener noreferrer" data-cat="runtime"> runtime7,198

MLX-LM

Run and quantize LLMs on Apple silicon through MLX. The default path on unified-memory Macs.

Loads
mlx · safetensors
Runs on
Metal
Multi-GPU
no
KV cache
quantizable
API
other
github.com/ml-explore/mlx-lm </a> <a class="card rt" href="https://github.com/janhq/jan" rel="noopener noreferrer" data-cat="runtime"> runtime44,759

Jan

Offline desktop assistant that can also expose a local OpenAI-compatible server.

Loads
gguf
Runs on
CPU · CUDA · Vulkan · Metal
Multi-GPU
no
KV cache
quantizable
API
OpenAI-compatible
github.com/janhq/jan </a> <a class="card rt" href="https://github.com/AtomicBot-ai/Atomic-Chat" rel="noopener noreferrer" data-cat="runtime"> runtime1,658

Atomic Chat

Local AI app with its own inference engine, built for agents and multi-model workflows.

Loads
gguf
Runs on
CPU · CUDA · Metal
Multi-GPU
no
KV cache
quantizable
API
OpenAI-compatible
github.com/AtomicBot-ai/Atomic-Chat </a> <a class="card rt" href="https://github.com/tetherto/qvac" rel="noopener noreferrer" data-cat="runtime"> runtime657

qvac

On-device AI SDK: GGUF models, RAG, images and music with no cloud and no API keys.

Loads
gguf
Runs on
CPU · CUDA · Metal
Multi-GPU
no
KV cache
quantizable
API
other
github.com/tetherto/qvac </a> <a class="card rt" href="https://github.com/cactus-compute/cactus" rel="noopener noreferrer" data-cat="runtime"> runtime6,084

Cactus

Quantization, kernels and runtime aimed at phones, wearables and robots rather than desktops.

Loads
gguf
Runs on
CPU · Metal · NPU
Multi-GPU
no
KV cache
quantizable
API
other
github.com/cactus-compute/cactus </a> <a class="card rt" href="https://github.com/vllm-project/vllm" rel="noopener noreferrer" data-cat="serving"> serving93,059

vLLM

The reference serving engine: paged KV cache, continuous batching, tensor parallelism. Strong on both CUDA and ROCm.

Loads
safetensors
Runs on
CUDA · ROCm · CPU
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/vllm-project/vllm </a> <a class="card rt" href="https://github.com/sgl-project/sglang" rel="noopener noreferrer" data-cat="serving"> serving36,716

SGLang

Serving framework with a prefix cache (RadixAttention) built for multi-turn, agents and structured output.

Loads
safetensors
Runs on
CUDA · ROCm
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/sgl-project/sglang </a> <a class="card rt" href="https://github.com/NVIDIA/TensorRT-LLM" rel="noopener noreferrer" data-cat="serving"> serving14,756

TensorRT-LLM

NVIDIA's engine: compiles the model into optimized kernels and graphs. Fastest on recent GeForce and datacenter GPUs.

Loads
safetensors
Runs on
CUDA
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/NVIDIA/TensorRT-LLM </a> <a class="card rt" href="https://github.com/huggingface/text-generation-inference" rel="noopener noreferrer" data-cat="serving"> serving10,884

Text Generation Inference

Hugging Face's serving stack. Still used in production, but development has slowed since early 2026.

Loads
safetensors
Runs on
CUDA · ROCm · CPU
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/huggingface/text-generation-inference </a> <a class="card rt" href="https://github.com/turboderp-org/exllamav3" rel="noopener noreferrer" data-cat="serving"> serving1,571

ExLlamaV3

Quantization and inference library tuned for one or two consumer GPUs, successor to ExLlamaV2.

Loads
exl3 · safetensors
Runs on
CUDA · ROCm
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/turboderp-org/exllamav3 </a> <a class="card rt" href="https://github.com/turboderp-org/exllamav2" rel="noopener noreferrer" data-cat="serving"> serving4,633

ExLlamaV2

The exl2 engine that made big models fit on 24 GB cards. Superseded by ExLlamaV3, quiet since March 2026.

Loads
exl2
Runs on
CUDA · ROCm
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/turboderp-org/exllamav2 </a> <a class="card rt" href="https://github.com/mlc-ai/mlc-llm" rel="noopener noreferrer" data-cat="serving"> serving23,202

MLC-LLM

Compiles models for a target device, from phones to servers, using TVM. Same stack as in-browser inference.

Loads
mlc
Runs on
CUDA · ROCm · Vulkan · Metal
Multi-GPU
no
KV cache
full precision
API
OpenAI-compatible
github.com/mlc-ai/mlc-llm </a> <a class="card rt" href="https://github.com/mlc-ai/web-llm" rel="noopener noreferrer" data-cat="serving"> serving19,214

WebLLM

Runs quantized models inside the browser on WebGPU. No server, no install, model cached locally.

Loads
mlc
Runs on
Metal
Multi-GPU
no
KV cache
full precision
API
other
github.com/mlc-ai/web-llm </a> <a class="card rt" href="https://github.com/bentoml/OpenLLM" rel="noopener noreferrer" data-cat="serving"> serving12,552

OpenLLM

BentoML's packaging layer: turn any open model into an OpenAI-compatible endpoint with a few lines.

Loads
safetensors · gguf
Runs on
CUDA · ROCm · CPU
Multi-GPU
yes
KV cache
full precision
API
OpenAI-compatible
github.com/bentoml/OpenLLM </a> <a class="card rt" href="https://github.com/ai-dynamo/dynamo" rel="noopener noreferrer" data-cat="serving"> serving8,206

Dynamo

Distributed serving framework that splits prefill and decode across machines. Datacenter scale, Rust core.

Loads
safetensors
Runs on
CUDA
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/ai-dynamo/dynamo </a> <a class="card rt" href="https://github.com/lightseekorg/tokenspeed" rel="noopener noreferrer" data-cat="serving"> serving2,189

TokenSpeed

Recent high-throughput serving engine, focused on squeezing latency out of the decode path.

Loads
safetensors
Runs on
CUDA
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/lightseekorg/tokenspeed </a> <a class="card rt" href="https://github.com/dphnAI/sonar" rel="noopener noreferrer" data-cat="serving"> serving1,869

Sonar

Large-scale LLM inference engine, aimed at clusters rather than single rigs.

Loads
safetensors
Runs on
CUDA
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/dphnAI/sonar </a> <a class="card rt" href="https://github.com/trymirai/uzu" rel="noopener noreferrer" data-cat="serving"> serving1,820

uzu

Rust inference engine for AI models, portable across CPU and GPU backends.

Loads
safetensors
Runs on
CPU · CUDA · Metal
Multi-GPU
yes
KV cache
quantizable
API
other
github.com/trymirai/uzu </a> <a class="card rt" href="https://github.com/lucasjinreal/Crane" rel="noopener noreferrer" data-cat="serving"> serving488

Crane

Pure Rust engine for LLM, VLM, TTS and OCR, built on Candle. Pitched as a simpler alternative to llama.cpp.

Loads
safetensors · gguf
Runs on
CPU · CUDA · Metal
Multi-GPU
yes
KV cache
quantizable
API
other
github.com/lucasjinreal/Crane </a> <a class="card rt" href="https://github.com/microsoft/sarathi-serve" rel="noopener noreferrer" data-cat="serving"> serving529

Sarathi-Serve

Research serving engine built around chunked prefill for lower tail latency. Dormant since January 2026.

Loads
safetensors
Runs on
CUDA
Multi-GPU
yes
KV cache
full precision
API
OpenAI-compatible
github.com/microsoft/sarathi-serve </a> <a class="card rt" href="https://github.com/Tiiny-AI/PowerInfer" rel="noopener noreferrer" data-cat="serving"> serving9,813

PowerInfer

Serves local models by keeping hot neurons on the GPU and offloading the rest. Quiet since May 2026.

Loads
gguf
Runs on
CUDA · CPU
Multi-GPU
yes
KV cache
full precision
API
other
github.com/Tiiny-AI/PowerInfer </a> <a class="card rt" href="https://github.com/openvinotoolkit/openvino" rel="noopener noreferrer" data-cat="serving"> serving10,943

OpenVINO

Intel's toolkit for CPU, integrated GPU and NPU inference. The practical path on non-NVIDIA hardware.

Loads
onnx · ir
Runs on
CPU · SYCL · NPU
Multi-GPU
no
KV cache
full precision
API
other
github.com/openvinotoolkit/openvino </a> <a class="card rt" href="https://github.com/alibaba/MNN" rel="noopener noreferrer" data-cat="serving"> serving16,166

MNN

Alibaba's lightweight engine for mobile and edge devices, with its own quantized format.

Loads
mnn
Runs on
CPU · Vulkan · Metal
Multi-GPU
no
KV cache
full precision
API
other
github.com/alibaba/MNN </a> <a class="card rt" href="https://github.com/kvcache-ai/ktransformers" rel="noopener noreferrer" data-cat="serving"> serving19,555

KTransformers

Runs large MoE models by keeping attention on the GPU and experts in CPU memory. Built for big sparse models on small VRAM.

Loads
safetensors · gguf
Runs on
CPU · CUDA
Multi-GPU
yes
KV cache
quantizable
API
OpenAI-compatible
github.com/kvcache-ai/ktransformers </a> <div class="card rt" data-cat="decision"> decision

JEV (System One)

TypeSafe AI's decision models: a typed question in, a typed decision out (choice, score, boolean) with a confidence. No text generation.

Loads
api
Runs on
Cloud · CUDA · CPU
Multi-GPU
no
KV cache
full precision
API
other

The category is decision, not chat: score candidate branches off one shared prefix.

typesafe.ai </div> <a class="card rt" href="https://github.com/feder-cr/jev" rel="noopener noreferrer" data-cat="decision"> decision1,175

jevos

Open-source alternative to Jev for yes/no decisions. C++, GGUF, runs on a laptop CPU behind a FastAPI server.

Loads
gguf
Runs on
CPU · CUDA
Multi-GPU
no
KV cache
full precision
API
other
github.com/feder-cr/jev </a> <a class="card rt" href="https://github.com/nokia-applied-research/AnyJev" rel="noopener noreferrer" data-cat="decision"> decision1,009

AnyJev

Turns any LLM into a Jev-style decision model - typed decisions with real probabilities, no training.

Loads
safetensors · gguf
Runs on
CPU · CUDA
Multi-GPU
yes
KV cache
full precision
API
other
github.com/nokia-applied-research/AnyJev </a> <a class="card rt" href="https://github.com/mizorewww/laya-mlx" rel="noopener noreferrer" data-cat="decision"> decision6,702

laya-mlx

Native MLX runtime for Laya typed decision models: 7-14 ms short decisions on an M3 Max, no generation.

Loads
mlx
Runs on
Metal
Multi-GPU
no
KV cache
full precision
API
other
github.com/mizorewww/laya-mlx </a> <a class="card rt" href="https://github.com/TypeLLM/TypeLLM" rel="noopener noreferrer" data-cat="decision"> decision910

TypeLLM

LLMs with type-safe generation: constrain the output to a declared type instead of parsing free text.

Loads
safetensors
Runs on
CUDA · CPU
Multi-GPU
no
KV cache
full precision
API
other
github.com/TypeLLM/TypeLLM </a> <a class="card rt" href="https://github.com/zwliJay/jev-forge" rel="noopener noreferrer" data-cat="decision"> decision119

jev-forge

Training and inference stack for Jev-style models: score dynamic candidate branches from a shared prefix, with calibration.

Loads
safetensors
Runs on
CUDA
Multi-GPU
no
KV cache
full precision
API
other
github.com/zwliJay/jev-forge </a> <a class="card rt" href="https://github.com/lyuyiqi/open-jev-fast" rel="noopener noreferrer" data-cat="decision"> decision88

open-jev-fast

Faster CUDA backend for Open-Jev-27B: fused kernels, prefix tree and CUDA graphs on bf16.

Loads
safetensors
Runs on
CUDA
Multi-GPU
no
KV cache
full precision
API
other
github.com/lyuyiqi/open-jev-fast </a> <a class="card rt" href="https://github.com/kikoncuo/jevfire" rel="noopener noreferrer" data-cat="decision"> decision71

jevfire

Parallel decisions on top of an existing vLLM server: one context, many decisions, vLLM API.

Loads
api
Runs on
CUDA
Multi-GPU
no
KV cache
full precision
API
other
github.com/kikoncuo/jevfire </a> <a class="card rt" href="https://github.com/Jwuthri/SelfJev" rel="noopener noreferrer" data-cat="decision"> decision61

SelfJev

Qwen3.5-4B plus a LoRA that answers with typed decisions and probabilities on a single GPU.

Loads
safetensors
Runs on
CUDA
Multi-GPU
no
KV cache
full precision
API
other
github.com/Jwuthri/SelfJev </a> <a class="card rt" href="https://github.com/SAGAR-TAMANG/sarvam-jev" rel="noopener noreferrer" data-cat="decision"> decision58

sarvam-jev

Decision models for Indic languages: constrained logit readout instead of generated JSON, runs in the browser.

Loads
api
Runs on
Cloud · CPU
Multi-GPU
no
KV cache
full precision
API
other
github.com/SAGAR-TAMANG/sarvam-jev </a> <a class="card rt" href="https://github.com/Comfy-Org/ComfyUI" rel="noopener noreferrer" data-cat="diffusion"> diffusion135,830

ComfyUI

Node-graph backend for image and video models: FLUX, Qwen Image, Wan and others, scriptable through an API.

Loads
safetensors · gguf
Runs on
CUDA · ROCm · CPU · Metal
Multi-GPU
yes
KV cache
full precision
API
other
github.com/Comfy-Org/ComfyUI </a> <a class="card rt" href="https://github.com/leejet/stable-diffusion.cpp" rel="noopener noreferrer" data-cat="diffusion"> diffusion7,498

stable-diffusion.cpp

Diffusion inference in plain C/C++ (SD, FLUX, Wan, Qwen Image), the GGUF-minded sibling of llama.cpp.

Loads
gguf · safetensors
Runs on
CPU · CUDA · Vulkan · Metal
Multi-GPU
no
KV cache
full precision
API
other
github.com/leejet/stable-diffusion.cpp </a> <a class="card rt" href="https://github.com/AUTOMATIC1111/stable-diffusion-webui" rel="noopener noreferrer" data-cat="diffusion"> diffusion165,182

Stable Diffusion web UI

The original A1111 interface. Enormous extension ecosystem, development quiet since March 2026.

Loads
safetensors
Runs on
CUDA · CPU · Metal
Multi-GPU
no
KV cache
full precision
API
OpenAI-compatible
github.com/AUTOMATIC1111/stable-diffusion-webui </a> <a class="card rt" href="https://github.com/lllyasviel/stable-diffusion-webui-forge" rel="noopener noreferrer" data-cat="diffusion"> diffusion13,044

Stable Diffusion webui Forge

Memory-optimized fork of A1111 for smaller cards. No commits since July 2025.

Loads
safetensors
Runs on
CUDA
Multi-GPU
no
KV cache
full precision
API
OpenAI-compatible
github.com/lllyasviel/stable-diffusion-webui-forge </a> <a class="card rt" href="https://github.com/invoke-ai/InvokeAI" rel="noopener noreferrer" data-cat="diffusion"> diffusion28,330

InvokeAI

Creative studio around diffusion models: canvas, layers and a professional workflow on top of the same weights.

Loads
safetensors
Runs on
CUDA · CPU · Metal
Multi-GPU
no
KV cache
full precision
API
OpenAI-compatible
github.com/invoke-ai/InvokeAI </a> <a class="card rt" href="https://github.com/huggingface/diffusers" rel="noopener noreferrer" data-cat="diffusion"> diffusion34,641

Diffusers

Hugging Face's Python library for diffusion pipelines. The base most other image tools build on.

Loads
safetensors · diffusers
Runs on
CUDA · ROCm · CPU · Metal
Multi-GPU
yes
KV cache
full precision
API
other
github.com/huggingface/diffusers </a> <a class="card rt" href="https://github.com/mflux-community/mflux" rel="noopener noreferrer" data-cat="diffusion"> diffusion2,433

mflux

Apple MLX implementations of current image and video models, optimized for unified-memory Macs.

Loads
mlx
Runs on
Metal
Multi-GPU
no
KV cache
full precision
API
other
github.com/mflux-community/mflux </a> <a class="card rt" href="https://github.com/drawthingsai/draw-things-community" rel="noopener noreferrer" data-cat="diffusion"> diffusion576

Draw Things

Local image generation app for Apple devices and desktop, with its own models and LoRA support.

Loads
safetensors · mlx
Runs on
Metal
Multi-GPU
no
KV cache
full precision
API
other
github.com/drawthingsai/draw-things-community </a> <a class="card rt" href="https://github.com/ggml-org/whisper.cpp" rel="noopener noreferrer" data-cat="speech"> speech54,095

whisper.cpp

Whisper speech recognition in C/C++, GGML-quantized. The default local transcriber.

Loads
ggml
Runs on
CPU · CUDA · Vulkan · Metal
Multi-GPU
no
KV cache
full precision
API
other
github.com/ggml-org/whisper.cpp </a> <a class="card rt" href="https://github.com/SYSTRAN/faster-whisper" rel="noopener noreferrer" data-cat="speech"> speech25,673

faster-whisper

CTranslate2 implementation of Whisper, several times faster than the reference at the same accuracy.

Loads
ctranslate2
Runs on
CPU · CUDA
Multi-GPU
yes
KV cache
full precision
API
OpenAI-compatible
github.com/SYSTRAN/faster-whisper </a> <a class="card rt" href="https://github.com/k2-fsa/sherpa-onnx" rel="noopener noreferrer" data-cat="speech"> speech15,077

sherpa-onnx

One ONNX runtime for speech recognition, synthesis, diarization and source separation, embeddable everywhere.

Loads
onnx
Runs on
CPU · CUDA · Vulkan · Metal
Multi-GPU
no
KV cache
full precision
API
other
github.com/k2-fsa/sherpa-onnx </a> <a class="card rt" href="https://github.com/speaches-ai/speaches" rel="noopener noreferrer" data-cat="speech"> speech3,695

Speaches

Self-hosted speech-to-text and text-to-speech behind OpenAI-compatible endpoints.

Loads
onnx
Runs on
CPU · CUDA
Multi-GPU
no
KV cache
full precision
API
OpenAI-compatible
github.com/speaches-ai/speaches </a> <a class="card rt" href="https://github.com/rhasspy/piper" rel="noopener noreferrer" data-cat="speech"> speech11,298

Piper

Fast local neural text-to-speech with small ONNX voices. Quiet upstream since August 2025.

Loads
onnx
Runs on
CPU
Multi-GPU
no
KV cache
full precision
API
other
github.com/rhasspy/piper </a> <a class="card rt" href="https://github.com/hexgrad/kokoro" rel="noopener noreferrer" data-cat="speech"> speech9,116

Kokoro

Small, high-quality text-to-speech model, usually run through ONNX or MLX wrappers.

Loads
onnx · mlx
Runs on
CPU · Metal
Multi-GPU
no
KV cache
full precision
API
other
github.com/hexgrad/kokoro </a> <a class="card rt" href="https://github.com/open-webui/open-webui" rel="noopener noreferrer" data-cat="ui"> ui153,791

Open WebUI

Self-hosted chat front end for Ollama, vLLM or any OpenAI-compatible endpoint, with users and RAG.

Loads
n/a
Runs on
n/a
Multi-GPU
no
KV cache
full precision
API
other
github.com/open-webui/open-webui </a> <a class="card rt" href="https://github.com/Mintplex-Labs/anything-llm" rel="noopener noreferrer" data-cat="ui"> ui66,668

AnythingLLM

Desktop and server app for chatting with your own documents, local models or APIs.

Loads
n/a
Runs on
n/a
Multi-GPU
no
KV cache
full precision
API
other
github.com/Mintplex-Labs/anything-llm </a> <a class="card rt" href="https://github.com/LibreChat-AI/LibreChat" rel="noopener noreferrer" data-cat="ui"> ui45,203

LibreChat

Multi-provider chat interface with agents and MCP support, deployable next to your own models.

Loads
n/a
Runs on
n/a
Multi-GPU
no
KV cache
full precision
API
other
github.com/LibreChat-AI/LibreChat </a> <a class="card rt" href="https://github.com/ome-projects/ome" rel="noopener noreferrer" data-cat="orchestrator"> orchestrator514

OME

Kubernetes operator for LLM serving: GPU scheduling and model lifecycle as cluster resources.

Loads
n/a
Runs on
CUDA
Multi-GPU
yes
KV cache
full precision
API
other
github.com/ome-projects/ome </a> <a class="card rt" href="https://github.com/kserve/kserve" rel="noopener noreferrer" data-cat="orchestrator"> orchestrator6,058

KServe

Standard inference platform for Kubernetes, covering both predictive and generative models.

Loads
n/a
Runs on
CUDA · CPU
Multi-GPU
yes
KV cache
full precision
API
other
github.com/kserve/kserve </a> <a class="card rt" href="https://github.com/ray-project/ray" rel="noopener noreferrer" data-cat="orchestrator"> orchestrator43,963

Ray Serve

Composes multi-model pipelines and scales them across machines; used as the layer under several engines.

Loads
n/a
Runs on
CUDA · CPU
Multi-GPU
yes
KV cache
full precision
API
other
github.com/ray-project/ray </a>

Four layers, not one thing

People say “the engine” and mean four different things. Keeping them apart makes every benchmark and every bug report easier to read.

  • Kernels. The math: cuBLAS, ROCm, Metal, Vulkan, SYCL, AVX-512. Everything above inherits their speed and their limits.
  • Engine. What loads the weights, allocates the KV cache and decides how to split work: llama.cpp, vLLM, SGLang, ExLlamaV3, MLX.
  • Server. What exposes the model: an OpenAI-compatible HTTP port, a task queue, a gRPC endpoint.
  • App. What a person actually touches: Open WebUI, ComfyUI, LM Studio, KoboldCpp.

A “slow model” is usually a slow choice at one of those four layers, not a slow GPU.

Which one, by what you have

  • One card up to 24 GB, or CPU only. GGUF and llama.cpp (or a UI built on it: Ollama, LM Studio, KoboldCpp). Quantized weights, KV cache in q8_0, partial GPU offload. ik_llama.cpp when you want the extra quant formats.
  • Two to four consumer cards. Still llama.cpp for GGUF, or vLLM/ExLlamaV3 if you can hold full-precision weights. This is the range where tensor parallel and layer split start to matter.
  • 96 GB and up on one machine. vLLM or SGLang with batching: several users, long contexts, an OpenAI endpoint. TensorRT-LLM if the stack is NVIDIA-only and you want the last 10 %.
  • Apple silicon, 128 GB and up. MLX / MLX-LM, or GGUF through llama.cpp with Metal.
  • No NVIDIA at all. ROCm with llama.cpp or vLLM on recent Radeons, OpenVINO on Intel, Vulkan as the fallback.
  • Image and video. ComfyUI, or stable-diffusion.cpp if you want one binary and GGUF quantization.
  • Speech. whisper.cpp (or faster-whisper) in, Piper/Kokoro out, sherpa-onnx if you need it embedded.
  • A whole cluster. OME or KServe on Kubernetes, Dynamo or Ray Serve underneath.

The decision engines: JEV and friends

JEV (TypeSafe AI’s System One models) is not a chat model and not a smaller vLLM. It takes a typed question and returns a typed decision — a choice, a score, a boolean — with a confidence, by reading logits from one forward pass over a shared prefix. No tokens are generated, so the VRAM and the latency are a different order of magnitude: a few milliseconds and a few gigabytes, on a laptop.

That makes it the right tool for classification, routing, rubric scoring, verification and agent guardrails — the places where people otherwise pay for a chat model and then try to parse its answer. Open implementations exist (jevos, AnyJev, laya-mlx), and the same idea shows up in TypeLLM for type-safe generation.

How to read this page

Cards are grouped by type and can be filtered. formats is what the engine can load, backends is what it really runs on today, multi-GPU says whether one model can be split across cards, and KV cache says whether the context cache can be quantized — the cheapest way to buy context on a small card. Star counts are a snapshot taken on 2026-10-02, kept only as a rough popularity signal; a quiet repository is not always a dead one, and the notes say when a project has gone silent.