#1919th of 20 in Model serving
TabbyAPI
OpenAI-compatible API server for running ExLlamaV3 models
- Stars
- 1.5k
- License
- AGPL-3.0
- Last commit
- Oct 2026
- Language
- Python
Overview
TabbyAPI is a FastAPI server that loads and serves LLMs through the ExLlamaV3 backend, exposing an OpenAI-compatible HTTP API. It supports runtime model loading and unloading, HuggingFace downloads, embedding models, constrained output (JSON schema, regex, EBNF), and tool calling. The README describes it as a hobby project for a small user base, not for production servers.
Who it is for: Self-hosters running EXL3 models on Nvidia or AMD GPUs
Strengths
- Continuous batching with paged attention on Nvidia Ampere and newer GPUs
- Speculative decoding with draft models; JSON schema, regex and EBNF constraints
- Load, unload and download models at runtime without restarting the server
- Published Docker images for CUDA 12.8, CUDA 13 and ROCm
Weaknesses
- README says it is not meant for production servers
- Only EXL3 and FP16/BF16 models; no GGUF support listed
- Rolling release with no tagged releases; dependencies may need reinstalling
- AGPL-3.0 license may restrict use in some network services
What it needs
- GPU required
- Docker + Compose
- Needs ExLlamaV3, NVIDIA container toolkit (Docker)
- Models: EXL3, FP16, BF16
- port 5000
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |