#1919th of 20 in Model serving

TabbyAPI

OpenAI-compatible API server for running ExLlamaV3 models

Stars
1.5k
License
AGPL-3.0
Last commit
Oct 2026
Language
Python

Overview

TabbyAPI is a FastAPI server that loads and serves LLMs through the ExLlamaV3 backend, exposing an OpenAI-compatible HTTP API. It supports runtime model loading and unloading, HuggingFace downloads, embedding models, constrained output (JSON schema, regex, EBNF), and tool calling. The README describes it as a hobby project for a small user base, not for production servers.

Who it is for: Self-hosters running EXL3 models on Nvidia or AMD GPUs

Strengths

  • Continuous batching with paged attention on Nvidia Ampere and newer GPUs
  • Speculative decoding with draft models; JSON schema, regex and EBNF constraints
  • Load, unload and download models at runtime without restarting the server
  • Published Docker images for CUDA 12.8, CUDA 13 and ROCm

Weaknesses

  • README says it is not meant for production servers
  • Only EXL3 and FP16/BF16 models; no GGUF support listed
  • Rolling release with no tagged releases; dependencies may need reinstalling
  • AGPL-3.0 license may restrict use in some network services

What it needs

  • GPU required
  • Docker + Compose
  • Needs ExLlamaV3, NVIDIA container toolkit (Docker)
  • Models: EXL3, FP16, BF16
  • port 5000

Also in Model serving

See all 20
Also in Model serving
RankProjectScore
1LocalAIOne OpenAI-compatible server for text, speech, image and video models49.5k stars, MIT82 out of 100
2llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API130.8k stars, MIT79 out of 100
3vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs93.5k stars, Apache-2.075 out of 100
#4OllamaRuns open-weight models locally behind a CLI and REST API182.7k stars, MIT74 out of 100
#5colibriC inference engine that runs huge MoE models by streaming experts from disk41k stars, Apache-2.070 out of 100