#1212th of 20 in Model serving
Text Generation Web UI
Local LLM chat UI and API with five switchable loader backends
- Stars
- 47.7k
- License
- AGPL-3.0
- Last commit
- Aug 2026
- Last release
- May 2026
Overview
TextGen runs local LLMs behind a chat UI and an OpenAI/Anthropic-compatible API with tool calling and MCP, with llama.cpp, ik_llama.cpp, Transformers, ExLlamaV3 or TensorRT-LLM loaders switchable without restart. Portable builds for Linux, Windows and macOS bundle CUDA, Vulkan, ROCm or CPU dependencies for GGUF; the full install adds LoRA training, image generation and extensions. Web UI on port 7860.
Who it is for: hobbyists running local models with a full-featured UI
Strengths
- Portable builds with all dependencies for CUDA, Vulkan, ROCm and CPU
- Five loaders switchable without restarting
- OpenAI and Anthropic-compatible API with tool calling and MCP servers
- LoRA training and diffusers image generation in the same app
Weaknesses
- Full install needs ~10 GB disk and PyTorch; portable build is GGUF only
- Multi-user mode does not save chat histories; meant for small trusted teams
- Docker needs per-GPU Dockerfile symlinks and manual .env edits
- AGPL-3.0 license
What it needs
- GPU optional
- Docker + Compose
- Models: GGUF via llama.cpp and ik_llama.cpp, Transformers safetensors, EXL3 via ExLlamaV3, TensorRT-LLM
- port 7860
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |