#1212th of 20 in Model serving

Text Generation Web UI

Local LLM chat UI and API with five switchable loader backends

Stars
47.7k
License
AGPL-3.0
Last commit
Aug 2026
Last release
May 2026

Overview

TextGen runs local LLMs behind a chat UI and an OpenAI/Anthropic-compatible API with tool calling and MCP, with llama.cpp, ik_llama.cpp, Transformers, ExLlamaV3 or TensorRT-LLM loaders switchable without restart. Portable builds for Linux, Windows and macOS bundle CUDA, Vulkan, ROCm or CPU dependencies for GGUF; the full install adds LoRA training, image generation and extensions. Web UI on port 7860.

Who it is for: hobbyists running local models with a full-featured UI

Strengths

  • Portable builds with all dependencies for CUDA, Vulkan, ROCm and CPU
  • Five loaders switchable without restarting
  • OpenAI and Anthropic-compatible API with tool calling and MCP servers
  • LoRA training and diffusers image generation in the same app

Weaknesses

  • Full install needs ~10 GB disk and PyTorch; portable build is GGUF only
  • Multi-user mode does not save chat histories; meant for small trusted teams
  • Docker needs per-GPU Dockerfile symlinks and manual .env edits
  • AGPL-3.0 license

What it needs

  • GPU optional
  • Docker + Compose
  • Models: GGUF via llama.cpp and ik_llama.cpp, Transformers safetensors, EXL3 via ExLlamaV3, TensorRT-LLM
  • port 7860

Also in Model serving

See all 20
Also in Model serving
RankProjectScore
1LocalAIOne OpenAI-compatible server for text, speech, image and video models49.5k stars, MIT82 out of 100
2llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API130.7k stars, MIT79 out of 100
3vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs93.5k stars, Apache-2.075 out of 100
#4OllamaRuns open-weight models locally behind a CLI and REST API182.6k stars, MIT74 out of 100
#5colibriC inference engine that runs huge MoE models by streaming experts from disk41k stars, Apache-2.070 out of 100