11st of 20 in Model serving
LocalAI
One OpenAI-compatible server for text, speech, image and video models
Documentation ↗
(opens in a new tab)Website ↗
(opens in a new tab)Repository on GitHub ↗
(opens in a new tab)
- Stars
- 49.5k
- License
- MIT
- Last commit
- Oct 2026
- Last release
- Oct 2026
Overview
LocalAI is a Go server on port 8080 with OpenAI, Anthropic, ElevenLabs and Ollama-compatible APIs for text, vision, speech, image and video. Backends (llama.cpp, vLLM, SGLang, whisper.cpp, diffusers, MLX, 60+ total) ship as separate OCI images pulled on demand; containers exist for CPU, CUDA, ROCm, Intel and Vulkan. It adds API keys, quotas and OIDC, agents with MCP, and a PostgreSQL/NATS distributed mode.
Who it is for: self-hosters wanting one API for LLM, speech and image models
Strengths
- Small core; 60+ backends installed on demand as OCI images
- OpenAI, Anthropic, ElevenLabs and Ollama API compatibility in one server
- Multi-user: API keys, per-user quotas, role-based access, OIDC
- Container images for CPU, CUDA 12/13, ROCm, Intel oneAPI, Vulkan, Jetson
Weaknesses
- First model load pulls backend images; needs network and disk space
- macOS DMG is unsigned and needs quarantine removal
- Distributed mode requires PostgreSQL and NATS
- Very wide scope (agents, biometrics, video) increases configuration surface
What it needs
- GPU optional
- Docker + Compose
- Needs PostgreSQL and NATS (distributed mode only)
- Models: GGUF via llama.cpp, vLLM, SGLang, transformers, MLX, diffusers, whisper.cpp backends, models from gallery, Hugging Face, Ollama registry, OCI images, YAML
- port 8080
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |
| #6 | SGLangInference server for LLM, vision-language and diffusion models | 68 out of 100 |