#1010th of 20 in Model serving
Xinference
Serves LLM, embedding, speech and image models behind one OpenAI-compatible API
Documentation ↗
(opens in a new tab)Website ↗
(opens in a new tab)Repository on GitHub ↗
(opens in a new tab)
- Stars
- 9.6k
- License
- Apache-2.0
- Last commit
- Oct 2026
- Last release
- Oct 2026
- Language
- Python
Overview
Xinference is a Python library and server that runs language, embedding, speech recognition, image and multimodal models from a single command, with built-in model definitions and support for custom ones. It exposes an OpenAI-compatible REST API with function calling, plus RPC, a CLI and a web UI, and can spread models across multiple workers. It runs on GPUs and CPUs through backends including vLLM and its own llama.cpp binding.
Who it is for: Self-hosters serving many model types through one OpenAI-style endpoint
Strengths
- One server covers LLM, embedding, audio, image and multimodal models
- OpenAI-compatible REST API with function calling
- Distributed inference across workers; Helm chart for Kubernetes
- Install via pip, one-line script, Docker or Helm
Weaknesses
- Enterprise, Cloud and managed Model API offerings are commercial; edition differences not listed
- 3.0.0 release notes mention breaking changes and migration steps
- RAM and VRAM requirements are not stated; they depend on the model
- README feature list is dense; hardware and backend support per model is unclear
What it needs
- GPU optional
- Models: vLLM, ggml, llama.cpp (xllamacpp), TensorRT, embedding models
- port 9997
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |