#2020th of 20 in Model serving
OpenLLM
One-command OpenAI-compatible endpoints for curated open LLMs
- Stars
- 12.6k
- License
- Apache-2.0
- Last commit
- May 2026
- Last release
- Apr 2025
Overview
OpenLLM serves open LLMs as OpenAI-compatible APIs with one command: pip install openllm, then openllm serve llama3.2:1b starts a vLLM-backed server on port 3000 with a chat UI at /chat. Models come as prebuilt Bentos from a curated repository (Llama 3.x and 4, Qwen2.5, Mistral, Phi-4, Gemma, DeepSeek R1) and the catalog states the GPU each needs. openllm deploy pushes the same Bento to BentoCloud.
Who it is for: developers wanting a quick OpenAI-compatible endpoint for curated models
Strengths
- One command gives an OpenAI API plus /chat UI on port 3000
- Catalog lists the required GPU per model tag (12 GB to 16x80 GB)
- vLLM backend for serving
- Same Bento deploys to Docker, Kubernetes or BentoCloud
Weaknesses
- Every catalog model requires a GPU; no CPU-only entries
- Custom model repositories must be public
- Catalog tops out around Llama 3.3 and Qwen2.5; last commit 2026-05-29
- Adding models means building BentoML Bentos
What it needs
- GPU required
- Models: Llama 3.1/3.2/3.3/4, Qwen2.5, Qwen2.5-Coder, QwQ, Mistral, Mistral Large, Pixtral, Phi-4, Gemma 2/3, Jamba 1.5, DeepSeek R1
- port 3000
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |