33rd of 20 in Model serving
vLLM
High-throughput LLM serving engine with OpenAI and Anthropic APIs
Documentation ↗
(opens in a new tab)Website ↗
(opens in a new tab)Repository on GitHub ↗
(opens in a new tab)
- Stars
- 93.5k
- License
- Apache-2.0
- Last commit
- Oct 2026
- Last release
- Oct 2026
Overview
vLLM is a Python serving engine for Hugging Face models that batches requests continuously with PagedAttention, prefix caching and speculative decoding, exposing an OpenAI-compatible API plus Anthropic Messages API and gRPC. It covers 200+ architectures (dense, MoE, multimodal, embedding) with FP8, INT8, GPTQ, AWQ and GGUF quantization and tensor, pipeline and expert parallelism.
Who it is for: teams serving open models to many concurrent users
Strengths
- Continuous batching with PagedAttention for high multi-user throughput
- 200+ Hugging Face architectures including MoE, multimodal and embedding models
- OpenAI, Anthropic Messages and gRPC endpoints with tool calling and structured output
- Runs on NVIDIA, AMD, Intel GPUs, CPUs, TPUs, Gaudi, Ascend via plugins
Weaknesses
- No web UI; API server only
- README gives no VRAM, port or model-size guidance
- Heavy Python, PyTorch and CUDA dependency chain; no single binary
- Most optimized kernels target NVIDIA and AMD GPUs; CPU path is secondary
What it needs
- GPU optional
- Docker
- Models: 200+ Hugging Face architectures: Llama, Qwen, Gemma, Mixtral, DeepSeek-V3, GPT-OSS, LLaVA, Qwen-VL, E5-Mistral
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |
| #6 | SGLangInference server for LLM, vision-language and diffusion models | 68 out of 100 |