#1616th of 20 in Model serving
LMDeploy
LLM and VLM serving toolkit with the TurboMind and PyTorch engines
- Stars
- 8.1k
- License
- Apache-2.0
- Last commit
- Oct 2026
- Last release
- Sep 2026
Overview
LMDeploy compresses and serves LLMs and VLMs with two engines: TurboMind (CUDA, persistent batching, blocked KV cache, AWQ W4A16, MXFP4) and a pure-Python PyTorch engine that also runs on Huawei Ascend. pip install lmdeploy adds an API server plus a proxy for multi-model, multi-machine serving. Models span Llama, Qwen3, DeepSeek-V3/V4, GLM-5, InternVL and Qwen3-VL.
Who it is for: teams serving LLMs and VLMs on NVIDIA or Ascend hardware
Strengths
- TurboMind engine with persistent batching, blocked KV cache and 4-bit AWQ inference
- Online INT8/INT4 KV cache quantization and prefix caching usable together
- Wide VLM list: InternVL 1 to 3.5, Qwen2/2.5/3-VL, LLaVA, Gemma3, Llama4
- PyTorch engine supports Huawei Ascend NPUs with graph mode
Weaknesses
- The two engines support different model sets and dtypes; check the matrix
- Prebuilt wheels target CUDA 12.8; other CUDA versions need source builds
- No port, VRAM or web UI details in the README
- Community channels are WeChat-centric alongside Discord
What it needs
- GPU required
- Docker
- Models: Llama 1-4, Qwen1.5-3.5, InternLM2/3, DeepSeek V2-V4, GLM-4/5, Mixtral, Gemma, Phi-3/4, gpt-oss, VLMs: InternVL, Qwen-VL, LLaVA, DeepSeek-VL, CogVLM, MiniCPM-V, Molmo, Gemma3, Llama4
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |