#1515th of 20 in Model serving
GPUStack
GPU cluster manager that deploys models on vLLM, SGLang and TensorRT-LLM
- Stars
- 5.8k
- License
- Apache-2.0
- Last commit
- Oct 2026
- Last release
- Jul 2026
Overview
GPUStack is a GPU cluster manager that deploys models across on-prem, Kubernetes and cloud workers, configuring vLLM, SGLang, TensorRT-LLM or custom engines behind OpenAI-compatible APIs with auth, API keys and token metering. The server is one Docker container on port 80 and can run CPU-only; Linux workers join with a privileged Docker command. It supports NVIDIA, AMD, Ascend and six Chinese accelerator families.
Who it is for: ops teams running shared GPU fleets as a model service
Strengths
- Multi-cluster: on-prem, Kubernetes and cloud GPUs under one server
- Auto-selects and tunes vLLM, SGLang or TensorRT-LLM per model
- Built-in auth, API keys, token metering, Grafana and Prometheus dashboards
- SSH-accessible GPU instances on demand for fine-tuning
Weaknesses
- Workers are Linux-only; macOS cannot be a worker, Windows needs WSL2
- Worker container runs privileged with the Docker socket mounted
- Cluster topology view is in the paid GPUStack Enterprise
- Quick start assumes an NVIDIA GPU; other vendors need extra steps
What it needs
- GPU required
- Compose
- Models: catalog models (e.g. Qwen3.5-0.8B) via vLLM, SGLang, TensorRT-LLM, LLM, voice, image and video models
- port 80
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |