#1515th of 20 in Model serving

GPUStack

GPU cluster manager that deploys models on vLLM, SGLang and TensorRT-LLM

Stars
5.8k
License
Apache-2.0
Last commit
Oct 2026
Last release
Jul 2026

Overview

GPUStack is a GPU cluster manager that deploys models across on-prem, Kubernetes and cloud workers, configuring vLLM, SGLang, TensorRT-LLM or custom engines behind OpenAI-compatible APIs with auth, API keys and token metering. The server is one Docker container on port 80 and can run CPU-only; Linux workers join with a privileged Docker command. It supports NVIDIA, AMD, Ascend and six Chinese accelerator families.

Who it is for: ops teams running shared GPU fleets as a model service

Strengths

  • Multi-cluster: on-prem, Kubernetes and cloud GPUs under one server
  • Auto-selects and tunes vLLM, SGLang or TensorRT-LLM per model
  • Built-in auth, API keys, token metering, Grafana and Prometheus dashboards
  • SSH-accessible GPU instances on demand for fine-tuning

Weaknesses

  • Workers are Linux-only; macOS cannot be a worker, Windows needs WSL2
  • Worker container runs privileged with the Docker socket mounted
  • Cluster topology view is in the paid GPUStack Enterprise
  • Quick start assumes an NVIDIA GPU; other vendors need extra steps

What it needs

  • GPU required
  • Compose
  • Models: catalog models (e.g. Qwen3.5-0.8B) via vLLM, SGLang, TensorRT-LLM, LLM, voice, image and video models
  • port 80

Also in Model serving

See all 20
Also in Model serving
RankProjectScore
1LocalAIOne OpenAI-compatible server for text, speech, image and video models49.5k stars, MIT82 out of 100
2llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API130.7k stars, MIT79 out of 100
3vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs93.5k stars, Apache-2.075 out of 100
#4OllamaRuns open-weight models locally behind a CLI and REST API182.6k stars, MIT74 out of 100
#5colibriC inference engine that runs huge MoE models by streaming experts from disk41k stars, Apache-2.070 out of 100