#66th of 20 in Model serving

SGLang

Inference server for LLM, vision-language and diffusion models

Stars
37k
License
Apache-2.0
Last commit
Oct 2026
Last release
Oct 2026
Language
Python

Overview

SGLang is a Python inference framework for serving large language, vision-language and diffusion models, aimed at agentic workloads, large-scale serving and RL rollouts. It runs on NVIDIA, AMD, Google TPU, Intel, Apple Silicon, Huawei Ascend and Moore Threads hardware, and includes a built-in engine for image and video generation. Install via the lmsysorg/sglang Docker image or uv pip.

Who it is for: Engineers serving LLMs at scale or running RL rollouts on GPU clusters

Strengths

  • Supports NVIDIA, AMD, TPU, Intel, Apple Silicon, Ascend and Moore Threads hardware
  • Image and video diffusion engine ships in the same package
  • Integrated by RL training frameworks such as verl and slime for rollouts
  • Apache-2.0 license, with a Docker image and a cookbook of launch commands

Weaknesses

  • Audio TTS/ASR serving lives in a separate project, SGLang Omni
  • Install requires --prerelease=allow with uv, suggesting prerelease dependencies
  • Several hardware backends (Trainium, Cambricon, MetaX) are still in progress
  • README gives no RAM or VRAM requirements; sizing depends on model and hardware

What it needs

  • Docker + Compose
  • Models: LLMs, vision-language models, diffusion models

Also in Model serving

See all 20
Also in Model serving
RankProjectScore
1LocalAIOne OpenAI-compatible server for text, speech, image and video models49.5k stars, MIT82 out of 100
2llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API130.7k stars, MIT79 out of 100
3vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs93.5k stars, Apache-2.075 out of 100
#4OllamaRuns open-weight models locally behind a CLI and REST API182.6k stars, MIT74 out of 100
#5colibriC inference engine that runs huge MoE models by streaming experts from disk41k stars, Apache-2.070 out of 100