#2020th of 20 in Model serving

OpenLLM

One-command OpenAI-compatible endpoints for curated open LLMs

Stars
12.6k
License
Apache-2.0
Last commit
May 2026
Last release
Apr 2025

Overview

OpenLLM serves open LLMs as OpenAI-compatible APIs with one command: pip install openllm, then openllm serve llama3.2:1b starts a vLLM-backed server on port 3000 with a chat UI at /chat. Models come as prebuilt Bentos from a curated repository (Llama 3.x and 4, Qwen2.5, Mistral, Phi-4, Gemma, DeepSeek R1) and the catalog states the GPU each needs. openllm deploy pushes the same Bento to BentoCloud.

Who it is for: developers wanting a quick OpenAI-compatible endpoint for curated models

Strengths

  • One command gives an OpenAI API plus /chat UI on port 3000
  • Catalog lists the required GPU per model tag (12 GB to 16x80 GB)
  • vLLM backend for serving
  • Same Bento deploys to Docker, Kubernetes or BentoCloud

Weaknesses

  • Every catalog model requires a GPU; no CPU-only entries
  • Custom model repositories must be public
  • Catalog tops out around Llama 3.3 and Qwen2.5; last commit 2026-05-29
  • Adding models means building BentoML Bentos

What it needs

  • GPU required
  • Models: Llama 3.1/3.2/3.3/4, Qwen2.5, Qwen2.5-Coder, QwQ, Mistral, Mistral Large, Pixtral, Phi-4, Gemma 2/3, Jamba 1.5, DeepSeek R1
  • port 3000

Also in Model serving

See all 20
Also in Model serving
RankProjectScore
1LocalAIOne OpenAI-compatible server for text, speech, image and video models49.5k stars, MIT82 out of 100
2llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API130.7k stars, MIT79 out of 100
3vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs93.5k stars, Apache-2.075 out of 100
#4OllamaRuns open-weight models locally behind a CLI and REST API182.6k stars, MIT74 out of 100
#5colibriC inference engine that runs huge MoE models by streaming experts from disk41k stars, Apache-2.070 out of 100