#55th of 20 in Model serving
colibri
C inference engine that runs huge MoE models by streaming experts from disk
Documentation ↗
(opens in a new tab)Website ↗
(opens in a new tab)Repository on GitHub ↗
(opens in a new tab)
- Stars
- 41k
- License
- Apache-2.0
- Last commit
- Oct 2026
- Last release
- Oct 2026
- Language
- C
Overview
colibri runs large mixture-of-experts models such as GLM-5.2 (744B) and Kimi K3 (2.8T) on ordinary hardware. It keeps the dense weights in RAM and reads routed experts from disk through a cache, with optional Vulkan or CUDA offload. It serves a browser dashboard plus OpenAI- and Anthropic-compatible HTTP endpoints, and ships guided setup scripts for Windows, Linux and macOS.
Who it is for: Self-hosters running 100B+ MoE models on consumer RAM, disk and optional GPU
Strengths
- Pure C engines, no GPU required; 8 GB RAM minimum for the smallest model
- Serves OpenAI and Anthropic-style APIs on port 8000, plus a web dashboard
- Vulkan works on AMD, Intel and NVIDIA; CUDA path for NVIDIA on Linux and Windows
- Setup script detects hardware, recommends a model, and resumes interrupted downloads
Weaknesses
- Large models stream from disk: 0.05-3 tok/s on typical machines, per README tables
- Discrete-GPU performance mostly unmeasured by the authors; relies on user reports
- Supports a fixed list of model families, one engine each; other architectures need new code
- Some models need manual conversion or preparation steps after download
What it needs
- RAM ≥ 8 GB
- GPU optional
- Docker + Compose
- Models: GLM-5.2, GLM-5.3, GLM-5.3-Flash, Inkling, Kimi K3
- port 8000
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #6 | SGLangInference server for LLM, vision-language and diffusion models | 68 out of 100 |