22nd of 20 in Model serving
llama.cpp
C/C++ inference engine serving GGUF models over an OpenAI-compatible API
- Stars
- 130.8k
- License
- MIT
- Last commit
- Oct 2026
- Last release
- Oct 2026
Overview
llama.cpp is a C/C++ inference engine for LLMs and VLMs with no dependencies, built on ggml. llama serve starts an OpenAI-compatible API server with a built-in web UI, pulling GGUF models straight from Hugging Face, with 1.5 to 8-bit quantization and CPU+GPU hybrid offload for models larger than VRAM. Backends cover CUDA, HIP, Metal, Vulkan, SYCL, OpenCL, CANN, MUSA and WebGPU.
Who it is for: engineers who want a lean local inference server
Strengths
- Plain C/C++ with no runtime dependencies; prebuilt binaries and Docker
- Backends for NVIDIA, AMD, Apple Metal, Intel SYCL, Vulkan, Ascend, Moore Threads
- Hybrid CPU+GPU offload runs models larger than available VRAM
- Built-in web UI and OpenAI-compatible server via llama serve
Weaknesses
- GGUF model format only
- README gives no port, auth or sizing guidance; see tools/server docs
- Install script is curl piped to sh; otherwise build from source
- OpenVINO backend still in progress
What it needs
- GPU optional
- Docker
- Models: GGUF models from Hugging Face (e.g. Qwen3.5-0.8B-GGUF)
Head to head
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |
| #6 | SGLangInference server for LLM, vision-language and diffusion models | 68 out of 100 |