22nd of 20 in Model serving

llama.cpp

C/C++ inference engine serving GGUF models over an OpenAI-compatible API

Stars
130.8k
License
MIT
Last commit
Oct 2026
Last release
Oct 2026

Overview

llama.cpp is a C/C++ inference engine for LLMs and VLMs with no dependencies, built on ggml. llama serve starts an OpenAI-compatible API server with a built-in web UI, pulling GGUF models straight from Hugging Face, with 1.5 to 8-bit quantization and CPU+GPU hybrid offload for models larger than VRAM. Backends cover CUDA, HIP, Metal, Vulkan, SYCL, OpenCL, CANN, MUSA and WebGPU.

Who it is for: engineers who want a lean local inference server

Strengths

  • Plain C/C++ with no runtime dependencies; prebuilt binaries and Docker
  • Backends for NVIDIA, AMD, Apple Metal, Intel SYCL, Vulkan, Ascend, Moore Threads
  • Hybrid CPU+GPU offload runs models larger than available VRAM
  • Built-in web UI and OpenAI-compatible server via llama serve

Weaknesses

  • GGUF model format only
  • README gives no port, auth or sizing guidance; see tools/server docs
  • Install script is curl piped to sh; otherwise build from source
  • OpenVINO backend still in progress

What it needs

  • GPU optional
  • Docker
  • Models: GGUF models from Hugging Face (e.g. Qwen3.5-0.8B-GGUF)

Head to head

Also in Model serving

See all 20
Also in Model serving
RankProjectScore
1LocalAIOne OpenAI-compatible server for text, speech, image and video models49.5k stars, MIT82 out of 100
3vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs93.6k stars, Apache-2.075 out of 100
#4OllamaRuns open-weight models locally behind a CLI and REST API182.7k stars, MIT74 out of 100
#5colibriC inference engine that runs huge MoE models by streaming experts from disk41.1k stars, Apache-2.070 out of 100
#6SGLangInference server for LLM, vision-language and diffusion models37k stars, Apache-2.068 out of 100