#44th of 20 in Model serving

Ollama

Runs open-weight models locally behind a CLI and REST API

Stars
182.7k
License
MIT
Last commit
Oct 2026
Last release
Oct 2026

Overview

Ollama runs open-weight models locally with a CLI and a REST API on port 11434, pulling models from its own library (for example gemma4) and using llama.cpp as the inference backend. Install scripts cover macOS, Windows and Linux, and an official Docker image exists. The ollama launch command wires it into coding agents such as Claude Code, Codex, Copilot CLI and OpenCode, or into OpenClaw as a chat assistant.

Who it is for: anyone wanting local models behind a simple API

Strengths

  • One command pulls and runs a model; REST API on 11434
  • Official Docker image plus Python and JavaScript libraries
  • ollama launch integrates with Claude Code, Codex, Copilot CLI, OpenCode
  • Broad ecosystem: dozens of web, desktop and IDE clients listed

Weaknesses

  • Single inference backend: llama.cpp
  • Install is a curl piped to sh script
  • README gives no RAM or VRAM guidance per model size
  • Models come from Ollama's own registry; others need import steps

What it needs

  • GPU optional
  • Docker
  • Models: Ollama library models (e.g. gemma4), GGUF via llama.cpp
  • port 11434

Also in Model serving

See all 20
Also in Model serving
RankProjectScore
1LocalAIOne OpenAI-compatible server for text, speech, image and video models49.5k stars, MIT82 out of 100
2llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API130.8k stars, MIT79 out of 100
3vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs93.5k stars, Apache-2.075 out of 100
#5colibriC inference engine that runs huge MoE models by streaming experts from disk41k stars, Apache-2.070 out of 100
#6SGLangInference server for LLM, vision-language and diffusion models37k stars, Apache-2.068 out of 100