#1717th of 20 in Model serving

llamafile

Single-file executables that bundle llama.cpp with model weights

Stars
26.2k
License
custom license
Last commit
Oct 2026
Last release
Sep 2026

Overview

llamafile packages llama.cpp and model weights into one executable using Cosmopolitan Libc, so a downloaded .llamafile runs on Linux, macOS, Windows and BSD across CPU architectures with no install and serves a local web UI and API. Since 0.10 it tracks upstream llama.cpp closely for newer models, and whisperfile applies the same packaging to speech-to-text. Maintained by Mozilla.ai.

Who it is for: people who want a model that runs with zero setup

Strengths

  • Single file, no installation, runs across OSes and CPU architectures
  • 0.10 build system tracks upstream llama.cpp for recent model support
  • Can run external GGUF weights with the bare llamafile binary
  • whisperfile gives single-file transcription and translation

Weaknesses

  • Windows cannot run executables above 4 GB; larger models need external weights
  • 0.10.x dropped some classic features; older releases remain for those
  • Pre-built llamafiles limited to Mozilla.ai's Hugging Face uploads
  • One model per file; not a multi-model server

What it needs

  • GPU optional
  • Models: GGUF (bundled or external), e.g. Qwen3.5-0.8B

Also in Model serving

See all 20
Also in Model serving
RankProjectScore
1LocalAIOne OpenAI-compatible server for text, speech, image and video models49.5k stars, MIT82 out of 100
2llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API130.7k stars, MIT79 out of 100
3vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs93.5k stars, Apache-2.075 out of 100
#4OllamaRuns open-weight models locally behind a CLI and REST API182.6k stars, MIT74 out of 100
#5colibriC inference engine that runs huge MoE models by streaming experts from disk41k stars, Apache-2.070 out of 100