#1717th of 20 in Model serving
llamafile
Single-file executables that bundle llama.cpp with model weights
- Stars
- 26.2k
- License
- custom license
- Last commit
- Oct 2026
- Last release
- Sep 2026
Overview
llamafile packages llama.cpp and model weights into one executable using Cosmopolitan Libc, so a downloaded .llamafile runs on Linux, macOS, Windows and BSD across CPU architectures with no install and serves a local web UI and API. Since 0.10 it tracks upstream llama.cpp closely for newer models, and whisperfile applies the same packaging to speech-to-text. Maintained by Mozilla.ai.
Who it is for: people who want a model that runs with zero setup
Strengths
- Single file, no installation, runs across OSes and CPU architectures
- 0.10 build system tracks upstream llama.cpp for recent model support
- Can run external GGUF weights with the bare llamafile binary
- whisperfile gives single-file transcription and translation
Weaknesses
- Windows cannot run executables above 4 GB; larger models need external weights
- 0.10.x dropped some classic features; older releases remain for those
- Pre-built llamafiles limited to Mozilla.ai's Hugging Face uploads
- One model per file; not a multi-model server
What it needs
- GPU optional
- Models: GGUF (bundled or external), e.g. Qwen3.5-0.8B
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |