#99th of 20 in Model serving
llama-swap
Go proxy that hot-swaps local model servers per request
- Stars
- 5.9k
- License
- MIT
- Last commit
- Oct 2026
- Last release
- Oct 2026
Overview
llama-swap is one Go binary that proxies OpenAI and Anthropic API calls to local servers (llama-server, vLLM, stable-diffusion.cpp, whisper.cpp) and starts, stops or swaps the right one per model ID from a YAML file. It adds a web UI, log streaming, Prometheus metrics, API keys, TTL unload and a matrix DSL for concurrent models. Unified Docker images bundle the servers for CUDA and Vulkan.
Who it is for: home-lab users juggling several local models on one GPU
Strengths
- One binary, one YAML file, zero dependencies
- Hot-swaps any OpenAI or Anthropic-compatible upstream per model ID, with ttl unload
- Unified images bundle llama-server, stable-diffusion.cpp, whisper.cpp, audio.cpp
- Web UI with playground, token metrics, request inspection and live logs
Weaknesses
- Basic mode runs one model at a time; concurrency needs the matrix DSL
- Python servers like vLLM or tabbyAPI should run in containers for clean SIGTERM
- nginx needs proxy_buffering off or SSE streaming breaks
- Container listen address must stay 0.0.0.0 when publishing ports
What it needs
- GPU optional
- Docker
- Needs an upstream inference server (llama-server, vLLM, etc.)
- Models: any model served by the configured upstream (GGUF via llama-server, etc.)
- port 8080
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |