#99th of 20 in Model serving

llama-swap

Go proxy that hot-swaps local model servers per request

Stars
5.9k
License
MIT
Last commit
Oct 2026
Last release
Oct 2026

Overview

llama-swap is one Go binary that proxies OpenAI and Anthropic API calls to local servers (llama-server, vLLM, stable-diffusion.cpp, whisper.cpp) and starts, stops or swaps the right one per model ID from a YAML file. It adds a web UI, log streaming, Prometheus metrics, API keys, TTL unload and a matrix DSL for concurrent models. Unified Docker images bundle the servers for CUDA and Vulkan.

Who it is for: home-lab users juggling several local models on one GPU

Strengths

  • One binary, one YAML file, zero dependencies
  • Hot-swaps any OpenAI or Anthropic-compatible upstream per model ID, with ttl unload
  • Unified images bundle llama-server, stable-diffusion.cpp, whisper.cpp, audio.cpp
  • Web UI with playground, token metrics, request inspection and live logs

Weaknesses

  • Basic mode runs one model at a time; concurrency needs the matrix DSL
  • Python servers like vLLM or tabbyAPI should run in containers for clean SIGTERM
  • nginx needs proxy_buffering off or SSE streaming breaks
  • Container listen address must stay 0.0.0.0 when publishing ports

What it needs

  • GPU optional
  • Docker
  • Needs an upstream inference server (llama-server, vLLM, etc.)
  • Models: any model served by the configured upstream (GGUF via llama-server, etc.)
  • port 8080

Also in Model serving

See all 20
Also in Model serving
RankProjectScore
1LocalAIOne OpenAI-compatible server for text, speech, image and video models49.5k stars, MIT82 out of 100
2llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API130.8k stars, MIT79 out of 100
3vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs93.5k stars, Apache-2.075 out of 100
#4OllamaRuns open-weight models locally behind a CLI and REST API182.7k stars, MIT74 out of 100
#5colibriC inference engine that runs huge MoE models by streaming experts from disk41k stars, Apache-2.070 out of 100