#77th of 20 in Model serving
Lemonade
Local AI server that targets GPUs and AMD NPUs with OpenAI-style APIs
- Stars
- 5.9k
- License
- Apache-2.0
- Last commit
- Oct 2026
- Last release
- Oct 2026
Overview
Lemonade is a local AI server with OpenAI, Anthropic and Ollama-compatible APIs on port 13305 that runs GGUF, FLM and ONNX models, Whisper transcription, Kokoro speech and Stable Diffusion images. It picks the backend for the hardware: llama.cpp on CPU, CUDA, Vulkan, ROCm or Metal, plus AMD XDNA2 NPU paths for Ryzen AI. Packages exist for Windows, macOS, Debian, Fedora, Ubuntu, Arch, Snap and Docker.
Who it is for: PC users with AMD or NVIDIA hardware wanting a local API
Strengths
- NPU backends for AMD XDNA2 (Ryzen AI) alongside CUDA, ROCm, Vulkan, Metal
- Chat, speech-to-text, text-to-speech, image and audio generation in one server
- Native packages: msi, pkg, deb, rpm, Arch, Snap, PPA, Docker
- Model aliases enable active-standby failover between models
Weaknesses
- Many engines (vllm, ds4, openmoss, trellis) are marked experimental
- NPU support covers AMD XDNA2 only
- macOS gets Metal only; several backends are Windows or Linux only
- Cloud offload to OpenAI-compatible providers is experimental
What it needs
- GPU optional
- Docker
- Models: GGUF, FLM and ONNX LLMs (e.g. Gemma 4, Qwen3), Whisper, Kokoro, SDXL-Turbo
- port 13305
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |