#77th of 20 in Model serving

Lemonade

Local AI server that targets GPUs and AMD NPUs with OpenAI-style APIs

Stars
5.9k
License
Apache-2.0
Last commit
Oct 2026
Last release
Oct 2026

Overview

Lemonade is a local AI server with OpenAI, Anthropic and Ollama-compatible APIs on port 13305 that runs GGUF, FLM and ONNX models, Whisper transcription, Kokoro speech and Stable Diffusion images. It picks the backend for the hardware: llama.cpp on CPU, CUDA, Vulkan, ROCm or Metal, plus AMD XDNA2 NPU paths for Ryzen AI. Packages exist for Windows, macOS, Debian, Fedora, Ubuntu, Arch, Snap and Docker.

Who it is for: PC users with AMD or NVIDIA hardware wanting a local API

Strengths

  • NPU backends for AMD XDNA2 (Ryzen AI) alongside CUDA, ROCm, Vulkan, Metal
  • Chat, speech-to-text, text-to-speech, image and audio generation in one server
  • Native packages: msi, pkg, deb, rpm, Arch, Snap, PPA, Docker
  • Model aliases enable active-standby failover between models

Weaknesses

  • Many engines (vllm, ds4, openmoss, trellis) are marked experimental
  • NPU support covers AMD XDNA2 only
  • macOS gets Metal only; several backends are Windows or Linux only
  • Cloud offload to OpenAI-compatible providers is experimental

What it needs

  • GPU optional
  • Docker
  • Models: GGUF, FLM and ONNX LLMs (e.g. Gemma 4, Qwen3), Whisper, Kokoro, SDXL-Turbo
  • port 13305

Also in Model serving

See all 20
Also in Model serving
RankProjectScore
1LocalAIOne OpenAI-compatible server for text, speech, image and video models49.5k stars, MIT82 out of 100
2llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API130.7k stars, MIT79 out of 100
3vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs93.5k stars, Apache-2.075 out of 100
#4OllamaRuns open-weight models locally behind a CLI and REST API182.6k stars, MIT74 out of 100
#5colibriC inference engine that runs huge MoE models by streaming experts from disk41k stars, Apache-2.070 out of 100