Voicebox vs Kokoro-FastAPI
Two of the top voice, side by side: score, setup, license, activity and what each review found.
Voicebox
Local voice studio for cloning, TTS, dictation and agent speech
Kokoro-FastAPI
OpenAI-compatible Kokoro-82M speech API in CPU and GPU images
| What we compare | Voicebox | Kokoro-FastAPI |
|---|---|---|
| Score parts, out of 100 | ||
| Adoption | 86, widely used | 31, known |
| Freshness | 100, active | 100, active |
| Maintenance | 54, fair | 98, healthy |
| Easy to run | 67, easy | 50, easy |
| Agent-ready | 0, none | 45, minimal |
| Facts from GitHub and the README | ||
| Stars | 56.9k | 5.5k |
| License | MIT (permissive) | Apache-2.0 (permissive) |
| Last commit | Oct 2026 | Oct 2026 |
| Last release | Apr 2026 | Sep 2026 |
| Language | Not stated | Not stated |
| Docker | Yes | Yes |
| GPU | Optional | Optional |
| arm64 or Apple Silicon | Mentioned | Mentioned |
Voicebox
Desktop app (Tauri) and Docker service that clones voices from a short sample and generates speech through eight TTS engines, including Qwen3-TTS, Chatterbox and Kokoro, in 23 languages. Adds Whisper dictation with a global hotkey, a REST API on port 17493 and an MCP server so coding agents can speak in a cloned voice. For individuals who want ElevenLabs-style voice I/O on their own machine.
Who it is for: Individuals wanting local voice cloning, TTS and dictation
Strengths
- Eight switchable TTS engines; Chatterbox Multilingual covers 23 languages
- REST API plus HTTP and stdio MCP server for Claude Code, Cursor, Windsurf
- Runs on MLX, CUDA, ROCm, DirectML, Intel Arc or CPU
- Auto-chunking with crossfade handles scripts up to 50,000 characters
Weaknesses
- No prebuilt Linux binaries; build from source or use Docker
- Only Chatterbox Turbo honors tags like [laugh]; other engines read them aloud
- Dictation auto-paste and the permission flow are macOS-specific
- Docker deployment gets one line in the README; details are in external docs
- GPU optional
- Docker + Compose
- Models: Qwen3-TTS 0.6B/1.7B, Qwen CustomVoice, Qwen VoiceDesign, LuxTTS, Chatterbox Multilingual
- port 17493
- README: alternative to ElevenLabs
Kokoro-FastAPI
Serves the Kokoro-82M model behind an OpenAI-compatible /v1/audio/speech endpoint on port 8880, streaming mp3, wav, opus, flac, aac or pcm. Covers English (US/GB), Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese and Mandarin, with weighted voice mixing, inline [voice:] and [pause:] tags, word timestamps and phoneme endpoints. Prebuilt images exist for CPU, CUDA (amd64 and arm64) and experimental ROCm.
Who it is for: Self-hosters wanting an OpenAI-style TTS endpoint
Strengths
- Drop-in for the OpenAI Python client; models baked into the images
- Weighted voice mixing and inline speaker, pause, rate and IPA tags
- Per-word timestamp captions and phoneme in/out endpoints
- First-token latency about 300 ms on GPU
Weaknesses
- CPU first-token latency: 3.5 s on an older i7, under 1 s on M3 Pro
- No true voice cloning; /dev/tune only nudges toward a reference clip
- ROCm image is experimental and amd64 only
- Apple Silicon GPU (MPS) only when run natively via uv, not in Docker
- GPU optional
- Docker + Compose
- Needs espeak-ng (optional fallback)
- Models: Kokoro-82M v1.0
- port 8880