11st of 11 in Voice
Voicebox
Local voice studio for cloning, TTS, dictation and agent speech
Documentation ↗
(opens in a new tab)Website ↗
(opens in a new tab)Repository on GitHub ↗
(opens in a new tab)
- Stars
- 56.8k
- License
- MIT
- Last commit
- Oct 2026
- Last release
- Apr 2026
Overview
Desktop app (Tauri) and Docker service that clones voices from a short sample and generates speech through eight TTS engines, including Qwen3-TTS, Chatterbox and Kokoro, in 23 languages. Adds Whisper dictation with a global hotkey, a REST API on port 17493 and an MCP server so coding agents can speak in a cloned voice. For individuals who want ElevenLabs-style voice I/O on their own machine.
Who it is for: Individuals wanting local voice cloning, TTS and dictation
Its own README calls it an alternative to ElevenLabs.
Strengths
- Eight switchable TTS engines; Chatterbox Multilingual covers 23 languages
- REST API plus HTTP and stdio MCP server for Claude Code, Cursor, Windsurf
- Runs on MLX, CUDA, ROCm, DirectML, Intel Arc or CPU
- Auto-chunking with crossfade handles scripts up to 50,000 characters
Weaknesses
- No prebuilt Linux binaries; build from source or use Docker
- Only Chatterbox Turbo honors tags like [laugh]; other engines read them aloud
- Dictation auto-paste and the permission flow are macOS-specific
- Docker deployment gets one line in the README; details are in external docs
What it needs
- GPU optional
- Docker + Compose
- Models: Qwen3-TTS 0.6B/1.7B, Qwen CustomVoice, Qwen VoiceDesign, LuxTTS, Chatterbox Multilingual
- port 17493
- README: alternative to ElevenLabs
Also in Voice
See all 11| Rank | Project | Score |
|---|---|---|
| 2 | Speech-to-SpeechModular voice-agent pipeline exposed through the OpenAI Realtime API | 67 out of 100 |
| 3 | Pocket TTS100M-parameter CPU text-to-speech with streaming and voice cloning | 64 out of 100 |
| #4 | Kokoro-FastAPIOpenAI-compatible Kokoro-82M speech API in CPU and GPU images | 64 out of 100 |
| #5 | IndexTTSZero-shot TTS with emotion, speed and pronunciation control | 61 out of 100 |
| #6 | GPT-SoVITSFew-shot voice cloning and TTS with a training web UI | 59 out of 100 |