Voicebox vs Kokoro-FastAPI

Two of the top voice, side by side: score, setup, license, activity and what each review found.

11st of 11 in Voice

Voicebox

Local voice studio for cloning, TTS, dictation and agent speech

72 out of 100
#44th of 11 in Voice

Kokoro-FastAPI

OpenAI-compatible Kokoro-82M speech API in CPU and GPU images

Voicebox vs Kokoro-FastAPI: score parts and facts
What we compareVoiceboxKokoro-FastAPI
Score parts, out of 100
Adoption86, widely used31, known
Freshness100, active100, active
Maintenance54, fair98, healthy
Easy to run67, easy50, easy
Agent-ready0, none45, minimal
Facts from GitHub and the README
Stars56.9k5.5k
LicenseMIT (permissive)Apache-2.0 (permissive)
Last commitOct 2026Oct 2026
Last releaseApr 2026Sep 2026
LanguageNot statedNot stated
DockerYesYes
GPUOptionalOptional
arm64 or Apple SiliconMentionedMentioned

Voicebox

Desktop app (Tauri) and Docker service that clones voices from a short sample and generates speech through eight TTS engines, including Qwen3-TTS, Chatterbox and Kokoro, in 23 languages. Adds Whisper dictation with a global hotkey, a REST API on port 17493 and an MCP server so coding agents can speak in a cloned voice. For individuals who want ElevenLabs-style voice I/O on their own machine.

Who it is for: Individuals wanting local voice cloning, TTS and dictation

Strengths

  • Eight switchable TTS engines; Chatterbox Multilingual covers 23 languages
  • REST API plus HTTP and stdio MCP server for Claude Code, Cursor, Windsurf
  • Runs on MLX, CUDA, ROCm, DirectML, Intel Arc or CPU
  • Auto-chunking with crossfade handles scripts up to 50,000 characters

Weaknesses

  • No prebuilt Linux binaries; build from source or use Docker
  • Only Chatterbox Turbo honors tags like [laugh]; other engines read them aloud
  • Dictation auto-paste and the permission flow are macOS-specific
  • Docker deployment gets one line in the README; details are in external docs
  • GPU optional
  • Docker + Compose
  • Models: Qwen3-TTS 0.6B/1.7B, Qwen CustomVoice, Qwen VoiceDesign, LuxTTS, Chatterbox Multilingual
  • port 17493
  • README: alternative to ElevenLabs

Kokoro-FastAPI

Serves the Kokoro-82M model behind an OpenAI-compatible /v1/audio/speech endpoint on port 8880, streaming mp3, wav, opus, flac, aac or pcm. Covers English (US/GB), Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese and Mandarin, with weighted voice mixing, inline [voice:] and [pause:] tags, word timestamps and phoneme endpoints. Prebuilt images exist for CPU, CUDA (amd64 and arm64) and experimental ROCm.

Who it is for: Self-hosters wanting an OpenAI-style TTS endpoint

Strengths

  • Drop-in for the OpenAI Python client; models baked into the images
  • Weighted voice mixing and inline speaker, pause, rate and IPA tags
  • Per-word timestamp captions and phoneme in/out endpoints
  • First-token latency about 300 ms on GPU

Weaknesses

  • CPU first-token latency: 3.5 s on an older i7, under 1 s on M3 Pro
  • No true voice cloning; /dev/tune only nudges toward a reference clip
  • ROCm image is experimental and amd64 only
  • Apple Silicon GPU (MPS) only when run natively via uv, not in Docker
  • GPU optional
  • Docker + Compose
  • Needs espeak-ng (optional fallback)
  • Models: Kokoro-82M v1.0
  • port 8880

More in Voice