Voicebox vs Pocket TTS

Two of the top voice, side by side: score, setup, license, activity and what each review found.

11st of 11 in Voice

Voicebox

Local voice studio for cloning, TTS, dictation and agent speech

72 out of 100
33rd of 11 in Voice

Pocket TTS

100M-parameter CPU text-to-speech with streaming and voice cloning

Voicebox vs Pocket TTS: score parts and facts
What we compareVoiceboxPocket TTS
Score parts, out of 100
Adoption86, widely used43, known
Freshness100, active100, active
Maintenance54, fair87, healthy
Easy to run67, easy50, easy
Agent-ready0, none30, minimal
Facts from GitHub and the README
Stars56.9k9.9k
LicenseMIT (permissive)MIT (permissive)
Last commitOct 2026Oct 2026
Last releaseApr 2026Sep 2026
LanguageNot statedNot stated
DockerYesYes
GPUOptionalNot needed
arm64 or Apple SiliconMentionedMentioned

Voicebox

Desktop app (Tauri) and Docker service that clones voices from a short sample and generates speech through eight TTS engines, including Qwen3-TTS, Chatterbox and Kokoro, in 23 languages. Adds Whisper dictation with a global hotkey, a REST API on port 17493 and an MCP server so coding agents can speak in a cloned voice. For individuals who want ElevenLabs-style voice I/O on their own machine.

Who it is for: Individuals wanting local voice cloning, TTS and dictation

Strengths

  • Eight switchable TTS engines; Chatterbox Multilingual covers 23 languages
  • REST API plus HTTP and stdio MCP server for Claude Code, Cursor, Windsurf
  • Runs on MLX, CUDA, ROCm, DirectML, Intel Arc or CPU
  • Auto-chunking with crossfade handles scripts up to 50,000 characters

Weaknesses

  • No prebuilt Linux binaries; build from source or use Docker
  • Only Chatterbox Turbo honors tags like [laugh]; other engines read them aloud
  • Dictation auto-paste and the permission flow are macOS-specific
  • Docker deployment gets one line in the README; details are in external docs
  • GPU optional
  • Docker + Compose
  • Models: Qwen3-TTS 0.6B/1.7B, Qwen CustomVoice, Qwen VoiceDesign, LuxTTS, Chatterbox Multilingual
  • port 17493
  • README: alternative to ElevenLabs

Pocket TTS

Generates speech on CPU with a 100M-parameter model: about 200 ms to the first audio chunk and roughly 6x real time on an M4 MacBook Air using two cores. Covers English, French, German, Portuguese, Italian, Spanish and Dutch, clones a voice from a WAV file, and runs as a CLI, a Python library or an HTTP server with a web UI on port 8000. For developers who want TTS without a GPU.

Who it is for: Developers adding TTS to apps without a GPU

Strengths

  • Runs on 2 CPU cores; no CUDA build of PyTorch needed
  • Streaming output with about 200 ms first-chunk latency
  • Voice cloning from any WAV; export voices to safetensors for fast loading
  • Training code released; community models load via --config

Weaknesses

  • Seven European languages; others depend on community-trained models
  • No pause or silence markup in text input
  • serve command and Docker image are CPU-only; GPU use is unsupported and manual
  • Linux pip pulls CUDA PyTorch (about 3 GB) unless the CPU index is set
  • no GPU
  • Docker + Compose
  • Models: Pocket TTS 100M, 24-layer language variants, community checkpoints via --config
  • port 8000

More in Voice