33rd of 11 in Voice
Pocket TTS
100M-parameter CPU text-to-speech with streaming and voice cloning
Live demo ↗
(opens in a new tab)Documentation ↗
(opens in a new tab)Repository on GitHub ↗
(opens in a new tab)
- Stars
- 9.9k
- License
- MIT
- Last commit
- Oct 2026
- Last release
- Sep 2026
Overview
Generates speech on CPU with a 100M-parameter model: about 200 ms to the first audio chunk and roughly 6x real time on an M4 MacBook Air using two cores. Covers English, French, German, Portuguese, Italian, Spanish and Dutch, clones a voice from a WAV file, and runs as a CLI, a Python library or an HTTP server with a web UI on port 8000. For developers who want TTS without a GPU.
Who it is for: Developers adding TTS to apps without a GPU
Strengths
- Runs on 2 CPU cores; no CUDA build of PyTorch needed
- Streaming output with about 200 ms first-chunk latency
- Voice cloning from any WAV; export voices to safetensors for fast loading
- Training code released; community models load via --config
Weaknesses
- Seven European languages; others depend on community-trained models
- No pause or silence markup in text input
- serve command and Docker image are CPU-only; GPU use is unsupported and manual
- Linux pip pulls CUDA PyTorch (about 3 GB) unless the CPU index is set
What it needs
- no GPU
- Docker + Compose
- Models: Pocket TTS 100M, 24-layer language variants, community checkpoints via --config
- port 8000
Also in Voice
See all 11| Rank | Project | Score |
|---|---|---|
| 1 | VoiceboxLocal voice studio for cloning, TTS, dictation and agent speech | 72 out of 100 |
| 2 | Speech-to-SpeechModular voice-agent pipeline exposed through the OpenAI Realtime API | 67 out of 100 |
| #4 | Kokoro-FastAPIOpenAI-compatible Kokoro-82M speech API in CPU and GPU images | 64 out of 100 |
| #5 | IndexTTSZero-shot TTS with emotion, speed and pronunciation control | 61 out of 100 |
| #6 | GPT-SoVITSFew-shot voice cloning and TTS with a training web UI | 59 out of 100 |