Pocket TTS vs Kokoro-FastAPI
Two of the top voice, side by side: score, setup, license, activity and what each review found.
Pocket TTS
100M-parameter CPU text-to-speech with streaming and voice cloning
Kokoro-FastAPI
OpenAI-compatible Kokoro-82M speech API in CPU and GPU images
| What we compare | Pocket TTS | Kokoro-FastAPI |
|---|---|---|
| Score parts, out of 100 | ||
| Adoption | 43, known | 31, known |
| Freshness | 100, active | 100, active |
| Maintenance | 87, healthy | 98, healthy |
| Easy to run | 50, easy | 50, easy |
| Agent-ready | 30, minimal | 45, minimal |
| Facts from GitHub and the README | ||
| Stars | 9.9k | 5.5k |
| License | MIT (permissive) | Apache-2.0 (permissive) |
| Last commit | Oct 2026 | Oct 2026 |
| Last release | Sep 2026 | Sep 2026 |
| Language | Not stated | Not stated |
| Docker | Yes | Yes |
| GPU | Not needed | Optional |
| arm64 or Apple Silicon | Mentioned | Mentioned |
Pocket TTS
Generates speech on CPU with a 100M-parameter model: about 200 ms to the first audio chunk and roughly 6x real time on an M4 MacBook Air using two cores. Covers English, French, German, Portuguese, Italian, Spanish and Dutch, clones a voice from a WAV file, and runs as a CLI, a Python library or an HTTP server with a web UI on port 8000. For developers who want TTS without a GPU.
Who it is for: Developers adding TTS to apps without a GPU
Strengths
- Runs on 2 CPU cores; no CUDA build of PyTorch needed
- Streaming output with about 200 ms first-chunk latency
- Voice cloning from any WAV; export voices to safetensors for fast loading
- Training code released; community models load via --config
Weaknesses
- Seven European languages; others depend on community-trained models
- No pause or silence markup in text input
- serve command and Docker image are CPU-only; GPU use is unsupported and manual
- Linux pip pulls CUDA PyTorch (about 3 GB) unless the CPU index is set
- no GPU
- Docker + Compose
- Models: Pocket TTS 100M, 24-layer language variants, community checkpoints via --config
- port 8000
Kokoro-FastAPI
Serves the Kokoro-82M model behind an OpenAI-compatible /v1/audio/speech endpoint on port 8880, streaming mp3, wav, opus, flac, aac or pcm. Covers English (US/GB), Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese and Mandarin, with weighted voice mixing, inline [voice:] and [pause:] tags, word timestamps and phoneme endpoints. Prebuilt images exist for CPU, CUDA (amd64 and arm64) and experimental ROCm.
Who it is for: Self-hosters wanting an OpenAI-style TTS endpoint
Strengths
- Drop-in for the OpenAI Python client; models baked into the images
- Weighted voice mixing and inline speaker, pause, rate and IPA tags
- Per-word timestamp captions and phoneme in/out endpoints
- First-token latency about 300 ms on GPU
Weaknesses
- CPU first-token latency: 3.5 s on an older i7, under 1 s on M3 Pro
- No true voice cloning; /dev/tune only nudges toward a reference clip
- ROCm image is experimental and amd64 only
- Apple Silicon GPU (MPS) only when run natively via uv, not in Docker
- GPU optional
- Docker + Compose
- Needs espeak-ng (optional fallback)
- Models: Kokoro-82M v1.0
- port 8880