#44th of 11 in Voice
Kokoro-FastAPI
OpenAI-compatible Kokoro-82M speech API in CPU and GPU images
- Stars
- 5.5k
- License
- Apache-2.0
- Last commit
- Oct 2026
- Last release
- Sep 2026
Overview
Serves the Kokoro-82M model behind an OpenAI-compatible /v1/audio/speech endpoint on port 8880, streaming mp3, wav, opus, flac, aac or pcm. Covers English (US/GB), Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese and Mandarin, with weighted voice mixing, inline [voice:] and [pause:] tags, word timestamps and phoneme endpoints. Prebuilt images exist for CPU, CUDA (amd64 and arm64) and experimental ROCm.
Who it is for: Self-hosters wanting an OpenAI-style TTS endpoint
Strengths
- Drop-in for the OpenAI Python client; models baked into the images
- Weighted voice mixing and inline speaker, pause, rate and IPA tags
- Per-word timestamp captions and phoneme in/out endpoints
- First-token latency about 300 ms on GPU
Weaknesses
- CPU first-token latency: 3.5 s on an older i7, under 1 s on M3 Pro
- No true voice cloning; /dev/tune only nudges toward a reference clip
- ROCm image is experimental and amd64 only
- Apple Silicon GPU (MPS) only when run natively via uv, not in Docker
What it needs
- GPU optional
- Docker + Compose
- Needs espeak-ng (optional fallback)
- Models: Kokoro-82M v1.0
- port 8880
Also in Voice
See all 11| Rank | Project | Score |
|---|---|---|
| 1 | VoiceboxLocal voice studio for cloning, TTS, dictation and agent speech | 72 out of 100 |
| 2 | Speech-to-SpeechModular voice-agent pipeline exposed through the OpenAI Realtime API | 67 out of 100 |
| 3 | Pocket TTS100M-parameter CPU text-to-speech with streaming and voice cloning | 64 out of 100 |
| #5 | IndexTTSZero-shot TTS with emotion, speed and pronunciation control | 61 out of 100 |
| #6 | GPT-SoVITSFew-shot voice cloning and TTS with a training web UI | 59 out of 100 |