#1111th of 11 in Voice
Speaches
OpenAI-compatible STT and TTS server with faster-whisper, Kokoro and Piper
Documentation ↗
(opens in a new tab)Website ↗
(opens in a new tab)Repository on GitHub ↗
(opens in a new tab)
- Stars
- 3.7k
- License
- MIT
- Last commit
- Apr 2026
- Last release
- Dec 2025
Overview
Exposes OpenAI-style audio endpoints: streaming transcription and translation through faster-whisper, speech generation through Kokoro and Piper, plus a Realtime API and audio chat completions. Models load on first request and unload after inactivity, on CPU or GPU, via Docker Compose. For self-hosters who want one container that OpenAI SDKs can talk to for speech.
Who it is for: Self-hosters replacing OpenAI audio endpoints
Strengths
- Works with any OpenAI SDK; transcription streams over SSE
- Dynamic model loading and unloading after idle time
- Supports the Realtime API and audio-in, audio-out chat completions
- CPU and GPU Docker images with Compose files
Weaknesses
- Last commit April 2026; development has slowed
- README is short; port, env vars and limits live only in the external docs
- TTS limited to Kokoro and Piper models
- Streaming transcription demo is marked TODO in the README
What it needs
- GPU optional
- Docker + Compose
- Models: faster-whisper (CTranslate2 Whisper), Kokoro, Piper
Also in Voice
See all 11| Rank | Project | Score |
|---|---|---|
| 1 | VoiceboxLocal voice studio for cloning, TTS, dictation and agent speech | 72 out of 100 |
| 2 | Speech-to-SpeechModular voice-agent pipeline exposed through the OpenAI Realtime API | 67 out of 100 |
| 3 | Pocket TTS100M-parameter CPU text-to-speech with streaming and voice cloning | 64 out of 100 |
| #4 | Kokoro-FastAPIOpenAI-compatible Kokoro-82M speech API in CPU and GPU images | 64 out of 100 |
| #5 | IndexTTSZero-shot TTS with emotion, speed and pronunciation control | 61 out of 100 |