#1111th of 11 in Voice

Speaches

OpenAI-compatible STT and TTS server with faster-whisper, Kokoro and Piper

Stars
3.7k
License
MIT
Last commit
Apr 2026
Last release
Dec 2025

Overview

Exposes OpenAI-style audio endpoints: streaming transcription and translation through faster-whisper, speech generation through Kokoro and Piper, plus a Realtime API and audio chat completions. Models load on first request and unload after inactivity, on CPU or GPU, via Docker Compose. For self-hosters who want one container that OpenAI SDKs can talk to for speech.

Who it is for: Self-hosters replacing OpenAI audio endpoints

Strengths

  • Works with any OpenAI SDK; transcription streams over SSE
  • Dynamic model loading and unloading after idle time
  • Supports the Realtime API and audio-in, audio-out chat completions
  • CPU and GPU Docker images with Compose files

Weaknesses

  • Last commit April 2026; development has slowed
  • README is short; port, env vars and limits live only in the external docs
  • TTS limited to Kokoro and Piper models
  • Streaming transcription demo is marked TODO in the README

What it needs

  • GPU optional
  • Docker + Compose
  • Models: faster-whisper (CTranslate2 Whisper), Kokoro, Piper

Also in Voice

See all 11
Also in Voice
RankProjectScore
1VoiceboxLocal voice studio for cloning, TTS, dictation and agent speech56.8k stars, MIT72 out of 100
2Speech-to-SpeechModular voice-agent pipeline exposed through the OpenAI Realtime API13.4k stars, Apache-2.067 out of 100
3Pocket TTS100M-parameter CPU text-to-speech with streaming and voice cloning Live demo ↗ (opens in a new tab)9.9k stars, MIT64 out of 100
#4Kokoro-FastAPIOpenAI-compatible Kokoro-82M speech API in CPU and GPU images Live demo ↗ (opens in a new tab)5.5k stars, Apache-2.064 out of 100
#5IndexTTSZero-shot TTS with emotion, speed and pronunciation control Live demo ↗ (opens in a new tab)24.4k stars, custom license61 out of 100