Speech-to-Speech vs Kokoro-FastAPI
Two of the top voice, side by side: score, setup, license, activity and what each review found.
Speech-to-Speech
Modular voice-agent pipeline exposed through the OpenAI Realtime API
Kokoro-FastAPI
OpenAI-compatible Kokoro-82M speech API in CPU and GPU images
| What we compare | Speech-to-Speech | Kokoro-FastAPI |
|---|---|---|
| Score parts, out of 100 | ||
| Adoption | 52, popular | 31, known |
| Freshness | 100, active | 100, active |
| Maintenance | 94, healthy | 98, healthy |
| Easy to run | 50, easy | 50, easy |
| Agent-ready | 30, minimal | 45, minimal |
| Facts from GitHub and the README | ||
| Stars | 13.4k | 5.5k |
| License | Apache-2.0 (permissive) | Apache-2.0 (permissive) |
| Last commit | Oct 2026 | Oct 2026 |
| Last release | Sep 2026 | Sep 2026 |
| Language | Python | Not stated |
| Docker | Yes | Yes |
| GPU | Optional | Optional |
| arm64 or Apple Silicon | Mentioned | Mentioned |
Speech-to-Speech
Runs a VAD, STT, LLM and TTS cascade, with each stage in its own thread and every backend swappable by CLI flag. It serves the core OpenAI Realtime event set over WebSocket and WebRTC, so existing Realtime clients can point at it. Defaults are Parakeet TDT for speech recognition and Qwen3-TTS for speech output, with the LLM running locally or through any OpenAI-compatible endpoint.
Who it is for: Developers building self-hosted voice agents or Realtime API backends
Strengths
- Implements core OpenAI Realtime events over WebSocket and WebRTC, so client swaps are easy
- Fully local on Apple Silicon (MLX) or NVIDIA CUDA, with no API key needed
- Many interchangeable STT and TTS backends, including Whisper, Kokoro, Pocket TTS and OmniVoice
- Apache-2.0, installable from PyPI, with a packaged microphone client
Weaknesses
- Qwen3-TTS GGML wheel targets CUDA 12.8 and glibc 2.39 by default
- Fully local NVIDIA setup budgets about 24 GB VRAM; the README calls this an estimate
- Only the core Realtime event set is implemented, not the full API
- Some extras conflict, e.g. DeepFilterNet needs numpy<2 while Pocket TTS needs numpy>=2
- RAM ≥ 16 GB
- GPU optional
- Docker + Compose
- Needs OpenAI-compatible LLM server (optional), PortAudio and libsndfile on Ubuntu
- Models: Parakeet TDT, Qwen3-TTS, Whisper, Kokoro-82M, Transformers LLMs
- port 8765
Kokoro-FastAPI
Serves the Kokoro-82M model behind an OpenAI-compatible /v1/audio/speech endpoint on port 8880, streaming mp3, wav, opus, flac, aac or pcm. Covers English (US/GB), Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese and Mandarin, with weighted voice mixing, inline [voice:] and [pause:] tags, word timestamps and phoneme endpoints. Prebuilt images exist for CPU, CUDA (amd64 and arm64) and experimental ROCm.
Who it is for: Self-hosters wanting an OpenAI-style TTS endpoint
Strengths
- Drop-in for the OpenAI Python client; models baked into the images
- Weighted voice mixing and inline speaker, pause, rate and IPA tags
- Per-word timestamp captions and phoneme in/out endpoints
- First-token latency about 300 ms on GPU
Weaknesses
- CPU first-token latency: 3.5 s on an older i7, under 1 s on M3 Pro
- No true voice cloning; /dev/tune only nudges toward a reference clip
- ROCm image is experimental and amd64 only
- Apple Silicon GPU (MPS) only when run natively via uv, not in Docker
- GPU optional
- Docker + Compose
- Needs espeak-ng (optional fallback)
- Models: Kokoro-82M v1.0
- port 8880