Speech-to-Speech vs Pocket TTS
Two of the top voice, side by side: score, setup, license, activity and what each review found.
Speech-to-Speech
Modular voice-agent pipeline exposed through the OpenAI Realtime API
Pocket TTS
100M-parameter CPU text-to-speech with streaming and voice cloning
| What we compare | Speech-to-Speech | Pocket TTS |
|---|---|---|
| Score parts, out of 100 | ||
| Adoption | 52, popular | 43, known |
| Freshness | 100, active | 100, active |
| Maintenance | 94, healthy | 87, healthy |
| Easy to run | 50, easy | 50, easy |
| Agent-ready | 30, minimal | 30, minimal |
| Facts from GitHub and the README | ||
| Stars | 13.4k | 9.9k |
| License | Apache-2.0 (permissive) | MIT (permissive) |
| Last commit | Oct 2026 | Oct 2026 |
| Last release | Sep 2026 | Sep 2026 |
| Language | Python | Not stated |
| Docker | Yes | Yes |
| GPU | Optional | Not needed |
| arm64 or Apple Silicon | Mentioned | Mentioned |
Speech-to-Speech
Runs a VAD, STT, LLM and TTS cascade, with each stage in its own thread and every backend swappable by CLI flag. It serves the core OpenAI Realtime event set over WebSocket and WebRTC, so existing Realtime clients can point at it. Defaults are Parakeet TDT for speech recognition and Qwen3-TTS for speech output, with the LLM running locally or through any OpenAI-compatible endpoint.
Who it is for: Developers building self-hosted voice agents or Realtime API backends
Strengths
- Implements core OpenAI Realtime events over WebSocket and WebRTC, so client swaps are easy
- Fully local on Apple Silicon (MLX) or NVIDIA CUDA, with no API key needed
- Many interchangeable STT and TTS backends, including Whisper, Kokoro, Pocket TTS and OmniVoice
- Apache-2.0, installable from PyPI, with a packaged microphone client
Weaknesses
- Qwen3-TTS GGML wheel targets CUDA 12.8 and glibc 2.39 by default
- Fully local NVIDIA setup budgets about 24 GB VRAM; the README calls this an estimate
- Only the core Realtime event set is implemented, not the full API
- Some extras conflict, e.g. DeepFilterNet needs numpy<2 while Pocket TTS needs numpy>=2
- RAM ≥ 16 GB
- GPU optional
- Docker + Compose
- Needs OpenAI-compatible LLM server (optional), PortAudio and libsndfile on Ubuntu
- Models: Parakeet TDT, Qwen3-TTS, Whisper, Kokoro-82M, Transformers LLMs
- port 8765
Pocket TTS
Generates speech on CPU with a 100M-parameter model: about 200 ms to the first audio chunk and roughly 6x real time on an M4 MacBook Air using two cores. Covers English, French, German, Portuguese, Italian, Spanish and Dutch, clones a voice from a WAV file, and runs as a CLI, a Python library or an HTTP server with a web UI on port 8000. For developers who want TTS without a GPU.
Who it is for: Developers adding TTS to apps without a GPU
Strengths
- Runs on 2 CPU cores; no CUDA build of PyTorch needed
- Streaming output with about 200 ms first-chunk latency
- Voice cloning from any WAV; export voices to safetensors for fast loading
- Training code released; community models load via --config
Weaknesses
- Seven European languages; others depend on community-trained models
- No pause or silence markup in text input
- serve command and Docker image are CPU-only; GPU use is unsupported and manual
- Linux pip pulls CUDA PyTorch (about 3 GB) unless the CPU index is set
- no GPU
- Docker + Compose
- Models: Pocket TTS 100M, 24-layer language variants, community checkpoints via --config
- port 8000