22nd of 11 in Voice

Speech-to-Speech

Modular voice-agent pipeline exposed through the OpenAI Realtime API

Stars
13.4k
License
Apache-2.0
Last commit
Oct 2026
Last release
Sep 2026
Language
Python

Overview

Runs a VAD, STT, LLM and TTS cascade, with each stage in its own thread and every backend swappable by CLI flag. It serves the core OpenAI Realtime event set over WebSocket and WebRTC, so existing Realtime clients can point at it. Defaults are Parakeet TDT for speech recognition and Qwen3-TTS for speech output, with the LLM running locally or through any OpenAI-compatible endpoint.

Who it is for: Developers building self-hosted voice agents or Realtime API backends

Strengths

  • Implements core OpenAI Realtime events over WebSocket and WebRTC, so client swaps are easy
  • Fully local on Apple Silicon (MLX) or NVIDIA CUDA, with no API key needed
  • Many interchangeable STT and TTS backends, including Whisper, Kokoro, Pocket TTS and OmniVoice
  • Apache-2.0, installable from PyPI, with a packaged microphone client

Weaknesses

  • Qwen3-TTS GGML wheel targets CUDA 12.8 and glibc 2.39 by default
  • Fully local NVIDIA setup budgets about 24 GB VRAM; the README calls this an estimate
  • Only the core Realtime event set is implemented, not the full API
  • Some extras conflict, e.g. DeepFilterNet needs numpy<2 while Pocket TTS needs numpy>=2

What it needs

  • RAM ≥ 16 GB
  • GPU optional
  • Docker + Compose
  • Needs OpenAI-compatible LLM server (optional), PortAudio and libsndfile on Ubuntu
  • Models: Parakeet TDT, Qwen3-TTS, Whisper, Kokoro-82M, Transformers LLMs
  • port 8765

Also in Voice

See all 11
Also in Voice
RankProjectScore
1VoiceboxLocal voice studio for cloning, TTS, dictation and agent speech56.8k stars, MIT72 out of 100
3Pocket TTS100M-parameter CPU text-to-speech with streaming and voice cloning Live demo ↗ (opens in a new tab)9.9k stars, MIT64 out of 100
#4Kokoro-FastAPIOpenAI-compatible Kokoro-82M speech API in CPU and GPU images Live demo ↗ (opens in a new tab)5.5k stars, Apache-2.064 out of 100
#5IndexTTSZero-shot TTS with emotion, speed and pronunciation control Live demo ↗ (opens in a new tab)24.4k stars, custom license61 out of 100
#6GPT-SoVITSFew-shot voice cloning and TTS with a training web UI Live demo ↗ (opens in a new tab)62.6k stars, MIT59 out of 100