Open-source alternatives to ElevenLabs

The highest-scoring open-source projects in Voice. They are picks, not exact replacements, so read the weaknesses before you switch.

Open-source alternatives to ElevenLabs
RankProjectScore
1VoiceboxLocal voice studio for cloning, TTS, dictation and agent speechVoice, 56.8k stars, MIT72 out of 100
2Speech-to-SpeechModular voice-agent pipeline exposed through the OpenAI Realtime APIVoice, 13.4k stars, Apache-2.067 out of 100
3Pocket TTS100M-parameter CPU text-to-speech with streaming and voice cloning Live demo ↗ (opens in a new tab)Voice, 9.9k stars, MIT64 out of 100
#4Kokoro-FastAPIOpenAI-compatible Kokoro-82M speech API in CPU and GPU images Live demo ↗ (opens in a new tab)Voice, 5.5k stars, Apache-2.064 out of 100
#5IndexTTSZero-shot TTS with emotion, speed and pronunciation control Live demo ↗ (opens in a new tab)Voice, 24.4k stars, custom license61 out of 100
#6GPT-SoVITSFew-shot voice cloning and TTS with a training web UI Live demo ↗ (opens in a new tab)Voice, 62.6k stars, MIT59 out of 100
#7F5-TTSFlow-matching TTS and voice cloning with Gradio and CLI Live demo ↗ (opens in a new tab)Voice, 15.4k stars, MIT58 out of 100
#8OpenReaderReads EPUB, PDF and DOCX aloud with synced word highlightingVoice, 540 stars, MIT56 out of 100
#9SpeakrTranscribe, summarize and search recordings with pluggable ASR and LLMsVoice, 4.1k stars, AGPL-3.052 out of 100
#10WhisperLiveNear-real-time Whisper transcription server over WebSocketVoice, 4.3k stars, MIT49 out of 100

Reviews

172 out of 100

Voicebox

Local voice studio for cloning, TTS, dictation and agent speech

56.8k stars, MIT, last commit Oct 2026

Its own README calls it an alternative to ElevenLabs.

Desktop app (Tauri) and Docker service that clones voices from a short sample and generates speech through eight TTS engines, including Qwen3-TTS, Chatterbox and Kokoro, in 23 languages. Adds Whisper dictation with a global hotkey, a REST API on port 17493 and an MCP server so coding agents can speak in a cloned voice. For individuals who want ElevenLabs-style voice I/O on their own machine.

Strengths

  • Eight switchable TTS engines; Chatterbox Multilingual covers 23 languages
  • REST API plus HTTP and stdio MCP server for Claude Code, Cursor, Windsurf
  • Runs on MLX, CUDA, ROCm, DirectML, Intel Arc or CPU
  • Auto-chunking with crossfade handles scripts up to 50,000 characters

Weaknesses

  • No prebuilt Linux binaries; build from source or use Docker
  • Only Chatterbox Turbo honors tags like [laugh]; other engines read them aloud
  • Dictation auto-paste and the permission flow are macOS-specific
  • Docker deployment gets one line in the README; details are in external docs
  • GPU optional
  • Docker + Compose
  • Models: Qwen3-TTS 0.6B/1.7B, Qwen CustomVoice, Qwen VoiceDesign, LuxTTS, Chatterbox Multilingual
  • port 17493
  • README: alternative to ElevenLabs
267 out of 100

Speech-to-Speech

Modular voice-agent pipeline exposed through the OpenAI Realtime API

13.4k stars, Apache-2.0, last commit Oct 2026

Runs a VAD, STT, LLM and TTS cascade, with each stage in its own thread and every backend swappable by CLI flag. It serves the core OpenAI Realtime event set over WebSocket and WebRTC, so existing Realtime clients can point at it. Defaults are Parakeet TDT for speech recognition and Qwen3-TTS for speech output, with the LLM running locally or through any OpenAI-compatible endpoint.

Strengths

  • Implements core OpenAI Realtime events over WebSocket and WebRTC, so client swaps are easy
  • Fully local on Apple Silicon (MLX) or NVIDIA CUDA, with no API key needed
  • Many interchangeable STT and TTS backends, including Whisper, Kokoro, Pocket TTS and OmniVoice
  • Apache-2.0, installable from PyPI, with a packaged microphone client

Weaknesses

  • Qwen3-TTS GGML wheel targets CUDA 12.8 and glibc 2.39 by default
  • Fully local NVIDIA setup budgets about 24 GB VRAM; the README calls this an estimate
  • Only the core Realtime event set is implemented, not the full API
  • Some extras conflict, e.g. DeepFilterNet needs numpy<2 while Pocket TTS needs numpy>=2
  • RAM ≥ 16 GB
  • GPU optional
  • Docker + Compose
  • Needs OpenAI-compatible LLM server (optional), PortAudio and libsndfile on Ubuntu
  • Models: Parakeet TDT, Qwen3-TTS, Whisper, Kokoro-82M, Transformers LLMs
  • port 8765
364 out of 100

Pocket TTS

100M-parameter CPU text-to-speech with streaming and voice cloning

Live demo ↗ (opens in a new tab)9.9k stars, MIT, last commit Oct 2026

Generates speech on CPU with a 100M-parameter model: about 200 ms to the first audio chunk and roughly 6x real time on an M4 MacBook Air using two cores. Covers English, French, German, Portuguese, Italian, Spanish and Dutch, clones a voice from a WAV file, and runs as a CLI, a Python library or an HTTP server with a web UI on port 8000. For developers who want TTS without a GPU.

Strengths

  • Runs on 2 CPU cores; no CUDA build of PyTorch needed
  • Streaming output with about 200 ms first-chunk latency
  • Voice cloning from any WAV; export voices to safetensors for fast loading
  • Training code released; community models load via --config

Weaknesses

  • Seven European languages; others depend on community-trained models
  • No pause or silence markup in text input
  • serve command and Docker image are CPU-only; GPU use is unsupported and manual
  • Linux pip pulls CUDA PyTorch (about 3 GB) unless the CPU index is set
  • no GPU
  • Docker + Compose
  • Models: Pocket TTS 100M, 24-layer language variants, community checkpoints via --config
  • port 8000
#464 out of 100

Kokoro-FastAPI

OpenAI-compatible Kokoro-82M speech API in CPU and GPU images

Live demo ↗ (opens in a new tab)5.5k stars, Apache-2.0, last commit Oct 2026

Serves the Kokoro-82M model behind an OpenAI-compatible /v1/audio/speech endpoint on port 8880, streaming mp3, wav, opus, flac, aac or pcm. Covers English (US/GB), Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese and Mandarin, with weighted voice mixing, inline [voice:] and [pause:] tags, word timestamps and phoneme endpoints. Prebuilt images exist for CPU, CUDA (amd64 and arm64) and experimental ROCm.

Strengths

  • Drop-in for the OpenAI Python client; models baked into the images
  • Weighted voice mixing and inline speaker, pause, rate and IPA tags
  • Per-word timestamp captions and phoneme in/out endpoints
  • First-token latency about 300 ms on GPU

Weaknesses

  • CPU first-token latency: 3.5 s on an older i7, under 1 s on M3 Pro
  • No true voice cloning; /dev/tune only nudges toward a reference clip
  • ROCm image is experimental and amd64 only
  • Apple Silicon GPU (MPS) only when run natively via uv, not in Docker
  • GPU optional
  • Docker + Compose
  • Needs espeak-ng (optional fallback)
  • Models: Kokoro-82M v1.0
  • port 8880
#561 out of 100

IndexTTS

Zero-shot TTS with emotion, speed and pronunciation control

Live demo ↗ (opens in a new tab)24.4k stars, custom license, last commit Sep 2026

Clones a voice from one reference clip and synthesizes speech in Chinese, English, Japanese, Spanish and Arabic (IndexTTS-2.5). Emotion comes from a second reference clip, an 8-value vector or the text itself; speed is set by duration_factor (0.5x to 2.0x) and pronunciation by inline Pinyin, CMU phonemes or Kana. Ships a Gradio web UI on port 7860 and a Python API; a vLLM recipe covers production serving.

Strengths

  • Emotion control via reference audio, an 8-value vector or a text description
  • Inline pronunciation overrides: Pinyin, CMU phonemes and Japanese Kana
  • BF16 inference with optional DeepSpeed and compiled CUDA kernels
  • Published vLLM recipe for production deployment

Weaknesses

  • No Dockerfile or compose file; install is uv plus CUDA Toolkit 12.8 or newer
  • Model weights (IndexTTS-2.5, IndexTTS-2) are separate multi-GB downloads
  • Five languages only; no streaming API is documented in the README
  • License is non-standard (GitHub reports NOASSERTION); check terms before commercial use
  • Needs uv
  • Models: IndexTTS-2.5, IndexTTS-2, IndexTTS-1.5 (legacy)
  • port 7860
#659 out of 100

GPT-SoVITS

Few-shot voice cloning and TTS with a training web UI

Live demo ↗ (opens in a new tab)62.6k stars, MIT, last commit Oct 2026

Clones a voice from a 5-second sample (zero-shot) or fine-tunes GPT and SoVITS models on about one minute of audio, then synthesizes speech in Chinese, English, Japanese, Korean and Cantonese. The Gradio web UI bundles dataset tools: UVR5 vocal separation, slicing, ASR and label proofreading. Aimed at hobbyists and studios building custom voices locally.

Strengths

  • Zero-shot cloning from 5 s of audio; few-shot fine-tune from about 1 minute
  • Cross-lingual synthesis across zh, en, ja, ko and yue
  • Compose services for CUDA 12.6 and 12.8, plus Lite images without ASR and UVR5 models
  • Reported RTF 0.028 on an RTX 4060 Ti for v2 ProPlus

Weaknesses

  • Pretrained weights are separate downloads from Hugging Face or ModelScope
  • Docker images lag the code; README says to pull latest source before using them
  • Training on Apple Silicon GPUs gives lower quality; macOS falls back to CPU
  • Five model generations (v1 to v5) with different tradeoffs to choose between
  • GPU optional
  • Docker + Compose
  • Needs ffmpeg
  • Models: GPT-SoVITS v1-v5 pretrained models, UVR5 vocal separation models, Faster Whisper large-v3 (ASR), FunASR Paraformer (Chinese ASR)
#758 out of 100

F5-TTS

Flow-matching TTS and voice cloning with Gradio and CLI

Live demo ↗ (opens in a new tab)15.4k stars, MIT, last commit Sep 2026

Synthesizes speech from a reference clip and its transcript using the F5-TTS diffusion transformer (plus an E2 TTS reproduction). Runs as a pip package with a Gradio web app on port 7860, a CLI and a Docker image; a Triton and TensorRT-LLM runtime reaches RTF 0.039 on an L20 GPU. Suited to researchers and builders who want a trainable open TTS model.

Strengths

  • pip install f5-tts; Gradio UI, CLI and a ghcr.io Docker image
  • Triton plus TensorRT-LLM runtime: 253 ms average latency at concurrency 2 on L20
  • Training and fine-tuning via Accelerate or a Gradio finetune app
  • PyTorch install documented for NVIDIA, AMD ROCm, Intel XPU and Apple Silicon

Weaknesses

  • Pretrained weights are CC-BY-NC (Emilia data); code is MIT, models are non-commercial
  • Reference audio needs a transcript, or an ASR model runs and uses more GPU memory
  • No compose file in the repo; the README's compose example assumes an NVIDIA GPU
  • Base checkpoints cover Chinese and English; other languages need community models
  • Docker
  • Needs ffmpeg
  • Models: F5-TTS v1 Base, E2 TTS, Vocos and BigVGAN vocoders
  • port 7860
#856 out of 100

OpenReader

Reads EPUB, PDF and DOCX aloud with synced word highlighting

540 stars, MIT, last commit Oct 2026

Next.js server that narrates EPUB, PDF, TXT, Markdown and DOCX files with synchronized read-along, generating audio ahead of playback through a self-hosted OpenAI-compatible TTS server (Kokoro-FastAPI, KittenTTS-FastAPI, Orpheus-FastAPI) or OpenAI, Replicate and DeepInfra. PDF layout is parsed with PP-DocLayoutV3 and words aligned with ONNX Whisper in a NATS JetStream worker. Exports M4B or MP3 audiobooks.

Strengths

  • Layout-aware PDF parsing and word-by-word highlighting
  • Audio cache reused across seeks, reloads and audiobook export
  • Storage on embedded SeaweedFS or S3; SQLite or Postgres; built-in auth
  • amd64 and arm64 Docker images with automatic startup migrations

Weaknesses

  • Needs a separate TTS server or cloud TTS API; nothing is bundled
  • Word alignment and DOCX conversion run in a NATS JetStream compute worker you deploy
  • Setup details (ports, env vars) are only in the external docs
  • no GPU
  • Docker
  • Needs OpenAI-compatible TTS server or cloud TTS API, NATS JetStream (compute worker), SQLite or PostgreSQL, SeaweedFS (embedded) or S3-compatible storage
  • Models: Kokoro-FastAPI, KittenTTS-FastAPI, Orpheus-FastAPI, OpenAI TTS, Replicate
#952 out of 100

Speakr

Transcribe, summarize and search recordings with pluggable ASR and LLMs

4.1k stars, AGPL-3.0, last commit Oct 2026

Web app that records or ingests audio, transcribes it through a connector (self-hosted WhisperX, OpenAI, Mistral Voxtral, AssemblyAI, OpenASR, FunASR), then writes summaries, action items and per-recording chat with an OpenAI-compatible LLM, OpenRouter or Ollama. Adds diarization, voice profiles, OIDC SSO, groups, a Swagger REST API and signed webhooks. Flask app on port 8899 with SQLite or PostgreSQL.

Strengths

  • Eight ASR connectors auto-detected from config; WhisperX enables voice profiles
  • Multi-user with OIDC SSO (Keycloak, Azure AD, Google, Auth0), groups and sharing
  • REST API v1 with Swagger UI, HMAC-signed webhooks, per-user token budgets
  • Lite image (about 725 MB) skips PyTorch; full image is about 4.4 GB

Weaknesses

  • Still alpha (v0.10.13-alpha) with frequent feature churn between releases
  • No bundled ASR; needs an API key or a separate GPU WhisperX container
  • Dual-licensed: AGPLv3, or a paid commercial license for proprietary use
  • Lite image downgrades Inquire semantic search to basic text search
  • no GPU
  • Docker
  • Needs ASR service or API (WhisperX, OpenAI, Mistral, AssemblyAI, OpenASR, FunASR), LLM API (OpenAI-compatible, OpenRouter or Ollama), SQLite or PostgreSQL
  • Models: WhisperX, OpenAI gpt-4o-transcribe-diarize, Mistral Voxtral, AssemblyAI, VibeVoice via vLLM
  • port 8899
#1049 out of 100

WhisperLive

Near-real-time Whisper transcription server over WebSocket

4.3k stars, MIT, last commit Oct 2026

Streams audio from a microphone, file, RTSP or HLS source to a server on port 9090 and returns partial and committed Whisper transcripts over WebSocket, with an optional OpenAI-compatible REST endpoint. Backends are faster-whisper (CPU, CUDA, ROCm), TensorRT-LLM and OpenVINO; extras include word timestamps, hotwords, pyannote diarization and translation. For teams embedding live captions or dictation.

Strengths

  • Three inference backends: faster-whisper, TensorRT-LLM, OpenVINO (Intel iGPU/dGPU)
  • Prebuilt GPU, CPU and OpenVINO Docker images; ROCm Dockerfile
  • Word-level timestamps, hotword boosting and batched multi-client inference
  • Chrome, Firefox and iOS clients; Python streaming client for raw PCM

Weaknesses

  • Defaults allow 4 clients and 600 s per connection; must be tuned for more
  • Without a fixed model, a new Whisper instance loads per client connection
  • TensorRT backend requires building engines and is recommended only via Docker
  • Diarization needs the optional pyannote.audio dependency
  • GPU optional
  • Docker
  • Needs PortAudio (client microphone input)
  • Models: Whisper via faster-whisper (CTranslate2), Whisper TensorRT-LLM engines, OpenVINO Whisper models
  • port 9090