#55th of 11 in Voice
IndexTTS
Zero-shot TTS with emotion, speed and pronunciation control
- Stars
- 24.4k
- License
- custom license
- Last commit
- Sep 2026
- Last release
- Aug 2026
Overview
Clones a voice from one reference clip and synthesizes speech in Chinese, English, Japanese, Spanish and Arabic (IndexTTS-2.5). Emotion comes from a second reference clip, an 8-value vector or the text itself; speed is set by duration_factor (0.5x to 2.0x) and pronunciation by inline Pinyin, CMU phonemes or Kana. Ships a Gradio web UI on port 7860 and a Python API; a vLLM recipe covers production serving.
Who it is for: Developers needing controllable multilingual voice cloning
Strengths
- Emotion control via reference audio, an 8-value vector or a text description
- Inline pronunciation overrides: Pinyin, CMU phonemes and Japanese Kana
- BF16 inference with optional DeepSpeed and compiled CUDA kernels
- Published vLLM recipe for production deployment
Weaknesses
- No Dockerfile or compose file; install is uv plus CUDA Toolkit 12.8 or newer
- Model weights (IndexTTS-2.5, IndexTTS-2) are separate multi-GB downloads
- Five languages only; no streaming API is documented in the README
- License is non-standard (GitHub reports NOASSERTION); check terms before commercial use
What it needs
- Needs uv
- Models: IndexTTS-2.5, IndexTTS-2, IndexTTS-1.5 (legacy)
- port 7860
Also in Voice
See all 11| Rank | Project | Score |
|---|---|---|
| 1 | VoiceboxLocal voice studio for cloning, TTS, dictation and agent speech | 72 out of 100 |
| 2 | Speech-to-SpeechModular voice-agent pipeline exposed through the OpenAI Realtime API | 67 out of 100 |
| 3 | Pocket TTS100M-parameter CPU text-to-speech with streaming and voice cloning | 64 out of 100 |
| #4 | Kokoro-FastAPIOpenAI-compatible Kokoro-82M speech API in CPU and GPU images | 64 out of 100 |
| #6 | GPT-SoVITSFew-shot voice cloning and TTS with a training web UI | 59 out of 100 |