#66th of 11 in Voice
GPT-SoVITS
Few-shot voice cloning and TTS with a training web UI
Live demo ↗
(opens in a new tab)Documentation ↗
(opens in a new tab)Repository on GitHub ↗
(opens in a new tab)
- Stars
- 62.6k
- License
- MIT
- Last commit
- Oct 2026
- Last release
- Jun 2025
Overview
Clones a voice from a 5-second sample (zero-shot) or fine-tunes GPT and SoVITS models on about one minute of audio, then synthesizes speech in Chinese, English, Japanese, Korean and Cantonese. The Gradio web UI bundles dataset tools: UVR5 vocal separation, slicing, ASR and label proofreading. Aimed at hobbyists and studios building custom voices locally.
Who it is for: Hobbyists and studios training custom voices locally
Strengths
- Zero-shot cloning from 5 s of audio; few-shot fine-tune from about 1 minute
- Cross-lingual synthesis across zh, en, ja, ko and yue
- Compose services for CUDA 12.6 and 12.8, plus Lite images without ASR and UVR5 models
- Reported RTF 0.028 on an RTX 4060 Ti for v2 ProPlus
Weaknesses
- Pretrained weights are separate downloads from Hugging Face or ModelScope
- Docker images lag the code; README says to pull latest source before using them
- Training on Apple Silicon GPUs gives lower quality; macOS falls back to CPU
- Five model generations (v1 to v5) with different tradeoffs to choose between
What it needs
- GPU optional
- Docker + Compose
- Needs ffmpeg
- Models: GPT-SoVITS v1-v5 pretrained models, UVR5 vocal separation models, Faster Whisper large-v3 (ASR), FunASR Paraformer (Chinese ASR)
Also in Voice
See all 11| Rank | Project | Score |
|---|---|---|
| 1 | VoiceboxLocal voice studio for cloning, TTS, dictation and agent speech | 72 out of 100 |
| 2 | Speech-to-SpeechModular voice-agent pipeline exposed through the OpenAI Realtime API | 67 out of 100 |
| 3 | Pocket TTS100M-parameter CPU text-to-speech with streaming and voice cloning | 64 out of 100 |
| #4 | Kokoro-FastAPIOpenAI-compatible Kokoro-82M speech API in CPU and GPU images | 64 out of 100 |
| #5 | IndexTTSZero-shot TTS with emotion, speed and pronunciation control | 61 out of 100 |