#77th of 11 in Voice
F5-TTS
Flow-matching TTS and voice cloning with Gradio and CLI
- Stars
- 15.4k
- License
- MIT
- Last commit
- Sep 2026
- Last release
- Jul 2026
Overview
Synthesizes speech from a reference clip and its transcript using the F5-TTS diffusion transformer (plus an E2 TTS reproduction). Runs as a pip package with a Gradio web app on port 7860, a CLI and a Docker image; a Triton and TensorRT-LLM runtime reaches RTF 0.039 on an L20 GPU. Suited to researchers and builders who want a trainable open TTS model.
Who it is for: Researchers and builders training or serving open TTS models
Strengths
- pip install f5-tts; Gradio UI, CLI and a ghcr.io Docker image
- Triton plus TensorRT-LLM runtime: 253 ms average latency at concurrency 2 on L20
- Training and fine-tuning via Accelerate or a Gradio finetune app
- PyTorch install documented for NVIDIA, AMD ROCm, Intel XPU and Apple Silicon
Weaknesses
- Pretrained weights are CC-BY-NC (Emilia data); code is MIT, models are non-commercial
- Reference audio needs a transcript, or an ASR model runs and uses more GPU memory
- No compose file in the repo; the README's compose example assumes an NVIDIA GPU
- Base checkpoints cover Chinese and English; other languages need community models
What it needs
- Docker
- Needs ffmpeg
- Models: F5-TTS v1 Base, E2 TTS, Vocos and BigVGAN vocoders
- port 7860
Also in Voice
See all 11| Rank | Project | Score |
|---|---|---|
| 1 | VoiceboxLocal voice studio for cloning, TTS, dictation and agent speech | 72 out of 100 |
| 2 | Speech-to-SpeechModular voice-agent pipeline exposed through the OpenAI Realtime API | 67 out of 100 |
| 3 | Pocket TTS100M-parameter CPU text-to-speech with streaming and voice cloning | 64 out of 100 |
| #4 | Kokoro-FastAPIOpenAI-compatible Kokoro-82M speech API in CPU and GPU images | 64 out of 100 |
| #5 | IndexTTSZero-shot TTS with emotion, speed and pronunciation control | 61 out of 100 |