#1010th of 11 in Voice
WhisperLive
Near-real-time Whisper transcription server over WebSocket
- Stars
- 4.3k
- License
- MIT
- Last commit
- Oct 2026
- Last release
- Sep 2026
Overview
Streams audio from a microphone, file, RTSP or HLS source to a server on port 9090 and returns partial and committed Whisper transcripts over WebSocket, with an optional OpenAI-compatible REST endpoint. Backends are faster-whisper (CPU, CUDA, ROCm), TensorRT-LLM and OpenVINO; extras include word timestamps, hotwords, pyannote diarization and translation. For teams embedding live captions or dictation.
Who it is for: Teams adding live transcription to apps or devices
Strengths
- Three inference backends: faster-whisper, TensorRT-LLM, OpenVINO (Intel iGPU/dGPU)
- Prebuilt GPU, CPU and OpenVINO Docker images; ROCm Dockerfile
- Word-level timestamps, hotword boosting and batched multi-client inference
- Chrome, Firefox and iOS clients; Python streaming client for raw PCM
Weaknesses
- Defaults allow 4 clients and 600 s per connection; must be tuned for more
- Without a fixed model, a new Whisper instance loads per client connection
- TensorRT backend requires building engines and is recommended only via Docker
- Diarization needs the optional pyannote.audio dependency
What it needs
- GPU optional
- Docker
- Needs PortAudio (client microphone input)
- Models: Whisper via faster-whisper (CTranslate2), Whisper TensorRT-LLM engines, OpenVINO Whisper models
- port 9090
Also in Voice
See all 11| Rank | Project | Score |
|---|---|---|
| 1 | VoiceboxLocal voice studio for cloning, TTS, dictation and agent speech | 72 out of 100 |
| 2 | Speech-to-SpeechModular voice-agent pipeline exposed through the OpenAI Realtime API | 67 out of 100 |
| 3 | Pocket TTS100M-parameter CPU text-to-speech with streaming and voice cloning | 64 out of 100 |
| #4 | Kokoro-FastAPIOpenAI-compatible Kokoro-82M speech API in CPU and GPU images | 64 out of 100 |
| #5 | IndexTTSZero-shot TTS with emotion, speed and pronunciation control | 61 out of 100 |