#1010th of 11 in Voice

WhisperLive

Near-real-time Whisper transcription server over WebSocket

Stars
4.3k
License
MIT
Last commit
Oct 2026
Last release
Sep 2026

Overview

Streams audio from a microphone, file, RTSP or HLS source to a server on port 9090 and returns partial and committed Whisper transcripts over WebSocket, with an optional OpenAI-compatible REST endpoint. Backends are faster-whisper (CPU, CUDA, ROCm), TensorRT-LLM and OpenVINO; extras include word timestamps, hotwords, pyannote diarization and translation. For teams embedding live captions or dictation.

Who it is for: Teams adding live transcription to apps or devices

Strengths

  • Three inference backends: faster-whisper, TensorRT-LLM, OpenVINO (Intel iGPU/dGPU)
  • Prebuilt GPU, CPU and OpenVINO Docker images; ROCm Dockerfile
  • Word-level timestamps, hotword boosting and batched multi-client inference
  • Chrome, Firefox and iOS clients; Python streaming client for raw PCM

Weaknesses

  • Defaults allow 4 clients and 600 s per connection; must be tuned for more
  • Without a fixed model, a new Whisper instance loads per client connection
  • TensorRT backend requires building engines and is recommended only via Docker
  • Diarization needs the optional pyannote.audio dependency

What it needs

  • GPU optional
  • Docker
  • Needs PortAudio (client microphone input)
  • Models: Whisper via faster-whisper (CTranslate2), Whisper TensorRT-LLM engines, OpenVINO Whisper models
  • port 9090

Also in Voice

See all 11
Also in Voice
RankProjectScore
1VoiceboxLocal voice studio for cloning, TTS, dictation and agent speech56.8k stars, MIT72 out of 100
2Speech-to-SpeechModular voice-agent pipeline exposed through the OpenAI Realtime API13.4k stars, Apache-2.067 out of 100
3Pocket TTS100M-parameter CPU text-to-speech with streaming and voice cloning Live demo ↗ (opens in a new tab)9.9k stars, MIT64 out of 100
#4Kokoro-FastAPIOpenAI-compatible Kokoro-82M speech API in CPU and GPU images Live demo ↗ (opens in a new tab)5.5k stars, Apache-2.064 out of 100
#5IndexTTSZero-shot TTS with emotion, speed and pronunciation control Live demo ↗ (opens in a new tab)24.4k stars, custom license61 out of 100