#99th of 11 in Voice
Speakr
Transcribe, summarize and search recordings with pluggable ASR and LLMs
- Stars
- 4.1k
- License
- AGPL-3.0
- Last commit
- Oct 2026
- Last release
- Oct 2026
Overview
Web app that records or ingests audio, transcribes it through a connector (self-hosted WhisperX, OpenAI, Mistral Voxtral, AssemblyAI, OpenASR, FunASR), then writes summaries, action items and per-recording chat with an OpenAI-compatible LLM, OpenRouter or Ollama. Adds diarization, voice profiles, OIDC SSO, groups, a Swagger REST API and signed webhooks. Flask app on port 8899 with SQLite or PostgreSQL.
Who it is for: Privacy-focused teams and individuals archiving meetings
Strengths
- Eight ASR connectors auto-detected from config; WhisperX enables voice profiles
- Multi-user with OIDC SSO (Keycloak, Azure AD, Google, Auth0), groups and sharing
- REST API v1 with Swagger UI, HMAC-signed webhooks, per-user token budgets
- Lite image (about 725 MB) skips PyTorch; full image is about 4.4 GB
Weaknesses
- Still alpha (v0.10.13-alpha) with frequent feature churn between releases
- No bundled ASR; needs an API key or a separate GPU WhisperX container
- Dual-licensed: AGPLv3, or a paid commercial license for proprietary use
- Lite image downgrades Inquire semantic search to basic text search
What it needs
- no GPU
- Docker
- Needs ASR service or API (WhisperX, OpenAI, Mistral, AssemblyAI, OpenASR, FunASR), LLM API (OpenAI-compatible, OpenRouter or Ollama), SQLite or PostgreSQL
- Models: WhisperX, OpenAI gpt-4o-transcribe-diarize, Mistral Voxtral, AssemblyAI, VibeVoice via vLLM
- port 8899
Also in Voice
See all 11| Rank | Project | Score |
|---|---|---|
| 1 | VoiceboxLocal voice studio for cloning, TTS, dictation and agent speech | 72 out of 100 |
| 2 | Speech-to-SpeechModular voice-agent pipeline exposed through the OpenAI Realtime API | 67 out of 100 |
| 3 | Pocket TTS100M-parameter CPU text-to-speech with streaming and voice cloning | 64 out of 100 |
| #4 | Kokoro-FastAPIOpenAI-compatible Kokoro-82M speech API in CPU and GPU images | 64 out of 100 |
| #5 | IndexTTSZero-shot TTS with emotion, speed and pronunciation control | 61 out of 100 |