Pocket TTS vs Kokoro-FastAPI

Two of the top voice, side by side: score, setup, license, activity and what each review found.

33rd of 11 in Voice

Pocket TTS

100M-parameter CPU text-to-speech with streaming and voice cloning

#44th of 11 in Voice

Kokoro-FastAPI

OpenAI-compatible Kokoro-82M speech API in CPU and GPU images

Pocket TTS vs Kokoro-FastAPI: score parts and facts
What we comparePocket TTSKokoro-FastAPI
Score parts, out of 100
Adoption43, known31, known
Freshness100, active100, active
Maintenance87, healthy98, healthy
Easy to run50, easy50, easy
Agent-ready30, minimal45, minimal
Facts from GitHub and the README
Stars9.9k5.5k
LicenseMIT (permissive)Apache-2.0 (permissive)
Last commitOct 2026Oct 2026
Last releaseSep 2026Sep 2026
LanguageNot statedNot stated
DockerYesYes
GPUNot neededOptional
arm64 or Apple SiliconMentionedMentioned

Pocket TTS

Generates speech on CPU with a 100M-parameter model: about 200 ms to the first audio chunk and roughly 6x real time on an M4 MacBook Air using two cores. Covers English, French, German, Portuguese, Italian, Spanish and Dutch, clones a voice from a WAV file, and runs as a CLI, a Python library or an HTTP server with a web UI on port 8000. For developers who want TTS without a GPU.

Who it is for: Developers adding TTS to apps without a GPU

Strengths

  • Runs on 2 CPU cores; no CUDA build of PyTorch needed
  • Streaming output with about 200 ms first-chunk latency
  • Voice cloning from any WAV; export voices to safetensors for fast loading
  • Training code released; community models load via --config

Weaknesses

  • Seven European languages; others depend on community-trained models
  • No pause or silence markup in text input
  • serve command and Docker image are CPU-only; GPU use is unsupported and manual
  • Linux pip pulls CUDA PyTorch (about 3 GB) unless the CPU index is set
  • no GPU
  • Docker + Compose
  • Models: Pocket TTS 100M, 24-layer language variants, community checkpoints via --config
  • port 8000

Kokoro-FastAPI

Serves the Kokoro-82M model behind an OpenAI-compatible /v1/audio/speech endpoint on port 8880, streaming mp3, wav, opus, flac, aac or pcm. Covers English (US/GB), Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese and Mandarin, with weighted voice mixing, inline [voice:] and [pause:] tags, word timestamps and phoneme endpoints. Prebuilt images exist for CPU, CUDA (amd64 and arm64) and experimental ROCm.

Who it is for: Self-hosters wanting an OpenAI-style TTS endpoint

Strengths

  • Drop-in for the OpenAI Python client; models baked into the images
  • Weighted voice mixing and inline speaker, pause, rate and IPA tags
  • Per-word timestamp captions and phoneme in/out endpoints
  • First-token latency about 300 ms on GPU

Weaknesses

  • CPU first-token latency: 3.5 s on an older i7, under 1 s on M3 Pro
  • No true voice cloning; /dev/tune only nudges toward a reference clip
  • ROCm image is experimental and amd64 only
  • Apple Silicon GPU (MPS) only when run natively via uv, not in Docker
  • GPU optional
  • Docker + Compose
  • Needs espeak-ng (optional fallback)
  • Models: Kokoro-82M v1.0
  • port 8880

More in Voice