Kokoro: Local, CPU-Friendly TTS That Actually Sounds Good

TL;DR
An 82M parameter text-to-speech model that runs on CPU and produces high-quality speech across multiple languages - no cloud APIs or GPU required.
Local AI inference keeps getting more practical. Kokoro is an 82 million parameter text-to-speech model that runs entirely on CPU while producing surprisingly natural-sounding speech. It supports English, Mandarin, Hindi, and other languages with approximately 50 voice options.
What Makes Kokoro Interesting#
The key numbers:
- 82M parameters - small enough for CPU inference
- ~50 voices - predominantly English speakers
- 5GB Docker image - includes pre-downloaded voice models
- OpenAI-compatible API - drop-in replacement for existing integrations
Performance varies by hardware but stays practical:
- Intel Core i7-4770K: 4.7 seconds for a short paragraph
- Apple M2 Pro: 4.5 seconds
- AMD Ryzen 7 8745HS: 1.5 seconds
These benchmarks are for CPU-only inference. If you have an integrated GPU, you can go faster - there's a start-gpu_mac.sh script for Apple Silicon.
Quick Setup with Kokoro-FastAPI#
The easiest path is the containerized Kokoro-FastAPI wrapper:
podman run -p 8880:8880 ghcr.io/remsky/kokoro-fastapi-cpu
Once running, you get a web interface at localhost:8880/web for testing, plus an OpenAI-compatible speech API. Applications built for OpenAI's TTS can point at your local endpoint instead.
What HN Is Saying#
The Hacker News discussion (80+ comments) is largely positive, with developers sharing their real-world use cases.
Multiple commenters are using Kokoro for home automation and voice assistants:
"I use kokoro with home assistant and its great. I find its the most natural sounding and small too. I speak over sonos speakers when certain events happen."
Several developers have built article-to-podcast pipelines:
"About a month ago I setup Kokoro on my GTX1650 to do TTS for an article reader. A simple WebUI lets me paste a URL or a chunk of copy pasted text... Then for my morning drive I'll catch up on articles or blog posts I've gathered."
One developer ported it to iPhone's ANE (Apple Neural Engine) for mobile TTS with better battery life:
"Cool I actually got it ported to iPhone's ANE finally yesterday! So we can get both rt natural local TTS and 4x less battery drainage and thermals"
The comments also surface some practical limitations. Single words and short phrases can sound off:
"Try having it say simply 'six' and it almost always says something like 'ah-six-ah'. I found a way around that though. If you give it a longer sentence to say (eg 'The word is: six') it will say it fine."
The commenter notes you can crop out just the word you need using the timestamp data Kokoro returns with each generation.
Alternative models came up frequently. Pocket TTS from Kyutai Labs got several mentions for voice cloning. Supertonic 3 was praised for handling mixed-language text well:
"Supertonic 3 is the only one that can autodetect language and make a mix of different languages sound good."
Read the full thread at https://news.ycombinator.com/item?id=48821576.
Practical Applications#
The HN thread includes several specific use cases worth noting:
Accessibility tools: One developer uses Kokoro extensively for an accessibility product, appreciating the IPA pronunciation guides for handling homographs correctly.
Browser extensions: Someone built a Chrome extension that runs Kokoro on any webpage with sentence highlighting: Local Reader
Ebook audiobooks: Multiple commenters use Kokoro to generate audiobooks from EPUBs when no official audiobook exists.
Japanese language learning: Combined with an LLM, one developer built a local Japanese tutor with native-sounding speech.
Alternatives and Comparisons#
The discussion surfaced several other local TTS options:
| Model | Parameters | Voice Cloning | Notes |
|---|---|---|---|
| Kokoro | 82M | No | Best CPU efficiency |
| Pocket TTS | ~100M | Yes | Easy voice cloning |
| Chatterbox Turbo | Larger | Yes | Emotional control |
| Fish Audio S2 | Larger | Yes | Fine-grained tone control |
| Piper | Various | No | Lightweight, fast |
For pure CPU inference without voice cloning, Kokoro remains the standout choice. If you need voice cloning, Pocket TTS is the comparable-size option.
The Bigger Picture#
Local TTS has reached a practical inflection point. A 5GB download gets you production-quality speech synthesis that runs on consumer hardware. Combined with local STT (Parakeet, whisper.cpp) and local LLMs, you can build voice interfaces that never touch the cloud.
The quality is not quite ElevenLabs or Azure's DragonHD voices at peak performance. But it is good enough for most applications, and "good enough + completely private + zero marginal cost" is a compelling combination.
As one commenter put it:
"Both Text-to-Speech and Speech-to-Text now have local models that are good enough to get the job done. Kokoro for TTS, Parakeet for STT and Fluid-1 for text formatting. I hope this is a trend that continues for other applications."
Getting Started#
The fastest path to try Kokoro:
- Run the container:
podman run -p 8880:8880 ghcr.io/remsky/kokoro-fastapi-cpu - Open
localhost:8880/web - Paste some text and generate
For more control, check the Kokoro-82M Hugging Face page or the ONNX version at NeuML/kokoro-base-onnx for custom pipelines.
Continue Reading#
- DeepSeek R1 and V3: The Developer''s Guide to Open-Source AI
- The $44 Compiler: Persistent Projects Beat Persistent Agents
- Forge Shows the Local Agent Reliability Gap Is a Harness Problem
Sources#
- Local, CPU-Friendly, High-Quality TTS with Kokoro - Original article
- Hacker News Discussion
- Kokoro-82M on Hugging Face
- Kokoro-FastAPI
- Pocket TTS - Alternative with voice cloning
Get the next deep dive like this in your inbox
One email a week on News and the rest of the AI dev stack. Free.
Read next on local and open-weight models
Ornith-1.0: What an Open Source Self-Improving Coding Model Actually Means
DeepReinforce AI released Ornith-1.0, a family of open-source coding models claiming self-improvement. The HN thread reveals a mix of skepticism and genuine interest - here is what the model actually does and whether the hype holds up.
7 min readMesh LLM: Run 235B Models Across Your Home Lab with iroh
A new distributed inference system pools GPU resources across multiple machines and exposes them through a single OpenAI-compatible API. No RDMA, no NVLink - just QUIC and your existing hardware.
6 min readColibri: Run GLM 5.2 on a 32GB Laptop With Disk Streaming
A solo developer built a 1,300-line C inference engine that runs the 744B GLM 5.2 model on consumer hardware by streaming routed experts from disk. Here's how it works.
6 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.








