Kokoro TTS: How to Use It Online, Locally and in the Browser
Updated 2026-10-08
Kokoro TTS is a free, open-weight text-to-speech model with about 82 million parameters, published by the developer hexgrad on Hugging Face under the Apache 2.0 license. It is popular because it sounds natural while staying small enough to run on an ordinary laptop – even in a web browser – without a paid API. You can try it online in a hosted demo, install it with Python on your own computer, or run it in the browser through a JavaScript port.
This guide covers what Kokoro is, the three ways to use it, the voices, and what to watch out for.
What is Kokoro TTS?
Kokoro (model name Kokoro-82M) is a neural text-to-speech model: you give it text, it returns spoken audio. Key facts:
- Size: about 82 million parameters – tiny compared with many modern TTS models, so it is fast and cheap to run.
- License: Apache 2.0, which allows personal and commercial use of the model weights.
- Voices: dozens of preset voices with IDs such as
af_heart,af_bella(American English, female) andam_adam,am_michael(American English, male). The prefix shows language/accent and gender. - Languages: American and British English are the strongest; the model card also lists voices for several other languages, with varying quality.
- Where it lives: the official weights and model card are on Hugging Face (
hexgrad/Kokoro-82M), with the inference library on GitHub.
Because the weights are open, many third-party websites wrap Kokoro in a simple web form. Those sites are independent of the model's author.
Option 1: Try Kokoro TTS online
The quickest way is a hosted demo:
- Open the Kokoro demo on Hugging Face Spaces (search "Kokoro TTS" on huggingface.co).
- Paste your text.
- Pick a voice and speed, then generate and download the audio.
Shared demos can queue at busy times and usually limit text length. For long scripts, running it yourself is more reliable.
Option 2: Run Kokoro TTS locally with Python
Kokoro runs on CPU; a GPU makes it faster but isn't required.
- Install the speech front-end. Kokoro uses espeak-ng for pronunciation of some words – install it from your system's package manager (on Windows, the espeak-ng installer).
- Install the library:
pip install kokoro soundfile - Generate audio with a short script: create a pipeline for American English, pass your text and a voice ID such as
af_heart, and write the returned audio to a WAV file (Kokoro outputs 24 kHz audio).
The project's README on GitHub has the exact current code sample – check it, as the API can change between versions.
Option 3: Run Kokoro in the browser
A JavaScript port (kokoro-js, built on Transformers.js) runs the model directly in the browser with WebGPU or WebAssembly. Nothing is sent to a server: the model downloads once (tens to a few hundred MB depending on precision), then speech is generated on your own device. This is how most "Kokoro TTS online free, no sign-up" sites work. The first load is slow; after that, generation is quick on a modern computer.
Kokoro voices: which one to pick
| Voice ID | Accent | Good for |
|---|---|---|
| af_heart | American, female | Default; warm narration |
| af_bella | American, female | Bright, expressive reads |
| af_nicole | American, female | Calm, close-mic style |
| am_adam | American, male | Neutral narration |
| am_michael | American, male | Deeper voiceover |
| bf_emma / bm_george | British | British English content |
Voice quality differs – the model card grades the voices, and the top-graded English voices are the safest choice for long content.
Tips for better Kokoro output
- Write for the ear. Spell out unusual abbreviations and numbers ("2026" reads fine, but "approx." may not).
- Split long text. Generate by paragraph and join the files; it is faster and easier to fix one bad sentence.
- Use punctuation for pacing. Commas and full stops control pauses better than line breaks.
- Adjust speed slightly (0.9–1.1) rather than heavily; extreme speeds sound less natural.
- Check pronunciation of names and replace them with phonetic spellings if needed.
Kokoro TTS vs paid text-to-speech
Kokoro's strengths are cost (free to run), privacy (it can run fully offline) and speed. Paid cloud voices tend to offer more languages, voice cloning, emotion controls and a polished studio interface. For narration, YouTube voiceovers, accessibility and prototypes in English, Kokoro is often good enough; for voice cloning or many languages, a commercial service may fit better.
FAQ
Is Kokoro TTS free?
Yes. The model weights are released under the Apache 2.0 license and can be used without paying. Hosted websites that run it may add their own limits.
Can I use Kokoro TTS commercially?
The Apache 2.0 license permits commercial use of the model. Check the terms of any website or service you use to run it.
Does Kokoro TTS need a GPU?
No. It runs on a normal CPU; a GPU or WebGPU in the browser only makes it faster.
Does Kokoro support voice cloning?
No. Kokoro uses preset voices; it does not clone a voice from a sample.
What languages does Kokoro TTS support?
English (American and British) is the strongest. The model card lists voices for several more languages with varying quality.