Kokoro TTS

kokorottsai.com
On the map Visit site

Turn your text into natural, high-quality voices across many languages.

Description

Kokoro TTS · Kokoro TTS - Advanced AI text-to-speech model with only 82M parameters, delivers HQ and efficient speech synthesis.Turn text into natural, lifelike voices. · Listen to any webpage with VoiceRead · Top Benefits of Kokoro TTS · Try Kokoro TTS Online · Why Use Kokoro TTS?

"As a digital publisher, l always wanted toturn our e-book library into audiobooksespecially for niche genres. Kokoro TTS hasbeen a game-changer! The natural-sounding voices and fast conversion makeit so easy to offer audiobooks to ourreaders."

Features

Kokoro TTS is an efficient text-to-speech tool with multilingual and customizable voice support. Its 182M parameter architecture delivers high-quality audio, supporting languages like American English, British English, French, Korean, Japanese, and Mandarin. It features lifelike voice options, autom

FAQ

Kokoro TTS is an advanced, lightweight text-to-speech AI model that uses only 82 million parameters to generate high-quality, natural-sounding speech. It delivers performance comparable to much larger models while remaining highly efficient and resource-friendly.

Kokoro TTS is unique due to its small model size with high-quality output, its open-source and commercially friendly Apache 2.0 license, its efficient inference suitable for CPU and GPU deployment, its support for real-time audio generation with very low latency (40-70ms on GPU), and its flexible integration including ONNX compatibility and OpenAI API.

Kokoro TTS supports multiple languages including American English, British English, French, Korean, Japanese, and Mandarin. Various voice packs allow customization for different applications.

It can handle inputs up to 510 tokens per pass or about 500 characters per generation (5000 characters in streaming mode), allowing efficient synthesis of longer texts like articles or audiobooks.

Kokoro TTS runs efficiently on CPUs and GPUs, supports Docker, ONNX deployment, and can also be used as a local server for AI assistants. It aims for broad device compatibility, including setups without dedicated GPUs.

The model was trained on a carefully curated dataset of high-quality, permissively licensed audio, ensuring accurate and natural voice synthesis.

Key features include high quality, natural-sounding speech with minimal artifacts, fast processing suitable for real-time applications, multiple customizable voices and adjustable settings such as speed and silence trimming, automatic content segmentation for audiobook or chapter handling, and open-source code for unrestricted use and modification.

You can input text on the online demo or install the open-source model. After selecting your voice and adjusting parameters, you generate speech with instant playback and saving options. For developers, integration is possible via API or local server setup.

Yes, it is freely available under the Apache 2.0 license and also offers commercial use permissions. Community support is accessible via Discord or email.

Specs

Type Agent
SectionInfrastructure & MLOps
Pricing free (от $0/mo)
Platform Browser extension
Systems browser, api
Hostingself-hosted
Who forIndividual
Complexitydeveloper
Site languageen
Rating5.00 (1 reviews)
Views1 572
Launched2025-02-13

Can do

Platforms

Social

Found in sources

Similar in «Infrastructure & MLOps»

Submit a site to the catalog

Just send the link — we will work out the rest.

We will review what you send and add it to the catalog if it fits.