LIFEHUBBER
Choose theme

AI Resources

KittenTTS

GitHub stars: 15.6K GitHub forks: 897 Declared license: Apache-2.0: Apache-2.0 Last pushed October 6, 2026: Pushed 2d ago
Stats from GitHub

KittenTTS is a Python speech-generation library with KittenTTS 2 for reference voices and expressive narration, alongside small legacy ONNX models.

The current repository defaults to KittenTTS 2, a 1.7-billion-parameter speech model. The original 15M–80M ONNX models remain available for smaller CPU deployments; their features and requirements differ. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.

What it is

Text-to-speech library

A Python application supplies text and a built-in voice or, with KittenTTS 2, a reference recording. It receives audio samples or writes a WAV file.

Why it stands out

Two speech-model families

KittenTTS 2 adds reference voice cloning and expression controls. Legacy ONNX models offer eight built-in voices and adjustable speed without those features.

Availability

Local library and browser try-out

GitHub provides installation and API documentation. KittenML's platform offers a browser route to try speech generation before integrating the library.

Why it matters

What makes it useful

A developer can turn a script into narration using a built-in voice, or supply a reference recording to KittenTTS 2 for voice cloning. That newer family also documents multilingual voices and markup for emotion, vocal events and emphasis. Expression support is early beta: it steers delivery rather than guaranteeing the requested performance.

Notable points

What stands out

The legacy nano variants have the same 15 million parameter count but different listed disk sizes: 56 MB for fp32 and 25 MB for int8. A smaller download can come from quantization rather than fewer parameters; those sizes are not complete runtime memory estimates and do not describe KittenTTS 2.

Before using

What to review

KittenTTS 2's current quick start calls for Python 3.9+ and about 6 GB of CPU RAM or about 8 GB of free CUDA memory. Check the chosen weights and decoder; download size is not working memory, and CPU support does not promise real-time playback.

Installing kittenml normally brings in PyTorch as well as ONNX Runtime. The legacy documentation gives a separate dependency path when you only want the small ONNX models.

The current branch documents KittenTTS 2, while GitHub's latest numbered release is 0.8.1. Match the installed package and selected model to the relevant API documentation rather than assuming that older tag contains the newer family.

For legacy ONNX output, generate leaves text preprocessing off while generate_to_file enables it. Match that setting as well as the voice when comparing how prices, dates or abbreviations are spoken.

For KittenTTS 2's non-English voices, the documentation says to use normalize=False because its normalizer is English-tuned. Listen to the actual language and voice: the publisher notes that non-English output is less robust.

Code and model weights have separate terms. Read the original terms for the chosen model and consider what reference recordings you supply; cloning is a capability, not proof of permission or an exact voice match.

Reader fit

Who may find it relevant

Developers adding local narration to an application, with a choice between the newer voice controls and the smaller legacy models.

Someone who wants to hear it first can use the browser try-out; application integration still requires programming, adequate memory and listening checks.

Editorial note

Why LifeHubber lists it

For longer narration, KittenTTS 2 can return audio one sentence chunk at a time. An application can begin playback before the entire script has finished generating, rather than waiting for one complete WAV file. That makes it worth considering for narration players that need incremental audio. The first chunk still has a generation delay, and streamed joins can sound different from the completed file.

Source links

Source materials

Reader note

Before relying on this entry

LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.

What to explore next

Choose the speech component for your audio task.

KittenTTS supplies speech from text. If your project also needs transcription, voice direction or live conversation, continue with the different audio jobs before choosing the rest of the stack.

Advertisements

Advertisements

For project maintainers

Listed here? You can use the badge.

If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.

See what’s moving