Choose theme
AI Resources
KittenTTS
KittenTTS is a Python speech-generation library with KittenTTS 2 for reference voices and expressive narration, alongside small legacy ONNX models.
The current repository defaults to KittenTTS 2, a 1.7-billion-parameter speech model. The original 15M–80M ONNX models remain available for smaller CPU deployments; their features and requirements differ. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
Text-to-speech library
A Python application supplies text and a built-in voice or, with KittenTTS 2, a reference recording. It receives audio samples or writes a WAV file.
Why it stands out
Two speech-model families
KittenTTS 2 adds reference voice cloning and expression controls. Legacy ONNX models offer eight built-in voices and adjustable speed without those features.
Availability
Local library and browser try-out
GitHub provides installation and API documentation. KittenML's platform offers a browser route to try speech generation before integrating the library.
Why it matters
What makes it useful
A developer can turn a script into narration using a built-in voice, or supply a reference recording to KittenTTS 2 for voice cloning. That newer family also documents multilingual voices and markup for emotion, vocal events and emphasis. Expression support is early beta: it steers delivery rather than guaranteeing the requested performance.
What to know
Where it fits
KittenTTS is the speech-generation component inside a Python workflow. The application still handles text selection, playback and any larger narration or assistant interface. The legacy ONNX path keeps named voices and adjustable speech speed for applications that do not need cloning or expression control.
Notable points
What stands out
The legacy nano variants have the same 15 million parameter count but different listed disk sizes: 56 MB for fp32 and 25 MB for int8. A smaller download can come from quantization rather than fewer parameters; those sizes are not complete runtime memory estimates and do not describe KittenTTS 2.
Before using
What to review
KittenTTS 2's current quick start calls for Python 3.9+ and about 6 GB of CPU RAM or about 8 GB of free CUDA memory. Check the chosen weights and decoder; download size is not working memory, and CPU support does not promise real-time playback.
Installing kittenml normally brings in PyTorch as well as ONNX Runtime. The legacy documentation gives a separate dependency path when you only want the small ONNX models.
The current branch documents KittenTTS 2, while GitHub's latest numbered release is 0.8.1. Match the installed package and selected model to the relevant API documentation rather than assuming that older tag contains the newer family.
For legacy ONNX output, generate leaves text preprocessing off while generate_to_file enables it. Match that setting as well as the voice when comparing how prices, dates or abbreviations are spoken.
For KittenTTS 2's non-English voices, the documentation says to use normalize=False because its normalizer is English-tuned. Listen to the actual language and voice: the publisher notes that non-English output is less robust.
Code and model weights have separate terms. Read the original terms for the chosen model and consider what reference recordings you supply; cloning is a capability, not proof of permission or an exact voice match.
Reader fit
Who may find it relevant
Developers adding local narration to an application, with a choice between the newer voice controls and the smaller legacy models.
Someone who wants to hear it first can use the browser try-out; application integration still requires programming, adequate memory and listening checks.
Editorial note
Why LifeHubber lists it
For longer narration, KittenTTS 2 can return audio one sentence chunk at a time. An application can begin playback before the entire script has finished generating, rather than waiting for one complete WAV file. That makes it worth considering for narration players that need incremental audio. The first chunk still has a generation delay, and streamed joins can sound different from the completed file.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
What to explore next
Choose the speech component for your audio task.
KittenTTS supplies speech from text. If your project also needs transcription, voice direction or live conversation, continue with the different audio jobs before choosing the rest of the stack.
More in Speech Models
Keep browsing this category
Explore more speech model resources.
Fish Audio S2 Pro
fishaudio/s2-pro
A text-to-speech model with in-script control over pauses, emphasis and mood.
AuK
Tencent-Hunyuan/AuK
A 1.5B speech model for instruction-guided text-to-speech, content and acoustic editing, paralinguistic changes, speech enhancement, and source separation, with public code, weights, demos, ComfyUI nodes, and fine-tuning materials.
TADA
HumeAI/tada
A speech-language model that aligns speech and text into a single synchronized stream.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.