Choose theme
AI Resources
Audio8 TTS 0.1B ONNX INT8
Audio8 TTS 0.1B ONNX INT8 packages a compact multilingual text-to-speech model for CPU inference through ONNX Runtime, including zero-shot voice cloning.
The downloadable package combines INT8 speech-generation models with an FP16 audio codec. Audio8's separate runtime adds local voice registration, command-line use, streaming, and HTTP and OpenAI-compatible speech endpoints. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
Compact multilingual speech generation
The package turns text into 44.1 kHz mono speech and supports a packaged default voice or a registered reference voice across 11 recommended languages.
Why it stands out
A small CPU service, not only model files
The matching runtime exposes the model through a local web interface, HTTP API, or OpenAI-compatible speech endpoint. It selects ONNX Runtime's CPUExecutionProvider; GPU execution is outside this export's scope.
Availability
Model files and runtime are separate
Hugging Face hosts the ONNX model package. Its matching code is in the onnx_runtime_0_1b_int8 directory. The older onnx_runtime directory targets incompatible 0.6B INT4 graphs.
Why it matters
What makes it useful
For an app that already sends text to a speech endpoint, the local service offers an OpenAI-compatible /v1/audio/speech route. A builder can request speech from the CPU model through that familiar interface, while the model package supplies generation rather than transcription or a complete voice assistant.
What to know
Where it fits
Download the model files and use the matching runtime to prepare a voice, then send text from your app. Its local browser interface lets you select a voice, play or download WAV output, and inspect memory. The surrounding app still owns text handling, user controls, storage, and disclosure.
Notable points
What stands out
Audio8 reports about 0.6 GB after loading the sessions used for normal speech generation in its test setup. The complete model repository is about 1 GB because it also includes the encoder used to register reference voices. Actual memory use varies by platform and runtime.
Before using
What to review
The documented ONNX path requires Python 3.10 or newer. The model card names macOS arm64 and Linux x86_64 as tested platforms; the runtime also documents Windows PowerShell wrappers, which does not establish the same performance on Windows.
Audio8 recommends a 0.5 to 30 second reference recording with an exact transcript for voice registration. Noisy, long, or incorrectly transcribed references can reduce stability and speaker similarity.
INT8 quantization can change generated token sequences. Test the exact language, voice, text length, and deployment settings instead of assuming the same output as the source checkpoint.
The linked ONNX package and base checkpoint do not currently present one consistent license label. Review the current terms at the source and ask Audio8 if the intended use depends on them.
Obtain consent before cloning another person's voice, and clearly disclose synthetic speech where the context calls for it.
Reader fit
Who may find it relevant
Builders who want a compact local TTS service that can run through ONNX Runtime on a CPU.
People testing multilingual speech or reference-voice registration behind a familiar HTTP API.
Less relevant for readers who need a polished consumer app, verified performance on Windows, or settled license terms before evaluation.
Editorial note
Why LifeHubber lists it
Audio8 is useful for builders who want to try a supplied voice before preparing reference-voice registration. Its default-voice runtime path lets a first CPU speech trial start without collecting a reference recording or loading the optional registration encoder. Choose the voice before downloading extra registration files. The bundled default voice uses reference_codes.npy, and register_default_voice.py does not load the optional encoder. To register your own reference recording through the service, download the optional registration files as well. Trying the supplied voice and preparing a new voice therefore need different files.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
What to explore next
Compare other compact local voice paths.
Audio8 packages voice cloning behind a small CPU service. Continue with a wider set of lightweight entry points or a device-spanning ONNX toolkit when those tradeoffs matter more.
More in Speech Models
Keep browsing this category
Explore more speech model resources.
Fish Audio S2 Pro
fishaudio/s2-pro
A text-to-speech model with in-script control over pauses, emphasis and mood.
AuK
Tencent-Hunyuan/AuK
A 1.5B speech model for instruction-guided text-to-speech, content and acoustic editing, paralinguistic changes, speech enhancement, and source separation, with public code, weights, demos, ComfyUI nodes, and fine-tuning materials.
NVIDIA NemotronLabs VoiceChat 11B
nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
An NVIDIA 11B full-duplex conversational speech model with streaming speech understanding and generation, realtime interruption handling, a separate tool-call channel, and official offline and container deployment paths.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.