First question
What happens to the recording?
Separate transcription, voice generation, realtime agents, and meeting notes before comparing the tools that touch spoken data.
See this starting pointAI Resources
A focused map for turning speech into text, text into voices, and live conversations into AI workflows.
Voice tools can involve sensitive recordings or speaker likeness. Check consent, storage, source terms, and data handling before using real conversations.
Choose by situation
These paths organize source-linked Resources by the question they can help you investigate. They do not rank products or cover every option.
First question
Separate transcription, voice generation, realtime agents, and meeting notes before comparing the tools that touch spoken data.
See this starting pointVoice boundary
A tool that transcribes audio is not the same as one that generates, clones, or performs a voice. Check the source terms for the actual use.
See this starting pointTesting path
Use sample clips or non-sensitive meetings first, then check retention, export, deletion, and review controls before using important recordings.
See this starting pointCoverage and freshness
These groups are selective starting points, not a complete directory. The date reflects the newest included Resource’s LifeHubber added date, not a recheck of every linked source. Check the original source for current setup, terms, limits, privacy, access, costs, and behaviour.
Fresh in this topic
Recently added Resources from the groups below.
Transcribe and understand audio
Use this group when the job is turning speech, noisy audio, calls, or recordings into text or structured context.
CohereLabs/cohere-transcribe-03-2026
A compact 2B audio-to-text scope with 14-language coverage centres multilingual transcription rather than voice generation or meeting-agent features.
xzf-thu/Mega-ASR
Training and inference work on difficult real-world recordings focus on transcription under noise and uncontrolled acoustic conditions.
XiaomiMiMo/MiMo-V2.5-ASR
Coverage of Mandarin, English, Chinese dialects, code-switching, songs, and multiple speakers addresses ASR shaped by language mixing and speaker complexity.
OpenMOSS/MOSS-Audio
Speech, sound, music, captioning, time-aware questions, and ASR give MOSS-Audio a broader audio-understanding scope beyond transcription alone.
nvidia/nemotron-3.5-asr-streaming-0.6b
A 600M streaming design with multilingual locales and latency and throughput tables exposes the balance between performance and language coverage in live transcription.
Vaibhavs10/insanely-fast-whisper
An on-device CLI focused on fast Whisper transcription prioritises local processing speed over a complete voice application.
Generate or run voices
Open this group when the output is a spoken voice, a compact TTS model, or a voice workflow that needs source and consent checks.
fishaudio/s2-pro
Detailed prosody and emotional-delivery controls support expressive direction beyond simply converting text to speech.
KittenML/KittenTTS
A very small footprint targets voice generation under lightweight deployment constraints.
hexgrad/Kokoro-82M
With 82M parameters, voice materials, samples, and an inference library, this model offers a compact path from evaluation to implementation.
OpenMOSS/MOSS-TTS
Voice design, dialogue, realtime speech, compact generation, and sound effects place the MOSS-TTS family across several audio-output jobs.
OpenMOSS/MOSS-TTS-Nano
A tiny multilingual model with CPU-friendly realtime positioning targets responsive local speech on modest hardware.
sbintuitions/sarashina2.2-tts
Japanese-centred generation, English support, style transfer, and a zero-shot voice path cover Japanese delivery and voice adaptation.
supertone-inc/supertonic
Local ONNX inference, 31-language coverage, and expression tags show multilingual speech with controllable delivery across device types.
openbmb/VoxCPM2
Multilingual generation, voice design, controllable cloning, and streaming support combine tailored voices with live TTS output.
jamiepine/voicebox
A local-first studio combines cloning, synthesis, effects, and app workflows for work that needs an end-user production environment rather than only a model.
Meetings and realtime agents
Use this group when timing, rooms, calls, summaries, or meeting records are the hard part.
livekit/agents
WebRTC rooms, telephony, tools, and deployment paths show voice agents joining live calls and interacting during the session.
pipecat-ai/pipecat
Audio and video pipeline stages, transports, client SDKs, and structured flows show realtime conversation components being assembled and swapped.
open-software-network/os-june
Desktop meeting notes, dictation, projects, and local app state carry spoken work into organised desktop tasks rather than ending at transcription.
Zackriya-Solutions/meetily
Live transcription with local or external summary choices makes a meeting assistant’s processing boundary explicit.
yazinsai/OpenOats
Conversational note-taking distinguishes a responsive meeting companion from a passive recording pipeline.
NVIDIA/personaplex
A full-duplex speech-to-speech design with role prompts and voice conditioning addresses interruption handling and conversational persona in live agents.
Also in AI
Keep the thread going with AI Guides for decision habits for messy AI choices, AI Access for free and low-cost ways to compare AI model access, AI Ballot for a clearer view of what readers are leaning toward.