LIFEHUBBER
Theme

AI Resources

Voice, speech, and meeting AI

A focused map for turning speech into text, text into voices, and live conversations into AI workflows.

Voice tools can involve sensitive recordings or speaker likeness. Check consent, storage, source terms, and data handling before using real conversations.

Choose by situation

Start with the job or constraint that matters now.

These paths organize source-linked Resources by the question they can help you investigate. They do not rank products or cover every option.

First question

What happens to the recording?

Separate transcription, voice generation, realtime agents, and meeting notes before comparing the tools that touch spoken data.

See this starting point

Voice boundary

Separate speech from speaker likeness

A tool that transcribes audio is not the same as one that generates, clones, or performs a voice. Check the source terms for the actual use.

See this starting point

Testing path

Start with low-risk audio

Use sample clips or non-sensitive meetings first, then check retention, export, deletion, and review controls before using important recordings.

See this starting point

Coverage and freshness

Newest LifeHubber addition included here: July 6, 2026

These groups are selective starting points, not a complete directory. The date reflects the newest included Resource’s LifeHubber added date, not a recheck of every linked source. Check the original source for current setup, terms, limits, privacy, access, costs, and behaviour.

Fresh in this topic

Newer Resources already included in this map

3

Recently added Resources from the groups below.

Transcribe and understand audio

Speech-to-text and audio understanding projects

6

Use this group when the job is turning speech, noisy audio, calls, or recordings into text or structured context.

Cohere Transcribe

CohereLabs/cohere-transcribe-03-2026

Hugging Face
Why it fits this starting point

A compact 2B audio-to-text scope with 14-language coverage centres multilingual transcription rather than voice generation or meeting-agent features.

STT, ASR

Mega-ASR

xzf-thu/Mega-ASR

GitHub
Why it fits this starting point

Training and inference work on difficult real-world recordings focus on transcription under noise and uncontrolled acoustic conditions.

Robust ASR, real-world audio Added to LifeHubber: May 22, 2026

MiMo-V2.5-ASR

XiaomiMiMo/MiMo-V2.5-ASR

GitHub
Why it fits this starting point

Coverage of Mandarin, English, Chinese dialects, code-switching, songs, and multiple speakers addresses ASR shaped by language mixing and speaker complexity.

ASR, dialects, code-switching Added to LifeHubber: April 24, 2026

MOSS-Audio

OpenMOSS/MOSS-Audio

GitHub
Why it fits this starting point

Speech, sound, music, captioning, time-aware questions, and ASR give MOSS-Audio a broader audio-understanding scope beyond transcription alone.

Audio understanding, ASR, reasoning Added to LifeHubber: April 27, 2026

NVIDIA Nemotron 3.5 ASR Streaming 0.6B

nvidia/nemotron-3.5-asr-streaming-0.6b

Hugging Face
Why it fits this starting point

A 600M streaming design with multilingual locales and latency and throughput tables exposes the balance between performance and language coverage in live transcription.

Streaming ASR, multilingual transcription Added to LifeHubber: June 21, 2026

insanely-fast-whisper

Vaibhavs10/insanely-fast-whisper

GitHub
Why it fits this starting point

An on-device CLI focused on fast Whisper transcription prioritises local processing speed over a complete voice application.

Transcription, local inference

Generate or run voices

Text-to-speech and voice generation projects

9

Open this group when the output is a spoken voice, a compact TTS model, or a voice workflow that needs source and consent checks.

Fish Audio S2 Pro

fishaudio/s2-pro

Hugging Face
Why it fits this starting point

Detailed prosody and emotional-delivery controls support expressive direction beyond simply converting text to speech.

TTS, expressive speech

KittenTTS

KittenML/KittenTTS

GitHub
Why it fits this starting point

A very small footprint targets voice generation under lightweight deployment constraints.

Compact TTS

Kokoro-82M

hexgrad/Kokoro-82M

Hugging Face
Why it fits this starting point

With 82M parameters, voice materials, samples, and an inference library, this model offers a compact path from evaluation to implementation.

Compact TTS, voice generation Added to LifeHubber: May 31, 2026

MOSS-TTS Family

OpenMOSS/MOSS-TTS

GitHub
Why it fits this starting point

Voice design, dialogue, realtime speech, compact generation, and sound effects place the MOSS-TTS family across several audio-output jobs.

TTS family, voice generation Added to LifeHubber: May 27, 2026

MOSS-TTS-Nano

OpenMOSS/MOSS-TTS-Nano

GitHub
Why it fits this starting point

A tiny multilingual model with CPU-friendly realtime positioning targets responsive local speech on modest hardware.

TTS, realtime speech

sarashina2.2-tts

sbintuitions/sarashina2.2-tts

Hugging Face
Why it fits this starting point

Japanese-centred generation, English support, style transfer, and a zero-shot voice path cover Japanese delivery and voice adaptation.

Japanese TTS, voice generation Added to LifeHubber: April 30, 2026

Supertonic

supertone-inc/supertonic

GitHub
Why it fits this starting point

Local ONNX inference, 31-language coverage, and expression tags show multilingual speech with controllable delivery across device types.

On-device multilingual TTS Added to LifeHubber: May 16, 2026

VoxCPM2

openbmb/VoxCPM2

Hugging Face
Why it fits this starting point

Multilingual generation, voice design, controllable cloning, and streaming support combine tailored voices with live TTS output.

TTS, voice cloning

Voicebox

jamiepine/voicebox

GitHub
Why it fits this starting point

A local-first studio combines cloning, synthesis, effects, and app workflows for work that needs an end-user production environment rather than only a model.

Voice tooling Added to LifeHubber: April 19, 2026

Meetings and realtime agents

Live conversation and meeting workflows

6

Use this group when timing, rooms, calls, summaries, or meeting records are the hard part.

LiveKit Agents

livekit/agents

GitHub
Why it fits this starting point

WebRTC rooms, telephony, tools, and deployment paths show voice agents joining live calls and interacting during the session.

Realtime voice and multimodal agents Added to LifeHubber: May 9, 2026

Pipecat

pipecat-ai/pipecat

GitHub
Why it fits this starting point

Audio and video pipeline stages, transports, client SDKs, and structured flows show realtime conversation components being assembled and swapped.

Voice agents, multimodal pipelines Added to LifeHubber: May 6, 2026

June

open-software-network/os-june

GitHub
Why it fits this starting point

Desktop meeting notes, dictation, projects, and local app state carry spoken work into organised desktop tasks rather than ending at transcription.

Meeting notes, dictation, desktop agent work Added to LifeHubber: July 1, 2026

Meetily

Zackriya-Solutions/meetily

GitHub
Why it fits this starting point

Live transcription with local or external summary choices makes a meeting assistant’s processing boundary explicit.

Local meeting transcription and summaries Added to LifeHubber: July 6, 2026

OpenOats

yazinsai/OpenOats

GitHub
Why it fits this starting point

Conversational note-taking distinguishes a responsive meeting companion from a passive recording pipeline.

Meetings, note-taking

PersonaPlex

NVIDIA/personaplex

GitHub
Why it fits this starting point

A full-duplex speech-to-speech design with role prompts and voice conditioning addresses interruption handling and conversational persona in live agents.

STS, conversational speech

Also in AI

Follow the next layer.

Keep the thread going with AI Guides for decision habits for messy AI choices, AI Access for free and low-cost ways to compare AI model access, AI Ballot for a clearer view of what readers are leaning toward.