Choose theme
AI Resources
WaxalNLP
WaxalNLP is a collection of African-language audio and text for speech-recognition and text-to-speech research.
Its dataset card separates natural speech with transcriptions from scripted recordings for voice synthesis. Choose the language and speech task before loading a subset; this is data for model work, not a ready-made transcription or voice app. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
Paired audio and text
The card describes ASR data in 19 languages and TTS data in 17, with task-specific configurations and fields.
Why it stands out
A second ASR split design
ASR v2 rearranges existing utterances into speaker-disjoint splits for benchmarking; it does not add new recordings.
Availability
Language subsets on Hugging Face
The dataset page provides loading examples, provider-specific terms and the metadata used for its newer ASR splits.
Why it matters
What makes it useful
Load a named language configuration with the datasets library: sna_asr selects Shona speech-recognition data, while swa_tts selects Swahili synthesis data. ASR rows pair audio with a transcription; TTS rows pair audio with the text script. These pairs give a model-training or evaluation pipeline recordings and their associated words.
What to know
Where it fits
The recognition collection contains natural speech from varied voices. The synthesis collection uses single-speaker recordings of phonetically balanced scripts. Select the side matching recognition or voice-generation work; a shared language name does not make the recording styles or fields interchangeable.
Notable points
What stands out
For an ASR benchmark, the v2 design separates speakers between training and evaluation. The card publishes per-language split comparisons and a manifest recording the source revision and split parameters. Those records help identify which evaluation you are running; they do not establish a model's accuracy or eliminate every possible data overlap.
Before using
What to review
The card requires FFmpeg and datasets[audio] for audio loading, with compatible torch, torchaudio and torchcodec versions when decoding fails.
Its provider tables assign different CC-BY or CC-BY-SA terms to subsets, and the card asks users to check the chosen language's specific licence. Those dataset terms do not establish permission to reproduce someone's voice.
Streaming limits memory use but still transfers the selected language audio. Plan storage and decoding work around the subset you actually need.
Review collection provenance and representation gaps before selecting training or evaluation data. Decide whether identifiable voice recordings fit your consent, privacy and downstream-use rules; provider licence labels do not answer those questions.
Reader fit
Who may find it relevant
Researchers preparing African-language ASR training or evaluation data.
Speech-synthesis builders needing recorded audio paired with scripts.
Readers tracking how African languages are represented in speech datasets and research.
Less useful for someone who wants to upload a recording and immediately receive a transcript.
Editorial note
Why LifeHubber lists it
WaxalNLP is useful when a researcher needs to reconstruct the intended African-language evaluation set from the original recordings. Its v2 row map makes held-out utterance membership explicit across the old split boundaries, providing an assembly route beyond choosing a benchmark name. A v2 test set is not necessarily the old test set with a new label. For re-partitioned languages, the card's loading example reads all three original labelled splits—train, validation and test—then uses each row's id and the v2_split map to select the target split. Reading only the original test split silently misses moved utterances. The example excludes unlabeled and unmapped rows. Six already-disjoint languages have direct v2 configurations instead: Akan, Amharic, Oromo, Sidama, Tigrinya and Wolaytta.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
More in Datasets
Keep browsing this category
Explore more datasets.
ParseBench
run-llama/ParseBench
A document parsing benchmark for AI-agent workflows, focused on whether parsed PDFs preserve structure and meaning for downstream evaluation.
UltraData-SFT-2605
openbmb/UltraData-SFT-2605
An OpenBMB supervised fine-tuning dataset with 15,036,178 thinking and non-thinking samples across math, code, knowledge, Chinese, instruction-following, and multilingual configurations, used in MiniCPM5-1B-SFT post-training.
Terminal-Bench 2.0
harbor-framework/terminal-bench-2
A terminal-agent benchmark for evaluating AI agents on hard containerized command-line tasks, with Harbor run commands, task-level registry pages, GitHub and Hugging Face materials, docs, and paper links.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.