Theme
AI Resources
WaxalNLP
WaxalNLP is a multilingual African-language speech collection with separate datasets for automatic speech recognition and text-to-speech research.
The current card describes transcribed natural speech for ASR and single-speaker scripted recordings for TTS, supplied through several collection partners. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
Multilingual speech dataset
The ASR side covers 19 languages and the TTS side covers 17, with different recording styles, data structures, and intended speech tasks.
Why it stands out
African-language speech focus
The card reports about 1,250 hours of transcribed ASR speech and more than 180 hours of TTS recordings, while the newer speaker-disjoint ASR splits support cleaner evaluation across voices.
Availability
Hugging Face dataset page
Public materials are available through a Hugging Face dataset page with dataset-card details, usage information, and linked research context.
Why it matters
What makes it useful
Speech model progress depends on which languages have usable public data. Its African-language speech focus gives readers a dataset source to inspect for multilingual coverage beyond the most commonly represented benchmark languages.
What to know
Where it fits
Read it as part of the dataset and speech-research layer rather than the model or chatbot layer. It is most relevant to readers following language coverage, speech resources, and multilingual benchmarks.
Notable points
What stands out
The collection is not one uniform pool: language coverage, provider, recording conditions, and license differ between ASR and TTS subsets.
Before using
What to review
Which ASR or TTS language subset fits the work, including its recording conditions, data fields, split design, provider, and current license.
What the dataset card says about collection provenance and representation gaps, and whether identifiable voice data fits your own consent, privacy, and downstream-use rules.
The storage and tooling burden: the full collection is large, audio loading needs additional dependencies, and streaming still transfers the selected language data.
Reader fit
Who may find it relevant
Readers tracking multilingual speech datasets and language representation in AI.
Builders working on speech systems or research with African-language coverage in mind.
Less relevant for readers mainly focused on consumer assistants or non-speech tooling.
Editorial note
Why LifeHubber lists it
WaxalNLP helps readers choose by language and speech task instead of treating African-language data as one category; the ASR/TTS split, provider-specific terms, recording conditions, and evaluation design all affect fit.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
More in Datasets
Keep browsing this category
Explore more datasets.
ParseBench
run-llama/ParseBench
A document parsing benchmark for AI-agent workflows, focused on whether parsed PDFs preserve structure and meaning for downstream evaluation.
UltraData-SFT-2605
openbmb/UltraData-SFT-2605
An OpenBMB supervised fine-tuning dataset with 15,036,178 thinking and non-thinking samples across math, code, knowledge, Chinese, instruction-following, and multilingual configurations, used in MiniCPM5-1B-SFT post-training.
Terminal-Bench 2.0
harbor-framework/terminal-bench-2
A terminal-agent benchmark for evaluating AI agents on hard containerized command-line tasks, with Harbor run commands, task-level registry pages, GitHub and Hugging Face materials, docs, and paper links.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.