LIFEHUBBER
Choose theme

AI Resources

WaxalNLP

Hugging Face likes: 289 Hugging Face downloads, last 30 days: 6.8K Declared license: CC-BY-SA-4.0: CC-BY-SA-4.0 Last modified September 1, 2026: Modified 1mo ago
Stats from Hugging Face

WaxalNLP is a collection of African-language audio and text for speech-recognition and text-to-speech research.

Its dataset card separates natural speech with transcriptions from scripted recordings for voice synthesis. Choose the language and speech task before loading a subset; this is data for model work, not a ready-made transcription or voice app. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.

What it is

Paired audio and text

The card describes ASR data in 19 languages and TTS data in 17, with task-specific configurations and fields.

Why it stands out

A second ASR split design

ASR v2 rearranges existing utterances into speaker-disjoint splits for benchmarking; it does not add new recordings.

Availability

Language subsets on Hugging Face

The dataset page provides loading examples, provider-specific terms and the metadata used for its newer ASR splits.

Why it matters

What makes it useful

Load a named language configuration with the datasets library: sna_asr selects Shona speech-recognition data, while swa_tts selects Swahili synthesis data. ASR rows pair audio with a transcription; TTS rows pair audio with the text script. These pairs give a model-training or evaluation pipeline recordings and their associated words.

Notable points

What stands out

For an ASR benchmark, the v2 design separates speakers between training and evaluation. The card publishes per-language split comparisons and a manifest recording the source revision and split parameters. Those records help identify which evaluation you are running; they do not establish a model's accuracy or eliminate every possible data overlap.

Before using

What to review

The card requires FFmpeg and datasets[audio] for audio loading, with compatible torch, torchaudio and torchcodec versions when decoding fails.

Its provider tables assign different CC-BY or CC-BY-SA terms to subsets, and the card asks users to check the chosen language's specific licence. Those dataset terms do not establish permission to reproduce someone's voice.

Streaming limits memory use but still transfers the selected language audio. Plan storage and decoding work around the subset you actually need.

Review collection provenance and representation gaps before selecting training or evaluation data. Decide whether identifiable voice recordings fit your consent, privacy and downstream-use rules; provider licence labels do not answer those questions.

Reader fit

Who may find it relevant

Researchers preparing African-language ASR training or evaluation data.

Speech-synthesis builders needing recorded audio paired with scripts.

Readers tracking how African languages are represented in speech datasets and research.

Less useful for someone who wants to upload a recording and immediately receive a transcript.

Editorial note

Why LifeHubber lists it

WaxalNLP is useful when a researcher needs to reconstruct the intended African-language evaluation set from the original recordings. Its v2 row map makes held-out utterance membership explicit across the old split boundaries, providing an assembly route beyond choosing a benchmark name. A v2 test set is not necessarily the old test set with a new label. For re-partitioned languages, the card's loading example reads all three original labelled splits—train, validation and test—then uses each row's id and the v2_split map to select the target split. Reading only the original test split silently misses moved utterances. The example excludes unlabeled and unmapped rows. Six already-disjoint languages have direct v2 configurations instead: Akan, Amharic, Oromo, Sidama, Tigrinya and Wolaytta.

Source links

Source materials

Reader note

Before relying on this entry

LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.

Advertisements

Advertisements

For project maintainers

Listed here? You can use the badge.

If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.

See what’s moving