Open Source arXiv cs.CL

YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech

YODAS v3speech corpusmultilingualopen dataset

YODAS v3 is a weakly-labeled speech corpus containing over 1.1 million hours of 48 kHz multi-channel audio in 147 languages, released under a CC BY 3.0 license. It is positioned as not only the largest open speech dataset to date, but also the first truly large-scale speech corpus with high-fidelity stereo audio.

The release provides a collection methodology that introduces new techniques for gathering language-balanced speech data. The effectiveness of this approach is reflected in the corpus's language distribution: 22 languages have over 10,000 hours of data, and 73 languages have over 5,000 hours.

The authors also conduct extensive analyses of the data composition, including the distribution of languages, audio quality, and transcription quality. To show the dataset's effectiveness, they train baseline speech recognition and neural codec models.

The corpus is available for download at https://huggingface.co/datasets/espnet/yodas3. Its scale, multilingual coverage, and high-fidelity stereo audio make it a notable resource for speech and audio research.

Read original →

← Back to home