The Open ASR Leaderboard Adds Its First Global South Language
Leaderboards shape what gets built, and the Open ASR Leaderboard has worked to make its metrics more trustworthy via held-out private splits, benchmark-fitting analysis, and improved normalizers. Yet aggregate word error rate (WER) hides known disparities: prior research found commercial ASR systems roughly twice as bad for Black speakers than white speakers, with further differences by gender, age, and accent. The leaderboard's test sets record what was said but almost nothing about who said it, so these biases were invisible.
To close this gap, the new evaluation sets — Monsoon en-IN and Monsoon hi-IN — are introduced. Hindi, spoken by over half a billion people, is the first Indic language on the leaderboard's multilingual tab, which previously covered only European languages. Each set has a public split for self-scoring and a private split withheld to limit benchmark-specific optimization; the four splits are speaker-disjoint, comprising 4,888 speakers total, with 12 speaker attributes recorded for each.
The dataset was built to vary along nine axes: geography, age, gender, vocabulary, devices, acoustic environments, speech type, speech rate, and the existence of multiple valid transcripts for the same audio. This is a deliberate contrast to most benchmarks built from readily available audio. Collectors recruited across hundreds of districts rather than fewer locations, contributors used their own handsets and connections indoors and out, and prompts were designed to push toward opinion, disagreement, narration, and recall — contexts where named entities, numbers, and unrehearsed phrasing appear. Age and gender are recorded per speaker and verified. Multiple valid transcripts are a property of the reference rather than the audio, addressed in a later section.
By surfacing performance across these varied axes, the new splits let researchers and practitioners see where WER is not evenly distributed, encouraging more inclusive ASR development rather than optimizing only an average number.