BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models
Large end-to-end Transformer-based ASR models such as Whisper have transformed speech recognition, but most prominent models are highly multilingual and often perform poorly on languages that are less well represented in their training data. Effective language adaptation through simple fine-tuning on monolingual data has long been known, but it had only been applied to a small number of languages.
BuzzASR scales up this simple approach to all 102 languages covered in the FLEURS dataset. It also implements a more complex language adaptation strategy that integrates monolingual tokenizer replacement and data augmentation using text-only fine-tuning.
The models outperform Whisper-large-v3 on 77 out of 102 languages, reducing character error rates by a factor of over 2.8 on average. BuzzASR achieves state-of-the-art CER among open-source systems on 27 of 102 languages on the combined FLEURS and Common Voice test set. Its tokenizer replacement strategy yields an average 3.3x improvement in compression rate (characters per token) over Whisper's multilingual BPE, with gains of up to 21.7x.
All models, code, and detailed results are released at https://lemn-lab.github.io/buzz-asr, making broad multilingual ASR adaptation resources publicly available.