Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>While automatic speech recognition (ASR) systems have made remarkable progress in many high-resource languages [1–3], most of the world’s 7,000+ languages remain unsupported. Expanding ASR coverage has long been regarded as prohibitively expensive and of lim- ited value for benchmarking, further hampered by architectures that restrict language coverage to a fixed set. To transcend these limita- tions, this article introduces Omnilingual ASR, the first large-scale ASR system (1,660+ languages) designed with extensibility in mind. More specifically, OmniASR enables communities to add previously unserved languages with a small number of their own data samples. On the modeling side, Omnilingual ASR scales self-supervised pre- training to 7B parameters to learn robust speech representations and introduces an encoder-decoder architecture designed for zero-shot gen- eralization, leveraging a decoder inspired by large language models to effectively exploit these representations. This capability is enabled by a massive and diverse training corpus that combines breadth of cov- erage with linguistic variety, helping the model learn representations robust enough to adapt to previously unseen languages. The corpus incorporates publicly available resources with new community-sourced recordings. Omnilingual ASR expands coverage to more than 1,660 languages, the largest such effort to date, including over 500 languages never before served by any ASR system. Automatic evaluations show substantial gains over prior systems, especially in extreme low-resource conditions, and strong generalization to languages never encountered during training. Omnilingual ASR is released∗ as a family of models ranging from compact 300M variants for low-power devices to large 7B models for maximum accuracy. We highlight how open-sourcing these models and associated artifacts can lower barriers for researchers and communities alike, inviting new forms of participation without onerous requirements.</p>

Show More

Keywords

languages omnilingual models coverage training

Related Articles

PORE

About

Connect