10 months ago
Meta Unveils Omnilingual ASR Supporting Over 1,600 Languages
Imagine a special computer program that can understand and write down what people are saying in many different languages!
Meta has created a new one called Omnilingual ASR that can handle over 1,600 languages.
This is super exciting because it includes about 500 languages that computers haven't been very good at understanding before, like some languages spoken by smaller groups of people.
Meta wants to make sure everyone can use technology, no matter what language they speak.
They trained this program using a lot of real-world talking.
While it's really good, it's not perfect for every single language yet, but Meta is sharing it openly.
This means scientists and people who speak these languages can work together to make it even better.
They also released a big collection of recorded voices and text in 350 languages to help others build similar helpful tools.
Meta has launched Omnilingual ASR, an AI system capable of speech recognition in over 1,600 languages.
The system includes support for approximately 500 low-resource languages, which are being transcribed using AI for the first time.
Meta has also released the Omnilingual ASR Corpus, featuring transcribed speech in 350 underserved languages, to aid AI development.
The initiative aims to democratize digital speech technology and reduce the digital divide for underrepresented linguistic communities.
While achieving progress, Meta acknowledges that accuracy varies, with a significant portion of low-resource languages still facing challenges in transcription accuracy.
- Who
- Meta's Fundamental AI Research (FAIR) team
- What
- Unveiled Omnilingual ASR, a suite of open-weight AI models with automatic speech recognition (ASR) capabilities supporting over 1,600 languages, including 500 low-resource languages for the first time. Also released the Omnilingual ASR Corpus.
- Where
- Globally, with a focus on supporting diverse languages, including Indian dialects.
- When
- Monday, November 10
- Why
- To support a vast number of languages, bridge the digital divide for underrepresented linguistic communities, and enable developers to build a wide range of AI-driven speech applications.
Democratization and Inclusivity
Challenges and Competition
Supporting Low-Resource Languages
Democratization and Inclusivity
Meta's Omnilingual ASR supports 500 low-resource languages, many for the first time with AI, aiming to reduce the digital divide. The community-driven framework allows users to add new languages with minimal samples.
Challenges and Competition
Despite broad support, only 78% of the 1,600+ languages achieve a character error rate below 10%. This indicates persistent challenges in achieving high accuracy for many low-resource languages.
AI Innovation Landscape in India
Democratization and Inclusivity
Meta's release of open-weight models and corpora provides tools for developers and researchers, potentially boosting local AI innovation.
Challenges and Competition
Indian AI startups face stiff competition from tech giants like Meta and OpenAI, who are expanding into India, a key growth market. Government initiatives like Mission Bhashini aim to foster local AI but face challenges from these larger players.
Key facts
- Model Name
- Omnilingual ASR, Omnilingual wav2vec 2.0
- Developer
- Meta's Fundamental AI Research (FAIR)
- Languages Supported
- Over 1,600
- Low-Resource Languages Supported
- Approximately 500
- Model Architecture
- Omnilingual wav2vec 2.0 (scaled up to 7 billion parameters)
- Corpus Availability
- Omnilingual ASR Corpus of transcribed speech in 350 underserved languages, released under CC-BY license
- License
- Apache 2.0 license for LLM-ASR
Quotes
Meta
The company
“We then built two decoder variants to map those into character tokens. The first decoder relies on a traditional connectionist temporal classification (CTC) objective, while the second leverages a traditional transformer decoder, commonly used in LLMs.”
indianexpress.com
“In practice, this means that a speaker of an unsupported language can provide only a handful of paired audio-text samples and obtain usable transcription quality — without training data at scale, onerous expertise, or access to high-end compute.”
indianexpress.com





