Omnilingual ASR
Open (Apache 2.0) suite: CTC and LLM-ASR models at 300M/1B/3B/7B, v2 checkpoints and 'Unlimited' long-audio LLM-ASR variants added December 2025, plus a 7B wav2vec 2.0 speech encoder and a corpus covering 350+ underserved languages. Checkpoints download via fairseq2 (e.g. https://dl.fbaipublicfiles.com/mms/omniASR-LLM-7B-v2.pt). Successor to MMS. The 'first' claim is Meta's ('never previously supported by any ASR model').
- Input
- audio
- Output
- text
- License
- apache-2.0
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| GitHub (fairseq2 checkpoints) | omniASR_LLM_7B_v2 | github.com/facebookresearch/omnilingual-asr | — |
| Hugging Face (demo space and dataset) | — | huggingface.co/facebook | — |
Notable capabilities (2)
- FIRST ASR for 1,600+ languages: Transcribes 1,600+ languages, ~500 of them never before supported by any ASR system (Whisper covers 99). source
- Zero-shot in-context language extension: omniASR_LLM_7B_ZS transcribes new languages from a few paired audio-text examples at inference, extending potential coverage to 5,400+ languages. source
Meta's open massively multilingual speech recognition (Nov 2025, still Meta's current open ASR).
Sources: https://github.com/facebookresearch/omnilingual-asr · https://ai.meta.com/blog/omnilingual-asr-advancing-automatic-speech-recognition/
Other Meta models
Muse Voice Transcribe 1.0 · Muse Spark 1.3 · Muse Glimmer 30B · Llama 4 Maverick (17B-128E) · Llama 4 Scout (17B-16E)