Models / Meta/ omniASR LLM 7B v2

omniASR LLM 7B v2

Meta

Speech to textTranscribes a recording into words

Input: audio. Output: text.InputOutput
Type
Closed
Input
None held
Output
None held
Cached
None held

We don't hold a list price for this model yet · hosted only — we have no record of published weights

Our take

Written Sep 4, 2026

omniASR LLM 7B v2 is Meta's proprietary speech-to-text model with 7.8 billion parameters and broad language coverage. It excels on clean read-aloud English audio and processes very fast, but no commercial hosts currently offer it and its accuracy drops sharply on natural or accented recordings.

Who should pick it

Pick this for high-throughput batch transcription of clean, read-aloud English where speed matters — it processes an hour of audio in roughly 25 seconds on benchmark hardware. Consider it when you need coverage across 1,676 languages and can accept English-only accuracy validation. Skip it if you are transcribing podcasts, meetings, accented speech, or need a hosted API with clear pricing.

The case for it

  • Among the most accurate on clean read-aloud English audio, with a word error rate below the field median.
  • Very fast transcription: benchmark hardware processes an hour of audio in roughly 25 seconds.
  • Extremely broad language coverage, with 1,676 languages listed.
  • Strong on structured financial speech, with low error on earnings calls.

The case against it

  • Worse than most on podcasts, video, accented speech, and meetings — error rates rise well above field medians.
  • No commercial hosting available; zero offers listed and no inference pricing in our data.
  • Proprietary weights with no licence terms disclosed.
00

How good is it?

A speech-to-text model for turning recordings into written text, though it trails most others on accuracy.

Less good at
  • turning spoken English into written textOpen ASR WER · 63rd of 76
  • transcribing recordings of meetings in a roomRecorded meetings · 74th of 92
  • transcribing speakers with a range of accentsAccented speech · 64th of 76
  • transcribing podcasts and video audioPodcasts and video · 72nd of 92

TranscriptionTurning speech into text2.5 of 5Open ASR WER · 63rd of 76

Words it gets right

93.6%

Misses roughly one word in 16, averaged over nine English test sets.

How fast it listens

129×67th of 74

an hour of audio in 28 seconds, on the board's own hardware. Your machine will differ.

Languages

1676

Stated by the leaderboard; we do not hold the list itself.

Where it struggles
Read aloudaudiobooks, clean recording1.4%41st of 92
Podcasts and videoeveryday internet audio9.2%72nd of 92
Accented speechspeakers from many countries11.4%64th of 76
Meetingsa room, several people, far microphone13.9%74th of 92

Percentage of words wrong on each set, lower better. Bars are scaled to this model's own worst case; the placing beneath each rate is against every model measured on that set.

The figures above come from the Open ASR Leaderboard, an independent public test that runs every model on the same recordings. It is the only measurement of transcription quality we know of, so there are no other scores to show.

Other boards it appears on
Clean read speech 41st of 92Harder read speech 47th of 92European-accented speech 53rd of 92Financial calls 66th of 92Podcasts and video 72nd of 92Recorded meetings 74th of 92Accented speech 64th of 76

Each of these is the same transcription job on a different kind of recording, so together they say where it holds up and where it slips — not how closely it follows an instruction.

Every published score for this model9 scoresEvery figure we hold, from 9 boards, with who ran it and a link to the source — including the boards no rating above is built on.
128.8source ↗
6.4source ↗
11.44source ↗
3.22source ↗
13.86source ↗
4.03source ↗
9.2source ↗
1.35source ↗
3.25source ↗
01

Where to get it

We hold no priced listing for omniASR LLM 7B v2.

There is no copy to download and no host in our price data, so Meta is where to look. We watch OpenRouter, the provider APIs we track and the LiteLLM price set; this version appears in none of them, which is a gap in what we collect rather than a statement about what Meta sells.

02

When we formed this view

Recent changes

Sep 11, 2026BenchmarkScored 128.8 on Open ASR RTFx
What movedleaderboard
Sep 11, 2026BenchmarkScored 6.4 on Open ASR WER
What movedleaderboard
Sep 11, 2026BenchmarkScored 11.44 on Accented speech
What movedleaderboard
Sep 11, 2026BenchmarkScored 9.2 on Podcasts and video
What movedleaderboard
Aug 2, 2026BenchmarkScored 3.22 on Financial calls
What movedleaderboard
Aug 2, 2026BenchmarkScored 13.86 on Recorded meetings
What movedleaderboard
Aug 2, 2026BenchmarkScored 4.03 on European-accented speech
What movedleaderboard
Aug 2, 2026BenchmarkScored 1.35 on Clean read speech
What movedleaderboard
Aug 2, 2026BenchmarkScored 3.25 on Harder read speech
What movedleaderboard
Aug 1, 2026ListedListed on LLMap
What movedfirst indexed by our pipeline

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • We don't hold a list price for this model yet — the gap is ours, not the lab's.
03

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.

Identifiers

Takes in, gives back
Audio in, text out
Catalogue slug
facebook-omniasr-llm-7b-v2

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us