Models / Zoom/ Scribe v2 Pro

Scribe v2 Pro

Zoom

Speech to textTranscribes a recording into words

Input: audio. Output: text.InputOutput
Type
Closed
Input
None held
Output
None held
Cached
None held

We don't hold a list price for this model yet · hosted only — no weights published

Our take

Written Sep 30, 2026

Scribe v2 Pro is a speech-to-text model from Zoom that tops the Open ASR word error rate board, and it is among the best any model achieves on accented speech and on meeting recordings. We list no download and no host for it, so there is no route we can point you to for running it.

Who should pick it

Reach for it when the audio is a meeting room or a clearly recorded earnings call, where its word error rates are among the best any model achieves on those conditions. It is also 1st of 76 on Accented speech as of 28 Sep 2026, so non-native speakers are a strength rather than a risk. Skip it if you need a route to run it, if you need speaker labelling, timestamps or streaming, or if your audio is not English.

The case for it

  • 6.1% of words wrong on meeting recordings, among the best any model achieves on that condition, so meeting-room audio is the job to bring it.
  • 4.1% of words wrong on accented speech, among the best any model achieves there, which makes it a reasonable first trial for non-native speakers.
  • 1st of 92 on Financial calls as of 28 Sep 2026, so clearly recorded earnings calls are the safest place to start.

The case against it

  • We list no download and no host for it, so there is no way to run it yourself and no hosted offer to reach it through.
  • 26th of 92 on European-accented speech as of 28 Sep 2026, against 1st of 76 on Accented speech as of 28 Sep 2026, so a real meeting with non-native speakers is a qualified case.
  • Every accuracy figure is English only: 11 languages are listed, but none of the others is measured, so accuracy in them is unverified in our data.
00

How good is it?

A speech-to-text model for turning recordings of meetings, accented speakers and read-aloud audio into written text.

Good at
  • turning spoken English into written textOpen ASR WER · 1st of 76
  • transcribing recordings of meetings in a roomRecorded meetings · 1st of 92
  • transcribing speakers with a range of accentsAccented speech · 1st of 76
  • transcribing clear recordings of people reading aloudClean read speech · 17th of 92

TranscriptionTurning speech into text5 of 5Open ASR WER · 1st of 76

Words it gets right

96.4%

Misses roughly one word in 28, averaged over nine English test sets.

Languages

11

Stated by the leaderboard; we do not hold the list itself.

Where it struggles
Read aloudaudiobooks, clean recording1.1%17th of 92
Podcasts and videoeveryday internet audio7.8%24th of 92
Accented speechspeakers from many countries4.1%1st of 76
Meetingsa room, several people, far microphone6.1%1st of 92

Percentage of words wrong on each set, lower better. Bars are scaled to this model's own worst case; the placing beneath each rate is against every model measured on that set.

The figures above come from the Open ASR Leaderboard, an independent public test that runs every model on the same recordings. It is the only measurement of transcription quality we know of, so there are no other scores to show.

Other boards it appears on
Financial calls 1st of 92Recorded meetings 1st of 92Harder read speech 13th of 92Clean read speech 17th of 92Podcasts and video 24th of 92European-accented speech 26th of 92Accented speech 1st of 76

Each of these is the same transcription job on a different kind of recording, so together they say where it holds up and where it slips — not how closely it follows an instruction.

Every published score for this model8 scoresEvery figure we hold, from 8 boards, with who ran it and a link to the source — including the boards no rating above is built on.
3.59source ↗
4.12source ↗
1.36source ↗
6.08source ↗
3.08source ↗
7.75source ↗
1.11source ↗
2.24source ↗
01

Where to get it

We hold no priced listing for Scribe v2 Pro.

There is no copy to download and no host in our price data, so Zoom is where to look. We watch OpenRouter, the provider APIs we track and the LiteLLM price set; this version appears in none of them, which is a gap in what we collect rather than a statement about what Zoom sells.

02

When we formed this view

Recent changes

Sep 11, 2026ListedListed on LLMap
What movedfirst indexed by our pipeline
Sep 11, 2026BenchmarkScored 3.59 on Open ASR WER
What movedleaderboard
Sep 11, 2026BenchmarkScored 4.12 on Accented speech
What movedleaderboard
Sep 11, 2026BenchmarkScored 1.36 on Financial calls
What movedleaderboard
Sep 11, 2026BenchmarkScored 6.08 on Recorded meetings
What movedleaderboard
Sep 11, 2026BenchmarkScored 3.08 on European-accented speech
What movedleaderboard
Sep 11, 2026BenchmarkScored 7.75 on Podcasts and video
What movedleaderboard
Sep 11, 2026BenchmarkScored 1.11 on Clean read speech
What movedleaderboard
Sep 11, 2026BenchmarkScored 2.24 on Harder read speech
What movedleaderboard

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • We don't hold a list price for this model yet — the gap is ours, not the lab's.
03

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence row of this model's own. A source states its weights are not published, so the determination that governs it is the one for closed weights, API access only.

Identifiers

Takes in, gives back
Audio in, text out
Catalogue slug
zoom-scribe-v2-pro

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us