Scribe v1
Zoom
Speech to textTranscribes a recording into words
- Type
- Closed
- Input
- None held
- Output
- None held
- Cached
- None held
We don't hold a list price for this model yet · hosted only — no weights published
Our take
Written Sep 5, 2026Scribe v1 is Zoom's proprietary speech-to-text model, available only through ElevenLabs, that turns audio into written words. It excels at meeting-room transcription and clean read-aloud audio, though it supports just one language and offers no self-hosting path.
Pick this for meeting transcription where accuracy matters — it is among the best any model achieves on recorded meeting-room audio. Use it for clean read-aloud or financial call work, or for budget-conscious short-form transcription via its single hosted provider. Skip it if you need multiple languages, want to self-host, or are transcribing heavily accented speech where it lags the field's best.
The case for it
- Among the most accurate models we list on meeting-room audio, with a word error rate well below the field middle.
- Strong on clean and semi-clean read speech, beating the typical model on both benchmarks.
- Respectable on everyday internet audio such as podcasts and video, above the field middle.
- Very low cost per minute for hosted transcription via its single provider.
The case against it
- Only one language supported; every accuracy figure is English-only with no measurement for other languages.
- Accented speech accuracy lags its own meeting-room performance and the field's best.
- Locked to a single provider with no open-weight or self-host option.
How good is it?
A speech-to-text model for turning recordings into written text, with meeting audio as its strongest ground.
- turning spoken English into written textOpen ASR WER · 7th of 76
- transcribing recordings of meetings in a roomRecorded meetings · 4th of 92
- transcribing speakers with a range of accentsAccented speech · 17th of 76
- transcribing clear recordings of people reading aloudClean read speech · 13th of 92
TranscriptionTurning speech into text4 of 5Open ASR WER · 7th of 76
95.9%
Misses roughly one word in 24, averaged over nine English test sets.
1
Stated by the leaderboard; we do not hold the list itself.
Percentage of words wrong on each set, lower better. Bars are scaled to this model's own worst case; the placing beneath each rate is against every model measured on that set.
The figures above come from the Open ASR Leaderboard, an independent public test that runs every model on the same recordings. It is the only measurement of transcription quality we know of, so there are no other scores to show.
Each of these is the same transcription job on a different kind of recording, so together they say where it holds up and where it slips — not how closely it follows an instruction.
Every published score for this model8 scoresEvery figure we hold, from 8 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Where to rent it
Prices checked 15 days ago
- per minute of audio
- $0.004
- Context served
- —
- Throughput
- Not measured
| Provider | Price per minute of audio | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| ElevenLabsDirect | $0.004checked 15 days ago | not reported | not measured | Unknown | Unknown | Unknown |
Across the 1 listings we hold: 0 say they do not train on prompts, 0 say they do and 1 does not say. 0 appear in the zero-retention registry we check; the rest are unknown to us.
When we formed this view
Recent changes
What moved
first indexed by our pipelineEach date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- 1 of 1 listings publishes no parameter list, so what its API accepts is unknown to us.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 1 of 1 listings does not say whether it trains on prompts.
- We don't hold a list price for this model yet — the gap is ours, not the lab's.
- We hold no cached-input rate for any of its listings.
- We hold no batch or off-peak rate for any of its listings.
Licence and identifiers
What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.
Licence
We hold no licence row of this model's own. A source states its weights are not published, so the determination that governs it is the one for closed weights, API access only.
Identifiers
- Takes in, gives back
- Audio in, text out
- Catalogue slug
- zoom-scribe-v1