Scribe v2
ElevenLabs
Speech to textTranscribes a recording into words
- Type
- Proprietary
- Input
- None held
- Output
- None held
- Cached
- None held
We don't hold a list price for this model yet · hosted only — no weights published
Our take
Written Aug 2, 2026Scribe v2 is ElevenLabs' proprietary speech-to-text model that turns audio into written words across 90 languages. It excels on clean, prepared English audio with near-perfect accuracy, but degrades sharply in meetings and with accented speech where no pricing or hosted access is currently available.
Pick this for high-stakes transcription of clean, prepared English speech, or financial calls with controlled audio. Consider it for multilingual projects needing 90-language support, though accuracy is unverified beyond English. Skip it if you are transcribing meetings, podcasts, or accented speech, where error rates rise sevenfold to eightfold, or if you need a hosted option today.
The case for it
- Near-perfect on clean, read-aloud English: about one word in a hundred wrong, and only slightly more on European-accented read speech.
- Handles controlled difficult audio better than chaotic real-world audio: harder read speech at roughly a third the error rate of podcasts, video, or meetings.
- Broad language support with 90 languages listed, though accuracy is measured only in English.
The case against it
- Dramatically worse in uncontrolled real-world settings: meetings and accented speech see error rates roughly seven to eight times higher than clean read speech.
- Accented speech is the worst-performing condition tested, with the highest error rate of any scenario.
- No available pricing or hosted access in our data, and no release date or parameter count disclosed.
How good is it?
TranscriptionTurning speech into text4.5 of 5Open ASR WER · 7th of 74
95.4%
Misses roughly one word in 22, averaged over nine English test sets.
90
Stated by the leaderboard; we do not hold the list itself.
Percentage of words wrong on each set, lower better. Bars are scaled to this model's own worst case, not to the board.
The figures above come from the Open ASR Leaderboard, an independent public test that runs every model on the same recordings. It is the only measurement of transcription quality we know of, so there are no other scores to show.
These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done, which is why they get no rating.
Every published score for this model8 scoresEvery figure we hold, from 8 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Where to get it
We hold no priced listing for Scribe v2.
There is no copy to download and no host in our price data, so ElevenLabs is where to look. We watch OpenRouter, the provider APIs we track and the LiteLLM price set; this version appears in none of them, which is a gap in what we collect rather than a statement about what ElevenLabs sells.
When we formed this view
Dates behind this page
Prices last checked 6h ago
What we do not know about this model yet
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- We don't hold a list price for this model yet — the gap is ours, not the lab's.
Licence and identifiers
What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.
Licence
Commercial API terms. We hold no licence record for this model, so there is nothing to summarise here.
Identifiers
- Modality record
- audio->text
- Catalogue slug
- elevenlabs-scribe-v2