Models / ElevenLabs/ Scribe v2

Scribe v2

ElevenLabs

Speech to textTranscribes a recording into words

Input: audio. Output: text.InputOutput
Type
Closed
Input
None held
Output
None held
Cached
None held

We don't hold a list price for this model yet · hosted only — no weights published

Our take

Written Sep 14, 2026

Scribe v2 is ElevenLabs' proprietary speech-to-text service that handles 90 languages and charges per minute of audio. It is a strong value pick for high-volume transcription, with better-than-most accuracy on several English audio conditions.

Who should pick it

Pick this for high-volume transcription where per-minute cost matters, or for multilingual pipelines needing broad language coverage. Use it for clean read-aloud or European-accented English audio, podcasts, video, and meeting-room audio, where it outperforms most models on each condition. Skip it if you need open weights, alternative hosts for price competition, or verified accuracy outside English.

The case for it

  • Among the lowest per-minute rates we hold for proprietary speech-to-text.
  • Better than most models on read-aloud audio, with a word error rate below the median across 85 models.
  • Better than most on European-accented parliamentary speech, podcasts, video, and real meetings.
  • Better than most on accented earnings-call audio, well under the median across 69 models.

The case against it

  • Never 'among the best' on any condition we measure — always 'better than most' with a gap to the top.
  • Only one tracked offer, from ElevenLabs, so no alternative hosts or price competition.
  • All accuracy figures are English; the 90-language claim is unverified in our data.
00

How good is it?

A speech-to-text model for turning recordings of meetings, podcasts and accented speakers into written text.

Good at
  • turning spoken English into written textOpen ASR WER · 4th of 76
  • transcribing recordings of meetings in a roomRecorded meetings · 22nd of 92
  • transcribing speakers with a range of accentsAccented speech · 3rd of 76
  • transcribing podcasts and video audioPodcasts and video · 21st of 92
  • transcribing clear recordings of people reading aloudClean read speech · 12th of 92

TranscriptionTurning speech into text4.5 of 5Open ASR WER · 4th of 76

Words it gets right

96%

Misses roughly one word in 25, averaged over nine English test sets.

Languages

90

Stated by the leaderboard; we do not hold the list itself.

Where it struggles
Read aloudaudiobooks, clean recording1.1%12th of 92
Podcasts and videoeveryday internet audio7.6%21st of 92
Accented speechspeakers from many countries4.8%3rd of 76
Meetingsa room, several people, far microphone8.2%22nd of 92

Percentage of words wrong on each set, lower better. Bars are scaled to this model's own worst case; the placing beneath each rate is against every model measured on that set.

The figures above come from the Open ASR Leaderboard, an independent public test that runs every model on the same recordings. It is the only measurement of transcription quality we know of, so there are no other scores to show.

Other boards it appears on
European-accented speech 1st of 92Clean read speech 12th of 92Harder read speech 15th of 92Podcasts and video 21st of 92Recorded meetings 22nd of 92Financial calls 63rd of 92Accented speech 3rd of 76

Each of these is the same transcription job on a different kind of recording, so together they say where it holds up and where it slips — not how closely it follows an instruction.

Every published score for this model8 scoresEvery figure we hold, from 8 boards, with who ran it and a link to the source — including the boards no rating above is built on.
3.97source ↗
4.8source ↗
3.01source ↗
8.19source ↗
1.41source ↗
7.63source ↗
1.06source ↗
2.33source ↗
01

Where to rent it

Prices checked 2 hours ago

Cheapest published offer

ElevenLabs, direct

The only live listing we hold.

per minute of audio
$0.004
Context served
—
Throughput
Not measured
Current provider offers with price, context and prompt-privacy answers
ProviderPrice per minute of audioContextThroughputTrains on promptsLogs promptsZero retention
ElevenLabsDirect$0.004checked 2 hours agonot reportednot measuredUnknownUnknownUnknown

Across the 1 listings we hold: 0 say they do not train on prompts, 0 say they do and 1 does not say. 0 appear in the zero-retention registry we check; the rest are unknown to us.

02

When we formed this view

Recent changes

Sep 11, 2026BenchmarkScored 3.97 on Open ASR WER
What movedleaderboard
Sep 11, 2026BenchmarkScored 4.8 on Accented speech
What movedleaderboard
Sep 11, 2026BenchmarkScored 8.19 on Recorded meetings
What movedleaderboard
Sep 11, 2026BenchmarkScored 7.63 on Podcasts and video
What movedleaderboard
Sep 11, 2026BenchmarkScored 1.06 on Clean read speech
What movedleaderboard
Sep 11, 2026BenchmarkScored 2.33 on Harder read speech
What movedleaderboard
Aug 2, 2026BenchmarkScored 3.01 on Financial calls
What movedleaderboard
Aug 2, 2026BenchmarkScored 1.41 on European-accented speech
What movedleaderboard
Aug 1, 2026ListedListed on LLMap
What movedfirst indexed by our pipeline

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • 1 of 1 listings publishes no parameter list, so what its API accepts is unknown to us.
  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 1 of 1 listings does not say whether it trains on prompts.
  • We don't hold a list price for this model yet — the gap is ours, not the lab's.
  • We hold no cached-input rate for any of its listings.
  • We hold no batch or off-peak rate for any of its listings.
03

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence row of this model's own. A source states its weights are not published, so the determination that governs it is the one for closed weights, API access only.

Identifiers

Takes in, gives back
Audio in, text out
Catalogue slug
elevenlabs-scribe-v2

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us