Speechmatics Enhanced
Speechmatics
Speech to textTranscribes a recording into words
- Type
- Closed
- Input
- None held
- Output
- None held
- Cached
- None held
We don't hold a list price for this model yet · hosted only — no weights published
Our take
Written Sep 4, 2026Speechmatics Enhanced is a hosted-only speech-to-text engine covering 55 languages. Its English accuracy sits near the middle of the field overall, but it outperforms most rivals on podcasts, video and meeting-room audio with overlapping speakers.
Choose this for transcribing podcasts, YouTube and internet video, or meeting-room audio with overlapping speakers, where it beats most alternatives. It also suits multilingual deployments needing 55 languages where English accuracy is acceptable. Skip it if you need clean read-aloud transcription or heavily accented earnings calls, where it lags the field.
The case for it
- Better than most on everyday internet audio, with about 7.8% of words wrong on podcasts and video.
- Strong on meeting-room audio despite overlapping speech and far-field recording, at roughly 10% word error.
- 55 languages supported, broader than most transcription engines.
The case against it
- Near the middle of the field overall on English accuracy.
- Weak on clean read-aloud audio — the easiest test case — at 1.73% word error, worse than most.
- Struggles with heavily accented earnings calls, at 11.86% word error, worse than most.
How good is it?
TranscriptionTurning speech into text3 of 5Open ASR WER · 40th of 76
94.7%
Misses roughly one word in 19, averaged over nine English test sets.
55
Stated by the leaderboard; we do not hold the list itself.
Percentage of words wrong on each set, lower better. Bars are scaled to this model's own worst case; the placing beneath each rate is against every model measured on that set.
The figures above come from the Open ASR Leaderboard, an independent public test that runs every model on the same recordings. It is the only measurement of transcription quality we know of, so there are no other scores to show.
Each of these is the same transcription job on a different kind of recording, so together they say where it holds up and where it slips — not how closely it follows an instruction.
Every published score for this model8 scoresEvery figure we hold, from 8 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Where to get it
We hold no priced listing for Speechmatics Enhanced.
There is no copy to download and no host in our price data, so Speechmatics is where to look. We watch OpenRouter, the provider APIs we track and the LiteLLM price set; this version appears in none of them, which is a gap in what we collect rather than a statement about what Speechmatics sells.
When we formed this view
Recent changes
What moved
first indexed by our pipelineEach date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- We don't hold a list price for this model yet — the gap is ours, not the lab's.
Licence and identifiers
What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.
Licence
We hold no licence row of this model's own. A source states its weights are not published, so the determination that governs it is the one for closed weights, API access only.
Identifiers
- Takes in, gives back
- Audio in, text out
- Catalogue slug
- speechmatics-enhanced