Modulate Multilingual
Modulate
Speech to textTranscribes a recording into words
- Type
- Closed
- Input
- None held
- Output
- None held
- Cached
- None held
We don't hold a list price for this model yet · hosted only — no weights published
Our take
Written Sep 30, 2026Modulate Multilingual is a speech-to-text model with measured accuracy among the best we list on clean read-aloud and meeting audio, and 3rd of 76 on Open ASR WER as of 28 Sep 2026. We list no download and no host for it, so there is no route here to run it.
Reach for it when the recordings are clear — read-aloud audio, podcasts, video — or when you are transcribing meeting-room audio, where its measured error rates are among the best we list. Every accuracy figure we hold is English only, so treat the 99 listed languages as unmeasured. Skip it if you need a route to run it, or if your audio is heavily accented.
The case for it
- 0.9% of words wrong on clean read-aloud recordings, among the best any model achieves on this condition, and 1st of 92 on Clean read speech as of 28 Sep 2026.
- 6.3% of words wrong on meeting recordings, among the best any model achieves on this condition, and 2nd of 92 on Recorded meetings as of 28 Sep 2026.
- Better than most on accented speech and everyday internet audio: 6% of words wrong on accented recordings and 6.9% on podcasts and video.
- Strong on corporate earnings calls, at 2.06% of words wrong with professional reference transcripts.
The case against it
- We list no download for it and no host for it, so there is nothing here to point you to.
- All accuracy figures we hold are English only; the 99 languages listed are not measured by any source we hold.
- Accented speakers cost it several times its clean-speech error rate, so try it on your own recordings first.
How good is it?
A speech-to-text model for turning recordings of meetings, podcasts and read-aloud speech into written text.
- turning spoken English into written textOpen ASR WER · 3rd of 76
- transcribing recordings of meetings in a roomRecorded meetings · 2nd of 92
- transcribing speakers with a range of accentsAccented speech · 12th of 76
- transcribing podcasts and video audioPodcasts and video · 3rd of 92
- transcribing clear recordings of people reading aloudClean read speech · 1st of 92
TranscriptionTurning speech into text4.5 of 5Open ASR WER · 3rd of 76
96.2%
Misses roughly one word in 26, averaged over nine English test sets.
99
Stated by the leaderboard; we do not hold the list itself.
Percentage of words wrong on each set, lower better. Bars are scaled to this model's own worst case; the placing beneath each rate is against every model measured on that set.
The figures above come from the Open ASR Leaderboard, an independent public test that runs every model on the same recordings. It is the only measurement of transcription quality we know of, so there are no other scores to show.
Each of these is the same transcription job on a different kind of recording, so together they say where it holds up and where it slips — not how closely it follows an instruction.
Every published score for this model8 scoresEvery figure we hold, from 8 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Where to get it
We hold no priced listing for Modulate Multilingual.
There is no copy to download and no host in our price data, so Modulate is where to look. We watch OpenRouter, the provider APIs we track and the LiteLLM price set; this version appears in none of them, which is a gap in what we collect rather than a statement about what Modulate sells.
When we formed this view
Recent changes
What moved
first indexed by our pipelineEach date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- We don't hold a list price for this model yet — the gap is ours, not the lab's.
Licence and identifiers
What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.
Licence
We hold no licence row of this model's own. A source states its weights are not published, so the determination that governs it is the one for closed weights, API access only.
Identifiers
- Takes in, gives back
- Audio in, text out
- Catalogue slug
- modulate-multilingual