Zipformer cr CTC Transducer XL 290M (fast Gpu Asr, ctc Greedy Search)
Sounds Good AI
Speech to textTranscribes a recording into words
- Type
- Closed
- Input
- None held
- Output
- None held
- Cached
- None held
We don't hold a list price for this model yet · hosted only — we have no record of published weights
Our take
Written Sep 29, 2026Zipformer cr CTC Transducer XL 290M is built for bulk transcription where throughput matters more than accuracy, and on the leaderboard's own hardware it turns an hour of audio into text in under a second. We list no download and no host for it, so there is no route we can point you to.
Reach for it when you have a large backlog of clearly recorded audio and the volume matters more than the last point of accuracy. It is 62nd of 76 on Open ASR WER as of 28 Sep 2026, and worse than most models on clean read-aloud recordings, podcasts, accented speech and meetings alike. Skip it if you need accurate transcription of everyday internet audio, or if you need to run it yourself or pick a host.
The case for it
- Among the fastest speech-to-text models we list: 20,628 times real time on the leaderboard's own hardware, an hour of audio in under a second, so a long backlog becomes practical to work through.
- Respectable on clearly recorded corporate earnings calls, at 2.75% word error rate on SPGISpeech, though accented speakers cost it several times that error rate.
The case against it
- Worse than most models on every condition we hold: 1.7% on clean read-aloud recordings against a field middle of 1.5%, 9.5% on podcasts and video against 8.3%, and 9.3% on accented speech against 7.9%.
- No route to use it: we list no download for it and no host either, so there is no way we can point you to run it.
- English only: every accuracy figure we hold is English, and no source we hold measures other languages.
How good is it?
A speech-to-text model for turning recordings into written text, though it trails most models on transcription quality.
- turning spoken English into written textOpen ASR WER · 62nd of 76
- transcribing podcasts and video audioPodcasts and video · 75th of 92
TranscriptionTurning speech into text2.5 of 5Open ASR WER · 62nd of 76
93.6%
Misses roughly one word in 16, averaged over nine English test sets.
20,628×1st of 74
an hour of audio in under a second, on the board's own hardware. Your machine will differ.
Percentage of words wrong on each set, lower better. Bars are scaled to this model's own worst case; the placing beneath each rate is against every model measured on that set.
The figures above come from the Open ASR Leaderboard, an independent public test that runs every model on the same recordings. It is the only measurement of transcription quality we know of, so there are no other scores to show.
Each of these is the same transcription job on a different kind of recording, so together they say where it holds up and where it slips — not how closely it follows an instruction.
Every published score for this model9 scoresEvery figure we hold, from 9 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Where to get it
We hold no priced listing for Zipformer cr CTC Transducer XL 290M (fast Gpu Asr, ctc Greedy Search).
There is no copy to download and no host in our price data, so Sounds Good AI is where to look. We watch OpenRouter, the provider APIs we track and the LiteLLM price set; this version appears in none of them, which is a gap in what we collect rather than a statement about what Sounds Good AI sells.
When we formed this view
Recent changes
What moved
first indexed by our pipelineEach date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- We don't hold a list price for this model yet — the gap is ours, not the lab's.
Licence and identifiers
What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.
Licence
We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.
Identifiers
- Takes in, gives back
- Audio in, text out
- Catalogue slug
- soundsgoodai-zipformer-cr-ctc-transducer-xl-290m-fast-gpu-asr-ctc-greedy-search