Models / Sounds Good AI/ Zipformer cr CTC Transducer XL 290M (fast Gpu Asr, transducer Modified Beam Search)

Zipformer cr CTC Transducer XL 290M (fast Gpu Asr, transducer Modified Beam Search)

Sounds Good AI

Speech to textTranscribes a recording into words

Input: audio. Output: text.InputOutput
Type
Closed
Input
None held
Output
None held
Cached
None held

We don't hold a list price for this model yet · hosted only — we have no record of published weights

Our take

Written Sep 28, 2026

Zipformer cr CTC Transducer XL 290M is a speech-to-text model with a measured speed that is hard to argue with, and no route we can point you to for using it. We list no download and no host, so the numbers below are what it does on the leaderboard's own rig rather than something you can act on today.

Who should pick it

Worth knowing about if you transcribe large backlogs of clearly recorded audio, or professionally recorded calls, where it is among the most accurate models we list. The catch is availability: we list no download for it and no host offers it, so there is no way to run it from here. Skip it if you need something you can actually deploy, or if your audio is podcasts and everyday internet recordings, where it sits below the middle of the field.

The case for it

  • Among the most accurate models we list on clearly recorded corporate speech: 1.6% of words wrong on earnings calls with professional reference transcripts, 5th of 92 on Financial calls as of 28 Sep 2026.
  • Fast enough to make a long backlog practical: 19,047 times real time on the leaderboard's own hardware, so an hour of audio is transcribed in under a second there — a figure measured on that rig, not on yours.
  • Holds up on meeting-room audio better than most: 10.2% of words wrong on real meetings with crosstalk and a distant microphone, against a field middle of 10.3%.

The case against it

  • There is no route we can point you to: we list no download for it and no host offers it, so neither running it yourself nor reaching it through a host is available from our data.
  • Everyday internet audio is its weakest measured condition: 8.3% of words wrong on podcasts and video, worse than most models on that condition.
  • No language coverage is held, and every accuracy figure we hold is English only, so nothing here says how it handles any other language.
00

How good is it?

TranscriptionTurning speech into text3 of 5Open ASR WER · 42nd of 76

Words it gets right

94.7%

Misses roughly one word in 19, averaged over nine English test sets.

How fast it listens

19,047×2nd of 74

an hour of audio in under a second, on the board's own hardware. Your machine will differ.

Where it struggles
Read aloudaudiobooks, clean recording1.3%38th of 92
Podcasts and videoeveryday internet audio8.3%47th of 92
Accented speechspeakers from many countries7.9%38th of 76
Meetingsa room, several people, far microphone10.2%45th of 92

Percentage of words wrong on each set, lower better. Bars are scaled to this model's own worst case; the placing beneath each rate is against every model measured on that set.

The figures above come from the Open ASR Leaderboard, an independent public test that runs every model on the same recordings. It is the only measurement of transcription quality we know of, so there are no other scores to show.

Other boards it appears on
Financial calls 5th of 92Clean read speech 38th of 92Harder read speech 41st of 92Recorded meetings 45th of 92Podcasts and video 47th of 92European-accented speech 67th of 92Accented speech 38th of 76

Each of these is the same transcription job on a different kind of recording, so together they say where it holds up and where it slips — not how closely it follows an instruction.

Every published score for this model9 scoresEvery figure we hold, from 9 boards, with who ran it and a link to the source — including the boards no rating above is built on.
19047source ↗
5.29source ↗
7.9source ↗
1.64source ↗
10.24source ↗
4.31source ↗
8.3source ↗
1.31source ↗
3.03source ↗
01

Where to get it

We hold no priced listing for Zipformer cr CTC Transducer XL 290M (fast Gpu Asr, transducer Modified Beam Search).

There is no copy to download and no host in our price data, so Sounds Good AI is where to look. We watch OpenRouter, the provider APIs we track and the LiteLLM price set; this version appears in none of them, which is a gap in what we collect rather than a statement about what Sounds Good AI sells.

02

When we formed this view

Recent changes

Sep 28, 2026ListedListed on LLMap
What movedfirst indexed by our pipeline
Sep 28, 2026BenchmarkScored 19047 on Open ASR RTFx
What movedleaderboard
Sep 28, 2026BenchmarkScored 5.29 on Open ASR WER
What movedleaderboard
Sep 28, 2026BenchmarkScored 7.9 on Accented speech
What movedleaderboard
Sep 28, 2026BenchmarkScored 1.64 on Financial calls
What movedleaderboard
Sep 28, 2026BenchmarkScored 10.24 on Recorded meetings
What movedleaderboard
Sep 28, 2026BenchmarkScored 4.31 on European-accented speech
What movedleaderboard
Sep 28, 2026BenchmarkScored 8.3 on Podcasts and video
What movedleaderboard
Sep 28, 2026BenchmarkScored 1.31 on Clean read speech
What movedleaderboard
Sep 28, 2026BenchmarkScored 3.03 on Harder read speech
What movedleaderboard

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • We don't hold a list price for this model yet — the gap is ours, not the lab's.
03

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.

Identifiers

Takes in, gives back
Audio in, text out
Catalogue slug
soundsgoodai-zipformer-cr-ctc-transducer-xl-290m-fast-gpu-asr-transducer-modified-beam-search

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us