Models / NVIDIA/ Parakeet TDT 0.6b v3 (fast Gpu Asr)

Parakeet TDT 0.6b v3 (fast Gpu Asr)

NVIDIA

Speech to textTranscribes a recording into words

Input: audio. Output: text.InputOutput
Type
Closed
Input
None held
Output
None held
Cached
None held

We don't hold a list price for this model yet · hosted only — we have no record of published weights

Our take

Written Sep 28, 2026

Parakeet TDT 0.6b v3 is a speech-to-text model with a measured turn of speed: on the leaderboard's own hardware an hour of audio finishes in under a second. The catch is access — we list no download for it and no host offers it, so there is no route we can point you to.

Who should pick it

Reach for it when you are working through a large backlog of clearly recorded English audio and throughput matters more than the last point of accuracy, and when the recording is accented speech, podcasts or meeting-room audio, where it sits better than most of the field. Skip it if you need corporate earnings-call transcription, where it ranks near the bottom, or if you need to run it yourself or reach it through a host — we list neither.

The case for it

  • 13,155 times real time on the leaderboard's own hardware, so a long backlog becomes practical to work through; your own machine may differ.
  • 5.8% of words wrong on accented speech, better than most models on that condition, so accented recordings are worth trying here.
  • 7.8% of words wrong on podcasts and video, better than most on that condition, which makes everyday internet audio a reasonable fit.

The case against it

  • We list no download for it and no host offers it, so there is no way we can point you to run it.
  • 3.56% of words wrong on financial calls, 78th of 92 on that condition — near the bottom of the field, so earnings-call work is the wrong job for it.
  • Every accuracy figure we hold is English; the 26 listed languages are unverified in our data.
00

How good is it?

A speech-to-text model for turning recordings of accented English speech into written text.

Good at
  • transcribing speakers with a range of accentsAccented speech · 8th of 76

TranscriptionTurning speech into text3.5 of 5Open ASR WER · 26th of 76

Words it gets right

95.3%

Misses roughly one word in 21, averaged over nine English test sets.

How fast it listens

13,155×4th of 74

an hour of audio in under a second, on the board's own hardware. Your machine will differ.

Languages

26

Stated by the leaderboard; we do not hold the list itself.

Where it struggles
Read aloudaudiobooks, clean recording1.5%46th of 92
Podcasts and videoeveryday internet audio7.8%24th of 92
Accented speechspeakers from many countries5.8%8th of 76
Meetingsa room, several people, far microphone9%32nd of 92

Percentage of words wrong on each set, lower better. Bars are scaled to this model's own worst case; the placing beneath each rate is against every model measured on that set.

The figures above come from the Open ASR Leaderboard, an independent public test that runs every model on the same recordings. It is the only measurement of transcription quality we know of, so there are no other scores to show.

Other boards it appears on
European-accented speech 24th of 92Podcasts and video 24th of 92Recorded meetings 32nd of 92Harder read speech 43rd of 92Clean read speech 46th of 92Financial calls 78th of 92Accented speech 8th of 76

Each of these is the same transcription job on a different kind of recording, so together they say where it holds up and where it slips — not how closely it follows an instruction.

Every published score for this model9 scoresEvery figure we hold, from 9 boards, with who ran it and a link to the source — including the boards no rating above is built on.
13155source ↗
4.71source ↗
5.81source ↗
3.56source ↗
9.02source ↗
3.04source ↗
7.75source ↗
1.46source ↗
3.08source ↗
01

Where to get it

We hold no priced listing for Parakeet TDT 0.6b v3 (fast Gpu Asr).

There is no copy to download and no host in our price data, so NVIDIA is where to look. We watch OpenRouter, the provider APIs we track and the LiteLLM price set; this version appears in none of them, which is a gap in what we collect rather than a statement about what NVIDIA sells.

02

Models people weigh against Parakeet TDT 0.6b v3 (fast Gpu Asr)

03

When we formed this view

Recent changes

Sep 28, 2026ListedListed on LLMap
What movedfirst indexed by our pipeline
Sep 28, 2026BenchmarkScored 13155 on Open ASR RTFx
What movedleaderboard
Sep 28, 2026BenchmarkScored 4.71 on Open ASR WER
What movedleaderboard
Sep 28, 2026BenchmarkScored 5.81 on Accented speech
What movedleaderboard
Sep 28, 2026BenchmarkScored 3.56 on Financial calls
What movedleaderboard
Sep 28, 2026BenchmarkScored 9.02 on Recorded meetings
What movedleaderboard
Sep 28, 2026BenchmarkScored 3.04 on European-accented speech
What movedleaderboard
Sep 28, 2026BenchmarkScored 7.75 on Podcasts and video
What movedleaderboard
Sep 28, 2026BenchmarkScored 1.46 on Clean read speech
What movedleaderboard
Sep 28, 2026BenchmarkScored 3.08 on Harder read speech
What movedleaderboard

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • We don't hold a list price for this model yet — the gap is ours, not the lab's.
04

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence record for this model, and no record of published weights either — so we can neither summarise its terms nor point you at the weights.

Identifiers

Takes in, gives back
Audio in, text out
Catalogue slug
nvidia-parakeet-tdt-0-6b-v3-fast-gpu-asr

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us