Catalogue
Models
455 models, and deliberately not every model there is: one earns a place by being sold by a provider we track, measured by a leaderboard we follow, or published as open weights by a lab we follow. Pointer names like “-latest” stay out, because they answer for whatever the vendor aimed them at that week.
How we choose what to list96 models
sprag SymphonyspragSpeech to text95.6%—19——no live host
Parakeet10 sizes here
Parakeet TDT 0.6b v3 (fast Gpu Asr)NVIDIASpeech to textParakeet TDT 0.6b v3 is a speech-to-text model with a measured turn of speed: on the leaderboard's own hardware an hour of audio finishes in under a second.95.3%13,155×260.6B—no live hostParakeet TDT 0.6b v2 (fast Gpu Asr)NVIDIASpeech to textParakeet TDT 0.6b v2 is a speech-to-text model with a measured speed advantage: on the leaderboard's own hardware it transcribes an hour of audio in under a second.95.3%14,199×—0.6B—no live hostParakeet TDT 0.6b v3NVIDIASpeech to textParakeet is a tiny downloadable speech-to-text model from NVIDIA that turns audio into written words.95.1%6,076×250.6B—no live hostParakeet TDT 0.6b v2NVIDIASpeech to textParakeet is NVIDIA's tiny downloadable speech-to-text model, small enough to run on almost any device and fast enough to process an hour of audio in under a second.95.3%6,025×10.6B—no live hostParakeet TDT CTC 110mNVIDIASpeech to textParakeet TDT CTC is a tiny English-only speech-to-text model from NVIDIA built for raw speed over accuracy.93.9%6,084×10.1B—no live hostParakeet TDT 1.1bNVIDIASpeech to textParakeet TDT is a tiny downloadable speech-to-text model from NVIDIA that turns audio into written words at extreme speed.—4,529×11.1B—no live hostParakeet RNNT 0.6bNVIDIASpeech to textParakeet RNNT is a 0.6-billion-parameter speech-to-text model from NVIDIA built for extreme speed on clean audio.93.8%5,431×10.6B—no live hostParakeet CTC 1.1bNVIDIASpeech to textParakeet CTC is a 1.1-billion-parameter speech-to-text model from NVIDIA built for extreme speed on modest hardware.94.1%5,023×11.1B—no live hostParakeet CTC 0.6bNVIDIASpeech to textParakeet CTC is a 0.6-billion-parameter speech-to-text model from NVIDIA that processes audio faster than real time.93.8%5,870×10.6B—no live hostParakeet RNNT 1.1bNVIDIASpeech to textParakeet RNNT is a compact downloadable speech-to-text model from NVIDIA built for extreme speed on English audio.94.2%4,139×11.1B—no live hostZipformer cr3 sizes here
Zipformer cr CTC Transducer XL 290M (fast Gpu Asr, ctc Greedy Search)Sounds Good AISpeech to textZipformer cr CTC Transducer XL 290M is built for bulk transcription where throughput matters more than accuracy, and on the leaderboard's own hardware it turns an hour of audio into text in under a second. We list no download and no host for it, so there is no route we can point you to.93.6%20,628×—0.3B—no live hostZipformer cr CTC Transducer XL 290M (fast Gpu Asr, transducer Modified Beam Search)Sounds Good AISpeech to textZipformer cr CTC Transducer XL 290M is a speech-to-text model with a measured speed that is hard to argue with, and no route we can point you to for using it.94.7%19,047×—0.3B—no live hostZipformer cr CTC Transducer XL 290MSounds Good AISpeech to textZipformer cr CTC Transducer XL is a 290-million-parameter speech-to-text model released in July 2026.94.7%158×10.3B—no live hostASR K1 (preview)sopheaSpeech to textASR K1 (preview) is a speech-to-text model with solid measured accuracy on English audio, better than most models on every condition we hold figures for.95.7%—2——no live hostMuse Voice TranscribeMetaSpeech to textMuse Voice Transcribe is Meta's proprietary speech-to-text model covering 25 languages.95.2%—25——no live hostAzure Speech2 sizes here
Azure Speech 07 2026MicrosoftSpeech to textAzure Speech 07 2026 is Microsoft's proprietary speech-to-text service that turns audio into written words across 25 languages.96.2%—25——no live hostAzure Speech 06 2026MicrosoftSpeech to textMicrosoft's Azure Speech 06 2026 is a speech-to-text model with solid measured accuracy on English audio, better than most models on every recording condition we hold.95.7%—25——no live hostNiagara4 sizes here
Niagara 84m Batch.enApplied Brain ResearchSpeech to textNiagara 84m Batch.en is a speech-to-text model you can download and run yourself, built for bulk transcription where speed matters more than word accuracy.92.5%1,928×184M—no live hostNiagara 9m Batch.enApplied Brain ResearchSpeech to textNiagara 9m Batch.en is a speech-to-text model you can download and run yourself, and it is built for one thing: turning a large audio backlog into rough text as fast as possible.86.5%7,158×19M—no live hostNiagara 38m Batch.enApplied Brain ResearchSpeech to textNiagara is a 38-million-parameter speech-to-text model built for extreme speed on English audio.91.3%4,529×138M—no live hostNiagara 19m Batch.enApplied Brain ResearchSpeech to textNiagara 19m Batch.en is a 20-million-parameter English speech-to-text model built for extreme speed.89.6%4,324×120M—no live hostModulate MultilingualModulateSpeech to textModulate Multilingual is a speech-to-text model with measured accuracy among the best we list on clean read-aloud and meeting audio, and 3rd of 76 on Open ASR WER as of 28 Sep 2026.96.2%—99——no live hostScribe2 sizes here
Scribe v2 ProZoomSpeech to textScribe v2 Pro is a speech-to-text model from Zoom that tops the Open ASR word error rate board, and it is among the best any model achieves on accented speech and on meeting recordings. We list no download and no host for it, so there is no route we can point you to for running it.96.4%—11——no live hostScribe v1ZoomSpeech to textScribe v1 is Zoom's proprietary speech-to-text model, available only through ElevenLabs, that turns audio into written words.95.9%—1—$0.004/minute of audiovia ElevenLabsGranite Speech6 sizes here
Granite Speech 5.0 470m Turboctc ncIBMSpeech to textGranite Speech 5.0 is a tiny English-only speech-to-text model from IBM that processes audio faster than almost anything else in its class.95.2%12,762×10.5B—no live hostGranite Speech 5.0 470m TurboctcIBMSpeech to textGranite Speech 5.0 Turboctc is a tiny downloadable speech-to-text model from IBM that turns English audio into text at extreme speed.95%12,946×10.5B—no live hostGranite Speech 4.1 2bIBMSpeech to textGranite Speech 4.1 is a compact downloadable speech-to-text model from IBM that turns audio into written words.95.4%546×62.3B—no live hostGranite Speech 4.1 2b NARIBMSpeech to textGranite Speech 4.1 is a 2.3-billion-parameter speech-to-text model from IBM that turns audio into written words under a permissive Apache licence.95.3%2,074×52.3B—no live hostGranite Speech 3.3 2bIBMSpeech to textGranite Speech 3.3 2b is a compact downloadable speech-to-text model from IBM with a permissive Apache licence.94.6%509×53B—no live hostGranite Speech 3.3 8bIBMSpeech to textGranite Speech 3.3 is an Apache-licensed speech-to-text model from IBM that turns audio into written words.94.7%263×58.6B—no live hostWhisper3 sizes here
Whisper 1 APIOpenAISpeech to textWhisper 1 API is OpenAI's hosted-only speech-to-text service, available through two cloud providers at identical per-minute rates.————$0.006/minute of audiovia OpenAIWhisper Large v3 TurboOpenAISpeech to textWhisper Large v3 Turbo is a tiny, MIT-licensed speech-to-text model built for speed over accuracy.93.6%797×990.8B$0.001/minute of audiovia GroqWhisper Large v3OpenAISpeech to textWhisper Large v3 turns recorded speech into written words, and is still the name most people reach for.94.2%470×991.5B$0.002/minute of audiovia GroqSpeechmatics EnhancedSpeechmaticsSpeech to textSpeechmatics Enhanced is a hosted-only speech-to-text engine covering 55 languages.94.7%—55——no live hostSmallest AI PulseSmallest AISpeech to textSmallest AI Pulse is a hosted-only speech-to-text model that turns audio into written words across 38 languages.95.6%—38——no live hostModulate VfastModulateSpeech to textModulate Vfast is a hosted-only speech-to-text model that ranks among the most accurate we list on both clean read-aloud audio and meeting-room recordings.—————no live hostSolaria 3GladiaSpeech to textSolaria 3 is a proprietary speech-to-text model from Gladia that turns audio into written words.95.4%—5——no live hostomniASR LLM 7B v2MetaSpeech to textomniASR LLM 7B v2 is Meta's proprietary speech-to-text model with 7.8 billion parameters and broad language coverage.93.6%129×16767.8B—no live hostomniASR CTC 7B v2MetaSpeech to textomniASR CTC 7B v2 is Meta's proprietary speech-to-text model with 6.5 billion parameters and 1,676 languages claimed.90.9%519×16766.5B—no live hostScribe v2ElevenLabsSpeech to textScribe v2 is ElevenLabs' proprietary speech-to-text service that handles 90 languages and charges per minute of audio.96%—90—$0.004/minute of audiovia ElevenLabsAvalon v1 enAqua VoiceSpeech to textAvalon v1 en is Aqua Voice's proprietary speech-to-text model for English audio.—————no live hostResonant2 sizes here
Resonant 1 FlashReson8Speech to textResonant 1 Flash is a proprietary speech-to-text model from Reson8 that turns audio into written words.96%—9——no live hostResonant 1Reson8Speech to textResonant 1 is a speech-to-text model with strong measured accuracy on clean and accented English, but we list no download and no host for it, so there is no route to run it that our data supports. Its weak spot is corporate earnings calls, where it sits near the bottom of the field.96%—9——no live hostUniversal2 sizes here
Universal 3 ProAssemblyAISpeech to textUniversal 3 Pro turns recorded speech into written text across 99 languages.——99——no live hostUniversal 3 5 ProAssemblyAISpeech to textUniversal 3 5 Pro is a speech-to-text model with strong measured accuracy on clearly recorded English, and a real weak spot in meeting rooms.95.7%—18——no live hostAudio8 ASR 0.1BAutoArk AISpeech to textAudio8 ASR is a tiny downloadable speech-to-text model from AutoArk AI that processes audio faster than almost anything we track — an hour of audio in about five seconds.93%719×70.3B—no live hostMOSS Transcribe Preview 2BOpenMOSSSpeech to textMOSS Transcribe Preview is a 2.4-billion-parameter speech-to-text model from OpenMOSS with a permissive Apache licence.95.1%151×12.4B—no live hostQwen3 ASR2 sizes here
Qwen3 ASR 1.7B HFQwenSpeech to textQwen3 ASR 1.7B HF is a speech-to-text model you can download and run yourself, with a licence that allows commercial use, changes and redistribution.95.7%820×302B—no live hostQwen3 ASR 0.6B HFQwenSpeech to textQwen3 ASR is a tiny downloadable speech-to-text model with a permissive Apache licence and support for 30 languages.95%744×300.8B—no live host