Hardware / NVIDIA /GeForce RTX 4090
NVIDIA · GPU

GeForce RTX 4090

Runs 125 models fully in memory, 9 more with CPU offload. 9 of those sit close enough to 24 GB that our size estimate is what decides them.

Memory
24 GB
Bandwidth
1008 GB/s
Runs
125
01

Specifications

Memory
24 GB

Decides which models fit at all.

Memory bandwidth
1008 GB/s

How fast it can read the model. Sets the speed of the reply.

Compute
82.6 TFLOPS

Trillions of half-precision sums a second. Sets how fast it reads a long prompt, not how fast it replies.

Power draw
450 W

At full load, for sizing a power supply.

MSRP
$1,599

What the maker asked for it new. We hold no used or street prices, so this is not what one costs today.

Released
Oct 12, 2022
Architecture
Ada Lovelace
02

What it runs

What is quantisation? →
Granite 4.1 8B8.8B66K129 tok/sFits in memoryQwen3.5-9B9.7B66K137 tok/sFits in memoryMinistral 3 8B 25128.9B66K127 tok/sFits in memoryMinistral 3 3B 25123.8B131K200 tok/sFits in memorygpt-oss-safeguard-20b21.5B131K314 tok/sFits in memoryGranite 4.0 Micro3.2B66K416 tok/sFits in memoryQwen3 VL 8B Thinking8.8B66K129 tok/sFits in memoryUI-TARS 7B8.3B66K160 tok/sFits in memoryLlama 3.2 3B Instruct3.2B131K416 tok/sFits in memoryQwen3 VL 8B Instruct8.8B66K129 tok/sFits in memoryQwen3 8B8.2B66K162 tok/sFits in memoryGemma 3 4B4.3B131K264 tok/sFits in memoryLlama 3.1 8B Instruct8B131K142 tok/sFits in memoryGranite 4.2 8B8.8B66K129 tok/sFits in memoryNemotron 3.5 Content Safety4.3B131K264 tok/sFits in memoryLlama Guard 4 12B12B131K111 tok/sFits in memoryGemma 3n 4B7.8B33K171 tok/sFits in memoryGemma 3 12B12.2B131K109 tok/sFits in memoryRocinante 12B12.2B66K109 tok/sFits in memoryLlama 3.2 1B Instruct1.2B33K1109 tok/sFits in memoryMistral Nemo12.2B66K109 tok/sFits in memoryQwen2.5 7B Instruct7.6B33K175 tok/sFits in memoryGemma 4 26B A4B25.8B16K283 tok/sFits in memoryReka Edge7.1B16K160 tok/sFits in memoryNemotron 3 Nano 30B A3B31.6B16K377 tok/sFits in memoryestMinistral 3 14B 251213.9B66K96 tok/sFits in memoryUnslopNemo 12B12.2B33K109 tok/sFits in memoryQwen3 30B A3B Thinking 250730.5B16K377 tok/sFits in memoryestQwen3 Coder 30B A3B Instruct30.5B16K377 tok/sFits in memoryestQwen3 30B A3B Instruct 250730.5B16K377 tok/sFits in memoryestQwen3 14B14.8B66K90 tok/sFits in memoryQwen3 VL 30B A3B Thinking31.1B8K377 tok/sFits in memoryestLlama 3 8B Lunaris8B8K142 tok/sFits in memoryQwen3 VL 30B A3B Instruct31.1B8K377 tok/sFits in memoryestHy-MT2-1.8B2B8K565 tok/sFits in memoryHy-MT2-30B-A3B30.1B8K377 tok/sFits in memoryestHy-MT2-7B8B8K121 tok/sFits in memoryPhi 414.7B16K91 tok/sFits in memoryMythoMax 13B13B8K102 tok/sFits in memoryReMM SLERP 13B13B4K102 tok/sFits in memoryReka Flash 320.9B33K64 tok/sFits in memoryCydonia 24B V4.123.6B33K56 tok/sFits in memoryUncensored24B33K55 tok/sFits in memoryMistral Small 3.2 24B24B33K55 tok/sFits in memoryMistral Small 3.1 24B24B33K55 tok/sFits in memoryMistral Small 323.6B33K56 tok/sFits in memoryGemma 3 27B27.4B33K49 tok/sFits in memoryTernary Bonsai 2 27B27B33K49 tok/sFits in memoryVoxtral Small 24B 250724.3B16K55 tok/sFits in memoryMuse Glimmer 30B29.8B33K45 tok/sFits in memoryQwen3.5-27B27.8B8K48 tok/sFits in memoryQwen3.6 27B27.8B8K48 tok/sFits in memoryGemma 2 27B27.2B8K49 tok/sFits in memoryNemotron 3.5 Lightning31.6B16K36 tok/sFits in memoryestQwen3.8 27B27.8B8K48 tok/sFits in memoryGLM 4.7 Flash31.2B4K36 tok/sFits in memoryLaguna XS 2.133.4B4K34 tok/sFits in memoryOlmo 3 32B Think32.2B2K41 tok/sFits in memoryest
Speech models it runs67 modelsTranscription and speech models that fit in memory. Throughput figures for a transcription model come from its text decoder, so treat this as a fit answer rather than a speed one.
Niagara 19m Batch.en20M37948 tok/sFits in memoryNiagara 38m Batch.en38M19973 tok/sFits in memoryHiggs Audio v3 8b STT v28.9B150 tok/sFits in memoryHiggs Audio v3 STT2.7B281 tok/sFits in memoryCohere Transcribe 03 20262.1B361 tok/sFits in memoryDistil Large v3.50.8B949 tok/sFits in memoryLite Whisper Large v3 Acc1.4B810 tok/sFits in memoryOwsm CTC v3.1 1B1.1B1209 tok/sFits in memoryOwsm CTC v3.2 ft 1B1B759 tok/sFits in memoryOwsm CTC v4 1B1B1330 tok/sFits in memoryHubert Large Ls960 ft0.3B2372 tok/sFits in memoryMMS 1b All1B1134 tok/sFits in memoryWav2vec2 Large 960h Lv60 Self0.3B2372 tok/sFits in memoryHojo ASR V15.2B219 tok/sFits in memoryGranite 4.0 1b Speech2.3B330 tok/sFits in memoryGranite Speech 3.3 2b3B443 tok/sFits in memoryGranite Speech 3.3 8b8.6B155 tok/sFits in memoryGranite Speech 4.1 2b2.3B578 tok/sFits in memoryGranite Speech 4.1 2b NAR2.3B578 tok/sFits in memorySTT 2.6b en2.6B512 tok/sFits in memoryPhi 4 Multimodal Instruct5.6B66K136 tok/sFits in memoryVibeVoice ASR HF8.3B137 tok/sFits in memoryVoxtral Mini 3B 25074.7B162 tok/sFits in memoryVoxtral Mini 4B Realtime 26024.4B173 tok/sFits in memoryCanary 180m Flash0.2B6299 tok/sFits in memoryCanary 1b1B759 tok/sFits in memoryCanary 1b Flash0.8B949 tok/sFits in memoryCanary 1b v21B759 tok/sFits in memoryCanary Qwen 2.5b2.6B292 tok/sFits in memoryNemotron 3.5 ASR Streaming 0.6b0.6B1265 tok/sFits in memoryNemotron Speech Streaming en 0.6b0.6B1265 tok/sFits in memoryParakeet CTC 0.6b0.6B2217 tok/sFits in memoryParakeet CTC 1.1b1.1B690 tok/sFits in memoryParakeet RNNT 0.6b0.6B1265 tok/sFits in memoryParakeet RNNT 1.1b1.1B1209 tok/sFits in memoryParakeet TDT CTC 110m0.1B6900 tok/sFits in memoryParakeet TDT 0.6b v20.6B1890 tok/sFits in memoryParakeet TDT 0.6b v30.6B1265 tok/sFits in memoryParakeet TDT 1.1b1.1B690 tok/sFits in memoryNyra Health CrisperWhisper1.6B474 tok/sFits in memoryWhisper Large v31.5B756 tok/sFits in memoryWhisper Large v3 Turbo0.8B1663 tok/sFits in memoryMOSS Transcribe Diarize0.9B843 tok/sFits in memoryQwen3 ASR 0.6B0.9B1478 tok/sFits in memoryQwen3 ASR 0.6B HF0.8B1663 tok/sFits in memoryQwen3 ASR 1.7B2.3B493 tok/sFits in memoryQwen3 ASR 1.7B HF2B380 tok/sFits in memoryZipformer cr CTC Transducer XL 290M0.3B3910 tok/sFits in memoryZipformer Transducer XL 290M0.3B2617 tok/sFits in memoryASR Conformer Largescaleasr0.5B2771 tok/sFits in memoryGLM ASR Nano 25122.3B493 tok/sFits in memoryGranite Speech 5.0 470m Turboctc0.5B2268 tok/sFits in memoryGranite Speech 5.0 470m Turboctc nc0.5B2660 tok/sFits in memoryFun ASR Nano 2512 HF0.8B1663 tok/sFits in memoryGemma 4 E2B it5.1B149 tok/sFits in memoryGemma 4 E4B it8B131K166 tok/sFits in memoryNiagara 9m Batch.en9M84329 tok/sFits in memoryNiagara 84m Batch.en84M9035 tok/sFits in memoryARK ASR 0.6B1.3B33K1023 tok/sFits in memoryARK ASR 3B4.1B33K324 tok/sFits in memoryAudio8 ASR 0.1B0.3B33K2530 tok/sFits in memoryMOSS Transcribe Preview 2B2.4B33K554 tok/sFits in memoryGemma 4 12B it12B111 tok/sFits in memoryMoonshine Streaming Medium0.3B4K4434 tok/sFits in memoryMoonshine Streaming Small0.1B4K7590 tok/sFits in memoryMoonshine Streaming Tiny0B4Kno speed estimateFits in memoryMoonshine Tiny0B194no speed estimateFits in memory
Runs only by spilling into system memory9 modelsThese load, but part of the weights sits in ordinary system memory, which is far slower than the chip.
Too large for this device90 modelsThe weights do not fit even with part of them offloaded to system memory. Closest calls first.
Llama 3.1 70B Instruct70.6B—no speed estimateToo largeLlama 3.3 Euryale 70B70.6B—no speed estimateToo largeLlama 3.3 70B Instruct70.6B—no speed estimateToo largeQwen2.5 72B Instruct72.7B—no speed estimateToo largeQwen3 Coder Next79.7B—no speed estimateToo largeQwen3 Next 80B A3B Instruct81.3B—no speed estimateToo largeGLM 4.6V108B—no speed estimateToo largeLing-2.6-flash107B—no speed estimateToo largeCommand A111B—no speed estimateToo largegpt-oss-120b120B—no speed estimateToo largeNemotron 3 Super124B—no speed estimateToo largeHermes 4 70B70.6B—no speed estimateToo largeLlama 3.1 Euryale 70B v2.270.6B—no speed estimateToo largeR1 Distill Llama 70B70.6B—no speed estimateToo largeHermes 3 70B Instruct70.6B—no speed estimateToo largeDevstral 2 2512125B—no speed estimateToo largeGLM 4.5V108B—no speed estimateToo largeLlama 4 Scout109B—no speed estimateToo largeMagnum v4 72B72.7B—no speed estimateToo largeQwen2.5 VL 72B Instruct73.4B—no speed estimateToo largeLaguna S 2.1118B—no speed estimateToo largeHunyuan A13B Instruct80.4B—no speed estimateToo largeQwen3 Next 80B A3B Thinking81.3B—no speed estimateToo largeMixtral 8x22B Instruct141B—no speed estimateToo largeGLM 4.5 Air111B—no speed estimateToo largeStep 3.7 Flash201B—no speed estimateToo largeMistral Small 4119B—no speed estimateToo largeQwen3.5-122B-A10B125B—no speed estimateToo largeMistral Medium 3.5128B—no speed estimateToo largeLing-3.0-flash128B—no speed estimateToo largeMiniMax M2229B—no speed estimateToo largeStep 3.5 Flash199B—no speed estimateToo largeQwen3 VL 235B A22B Thinking236B—no speed estimateToo largeMiniMax M2.5229B—no speed estimateToo largeQwen3 235B A22B Thinking 2507235B—no speed estimateToo largeQwen3 VL 235B A22B Instruct236B—no speed estimateToo largeDeepSeek V4 Flash291B—no speed estimateToo largeGLM 5.3 Flash321B—no speed estimateToo largeHy3299B—no speed estimateToo largeGLM 4.5358B—no speed estimateToo largeLaguna M.1226B—no speed estimateToo largeNex-N2-Pro397B—no speed estimateToo largeMiniMax M2.7229B—no speed estimateToo largeMiniMax M2.1229B—no speed estimateToo largeQwen3 235B A22B Instruct 2507235B—no speed estimateToo largeGLM 4.6357B—no speed estimateToo largeGLM 4.7358B—no speed estimateToo largeERNIE 4.5 VL 424B A47B424B—no speed estimateToo largeMiniMax M3427B—no speed estimateToo largeNex-N2.5-Pro397B—no speed estimateToo largeInkling Small266B—no speed estimateToo largeTrinity Large Thinking399B—no speed estimateToo largeJamba Large 1.7399B—no speed estimateToo largeHy3 preview299B—no speed estimateToo largeDeepSeek V4 Flash Vision Exp305B—no speed estimateToo largeMiniMax M1456B—no speed estimateToo largeMiniMax-01456B—no speed estimateToo largeMiMo-V2.5311B—no speed estimateToo largeQwen3 Coder 480B A35B480B—no speed estimateToo largeDeepSeek V3.1 Terminus685B—no speed estimateToo largeDeepSeek V3.2685B—no speed estimateToo largeDeepSeek V3.2 Exp685B—no speed estimateToo largeLlama 4 Maverick402B—no speed estimateToo largeQwen3.5 397B A17B403B—no speed estimateToo largeHermes 4 405B406B—no speed estimateToo largeHermes 3 405B Instruct406B—no speed estimateToo largeGLM 5.1754B—no speed estimateToo largeDeepSeek V3685B—no speed estimateToo largeGLM 5754B—no speed estimateToo largeDeepSeek V4.1 Flash763B—no speed estimateToo largeGLM 5.2753B—no speed estimateToo largeNemotron 3 Ultra561B—no speed estimateToo largeKimi K2 07111T—no speed estimateToo largeLing-2.6-1T1T—no speed estimateToo largeMiMo-V2.5-Pro1T—no speed estimateToo largeR1 0528685B—no speed estimateToo largeDeepSeek V3 0324685B—no speed estimateToo largeDeepSeek V3.1685B—no speed estimateToo largeRing-2.6-1T1T—no speed estimateToo largeKimi K2 Thinking1.1T—no speed estimateToo largeKimi K2.7 Code1.1T—no speed estimateToo largeKimi K2.61.1T—no speed estimateToo largeKimi K2.51.1T—no speed estimateToo largeHy4 preview780B—no speed estimateToo largeInkling952B—no speed estimateToo largeKimi K2 09051T—no speed estimateToo largeLongCat 2.01.8T—no speed estimateToo largeQwen3.8 2.4T A95B2.4T—no speed estimateToo largeDeepSeek V4 Pro1.6T—no speed estimateToo largeKimi K32.8T—no speed estimateToo large

224 models have a verdict on this device: 125 fit in memory, 9 spill into system memory and 90 are too large. Every one of them is on this page.

03

Other Ada Lovelace cards

Something wrong on this page? Tell us