Hardware / NVIDIA /GeForce RTX 4070 Ti SUPER
NVIDIA · GPU

GeForce RTX 4070 Ti SUPER

Runs 98 models fully in memory, 25 more with CPU offload.

Memory
16 GB
Bandwidth
672 GB/s
Runs
98
01

Specifications

Memory
16 GB

Decides which models fit at all.

Memory bandwidth
672 GB/s

How fast it can read the model. Sets the speed of the reply.

Compute
44.1 TFLOPS

Trillions of half-precision sums a second. Sets how fast it reads a long prompt, not how fast it replies.

Power draw
285 W

At full load, for sizing a power supply.

MSRP
$799

What the maker asked for it new. We hold no used or street prices, so this is not what one costs today.

Released
Jan 24, 2024
Architecture
Ada Lovelace
02

What it runs

What is quantisation? →
Ministral 3 3B 25123.8B66K233 tok/sFits in memoryGranite 4.0 Micro3.2B66K158 tok/sFits in memoryLlama 3.2 3B Instruct3.2B131K236 tok/sFits in memoryGemma 3 4B4.3B131K176 tok/sFits in memoryNemotron 3.5 Content Safety4.3B66K176 tok/sFits in memoryLlama 3.1 8B Instruct8B131K111 tok/sFits in memoryUI-TARS 7B8.3B66K107 tok/sFits in memoryLlama 3.2 1B Instruct1.2B33K739 tok/sFits in memoryQwen2.5 7B Instruct7.6B33K117 tok/sFits in memoryGemma 3n 4B7.8B33K114 tok/sFits in memoryReka Edge7.1B16K125 tok/sFits in memoryQwen3 8B8.2B33K108 tok/sFits in memoryGranite 4.1 8B8.8B33K101 tok/sFits in memoryMinistral 3 8B 25128.9B33K100 tok/sFits in memoryQwen3 VL 8B Thinking8.8B33K101 tok/sFits in memoryQwen3 VL 8B Instruct8.8B33K101 tok/sFits in memoryGranite 4.2 8B8.8B33K101 tok/sFits in memoryHy-MT2-1.8B2B8K215 tok/sFits in memoryQwen3.5-9B9.7B33K91 tok/sFits in memoryLlama 3 8B Lunaris8B8K111 tok/sFits in memoryLlama Guard 4 12B12B66K74 tok/sFits in memoryGemma 3 12B12.2B66K73 tok/sFits in memoryUnslopNemo 12B12.2B33K73 tok/sFits in memoryRocinante 12B12.2B33K73 tok/sFits in memoryMistral Nemo12.2B33K73 tok/sFits in memoryHy-MT2-7B8B8K94 tok/sFits in memoryMinistral 3 14B 251213.9B16K64 tok/sFits in memoryQwen3 14B14.8B16K60 tok/sFits in memoryPhi 414.7B16K60 tok/sFits in memoryReMM SLERP 13B13B4K68 tok/sFits in memoryMythoMax 13B13B4K68 tok/sFits in memory
Speech models it runs67 modelsTranscription and speech models that fit in memory. Throughput figures for a transcription model come from its text decoder, so treat this as a fit answer rather than a speed one.
Niagara 19m Batch.en20M25299 tok/sFits in memoryNiagara 38m Batch.en38M13315 tok/sFits in memoryHiggs Audio v3 STT2.7B187 tok/sFits in memoryCohere Transcribe 03 20262.1B241 tok/sFits in memoryDistil Large v3.50.8B633 tok/sFits in memoryLite Whisper Large v3 Acc1.4B540 tok/sFits in memoryOwsm CTC v3.1 1B1.1B806 tok/sFits in memoryOwsm CTC v3.2 ft 1B1B506 tok/sFits in memoryOwsm CTC v4 1B1B887 tok/sFits in memoryHubert Large Ls960 ft0.3B1581 tok/sFits in memoryMMS 1b All1B756 tok/sFits in memoryWav2vec2 Large 960h Lv60 Self0.3B1581 tok/sFits in memoryHojo ASR V15.2B146 tok/sFits in memoryGranite 4.0 1b Speech2.3B220 tok/sFits in memoryGranite Speech 3.3 2b3B296 tok/sFits in memoryGranite Speech 4.1 2b2.3B386 tok/sFits in memoryGranite Speech 4.1 2b NAR2.3B386 tok/sFits in memorySTT 2.6b en2.6B341 tok/sFits in memoryPhi 4 Multimodal Instruct5.6B66K158 tok/sFits in memoryVoxtral Mini 3B 25074.7B161 tok/sFits in memoryVoxtral Mini 4B Realtime 26024.4B172 tok/sFits in memoryCanary 180m Flash0.2B4199 tok/sFits in memoryCanary 1b1B506 tok/sFits in memoryCanary 1b Flash0.8B633 tok/sFits in memoryCanary 1b v21B506 tok/sFits in memoryCanary Qwen 2.5b2.6B195 tok/sFits in memoryNemotron 3.5 ASR Streaming 0.6b0.6B843 tok/sFits in memoryNemotron Speech Streaming en 0.6b0.6B843 tok/sFits in memoryParakeet CTC 0.6b0.6B1478 tok/sFits in memoryParakeet CTC 1.1b1.1B460 tok/sFits in memoryParakeet RNNT 0.6b0.6B843 tok/sFits in memoryParakeet RNNT 1.1b1.1B806 tok/sFits in memoryParakeet TDT CTC 110m0.1B4600 tok/sFits in memoryParakeet TDT 0.6b v20.6B1260 tok/sFits in memoryParakeet TDT 0.6b v30.6B843 tok/sFits in memoryParakeet TDT 1.1b1.1B460 tok/sFits in memoryNyra Health CrisperWhisper1.6B316 tok/sFits in memoryWhisper Large v31.5B504 tok/sFits in memoryWhisper Large v3 Turbo0.8B1109 tok/sFits in memoryMOSS Transcribe Diarize0.9B562 tok/sFits in memoryQwen3 ASR 0.6B0.9B985 tok/sFits in memoryQwen3 ASR 0.6B HF0.8B1109 tok/sFits in memoryQwen3 ASR 1.7B2.3B329 tok/sFits in memoryQwen3 ASR 1.7B HF2B253 tok/sFits in memoryZipformer cr CTC Transducer XL 290M0.3B2606 tok/sFits in memoryZipformer Transducer XL 290M0.3B1745 tok/sFits in memoryASR Conformer Largescaleasr0.5B1847 tok/sFits in memoryGLM ASR Nano 25122.3B329 tok/sFits in memoryGranite Speech 5.0 470m Turboctc0.5B1774 tok/sFits in memoryGranite Speech 5.0 470m Turboctc nc0.5B1512 tok/sFits in memoryFun ASR Nano 2512 HF0.8B1109 tok/sFits in memoryGemma 4 E2B it5.1B174 tok/sFits in memoryNiagara 9m Batch.en9M56220 tok/sFits in memoryNiagara 84m Batch.en84M10557 tok/sFits in memoryGemma 4 E4B it8B66K111 tok/sFits in memoryARK ASR 0.6B1.3B33K581 tok/sFits in memoryARK ASR 3B4.1B33K216 tok/sFits in memoryAudio8 ASR 0.1B0.3B33K1687 tok/sFits in memoryVibeVoice ASR HF8.3B107 tok/sFits in memoryMOSS Transcribe Preview 2B2.4B16K370 tok/sFits in memoryGranite Speech 3.3 8b8.6B103 tok/sFits in memoryHiggs Audio v3 8b STT v28.9B100 tok/sFits in memoryMoonshine Streaming Medium0.3B4K2956 tok/sFits in memoryMoonshine Streaming Small0.1B4K5060 tok/sFits in memoryGemma 4 12B it12B74 tok/sFits in memoryMoonshine Streaming Tiny0B4Kno speed estimateFits in memoryMoonshine Tiny0B194no speed estimateFits in memory
Runs only by spilling into system memory25 modelsThese load, but part of the weights sits in ordinary system memory, which is far slower than the chip.
Gemma 4 26B A4B25.8B—no speed estimateSpills to system RAMQwen3.5-27B27.8B—no speed estimateSpills to system RAMGLM 4.7 Flash31.2B—no speed estimateSpills to system RAMNemotron 3 Nano 30B A3B31.6B—no speed estimateSpills to system RAMestVoxtral Small 24B 250724.3B—no speed estimateSpills to system RAMgpt-oss-safeguard-20b21.5B—no speed estimateSpills to system RAMQwen3 VL 30B A3B Thinking31.1B—no speed estimateSpills to system RAMestCydonia 24B V4.123.6B—no speed estimateSpills to system RAMUncensored24B—no speed estimateSpills to system RAMMistral Small 3.2 24B24B—no speed estimateSpills to system RAMGemma 3 27B27.4B—no speed estimateSpills to system RAMQwen3.6 27B27.8B—no speed estimateSpills to system RAMQwen3 VL 30B A3B Instruct31.1B—no speed estimateSpills to system RAMestQwen3 30B A3B Thinking 250730.5B—no speed estimateSpills to system RAMestQwen3 Coder 30B A3B Instruct30.5B—no speed estimateSpills to system RAMestQwen3 30B A3B Instruct 250730.5B—no speed estimateSpills to system RAMestMistral Small 3.1 24B24B—no speed estimateSpills to system RAMReka Flash 320.9B—no speed estimateSpills to system RAMMistral Small 323.6B—no speed estimateSpills to system RAMGemma 2 27B27.2B—no speed estimateSpills to system RAMestMuse Glimmer 30B29.8B—no speed estimateSpills to system RAMestNemotron 3.5 Lightning31.6B—no speed estimateSpills to system RAMestQwen3.8 27B27.8B—no speed estimateSpills to system RAMHy-MT2-30B-A3B30.1B—no speed estimateSpills to system RAMestTernary Bonsai 2 27B27B—no speed estimateSpills to system RAMest
Too large for this device101 modelsThe weights do not fit even with part of them offloaded to system memory. Closest calls first.
Qwen3 32B32.8B—no speed estimateToo largeestQwen3.6 35B A3B36B—no speed estimateToo largeQwen3.5-35B-A3B36B—no speed estimateToo largeSkyfall 36B V236.9B—no speed estimateToo largeOlmo 3 32B Think32.2B—no speed estimateToo largeQwen3 VL 32B Instruct33.4B—no speed estimateToo largeNex-N2.5-Mini35.1B—no speed estimateToo largeLaguna XS 2.133.4B—no speed estimateToo largeGemma 4 31B31.3B—no speed estimateToo largeQwen2.5 Coder 32B Instruct32.8B—no speed estimateToo largeNex-N2-Mini35.1B—no speed estimateToo largeLlama 3.3 Euryale 70B70.6B—no speed estimateToo largeQwen2.5 VL 72B Instruct73.4B—no speed estimateToo largeQwen3 Next 80B A3B Thinking81.3B—no speed estimateToo largeQwen3 Coder Next79.7B—no speed estimateToo largeLlama 3.3 70B Instruct70.6B—no speed estimateToo largeQwen2.5 72B Instruct72.7B—no speed estimateToo largeQwen3 Next 80B A3B Instruct81.3B—no speed estimateToo largeCommand A111B—no speed estimateToo largegpt-oss-120b120B—no speed estimateToo largeLlama 3.1 70B Instruct70.6B—no speed estimateToo largeHermes 4 70B70.6B—no speed estimateToo largeLlama 3.1 Euryale 70B v2.270.6B—no speed estimateToo largeR1 Distill Llama 70B70.6B—no speed estimateToo largeHermes 3 70B Instruct70.6B—no speed estimateToo largeQwen3.5-122B-A10B125B—no speed estimateToo largeDevstral 2 2512125B—no speed estimateToo largeGLM 4.6V108B—no speed estimateToo largeLing-2.6-flash107B—no speed estimateToo largeMagnum v4 72B72.7B—no speed estimateToo largeMistral Medium 3.5128B—no speed estimateToo largeLaguna S 2.1118B—no speed estimateToo largeHunyuan A13B Instruct80.4B—no speed estimateToo largeMixtral 8x22B Instruct141B—no speed estimateToo largeGLM 4.5V108B—no speed estimateToo largeLlama 4 Scout109B—no speed estimateToo largeGLM 4.5 Air111B—no speed estimateToo largeStep 3.7 Flash201B—no speed estimateToo largeMistral Small 4119B—no speed estimateToo largeNemotron 3 Super124B—no speed estimateToo largeLing-3.0-flash128B—no speed estimateToo largeMiniMax M2.7229B—no speed estimateToo largeMiniMax M2229B—no speed estimateToo largeStep 3.5 Flash199B—no speed estimateToo largeQwen3 VL 235B A22B Instruct236B—no speed estimateToo largeMiniMax M2.5229B—no speed estimateToo largeQwen3 235B A22B Instruct 2507235B—no speed estimateToo largeQwen3 VL 235B A22B Thinking236B—no speed estimateToo largeDeepSeek V4 Flash291B—no speed estimateToo largeGLM 5.3 Flash321B—no speed estimateToo largeHy3299B—no speed estimateToo largeLaguna M.1226B—no speed estimateToo largeNex-N2-Pro397B—no speed estimateToo largeMiniMax M2.1229B—no speed estimateToo largeQwen3 235B A22B Thinking 2507235B—no speed estimateToo largeGLM 4.6357B—no speed estimateToo largeGLM 4.7358B—no speed estimateToo largeGLM 4.5358B—no speed estimateToo largeMiniMax M3427B—no speed estimateToo largeInkling Small266B—no speed estimateToo largeTrinity Large Thinking399B—no speed estimateToo largeJamba Large 1.7399B—no speed estimateToo largeLlama 4 Maverick402B—no speed estimateToo largeHermes 3 405B Instruct406B—no speed estimateToo largeERNIE 4.5 VL 424B A47B424B—no speed estimateToo largeHy3 preview299B—no speed estimateToo largeDeepSeek V4 Flash Vision Exp305B—no speed estimateToo largeMiniMax-01456B—no speed estimateToo largeMiMo-V2.5311B—no speed estimateToo largeQwen3 Coder 480B A35B480B—no speed estimateToo largeDeepSeek V3 0324685B—no speed estimateToo largeNex-N2.5-Pro397B—no speed estimateToo largeQwen3.5 397B A17B403B—no speed estimateToo largeHermes 4 405B406B—no speed estimateToo largeGLM 5.1754B—no speed estimateToo largeMiniMax M1456B—no speed estimateToo largeDeepSeek V3.1 Terminus685B—no speed estimateToo largeDeepSeek V3685B—no speed estimateToo largeGLM 5754B—no speed estimateToo largeInkling952B—no speed estimateToo largeNemotron 3 Ultra561B—no speed estimateToo largeMiMo-V2.5-Pro1T—no speed estimateToo largeKimi K2 07111T—no speed estimateToo largeLing-2.6-1T1T—no speed estimateToo largeKimi K2.7 Code1.1T—no speed estimateToo largeR1 0528685B—no speed estimateToo largeDeepSeek V3.1685B—no speed estimateToo largeDeepSeek V3.2685B—no speed estimateToo largeDeepSeek V3.2 Exp685B—no speed estimateToo largeKimi K2 09051T—no speed estimateToo largeRing-2.6-1T1T—no speed estimateToo largeKimi K2.61.1T—no speed estimateToo largeKimi K2.51.1T—no speed estimateToo largeGLM 5.2753B—no speed estimateToo largeDeepSeek V4.1 Flash763B—no speed estimateToo largeHy4 preview780B—no speed estimateToo largeKimi K2 Thinking1.1T—no speed estimateToo largeQwen3.8 2.4T A95B2.4T—no speed estimateToo largeDeepSeek V4 Pro1.6T—no speed estimateToo largeLongCat 2.01.8T—no speed estimateToo largeKimi K32.8T—no speed estimateToo large

224 models have a verdict on this device: 98 fit in memory, 25 spill into system memory and 101 are too large. Every one of them is on this page.

03

Other Ada Lovelace cards

Something wrong on this page? Tell us