Hardware / AMD /Radeon RX 9070
AMD · GPU
Radeon RX 9070
Runs 98 models fully in memory, 25 more with CPU offload.
Memory
16 GB
Bandwidth
640 GB/s
Runs
98
01
Specifications
Memory
16 GB
Decides which models fit at all.
Memory bandwidth
640 GB/s
How fast it can read the model. Sets the speed of the reply.
Compute
72.3 TFLOPS
Trillions of half-precision sums a second. Sets how fast it reads a long prompt, not how fast it replies.
Power draw
220 W
At full load, for sizing a power supply.
MSRP
$549
What the maker asked for it new. We hold no used or street prices, so this is not what one costs today.
Released
Mar 6, 2025
Architecture
RDNA 4
02
What is quantisation? →What it runs
Ministral 3 3B 25123.8B66K181 tok/sFits in memoryGranite 4.0 Micro3.2B66K214 tok/sFits in memoryLlama 3.2 3B Instruct3.2B131K183 tok/sFits in memoryGemma 3 4B4.3B131K136 tok/sFits in memoryNemotron 3.5 Content Safety4.3B66K136 tok/sFits in memoryLlama 3.2 1B Instruct1.2B33K572 tok/sFits in memoryUI-TARS 7B8.3B66K83 tok/sFits in memoryLlama 3.1 8B Instruct8B131K86 tok/sFits in memoryHy-MT2-1.8B2B8K166 tok/sFits in memoryQwen2.5 7B Instruct7.6B33K90 tok/sFits in memoryGemma 3n 4B7.8B33K88 tok/sFits in memoryReka Edge7.1B16K97 tok/sFits in memoryQwen3 8B8.2B33K84 tok/sFits in memoryGranite 4.1 8B8.8B33K78 tok/sFits in memoryQwen3 VL 8B Thinking8.8B33K78 tok/sFits in memoryQwen3 VL 8B Instruct8.8B33K78 tok/sFits in memoryGranite 4.2 8B8.8B33K78 tok/sFits in memoryMinistral 3 8B 25128.9B33K77 tok/sFits in memoryQwen3.5-9B9.7B33K71 tok/sFits in memoryLlama Guard 4 12B12B66K57 tok/sFits in memoryGemma 3 12B12.2B66K56 tok/sFits in memoryLlama 3 8B Lunaris8B8K86 tok/sFits in memoryUnslopNemo 12B12.2B33K56 tok/sFits in memoryRocinante 12B12.2B33K56 tok/sFits in memoryMistral Nemo12.2B33K56 tok/sFits in memoryHy-MT2-7B8B8K73 tok/sFits in memoryMinistral 3 14B 251213.9B16K49 tok/sFits in memoryQwen3 14B14.8B16K46 tok/sFits in memoryPhi 414.7B16K47 tok/sFits in memoryReMM SLERP 13B13B4K53 tok/sFits in memoryMythoMax 13B13B4K53 tok/sFits in memory
Speech models it runs67 modelsTranscription and speech models that fit in memory. Throughput figures for a transcription model come from its text decoder, so treat this as a fit answer rather than a speed one.
Niagara 19m Batch.en20M19577 tok/sFits in memoryNiagara 38m Batch.en38M10303 tok/sFits in memoryHiggs Audio v3 STT2.7B145 tok/sFits in memoryCohere Transcribe 03 20262.1B186 tok/sFits in memoryDistil Large v3.50.8B489 tok/sFits in memoryLite Whisper Large v3 Acc1.4B418 tok/sFits in memoryOwsm CTC v3.1 1B1.1B624 tok/sFits in memoryOwsm CTC v3.2 ft 1B1B392 tok/sFits in memoryOwsm CTC v4 1B1B686 tok/sFits in memoryHubert Large Ls960 ft0.3B1224 tok/sFits in memoryMMS 1b All1B585 tok/sFits in memoryWav2vec2 Large 960h Lv60 Self0.3B1224 tok/sFits in memoryHojo ASR V15.2B133 tok/sFits in memoryGranite 4.0 1b Speech2.3B170 tok/sFits in memoryGranite Speech 3.3 2b3B229 tok/sFits in memoryGranite Speech 4.1 2b2.3B298 tok/sFits in memoryGranite Speech 4.1 2b NAR2.3B298 tok/sFits in memorySTT 2.6b en2.6B264 tok/sFits in memoryPhi 4 Multimodal Instruct5.6B66K123 tok/sFits in memoryVoxtral Mini 3B 25074.7B124 tok/sFits in memoryVoxtral Mini 4B Realtime 26024.4B133 tok/sFits in memoryCanary 180m Flash0.2B3249 tok/sFits in memoryCanary 1b1B392 tok/sFits in memoryCanary 1b Flash0.8B489 tok/sFits in memoryCanary 1b v21B392 tok/sFits in memoryCanary Qwen 2.5b2.6B151 tok/sFits in memoryNemotron 3.5 ASR Streaming 0.6b0.6B653 tok/sFits in memoryNemotron Speech Streaming en 0.6b0.6B653 tok/sFits in memoryParakeet CTC 0.6b0.6B1144 tok/sFits in memoryParakeet CTC 1.1b1.1B356 tok/sFits in memoryParakeet RNNT 0.6b0.6B653 tok/sFits in memoryParakeet RNNT 1.1b1.1B624 tok/sFits in memoryParakeet TDT CTC 110m0.1B3559 tok/sFits in memoryParakeet TDT 0.6b v20.6B975 tok/sFits in memoryParakeet TDT 0.6b v30.6B653 tok/sFits in memoryParakeet TDT 1.1b1.1B356 tok/sFits in memoryNyra Health CrisperWhisper1.6B245 tok/sFits in memoryWhisper Large v31.5B390 tok/sFits in memoryWhisper Large v3 Turbo0.8B858 tok/sFits in memoryMOSS Transcribe Diarize0.9B435 tok/sFits in memoryQwen3 ASR 0.6B0.9B762 tok/sFits in memoryQwen3 ASR 0.6B HF0.8B858 tok/sFits in memoryQwen3 ASR 1.7B2.3B254 tok/sFits in memoryQwen3 ASR 1.7B HF2B196 tok/sFits in memoryZipformer cr CTC Transducer XL 290M0.3B2017 tok/sFits in memoryZipformer Transducer XL 290M0.3B1350 tok/sFits in memoryASR Conformer Largescaleasr0.5B1430 tok/sFits in memoryGLM ASR Nano 25122.3B254 tok/sFits in memoryGranite Speech 5.0 470m Turboctc0.5B1372 tok/sFits in memoryGranite Speech 5.0 470m Turboctc nc0.5B1170 tok/sFits in memoryFun ASR Nano 2512 HF0.8B858 tok/sFits in memoryGemma 4 E2B it5.1B135 tok/sFits in memoryNiagara 9m Batch.en9M43503 tok/sFits in memoryNiagara 84m Batch.en84M8169 tok/sFits in memoryARK ASR 0.6B1.3B33K528 tok/sFits in memoryARK ASR 3B4.1B33K167 tok/sFits in memoryAudio8 ASR 0.1B0.3B33K1305 tok/sFits in memoryMOSS Transcribe Preview 2B2.4B16K286 tok/sFits in memoryVibeVoice ASR HF8.3B83 tok/sFits in memoryGemma 4 E4B it8B66K86 tok/sFits in memoryMoonshine Streaming Medium0.3B4K2287 tok/sFits in memoryMoonshine Streaming Small0.1B4K3915 tok/sFits in memoryGranite Speech 3.3 8b8.6B80 tok/sFits in memoryHiggs Audio v3 8b STT v28.9B77 tok/sFits in memoryGemma 4 12B it12B57 tok/sFits in memoryMoonshine Streaming Tiny0B4Kno speed estimateFits in memoryMoonshine Tiny0B194no speed estimateFits in memory
Runs only by spilling into system memory25 modelsThese load, but part of the weights sits in ordinary system memory, which is far slower than the chip.
Gemma 4 26B A4B25.8B—no speed estimateSpills to system RAMQwen3.5-27B27.8B—no speed estimateSpills to system RAMGLM 4.7 Flash31.2B—no speed estimateSpills to system RAMNemotron 3 Nano 30B A3B31.6B—no speed estimateSpills to system RAMestVoxtral Small 24B 250724.3B—no speed estimateSpills to system RAMgpt-oss-safeguard-20b21.5B—no speed estimateSpills to system RAMestQwen3 VL 30B A3B Thinking31.1B—no speed estimateSpills to system RAMestCydonia 24B V4.123.6B—no speed estimateSpills to system RAMUncensored24B—no speed estimateSpills to system RAMMistral Small 3.2 24B24B—no speed estimateSpills to system RAMGemma 3 27B27.4B—no speed estimateSpills to system RAMQwen3.6 27B27.8B—no speed estimateSpills to system RAMQwen3 VL 30B A3B Instruct31.1B—no speed estimateSpills to system RAMestQwen3 30B A3B Thinking 250730.5B—no speed estimateSpills to system RAMestQwen3 Coder 30B A3B Instruct30.5B—no speed estimateSpills to system RAMestQwen3 30B A3B Instruct 250730.5B—no speed estimateSpills to system RAMestMistral Small 3.1 24B24B—no speed estimateSpills to system RAMReka Flash 320.9B—no speed estimateSpills to system RAMMistral Small 323.6B—no speed estimateSpills to system RAMGemma 2 27B27.2B—no speed estimateSpills to system RAMestMuse Glimmer 30B29.8B—no speed estimateSpills to system RAMestNemotron 3.5 Lightning31.6B—no speed estimateSpills to system RAMestQwen3.8 27B27.8B—no speed estimateSpills to system RAMHy-MT2-30B-A3B30.1B—no speed estimateSpills to system RAMestTernary Bonsai 2 27B27B—no speed estimateSpills to system RAMest
Too large for this device101 modelsThe weights do not fit even with part of them offloaded to system memory. Closest calls first.
Olmo 3 32B Think32.2B—no speed estimateToo largeestQwen3 32B32.8B—no speed estimateToo largeestQwen3.6 35B A3B36B—no speed estimateToo largeQwen3.5-35B-A3B36B—no speed estimateToo largeSkyfall 36B V236.9B—no speed estimateToo largeNex-N2.5-Mini35.1B—no speed estimateToo largeLaguna XS 2.133.4B—no speed estimateToo largeGemma 4 31B31.3B—no speed estimateToo largeQwen2.5 Coder 32B Instruct32.8B—no speed estimateToo largeQwen3 VL 32B Instruct33.4B—no speed estimateToo largeNex-N2-Mini35.1B—no speed estimateToo largeLlama 3.3 Euryale 70B70.6B—no speed estimateToo largeQwen2.5 VL 72B Instruct73.4B—no speed estimateToo largeQwen3 Next 80B A3B Thinking81.3B—no speed estimateToo largeQwen3 Coder Next79.7B—no speed estimateToo largeLlama 3.3 70B Instruct70.6B—no speed estimateToo largeQwen2.5 72B Instruct72.7B—no speed estimateToo largeQwen3 Next 80B A3B Instruct81.3B—no speed estimateToo largeCommand A111B—no speed estimateToo largegpt-oss-120b120B—no speed estimateToo largeLlama 3.1 70B Instruct70.6B—no speed estimateToo largeHermes 4 70B70.6B—no speed estimateToo largeLlama 3.1 Euryale 70B v2.270.6B—no speed estimateToo largeR1 Distill Llama 70B70.6B—no speed estimateToo largeHermes 3 70B Instruct70.6B—no speed estimateToo largeQwen3.5-122B-A10B125B—no speed estimateToo largeDevstral 2 2512125B—no speed estimateToo largeGLM 4.6V108B—no speed estimateToo largeLing-2.6-flash107B—no speed estimateToo largeMagnum v4 72B72.7B—no speed estimateToo largeMistral Medium 3.5128B—no speed estimateToo largeLaguna S 2.1118B—no speed estimateToo largeHunyuan A13B Instruct80.4B—no speed estimateToo largeMixtral 8x22B Instruct141B—no speed estimateToo largeGLM 4.5V108B—no speed estimateToo largeLlama 4 Scout109B—no speed estimateToo largeGLM 4.5 Air111B—no speed estimateToo largeStep 3.7 Flash201B—no speed estimateToo largeMistral Small 4119B—no speed estimateToo largeNemotron 3 Super124B—no speed estimateToo largeLing-3.0-flash128B—no speed estimateToo largeMiniMax M2.7229B—no speed estimateToo largeMiniMax M2229B—no speed estimateToo largeStep 3.5 Flash199B—no speed estimateToo largeQwen3 VL 235B A22B Thinking236B—no speed estimateToo largeQwen3 VL 235B A22B Instruct236B—no speed estimateToo largeMiniMax M2.5229B—no speed estimateToo largeQwen3 235B A22B Instruct 2507235B—no speed estimateToo largeDeepSeek V4 Flash291B—no speed estimateToo largeGLM 5.3 Flash321B—no speed estimateToo largeHy3299B—no speed estimateToo largeLaguna M.1226B—no speed estimateToo largeNex-N2-Pro397B—no speed estimateToo largeMiniMax M2.1229B—no speed estimateToo largeQwen3 235B A22B Thinking 2507235B—no speed estimateToo largeGLM 4.6357B—no speed estimateToo largeGLM 4.7358B—no speed estimateToo largeGLM 4.5358B—no speed estimateToo largeMiniMax M3427B—no speed estimateToo largeInkling Small266B—no speed estimateToo largeTrinity Large Thinking399B—no speed estimateToo largeJamba Large 1.7399B—no speed estimateToo largeLlama 4 Maverick402B—no speed estimateToo largeHermes 3 405B Instruct406B—no speed estimateToo largeERNIE 4.5 VL 424B A47B424B—no speed estimateToo largeHy3 preview299B—no speed estimateToo largeDeepSeek V4 Flash Vision Exp305B—no speed estimateToo largeMiniMax-01456B—no speed estimateToo largeMiMo-V2.5311B—no speed estimateToo largeQwen3 Coder 480B A35B480B—no speed estimateToo largeDeepSeek V3 0324685B—no speed estimateToo largeNex-N2.5-Pro397B—no speed estimateToo largeQwen3.5 397B A17B403B—no speed estimateToo largeHermes 4 405B406B—no speed estimateToo largeGLM 5.1754B—no speed estimateToo largeMiniMax M1456B—no speed estimateToo largeDeepSeek V3.1 Terminus685B—no speed estimateToo largeDeepSeek V3685B—no speed estimateToo largeDeepSeek V3.2685B—no speed estimateToo largeGLM 5754B—no speed estimateToo largeInkling952B—no speed estimateToo largeNemotron 3 Ultra561B—no speed estimateToo largeMiMo-V2.5-Pro1T—no speed estimateToo largeKimi K2 07111T—no speed estimateToo largeLing-2.6-1T1T—no speed estimateToo largeKimi K2.7 Code1.1T—no speed estimateToo largeR1 0528685B—no speed estimateToo largeDeepSeek V3.1685B—no speed estimateToo largeDeepSeek V3.2 Exp685B—no speed estimateToo largeKimi K2 09051T—no speed estimateToo largeRing-2.6-1T1T—no speed estimateToo largeKimi K2.61.1T—no speed estimateToo largeKimi K2.51.1T—no speed estimateToo largeGLM 5.2753B—no speed estimateToo largeDeepSeek V4.1 Flash763B—no speed estimateToo largeHy4 preview780B—no speed estimateToo largeKimi K2 Thinking1.1T—no speed estimateToo largeQwen3.8 2.4T A95B2.4T—no speed estimateToo largeDeepSeek V4 Pro1.6T—no speed estimateToo largeLongCat 2.01.8T—no speed estimateToo largeKimi K32.8T—no speed estimateToo large
224 models have a verdict on this device: 98 fit in memory, 25 spill into system memory and 101 are too large. Every one of them is on this page.
03