Hardware / NVIDIA /GeForce RTX 4060 8GB
NVIDIA · GPU
GeForce RTX 4060 8GB
Runs 86 models fully in memory, 9 more with CPU offload. 7 of those sit close enough to 8 GB that our size estimate is what decides them.
Memory
8 GB
Bandwidth
272 GB/s
Runs
86
01
Specifications
Memory
8 GB
Decides which models fit at all.
Memory bandwidth
272 GB/s
How fast it can read the model. Sets the speed of the reply.
Compute
15.1 TFLOPS
Trillions of half-precision sums a second. Sets how fast it reads a long prompt, not how fast it replies.
Power draw
115 W
At full load, for sizing a power supply.
MSRP
$299
What the maker asked for it new. We hold no used or street prices, so this is not what one costs today.
Released
Jun 29, 2023
Architecture
Ada Lovelace
02
What is quantisation? →What it runs
Llama 3.2 3B Instruct3.2B66K112 tok/sFits in memoryLlama 3.2 1B Instruct1.2B33K299 tok/sFits in memoryGranite 4.0 Micro3.2B33K112 tok/sFits in memoryMinistral 3 3B 25123.8B33K95 tok/sFits in memoryGemma 3 4B4.3B66K84 tok/sFits in memoryHy-MT2-1.8B2B8K153 tok/sFits in memoryNemotron 3.5 Content Safety4.3B16K84 tok/sFits in memoryGemma 3n 4B7.8B16K46 tok/sFits in memoryQwen2.5 7B Instruct7.6B16K47 tok/sFits in memoryLlama 3.1 8B Instruct8B16K45 tok/sFits in memoryReka Edge7.1B8K51 tok/sFits in memoryUI-TARS 7B8.3B8K43 tok/sFits in memoryLlama 3 8B Lunaris8B4K45 tok/sFits in memoryQwen3 8B8.2B4K44 tok/sFits in memoryHy-MT2-7B8B4K38 tok/sFits in memoryGranite 4.1 8B8.8B2K41 tok/sFits in memoryestMinistral 3 8B 25128.9B2K40 tok/sFits in memoryestQwen3 VL 8B Thinking8.8B2K41 tok/sFits in memoryestQwen3 VL 8B Instruct8.8B2K41 tok/sFits in memoryestGranite 4.2 8B8.8B2K41 tok/sFits in memoryest
Speech models it runs66 modelsTranscription and speech models that fit in memory. Throughput figures for a transcription model come from its text decoder, so treat this as a fit answer rather than a speed one.
Niagara 19m Batch.en20M17946 tok/sFits in memoryNiagara 38m Batch.en38M8051 tok/sFits in memoryCohere Transcribe 03 20262.1B146 tok/sFits in memoryDistil Large v3.50.8B449 tok/sFits in memoryLite Whisper Large v3 Acc1.4B256 tok/sFits in memoryOwsm CTC v3.1 1B1.1B278 tok/sFits in memoryOwsm CTC v3.2 ft 1B1B359 tok/sFits in memoryOwsm CTC v4 1B1B306 tok/sFits in memoryHubert Large Ls960 ft0.3B956 tok/sFits in memoryMMS 1b All1B205 tok/sFits in memoryWav2vec2 Large 960h Lv60 Self0.3B640 tok/sFits in memorySTT 2.6b en2.6B138 tok/sFits in memoryCanary 180m Flash0.2B1994 tok/sFits in memoryCanary 1b1B306 tok/sFits in memoryCanary 1b Flash0.8B256 tok/sFits in memoryCanary 1b v21B205 tok/sFits in memoryCanary Qwen 2.5b2.6B138 tok/sFits in memoryNemotron 3.5 ASR Streaming 0.6b0.6B341 tok/sFits in memoryNemotron Speech Streaming en 0.6b0.6B341 tok/sFits in memoryParakeet CTC 0.6b0.6B598 tok/sFits in memoryParakeet CTC 1.1b1.1B186 tok/sFits in memoryParakeet RNNT 0.6b0.6B341 tok/sFits in memoryParakeet RNNT 1.1b1.1B326 tok/sFits in memoryParakeet TDT CTC 110m0.1B1862 tok/sFits in memoryParakeet TDT 0.6b v20.6B510 tok/sFits in memoryParakeet TDT 0.6b v30.6B341 tok/sFits in memoryParakeet TDT 1.1b1.1B186 tok/sFits in memoryNyra Health CrisperWhisper1.6B128 tok/sFits in memoryWhisper Large v31.5B204 tok/sFits in memoryWhisper Large v3 Turbo0.8B449 tok/sFits in memoryQwen3 ASR 0.6B0.9B399 tok/sFits in memoryQwen3 ASR 1.7B2.3B156 tok/sFits in memoryZipformer cr CTC Transducer XL 290M0.3B1055 tok/sFits in memoryZipformer Transducer XL 290M0.3B706 tok/sFits in memoryASR Conformer Largescaleasr0.5B748 tok/sFits in memoryGLM ASR Nano 25122.3B156 tok/sFits in memoryGranite Speech 5.0 470m Turboctc0.5B612 tok/sFits in memoryGranite Speech 5.0 470m Turboctc nc0.5B410 tok/sFits in memoryNiagara 9m Batch.en9M22756 tok/sFits in memoryNiagara 84m Batch.en84M4273 tok/sFits in memoryARK ASR 0.6B1.3B33K276 tok/sFits in memoryAudio8 ASR 0.1B0.3B33K1196 tok/sFits in memoryHiggs Audio v3 STT2.7B133 tok/sFits in memoryGranite 4.0 1b Speech2.3B156 tok/sFits in memoryGranite Speech 3.3 2b3B120 tok/sFits in memoryGranite Speech 4.1 2b2.3B156 tok/sFits in memoryGranite Speech 4.1 2b NAR2.3B156 tok/sFits in memoryMOSS Transcribe Diarize0.9B228 tok/sFits in memoryQwen3 ASR 0.6B HF0.8B449 tok/sFits in memoryQwen3 ASR 1.7B HF2B180 tok/sFits in memoryFun ASR Nano 2512 HF0.8B449 tok/sFits in memoryMOSS Transcribe Preview 2B2.4B8K150 tok/sFits in memoryARK ASR 3B4.1B33K88 tok/sFits in memoryMoonshine Streaming Medium0.3B4K1196 tok/sFits in memoryMoonshine Streaming Small0.1B4K2048 tok/sFits in memoryGemma 4 E2B it5.1B70 tok/sFits in memoryHojo ASR V15.2B69 tok/sFits in memoryVoxtral Mini 4B Realtime 26024.4B82 tok/sFits in memoryVoxtral Mini 3B 25074.7B76 tok/sFits in memoryPhi 4 Multimodal Instruct5.6B16K64 tok/sFits in memoryVibeVoice ASR HF8.3B43 tok/sFits in memoryGemma 4 E4B it8B8K45 tok/sFits in memoryHiggs Audio v3 8b STT v28.9B40 tok/sFits in memoryestGranite Speech 3.3 8b8.6B42 tok/sFits in memoryestMoonshine Streaming Tiny0B4Kno speed estimateFits in memoryMoonshine Tiny0B194no speed estimateFits in memory
Runs only by spilling into system memory9 modelsThese load, but part of the weights sits in ordinary system memory, which is far slower than the chip.
Qwen3.5-9B9.7B—no speed estimateSpills to system RAMestMinistral 3 14B 251213.9B—no speed estimateSpills to system RAMestLlama Guard 4 12B12B—no speed estimateSpills to system RAMestGemma 3 12B12.2B—no speed estimateSpills to system RAMPhi 414.7B—no speed estimateSpills to system RAMUnslopNemo 12B12.2B—no speed estimateSpills to system RAMestRocinante 12B12.2B—no speed estimateSpills to system RAMestMistral Nemo12.2B—no speed estimateSpills to system RAMestGemma 4 12B it12B—no speed estimateSpills to system RAM
Too large for this device129 modelsThe weights do not fit even with part of them offloaded to system memory. Closest calls first.
ReMM SLERP 13B13B—no speed estimateToo largeQwen3 14B14.8B—no speed estimateToo largegpt-oss-safeguard-20b21.5B—no speed estimateToo largeCydonia 24B V4.123.6B—no speed estimateToo largeUncensored24B—no speed estimateToo largeReka Flash 320.9B—no speed estimateToo largeMythoMax 13B13B—no speed estimateToo largeGemma 3 27B27.4B—no speed estimateToo largeMistral Small 323.6B—no speed estimateToo largeMistral Small 3.1 24B24B—no speed estimateToo largeVoxtral Small 24B 250724.3B—no speed estimateToo largeGemma 4 26B A4B25.8B—no speed estimateToo largeGLM 4.7 Flash31.2B—no speed estimateToo largeTernary Bonsai 2 27B27B—no speed estimateToo largeGemma 2 27B27.2B—no speed estimateToo largeOlmo 3 32B Think32.2B—no speed estimateToo largeQwen3.8 27B27.8B—no speed estimateToo largeQwen3 32B32.8B—no speed estimateToo largeGemma 4 31B31.3B—no speed estimateToo largeMuse Glimmer 30B29.8B—no speed estimateToo largeHy-MT2-30B-A3B30.1B—no speed estimateToo largeQwen3 Coder 30B A3B Instruct30.5B—no speed estimateToo largeQwen3.5-35B-A3B36B—no speed estimateToo largeQwen3 VL 30B A3B Instruct31.1B—no speed estimateToo largeNemotron 3.5 Lightning31.6B—no speed estimateToo largeNex-N2.5-Mini35.1B—no speed estimateToo largeQwen3.6 35B A3B36B—no speed estimateToo largeMistral Small 3.2 24B24B—no speed estimateToo largeQwen3.5-27B27.8B—no speed estimateToo largeQwen3.6 27B27.8B—no speed estimateToo largeQwen3 30B A3B Thinking 250730.5B—no speed estimateToo largeQwen3 30B A3B Instruct 250730.5B—no speed estimateToo largeQwen3 VL 30B A3B Thinking31.1B—no speed estimateToo largeNemotron 3 Nano 30B A3B31.6B—no speed estimateToo largeLaguna XS 2.133.4B—no speed estimateToo largeQwen2.5 Coder 32B Instruct32.8B—no speed estimateToo largeQwen3 VL 32B Instruct33.4B—no speed estimateToo largeNex-N2-Mini35.1B—no speed estimateToo largeSkyfall 36B V236.9B—no speed estimateToo largeLlama 3.3 70B Instruct70.6B—no speed estimateToo largeLlama 3.3 Euryale 70B70.6B—no speed estimateToo largeR1 Distill Llama 70B70.6B—no speed estimateToo largeQwen3 Next 80B A3B Thinking81.3B—no speed estimateToo largeQwen2.5 VL 72B Instruct73.4B—no speed estimateToo largeHunyuan A13B Instruct80.4B—no speed estimateToo largeGLM 4.6V108B—no speed estimateToo largeCommand A111B—no speed estimateToo largegpt-oss-120b120B—no speed estimateToo largeMistral Small 4119B—no speed estimateToo largeLlama 3.1 70B Instruct70.6B—no speed estimateToo largeHermes 4 70B70.6B—no speed estimateToo largeLlama 3.1 Euryale 70B v2.270.6B—no speed estimateToo largeHermes 3 70B Instruct70.6B—no speed estimateToo largeGLM 4.5V108B—no speed estimateToo largeLing-2.6-flash107B—no speed estimateToo largeMagnum v4 72B72.7B—no speed estimateToo largeQwen2.5 72B Instruct72.7B—no speed estimateToo largeQwen3 Coder Next79.7B—no speed estimateToo largeLaguna S 2.1118B—no speed estimateToo largeMixtral 8x22B Instruct141B—no speed estimateToo largeQwen3 Next 80B A3B Instruct81.3B—no speed estimateToo largeNemotron 3 Super124B—no speed estimateToo largeQwen3.5-122B-A10B125B—no speed estimateToo largeMistral Medium 3.5128B—no speed estimateToo largeLlama 4 Scout109B—no speed estimateToo largeGLM 4.5 Air111B—no speed estimateToo largeStep 3.7 Flash201B—no speed estimateToo largeDevstral 2 2512125B—no speed estimateToo largeLing-3.0-flash128B—no speed estimateToo largeStep 3.5 Flash199B—no speed estimateToo largeQwen3 VL 235B A22B Thinking236B—no speed estimateToo largeQwen3 VL 235B A22B Instruct236B—no speed estimateToo largeInkling Small266B—no speed estimateToo largeMiniMax M2.7229B—no speed estimateToo largeMiniMax M2.5229B—no speed estimateToo largeMiniMax M2229B—no speed estimateToo largeQwen3 235B A22B Thinking 2507235B—no speed estimateToo largeQwen3 235B A22B Instruct 2507235B—no speed estimateToo largeDeepSeek V4 Flash291B—no speed estimateToo largeGLM 5.3 Flash321B—no speed estimateToo largeHy3299B—no speed estimateToo largeLaguna M.1226B—no speed estimateToo largeTrinity Large Thinking399B—no speed estimateToo largeMiniMax M2.1229B—no speed estimateToo largeGLM 4.6357B—no speed estimateToo largeGLM 4.7358B—no speed estimateToo largeGLM 4.5358B—no speed estimateToo largeMiniMax M3427B—no speed estimateToo largeJamba Large 1.7399B—no speed estimateToo largeLlama 4 Maverick402B—no speed estimateToo largeHermes 3 405B Instruct406B—no speed estimateToo largeERNIE 4.5 VL 424B A47B424B—no speed estimateToo largeHy3 preview299B—no speed estimateToo largeDeepSeek V4 Flash Vision Exp305B—no speed estimateToo largeMiniMax-01456B—no speed estimateToo largeMiMo-V2.5311B—no speed estimateToo largeQwen3 Coder 480B A35B480B—no speed estimateToo largeDeepSeek V3 0324685B—no speed estimateToo largeNex-N2-Pro397B—no speed estimateToo largeNex-N2.5-Pro397B—no speed estimateToo largeQwen3.5 397B A17B403B—no speed estimateToo largeHermes 4 405B406B—no speed estimateToo largeGLM 5.1754B—no speed estimateToo largeMiniMax M1456B—no speed estimateToo largeDeepSeek V3.1 Terminus685B—no speed estimateToo largeDeepSeek V3685B—no speed estimateToo largeDeepSeek V3.2685B—no speed estimateToo largeGLM 5754B—no speed estimateToo largeDeepSeek V4.1 Flash763B—no speed estimateToo largeInkling952B—no speed estimateToo largeNemotron 3 Ultra561B—no speed estimateToo largeMiMo-V2.5-Pro1T—no speed estimateToo largeKimi K2 07111T—no speed estimateToo largeLing-2.6-1T1T—no speed estimateToo largeKimi K2.7 Code1.1T—no speed estimateToo largeR1 0528685B—no speed estimateToo largeDeepSeek V3.1685B—no speed estimateToo largeDeepSeek V3.2 Exp685B—no speed estimateToo largeKimi K2 09051T—no speed estimateToo largeRing-2.6-1T1T—no speed estimateToo largeKimi K2.61.1T—no speed estimateToo largeKimi K2.51.1T—no speed estimateToo largeGLM 5.2753B—no speed estimateToo largeHy4 preview780B—no speed estimateToo largeDeepSeek V4 Pro1.6T—no speed estimateToo largeKimi K2 Thinking1.1T—no speed estimateToo largeQwen3.8 2.4T A95B2.4T—no speed estimateToo largeLongCat 2.01.8T—no speed estimateToo largeKimi K32.8T—no speed estimateToo large
224 models have a verdict on this device: 86 fit in memory, 9 spill into system memory and 129 are too large. Every one of them is on this page.
03