Provider guide
Where to run Gemma 4 31B
22 live offers tracked — output prices vary 3.4× between the cheapest and most expensive host, so the provider choice matters as much as the model choice.
FASTEST MEASURED
121 tok/s
measured throughput · $0.38 input / $1.15 output per million tokens
OpenRouterCHEAPEST$0.10 in/1M$0.34 out/1M—262K—13m agoDeepInfra$0.090 in/1M$0.34 out/1M—262Kfp46d agoChutes$0.12 in/1M$0.37 out/1M14 tok/s131Kfp414h agoDeepInfra$0.13 in/1M$0.38 out/1M—262Kfp86h agoNovita AI$0.14 in/1M$0.40 out/1M—262K—6h agoFriendli$0.14 in/1M$0.40 out/1M111 tok/s262K—12m agoSambaNova$0.38 in/1M$1.15 out/1M—131K—6h agoCoreWeavezero-retention$0.10 in/1M$0.34 out/1M39 tok/s262Kbf166h agoDeepInfraturbo tier$0.090 in/1M$0.34 out/1M50 tok/s262Kfp412m agoOpenInferencezero-retention$0.10 in/1M$0.35 out/1M54 tok/s262Kbf1612m agoVenice AIzero-retention$0.12 in/1M$0.36 out/1M38 tok/s256Kbf166h agoDeepInfrazero-retention$0.13 in/1M$0.38 out/1M49 tok/s262Kfp812m agoNovita AIzero-retention$0.14 in/1M$0.40 out/1M16 tok/s262Kbf1612m agoMorphzero-retention$0.14 in/1M$0.40 out/1M16 tok/s175Kfp46h agoSiliconFlowzero-retention$0.13 in/1M$0.40 out/1M26 tok/s262Kfp86h agoCrusoezero-retention$0.14 in/1M$0.40 out/1M34 tok/s262K—6h agoParasailzero-retention$0.15 in/1M$0.40 out/1M30 tok/s262Kfp812m agoPhalazero-retention$0.15 in/1M$0.46 out/1M21 tok/s262K—12m agoModelRunzero-retention$0.22 in/1M$0.55 out/1M61 tok/s262Kfp412m agoTogether AIzero-retention$0.28 in/1M$0.86 out/1M18 tok/s262K—8h agoSambaNovazero-retention$0.38 in/1M$1.15 out/1M121 tok/s131K—12m agoCerebraszero-retention$0.99 in/1M$1.49 out/1M29 tok/s131Kfp1612m ago
Full specs, hardware verdicts and benchmarks on theGemma 4 31B model page →