Models /Gemma 4 31B /Where to run
Provider guide

Where to run Gemma 4 31B

22 live offers tracked — output prices vary 3.4× between the cheapest and most expensive host, so the provider choice matters as much as the model choice.

CHEAPEST
$0.10 / $0.34
per 1M tokens in / out
FASTEST MEASURED
121 tok/s
measured throughput · $0.38 input / $1.15 output per million tokens
OpenRouterCHEAPEST$0.10 in/1M$0.34 out/1M262K13m agoDeepInfra$0.090 in/1M$0.34 out/1M262Kfp46d agoChutes$0.12 in/1M$0.37 out/1M14 tok/s131Kfp414h agoDeepInfra$0.13 in/1M$0.38 out/1M262Kfp86h agoNovita AI$0.14 in/1M$0.40 out/1M262K6h agoFriendli$0.14 in/1M$0.40 out/1M111 tok/s262K12m agoSambaNova$0.38 in/1M$1.15 out/1M131K6h agoCoreWeavezero-retention$0.10 in/1M$0.34 out/1M39 tok/s262Kbf166h agoDeepInfraturbo tier$0.090 in/1M$0.34 out/1M50 tok/s262Kfp412m agoOpenInferencezero-retention$0.10 in/1M$0.35 out/1M54 tok/s262Kbf1612m agoVenice AIzero-retention$0.12 in/1M$0.36 out/1M38 tok/s256Kbf166h agoDeepInfrazero-retention$0.13 in/1M$0.38 out/1M49 tok/s262Kfp812m agoNovita AIzero-retention$0.14 in/1M$0.40 out/1M16 tok/s262Kbf1612m agoMorphzero-retention$0.14 in/1M$0.40 out/1M16 tok/s175Kfp46h agoSiliconFlowzero-retention$0.13 in/1M$0.40 out/1M26 tok/s262Kfp86h agoCrusoezero-retention$0.14 in/1M$0.40 out/1M34 tok/s262K6h agoParasailzero-retention$0.15 in/1M$0.40 out/1M30 tok/s262Kfp812m agoPhalazero-retention$0.15 in/1M$0.46 out/1M21 tok/s262K12m agoModelRunzero-retention$0.22 in/1M$0.55 out/1M61 tok/s262Kfp412m agoTogether AIzero-retention$0.28 in/1M$0.86 out/1M18 tok/s262K8h agoSambaNovazero-retention$0.38 in/1M$1.15 out/1M121 tok/s131K12m agoCerebraszero-retention$0.99 in/1M$1.49 out/1M29 tok/s131Kfp1612m ago
Full specs, hardware verdicts and benchmarks on theGemma 4 31B model page →