Provider guide
Where to run gpt-oss-120b
23 live offers tracked — output prices vary 3.5× between the cheapest and most expensive host, so the provider choice matters as much as the model choice.
FASTEST MEASURED
919 tok/s
measured throughput · $0.35 input / $0.75 output per million tokens
OpenRouterCHEAPEST$0.037 in/1M$0.17 out/1M—131K—15m agoDeepInfra$0.037 in/1M$0.17 out/1M—131Kbfloat166h agoNovita AI$0.050 in/1M$0.25 out/1M—131K—6h agoSambaNova$0.22 in/1M$0.59 out/1M—131K—6h agoDeepInfrazero-retention$0.037 in/1M$0.17 out/1M40 tok/s131Kbf1613m agoCoreWeavezero-retention$0.030 in/1M$0.17 out/1M22 tok/s131Kfp413m agoNovita AIzero-retention$0.050 in/1M$0.25 out/1M102 tok/s131Kfp413m agoGoogle Vertex AIzero-retention$0.090 in/1M$0.36 out/1M210 tok/s131K—13m agoSiliconFlowzero-retention$0.050 in/1M$0.45 out/1M18 tok/s131Kfp813m agoDigitalOcean Gradientzero-retention$0.070 in/1M$0.49 out/1M34 tok/s128K—13m agoMancer 2zero-retention$0.10 in/1M$0.50 out/1M31 tok/s131Kfp86h agoBasetenzero-retention$0.10 in/1M$0.50 out/1M157 tok/s128Kfp413m agoPhalazero-retention$0.15 in/1M$0.60 out/1M80 tok/s131K—13m agoDeepInfraturbo tier$0.15 in/1M$0.60 out/1M129 tok/s131Kbf1613m agoTogether AIzero-retention$0.15 in/1M$0.60 out/1M148 tok/s131K—13m agoAmazon Bedrockzero-retention$0.15 in/1M$0.60 out/1M190 tok/s131K—13m agoAmazon Bedrockzero-retention$0.15 in/1M$0.60 out/1M122 tok/s131K—13m agoGroqzero-retention$0.15 in/1M$0.60 out/1M395 tok/s131K—13m agoNebius AI Studiozero-retention$0.15 in/1M$0.60 out/1M285 tok/s131Kfp46h agoCerebraszero-retention$0.35 in/1M$0.75 out/1M919 tok/s131Kfp1613m agoParasailzero-retention$0.10 in/1M$0.75 out/1M142 tok/s131Kfp413m agoMarazero-retention$0.15 in/1M$0.75 out/1M137 tok/s131K—13m agoSambaNovazero-retention$0.14 in/1M$0.95 out/1M325 tok/s131K—13m ago
Full specs, hardware verdicts and benchmarks on thegpt-oss-120b model page →