Models /Llama 3.3 70B Instruct /Where to run
Provider guide

Where to run Llama 3.3 70B Instruct

17 live offers tracked — output prices vary 7.0× between the cheapest and most expensive host, so the provider choice matters as much as the model choice.

CHEAPEST
$0.10 / $0.32
per 1M tokens in / out
FASTEST MEASURED
179 tok/s
measured throughput · $0.59 input / $0.79 output per million tokens
DeepInfra$0.10 in/1M$0.32 out/1M7 tok/s131Kfp86d agoOpenRouterCHEAPEST$0.10 in/1M$0.32 out/1M131K15m agoNovita AI$0.14 in/1M$0.40 out/1M6K6h agoSambaNova$0.45 in/1M$0.90 out/1M103 tok/s131Kbf166h agoSambaNova$0.60 in/1M$1.20 out/1M131K6h agoCloudflare Workers AI$0.29 in/1M$2.25 out/1M43 tok/s24Kfp815m agoDeepInfraturbo tier$0.10 in/1M$0.32 out/1M17 tok/s131Kfp815m agoAkashMLzero-retention$0.13 in/1M$0.40 out/1M21 tok/s131Kfp815m agoNovita AIzero-retention$0.14 in/1M$0.40 out/1M24 tok/s6Kbf1615m agoNebius AI Studiozero-retention$0.13 in/1M$0.40 out/1M16 tok/s131Kfp826h agoParasailzero-retention$0.22 in/1M$0.50 out/1M35 tok/s131Kfp820h agoCoreWeavezero-retention$0.71 in/1M$0.71 out/1M71 tok/s128Kfp1615m agoGoogle Vertex AIzero-retention$0.72 in/1M$0.72 out/1M53 tok/s128K15m agoGoogle Vertex AIzero-retention$0.72 in/1M$0.72 out/1M51 tok/s128K15m agoCrusoezero-retention$0.25 in/1M$0.75 out/1M60 tok/s131Kbf1615m agoGroqzero-retention$0.59 in/1M$0.79 out/1M179 tok/s131K15m agoTogether AIzero-retention$1.04 in/1M$1.04 out/1M42 tok/s131Kfp815m ago
Full specs, hardware verdicts and benchmarks on theLlama 3.3 70B Instruct model page →