Models /Llama 3.3 70B Instruct /Where to run
Provider guide

Where to run Llama 3.3 70B Instruct

13 live listings tracked — output prices vary 7.0× between the cheapest and most expensive host, so the provider choice matters as much as the model choice.

CHEAPEST
$0.10 / $0.32
per 1M tokens in / out
FASTEST MEASURED
156 tok/s
measured throughput · $0.59 input / $0.79 output per million tokens
OpenRouterCHEAPEST$0.10 in/1M$0.32 out/1M—131K—4 hours agoNovita AIzero-retention through OpenRouterDirect and through OpenRouter$0.14 in/1M$0.40 out/1M8 tok/sthrough OpenRouter12Kbf164 hours agoParasailzero-retention$0.22 in/1M$0.50 out/1M20 tok/s131Kfp84 hours agoAkashMLzero-retention$0.20 in/1M$0.52 out/1M20 tok/s131Kfp84 hours agoCoreWeavezero-retention$0.71 in/1M$0.71 out/1M51 tok/s128Kfp164 hours agoGoogle Vertex AIzero-retention$0.72 in/1M$0.72 out/1M42 tok/s128K—4 hours agoGoogle Vertex AIzero-retention$0.72 in/1M$0.72 out/1M81 tok/s128K—4 hours agoGroqzero-retention$0.59 in/1M$0.79 out/1M156 tok/s131K—4 hours agoSambaNovazero-retention$0.45 in/1M$0.90 out/1M114 tok/s131K—4 hours agoTogether AIzero-retention$1.04 in/1M$1.04 out/1M9 tok/s131K—4 hours agoSambaNova$0.60 in/1M$1.20 out/1M—131K—4 hours agoCloudflare Workers AI$0.29 in/1M$2.25 out/1M11 tok/s24Kfp84 hours agoDeepInfraturbo tier$0.10 in/1M$0.32 out/1M13 tok/s131Kfp84 hours ago
Full specs, hardware verdicts and benchmarks on the Llama 3.3 70B Instruct model page →