Provider guide
Where to run Llama 3.1 8B Instruct
6 live listings tracked — output prices vary 7.2× between the cheapest and most expensive host, so the provider choice matters as much as the model choice.
FASTEST MEASURED
106 tok/s
measured throughput · $0.22 input / $0.22 output per million tokens
DeepInfrazero-retention$0.020 in/1M$0.040 out/1M26 tok/s131Kfp84 hours agoNovita AIzero-retention through OpenRouterDirect and through OpenRouter$0.020 in/1M$0.050 out/1M49 tok/sthrough OpenRouter16Kfp84 hours agoOpenRouter$0.050 in/1M$0.080 out/1M—131K—4 hours agoGroqzero-retentionCHEAPEST$0.050 in/1M$0.080 out/1M73 tok/s131K—4 hours agoCoreWeavezero-retention$0.22 in/1M$0.22 out/1M106 tok/s131Kbf164 hours agoCloudflare Workers AI$0.15 in/1M$0.29 out/1M13 tok/s32Kfp84 hours ago
Full specs, hardware verdicts and benchmarks on the Llama 3.1 8B Instruct model page →