Providers /Groq
Models
10
Regions
—
00
Editorial by LLMap · updated Aug 2, 2026Our take
Groq is an inference provider that uses its own custom inference chips instead of standard graphics chips, designed for speed. It offers a small, curated menu of downloadable models at aggressive prices.
Pick this for latency-critical products such as voice agents or interactive interfaces, on supported models. Use it for cheap small-model bulk inference. Skip it if you need a broad model catalogue, an EU data-residency endpoint, or a single accountable host.
Strengths
- Purpose-built inference hardware designed for high tokens per second.
- Among the lowest small-model prices around.
Trade-offs
- Curated, limited catalogue; you adapt to its menu, not vice versa.
- EU endpoint is unverified in our data.
01
Privacy & data handling
- ✓
- Yes
- ✗
- No
- No answer on record
The colour says whether the answer favours you, not whether it is a yes — not training on your prompts earns a green cross, no zero-retention option earns a red one.
Trains on your prompts
✗No
Logs prompts
✗No
✓Offered
✓Attested
Prompt retention: none.
02
Models & pricing
gpt-oss-120b120B$0.15 / $0.60 /1M131K9m agogpt-oss-safeguard-20b21.5B$0.075 / $0.30 /1M131K9m agoKimi K2 09051T$1.00 / $3.00 /1M262K14h agoLlama 3.1 8B Instruct8B$0.050 / $0.080 /1M131K10m agoLlama 3.3 70B Instruct70.6B$0.59 / $0.79 /1M131K10m agoLlama 4 Scout109B$0.11 / $0.34 /1M131K10m agoMiniMax M2.7229B$0.60 / $1.80 /1M197K10m agoQwen3 32B32.8B$0.29 / $0.59 /1M131K9m agoWhisper Large v31.5B$0.002 /minute of audio—6h agoWhisper Large v3 Turbo0.8B$0.001 /minute of audio—6h ago