Our take
CoreWeave is a GPU-cloud inference provider with measured throughput strengths on select models and a zero-retention data option. Its compliance posture and geographic footprint remain largely unverified in our data.
Use this for high-throughput workloads on specific Qwen and IBM Granite models where 166–212 tps or 113 tps is available. Pick it when you need a contractual zero-retention guarantee paired with a no-training-on-prompts policy. Skip it if your procurement requires verified SOC 2, a confirmed headquarters location, or a guaranteed EU-region endpoint.
- Fastest measured throughput on qwen-qwen3-5-35b-a3b in the sample set: 212 tps, roughly 2.6× the throughput of qwen-qwen3-30b-a3b-instruct-2507 at 81 tps and 2.0× that of qwen-qwen3-6-35b-a3b at 166 tps.
- Strong throughput on a mid-size IBM Granite model: 113 tps on ibm-granite-granite-4-1-8b, nearly 6× the throughput of deepseek-deepseek-v4-flash at 19 tps and over 4× that of openai-gpt-oss-120b at 23 tps.
- Zero-retention data option available, paired with a commitment not to train on prompts.
- Lowest input price in the sample for openai-gpt-oss-safeguard-20b, matched only by openai-gpt-oss-120b at the same input tier.
- Compliance and geographic footprint are largely undocumented: SOC 2, headquarters country, and EU-region endpoint are all unverified in our data.
- Slow throughput on large and some popular models: deepseek-deepseek-v4-flash at 19 tps, openai-gpt-oss-120b at 23 tps, and meta-llama-llama-3-1-70b-instruct at 26 tps all sit below 40 tps.
- Premium pricing on Llama 3.3 70B relative to smaller models in the sample: meta-llama-llama-3-3-70b-instruct costs over 3× the input of qwen-qwen3-30b-a3b-instruct-2507 and more than 10× that of openai-gpt-oss-safeguard-20b.
Privacy & data handling
- ✓
- Yes
- ✗
- No
- No answer on record
The colour says whether the answer favours you, not whether it is a yes — not training on your prompts earns a green cross, no zero-retention option earns a red one.
Prompt retention: none.