Our take
Phala is a privacy-first, decentralized inference provider that guarantees zero data retention and does not train on prompts. It offers a modest catalogue of 20 tracked models, with competitive throughput on select mid-size options but significant slowdowns on its largest offerings.
Use this for privacy-sensitive workloads where zero-retention and no prompt training are hard requirements. Pick it when decentralized infrastructure matters more than centralization risk, or for latency-tolerant batch jobs that fit its 14–69 tokens-per-second range. Skip it if you need verified SOC 2, a confirmed EU-region endpoint, or guaranteed low-latency real-time chat, or if you require the largest models at usable speed.
- Strong privacy guarantees with verifiable zero-retention — does not train on prompts and explicitly offers a zero-retention option.
- Accessible entry pricing on small models, with Qwen2.5 7B Instruct and OpenAI GPT-OSS Safeguard 20B both at the lowest tier.
- Competitive throughput on select mid-size models: Qwen3 30B A3B Instruct at 69 tps, OpenAI GPT-OSS 120B at 66 tps, and OpenAI GPT-OSS Safeguard 20B at 63 tps.
- Throughput collapses on largest models with premium pricing: DeepSeek V3 2 at only 8 tps, and Qwen3 5 27B at only 7 tps — the latter's output price is 24× the cheapest sample while its throughput is roughly 10× worse than the best measured.
- Compliance and geographic transparency gaps: SOC 2, EU endpoint and headquarters country are all unverified or undisclosed in our data.
- One model lacks any throughput measurement: Qwen3 VL 30B A3B Instruct.
Privacy & data handling
- ✓
- Yes
- ✗
- No
- No answer on record
The colour says whether the answer favours you, not whether it is a yes — not training on your prompts earns a green cross, no zero-retention option earns a red one.
Prompt retention: none.