Our take
Modal is a privacy-first inference host with a zero-retention guarantee and a narrow catalogue of seven tracked offers. It suits workloads where prompt privacy matters more than breadth of choice.
Use this when zero-retention is a hard requirement and you are running GLM or DeepSeek Flash variants at modest scale. Pick it for low-latency, cost-conscious inference on those specific models. Skip it if you need a wide model catalogue, verified compliance attestations, or flagship-tier throughput.
- Strong privacy posture: a verified zero-retention option and a commitment not to train on prompts.
- Cheapest entry point in its own lineup for GLM-5-3-Flash.
- Best measured throughput on Qwen3-8-2-4t-A95B in its own data — 98 tps on one run versus 79 tps on another.
- Very small catalogue — only seven tracked offers.
- Mid-range throughput with wide variance: five of six measured runs fall between 77-87 tps.
- Premium pricing on its flagship-tier model and undocumented compliance: moonshotai-kimi-k3 costs many multiples more than its cheapest tier, while SOC 2, EU endpoint and headquarters country are all unverified.
Point your tools here
We have not recorded what a router needs for Modal yet — no base URL and no docs link on file. That is our gap, not a sign Modal has no API; its own documentation is the place to look until we close it.
Privacy & data handling
These answers cover requests sent to Modal through OpenRouter, as OpenRouter records them, last read 58 min ago. For Modal's own API we hold no answer: check its terms. SOC 2 and the data processing agreement below are hand-checked, not part of that sync.
Prompt retention: none.