Our take
ModelRun is a boutique inference host with a tiny, curated catalogue of five models and a zero-retention privacy option. It trades breadth for strong throughput numbers on the specific models it carries.
Use this when zero data retention is required, or for high-throughput inference on a model in its small catalogue — its fastest option handles 152 tokens per second. Pick it for budget-conscious Gemma 4 31B inference, which is the cheapest input in its own catalogue. Skip it if you need model choice, verified compliance credentials, or a documented EU endpoint.
- Strongest throughput in its own catalogue: its fastest model runs at 152 tokens per second, nearly a fifth quicker than the next fastest and three-quarters quicker than the slowest.
- Zero-retention privacy option available and it does not train on prompts.
- Lowest input price for Gemma 4 31B in its own catalogue.
- Very small catalogue limits model choice — only five tracked offers.
- Highest output price for its fastest model: that option costs more than three times the output of its mid-tier model and over ten times its cheapest.
- Compliance and geographic footprint are undocumented: SOC 2, EU endpoint and headquarters country are all unverified in our data.
Point your tools here
We have not recorded what a router needs for ModelRun yet — no base URL and no docs link on file. That is our gap, not a sign ModelRun has no API; its own documentation is the place to look until we close it.
Privacy & data handling
These answers cover requests sent to ModelRun through OpenRouter, as OpenRouter records them, last read 58 min ago. For ModelRun's own API we hold no answer: check its terms. SOC 2 and the data processing agreement below are hand-checked, not part of that sync.
Prompt retention: none.