Our take
A small, price-competitive host for downloadable models with measured throughput on most offers. It suits budget inference on specific models where its price beats larger catalogues, but its thin catalogue and unverified compliance documentation are real constraints.
Use this for budget inference on specific downloadable models where its price undercuts larger catalogues. Pick it for throughput-sensitive workloads on Qwen3 5-35B A3B — 107 tps is the fastest measured in its small set. Skip it if you need a broad model choice, verified SOC 2, a guaranteed EU endpoint, or a confirmed zero-retention policy.
- Measured throughput on 4 of 5 offers — Qwen3 5-35B A3B at 107 tps, Qwen3 6-35B A3B at 69 tps, Llama 3.3 70B at 24 tps, DeepSeek V4 Flash at 11 tps.
- Price-competitive on specific models — DeepSeek V4 Flash at $0.098/$0.196 undercuts some hosts; Llama 3.3 70B at $0.13/$0.4.
- Very small catalogue limits model choice — only 5 tracked offers versus 61-plus at larger hosts.
- No verified compliance or data-sovereignty documentation: SOC 2, EU endpoint, zero-retention, and trains-on-prompts are all unverified in our data.
- Some pricing is not the floor — Z AI GLM 5-2 at $0.77/$2.42 is relatively expensive, and Qwen3 variants at $1 output per million match or exceed competitors.
Privacy & data handling
- ✓
- Yes
- ✗
- No
- No answer on record
The colour says whether the answer favours you, not whether it is a yes — not training on your prompts earns a green cross, no zero-retention option earns a red one.
Prompt retention: none.