Our take
Baseten is a US-based inference host with a strong security posture and measured throughput for every model it serves. Its catalogue is narrow but includes several models we list nowhere else.
Use this for security-sensitive deployments where SOC 2, zero-retention and no-training matter more than price. Pick it when 37–159 tokens per second is sufficient, or for models we list no other host for: GPT-OSS 120B, Inkling variants, Kimi K2-6, Kimi K2-7-code, Nemotron 3 Ultra and GLM 5 series. Skip it if you need a verified EU endpoint, a deep catalogue, or budget pricing.
- Strong security and data-privacy posture: SOC 2 attestation, no training on prompts, and a zero-retention option.
- Throughput measured for every offer — all 13 have published tokens-per-second figures.
- Exclusive or rare model hosting: we list no other provider for GPT-OSS 120B, Inkling Small, Inkling, Kimi K2-6, Kimi K2-7-code, Nemotron 3 Ultra, and several GLM 5 variants.
- Premium pricing on several offers versus its own cheaper alternatives: DeepSeek V4 Pro costs many times more than its own DeepSeek V4 Flash, and GLM 5-3 is far pricier than its own GLM 5-3 Flash.
- Narrow catalogue depth — 13 tracked offers versus providers with 60 or more.
- No EU-region presence verified in our data.
Point your tools here
We have not recorded what a router needs for Baseten yet — no base URL and no docs link on file. That is our gap, not a sign Baseten has no API; its own documentation is the place to look until we close it.
Privacy & data handling
These answers cover requests sent to Baseten through OpenRouter, as OpenRouter records them, last read 58 min ago. For Baseten's own API we hold no answer: check its terms. SOC 2 and the data processing agreement below are hand-checked, not part of that sync.
We have not dated this attestation or linked the report yet.
Prompt retention: none.