Our take
Together AI is a US-headquartered inference provider with strong compliance credentials and a zero-retention option. It offers 20 tracked models with standout throughput on select options and competitive pricing on some smaller models.
Use this for compliance-sensitive deployments needing an independent security attestation and zero data retention. Pick it for high-throughput applications where ThinkingMachines Inkling Small or GPT-OSS 120B fit the model need. Skip it if you need a verified EU-region endpoint, a deep catalogue, or unmeasured models like Minimax M2-7.
- Strong compliance and data control posture: SOC 2 attestation, no training on prompts, and a zero-retention option.
- Exceptional throughput on specific models. Inkling Small reaches 189 tokens per second, 8.6× that of Llama Guard 4 12B. GPT-OSS 120B runs at 105 tokens per second, 3.5× that of Gemma 4 31B.
- Lowest input price among sampled models: GPT-OSS Safeguard 20B at five cents per million input tokens, one-quarter the input price of Llama Guard 4 12B.
- Limited catalogue depth versus the largest open-weight hosts: 20 tracked offers.
- EU-region presence is unverified in our data; throughput is also unmeasured for two of twelve sample offers.
- Several models carry high output prices relative to input. GPT-OSS 120B, Minimax M3 and M2-7 all charge four times as much for output as for input; Gemma 4 31B charges just over three times as much.
Privacy & data handling
- ✓
- Yes
- ✗
- No
- No answer on record
The colour says whether the answer favours you, not whether it is a yes — not training on your prompts earns a green cross, no zero-retention option earns a red one.
Prompt retention: none.