Our take
SiliconFlow is a China-headquartered host for downloadable models with a zero-retention data option and strong measured throughput on select models. Its 36 tracked offers include competitive picks for speed-sensitive workloads, though compliance documentation is thin and throughput varies widely across the catalogue.
Use this when zero data retention is required and the provider's terms are acceptable. Pick it for throughput-sensitive workloads on specific models where measured speed is strong, or for exploring Chinese-model access. Skip it if you need verified SOC 2, a guaranteed EU endpoint, or consistent performance across the full catalogue.
- Strong measured throughput on select models — DeepSeek V4 Flash at 69 TPS, StepFun Step 3.5 Flash at 55.5 TPS, Google Gemma 4 31B at 33 TPS, all above the sample median of 25.5 TPS.
- Zero-retention data option available and it does not train on prompts.
- Low entry price on some models: OpenAI GPT-OSS Safeguard 20B is the cheapest measured sample.
- Compliance and regional documentation gaps — SOC 2 and EU endpoint both unverified in our data.
- Inconsistent throughput across the catalogue: a 6.9x spread from 10 TPS to 69 TPS among measured samples.
- Output pricing can run high relative to input: seven of twelve measured models have output prices at least triple their input prices.
Point your tools here
https://api.siliconflow.com/v1Model IDs on this page are LLMap's own slugs, built for browsing and comparing — not what SiliconFlow's API expects. Its own model identifier lives in the docs above.
Privacy & data handling
These answers cover requests sent to SiliconFlow through OpenRouter, as OpenRouter records them, last read 56 min ago. For SiliconFlow's own API we hold no answer: check its terms. SOC 2 and the data processing agreement below are hand-checked, not part of that sync.
Prompt retention: none.