Providers /Venice AI
Inference provider · HQ US

Venice AI

Models
35
Data residency
No answer on record
00

Our take

Editorial by LLMap · updated Sep 5, 2026

Venice AI is a US-based host for downloadable models with a privacy-first stance: it does not train on prompts and offers zero-retention inference. Its 35 tracked offers include standout speed on a small NVIDIA model and competitive pricing on select entries, though throughput is uneven and compliance attestations remain unverified.

Use this for privacy-sensitive prototyping where zero-retention matters. Pick it for low-latency work on NVIDIA Nemotron 3.5 Lightning, or for budget small-model inference where Z AI GLM 4 7 Flash has the cheapest input in its own catalogue. Skip it if you need verified SOC 2, a disclosed EU endpoint, or predictable throughput across many models.

Strengths
  • Strongest privacy stance among tracked providers — does not train on prompts and offers a zero-retention option.
  • Highest measured throughput on a specific small model: NVIDIA Nemotron 3.5 Lightning at 232.5 tokens per second, roughly 9.7× Mistral Small 3.2 and 21× DeepSeek V3.2.
  • Lowest input price in its own catalogue: Z AI GLM 4 7 Flash at half the input price of Google Gemma 4 31B and well below Qwen3 5 9B.
Trade-offs
  • Compliance and regional documentation gaps: SOC 2 is unverified in our data and no EU endpoint is disclosed.
  • Inconsistent throughput with several slow offers — five models below 50 tps, and DeepSeek V3.2 slowest at 11 tps.
  • Output pricing on larger models is steep: Mistral Small 4 and Qwen3 235B A22B both charge many multiples above their cheapest entry.
01

Point your tools here

We have not recorded what a router needs for Venice AI yet — no base URL and no docs link on file. That is our gap, not a sign Venice AI has no API; its own documentation is the place to look until we close it.

02

Privacy & data handling

These answers cover requests sent to Venice AI through OpenRouter, as OpenRouter records them, last read 58 min ago. For Venice AI's own API we hold no answer: check its terms. SOC 2 and the data processing agreement below are hand-checked, not part of that sync.

Trains on prompts
✗No
Logs prompts
✗No
✓Offered
No answer on record
No answer on record

Prompt retention: none.

03

Models & pricing

DeepSeek V3.2Through OpenRouter685B$0.27 / $0.39 /1M160K58 min agoDeepSeek V4.1 Flashfp8Through OpenRouter763B$0.38 / $1.50 /1M1M58 min agoDeepSeek V4 FlashThrough OpenRouter291B$0.097 / $0.19 /1M1M59 min agoDeepSeek V4 ProThrough OpenRouter1.6T$1.65 / $3.30 /1M1M58 min agoGemma 4 26B A4Bbf16Through OpenRouter25.8B$0.13 / $0.40 /1M256K58 min agoGemma 4 31Bfp4Through OpenRouter31.3B$0.12 / $0.36 /1M256K58 min agoGLM 4.6fp4Through OpenRouter357B$0.43 / $1.75 /1M198K56 min agoGLM 4.7fp4Through OpenRouter358B$0.40 / $1.93 /1M198K13 hours agoGLM 4.7 Flashfp8Through OpenRouter31.2B$0.060 / $0.40 /1M128K56 min agoGLM 5fp8Through OpenRouter754B$1.00 / $3.20 /1M198K56 min agoGLM 5.1fp8Through OpenRouter754B$1.40 / $4.40 /1M200K56 min agoGLM 5.2fp8Through OpenRouter753B$1.40 / $4.40 /1M1M56 min agoGLM 5.3Through OpenRouter—$1.40 / $4.40 /1M1M56 min agoGLM 5.3 FlashThrough OpenRouter321B$0.15 / $0.50 /1M1M56 min agoKimi K2.5Through OpenRouter1.1T$0.53 / $3.33 /1M256K57 min agoKimi K2.6int4Through OpenRouter1.1T$0.75 / $3.50 /1M256K57 min agoKimi K2.7 Codeint4Through OpenRouter1.1T$0.75 / $3.50 /1M256K57 min agoMiMo-V2.5fp8Through OpenRouter311B$0.40 / $2.00 /1M1M56 min agoMiMo-V2.6-Flashfp8Through OpenRouter—$0.17 / $0.35 /1M1M56 min agoMiniMax M2.5Through OpenRouter229B$0.27 / $0.95 /1M198K57 min agoMiniMax M3fp8Through OpenRouter427B$0.30 / $1.20 /1M524K57 min agoMistral Small 3.2 24Bfp8Through OpenRouter24B$0.094 / $0.25 /1M256K57 min agoNemotron 3 Ultrafp8Through OpenRouter561B$0.63 / $3.13 /1M256K7 hours agoQwen3 235B A22B Instruct 2507fp8Through OpenRouter235B$0.15 / $0.75 /1M128K56 min agoQwen3 235B A22B Thinking 2507fp8Through OpenRouter235B$0.45 / $3.50 /1M128K56 min agoQwen3.5-35B-A3BThrough OpenRouter36B$0.31 / $1.25 /1M256K56 min agoQwen3.5 397B A17BThrough OpenRouter403B$0.75 / $4.50 /1M128K56 min agoQwen3.5-9Bfp8Through OpenRouter9.7B$0.10 / $0.15 /1M256K56 min agoQwen3.6 27Bfp8Through OpenRouter27.8B$0.33 / $3.25 /1M256K56 min agoQwen3.6 35B A3Bfp8Through OpenRouter36B$0.10 / $1.00 /1M256K56 min agoQwen3.8 2.4T A95BThrough OpenRouter2.4T$2.00 / $6.00 /1M262K56 min agoQwen3.8 27Bfp8Through OpenRouter27.8B$0.45 / $3.20 /1M262K56 min agoQwen3 Coder 480B A35Bfp8Through OpenRouter480B$0.35 / $1.50 /1M256K56 min agoQwen3 VL 235B A22B Instructfp8Through OpenRouter236B$0.21 / $1.90 /1M128K56 min agoUncensoredfp16Through OpenRouter24B$0.20 / $0.90 /1M128K58 min ago
Something wrong on this page? Tell us