DeepSeek V3.2
DeepSeek · released Dec 1, 2025 · deepseek-ai/DeepSeek-V3.2
- Type
- Open weightsMIT License
- Params
- 685B
- Context
- 164K
37B active per word · about 123K words of context
Our take
Written Sep 1, 2026DeepSeek V3.2 is a large downloadable text model built for coding and reasoning, with 685.4 billion total parameters and 37 billion active per word. It carries a permissive MIT licence and is available from 20 hosted providers.
Choose this for coding workloads where its measured arena strength matters, or for long-context tasks up to 163,840 tokens. Use it when you need wide provider choice with competitive floor pricing, or if Baidu's high-throughput option is available to you. Skip it if you need image or video input, or if web development coding is your main work — its score there lags well behind its general coding strength.
The case for it
- Arena Coding 1469.9 is its highest arena score, 44.8 points above its overall text score.
- Solves 70% of real GitHub issues end-to-end on SWE-bench Verified with mini-SWE-agent.
- 20 hosted offers give unusually wide provider choice, plus a permissive MIT licence.
- 53 tokens per second at Baidu, versus roughly half that at most measured alternatives.
The case against it
- Web development coding is a clear weak spot: 145.3 points below its general coding score.
- Creative writing is its lowest arena category, 24.1 points below its overall text score.
- Text-only; no image or video input, and throughput varies sharply by provider.
How good is it?
EverydayGeneral questions and everyday reasoning
Arena Text (overall)74th of 168 · 1425
Also on this board: 1423 (Sep 25, 2026). Read the pair, not the higher one.
CodingWriting and fixing code on its own
Arena Coding75th of 168 · 1469
Also on this board: 1476 (Sep 25, 2026). Read the pair, not the higher one.
AgenticPlanning, calling tools, staying on task
Not yet scored on Arena Agent. It is on SWE-bench Verified, in 10th of 42 with 70.
WritingDrafting and rewriting prose
Arena Creative Writing68th of 168 · 1400
Arena Creative Writing is the only board that has scored it for this.
Also on this board: 1391 (Sep 25, 2026). Read the pair, not the higher one.
These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done.
Every published score for this model8 scoresEvery figure we hold, from 8 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Can you run it yourself?
GeForce RTX 4090 · 24 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Apple M1 Pro (16-core GPU) · 32 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Radeon RX 7900 XT · 20 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Memory use by level
Against a 24 GB card.
What is quantisation? →This model on every device we track71 devicesThe Q4 build most people download, on each device: what the weights come to, how much context the memory leaves, and whether it runs. Smallest device that runs it first.
Check against your own machine → · Where to rent it hosted →
Or rent it from someone else
Prices checked between 1 hour and 4 days ago — each listing carries its own date.
Novita AI, direct
Cheapest of the 6 listings we can compare like for like — at 164K of context, out of 15 in the table below. 4 cheaper rows there are outside that comparison: a different quantisation or a different context length.
- per 1M tokens
- $0.27 in / $0.40 out
- Context served
- 164K
- Throughput
- Not measured
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| GMICloudfp8Through OpenRouter | $0.21 / $0.31checked 3 days ago | 164K147K max reply | 33 tok/s | No | Yesunknown period | Unknown |
| DeepInfrafp4Direct and through OpenRouter | $0.26 / $0.38checked 1 hour ago | 164K16K max reply through OpenRouter | 14 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| AtlasCloudfp8Through OpenRouter | $0.26 / $0.38checked 1 hour ago | 164K147K max reply | 34 tok/s | No | Yesunknown period | Unknown |
| Venice AIThrough OpenRouter | $0.27 / $0.39checked 1 hour ago | 160K33K max reply | 11 tok/s | No | No | Confirmed |
| Novita AIDirect | $0.27 / $0.40checked 4 days ago | 164K | not measured | Unknown | Unknown | Unknown |
| OpenRouterOpenRouter's own listing | $0.28 / $0.42checked 1 hour ago | 164K | not measured | Unknown | Unknown | Unknown |
| SiliconFlowfp8Through OpenRouter | $0.26 / $0.42checked 1 hour ago | 164K147K max reply | 14 tok/s | No | No | Confirmed |
| Baidufp8Through OpenRouter | $0.28 / $0.42checked 1 hour ago | 131K66K max reply | 30 tok/s | No | Yesunknown period | Unknown |
| DigitalOcean GradientThrough OpenRouter | $0.30 / $0.96checked 1 hour ago | 164K147K max reply | 17 tok/s | No | No | Confirmed |
| PhalaThrough OpenRouter | $1.00 / $1.00checked 1 hour ago | 164K64K max reply | 11 tok/s | No | No | Confirmed |
| Alibaba Cloudfp8Through OpenRouter | $0.37 / $1.11checked 1 hour ago | 131K66K max reply | 23 tok/s | No | Yesunknown period | Unknown |
| FriendliThrough OpenRouter | $0.50 / $1.50checked 1 hour ago | 164K147K max reply | 45 tok/s | No | Yesunknown period | Unknown |
| Google Vertex AIThrough OpenRouter | $0.56 / $1.68checked 1 hour ago | 164K66K max reply | 9 tok/s | No | No | Confirmed |
| MaraThrough OpenRouter | $3.00 / $4.50checked 7 hours ago | 33K7K max reply | 21 tok/s | No | No | Confirmed |
| SambaNovaDirect and through OpenRouter | $3.00 / $4.50checked 1 hour ago | 33K7K max reply through OpenRouter | 40 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
Across the 15 listings we hold: 13 say they do not train on prompts (2 of them only through OpenRouter), 0 say they do and 2 do not say. 8 appear in the zero-retention registry we check (2 of them only through OpenRouter); the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| GMICloudfp8Through OpenRouter | ✓ | ✗ | ✗ |
| DeepInfrafp4Direct and through OpenRouter | ✓ | ✓ | ✓ |
| AtlasCloudfp8Through OpenRouter | ✓ | ✓ | ✓ |
| Venice AIThrough OpenRouter | ✓ | ✓ | ✓ |
| Novita AIDirect | |||
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| SiliconFlowfp8Through OpenRouter | ✓ | ✓ | ✓ |
| Baidufp8Through OpenRouter | ✗ | ✓ | ✓ |
| DigitalOcean GradientThrough OpenRouter | ✗ | ✗ | ✗ |
| PhalaThrough OpenRouter | ✓ | ✓ | ✓ |
| Alibaba Cloudfp8Through OpenRouter | ✓ | ✓ | ✓ |
| FriendliThrough OpenRouter | ✓ | ✓ | ✓ |
| Google Vertex AIThrough OpenRouter | ✓ | ✓ | ✓ |
| MaraThrough OpenRouter | ✗ | ✓ | ✓ |
| SambaNovaDirect and through OpenRouter | ✗ | ✗ | ✗ |
Tool calling: 10 of 15 listings say yes, 4 say no, 1 publishes no parameter list. JSON output: 11 of 15 listings say yes, 3 say no, 1 publishes no parameter list. Strict schema: 11 of 15 listings say yes, 3 say no, 1 publishes no parameter list.
Models people weigh against DeepSeek V3.2
When we formed this view
Recent changes
What moved
input −48% ($0.260 → $0.134 per 1M tokens), output −47% ($0.38 → $0.20 per 1M tokens), cache read −48% ($0.130 → $0.067 per 1M tokens)What moved
input +20% ($0.25 → $0.30 per 1M tokens), output +20% ($0.80 → $0.96 per 1M tokens), cache read +20% ($0.075 → $0.090 per 1M tokens)What moved
input −41% ($0.425 → $0.250 per 1M tokens), output −41% ($1.36 → $0.80 per 1M tokens), cache read −50% ($0.150 → $0.075 per 1M tokens)Each date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
- 1 of 15 listings publishes no parameter list, so what its API accepts is unknown to us.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 2 of 15 listings do not say whether they train on prompts, and 2 answer only through OpenRouter, not for their own listing.
- We hold no batch or off-peak rate for any of its listings.
- We hold a decode speed for it, but no prompt-processing (prefill) figure, so how long the input side of a job takes is unknown to us.
Licence and identifiers
What the licence allowsMIT License, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.
Licence
MIT License
Fully permissive: do anything with attribution. No patent grant, unlike Apache-2.0.
Identifiers
- Hugging Face
- deepseek-ai/DeepSeek-V3.2
- Architecture
- Mixture of experts
- Takes in, gives back
- Text in, text out
- Catalogue slug
- deepseek-deepseek-v3-2