Models / DeepSeek/ DeepSeek V3

DeepSeek V3

DeepSeek · released Dec 25, 2024 · deepseek-ai/DeepSeek-V3

Input: text. Output: text.InputOutput
Type
Open weights
Params
685B
Context
164K

37B active per word · about 123K words of context

Our take

Written Aug 1, 2026

Built for coding and reasoning workloads, this large downloadable model from DeepSeek uses a mixture-of-experts design that keeps only 37 billion parameters active per token out of 684.5 billion total. It offers a wide context window and strong benchmarked coding performance, though its licence terms remain unverified and it handles text only.

Who should pick it

Pick this for coding workloads where benchmarked scores matter, or cost-sensitive high-volume input. Choose it when throughput is critical, or when downloadable weights matter but you can verify licence terms yourself. Skip it if you need multimodal input, require a verified permissive licence, or prioritise creative writing over coding.

The case for it

  • Strongest measured skill is coding, with a 29.1-point gap over its own overall text score.
  • Dramatic parameter efficiency: only 37 billion active per token from 684.5 billion total.
  • Wide context window of 163,840 tokens for long-document tasks.
  • Lowest input cost among its own tracked offers, with a 3.46× spread between cheapest and most expensive.

The case against it

  • Licence terms are unverified for commercial safety; not disclosed in our data.
  • Text-to-text only; no image, video, or audio handling.
  • Creative writing and instruction following lag its own coding score by 38.8 and 43.9 points.
00

How good is it?

IntelligencePuzzles, maths, exam questions

2 of 5

Arena Text (overall)98th of 143 · 1358.5

Arena Hard Prompts 107th of 143Arena Maths 109th of 139

CodingWriting and fixing code on its own

2 of 5

Arena Coding104th of 143 · 1387.7

LiveCodeBench 10th of 14

AgenticPlanning, calling tools, staying on task

not measured

Nobody we watch has scored DeepSeek V3 for this. We would take the rating from Arena Agent (IPS).

WritingWe do not rate this

Scored, not ratedThe placings are on the right.

Two boards come close and neither tests writing: Arena Creative Writing asks people which of two replies they prefer, and LiveBench Language tests whether a model understood a passage. So we show where DeepSeek V3 placed and give it no mark out of five.

Arena Creative Writing 80th of 143 · 1348.9
Also scored, on boards we give no mark for
Arena Instruction Following 98th of 143

These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done, which is why they get no rating.

Every published score for this model7 scoresEvery figure we hold, from 7 boards, with who ran it and a link to the source — including the boards no rating above is built on.
49.6independentsource ↗
1387.7independentsource ↗
1350.3independentsource ↗
1310.7independentsource ↗
1358.5independentsource ↗
01

Can you run it yourself?

Fits in memory
weights load entirely on the card
Spills to system RAM
some weights offload; much slower
Too large
will not load even with offload
est
size is calculated; the verdict could change by 10%
A card many people ownToo large

GeForce RTX 4090 · 24 GB

Weights at Q4_K_M404.4 / 24 GBmeasured
Usable contextNot calculated for spilled setups
Decode speedNot estimated for spilled setups

Too large for this card. The weights do not fit even with part of them offloaded to system memory.

One step upToo large

Apple M1 Pro (16-core GPU) · 32 GB

Weights at Q4_K_M404.4 / 32 GBmeasured
Usable contextNot calculated for spilled setups
Decode speedNot estimated for spilled setups

Too large for this card. The weights do not fit even with part of them offloaded to system memory.

One step downToo large

Radeon RX 7900 XT · 20 GB

Weights at Q4_K_M404.4 / 20 GBmeasured
Usable contextNot calculated for spilled setups
Decode speedNot estimated for spilled setups

Too large for this card. The weights do not fit even with part of them offloaded to system memory.

Your hardware
Checking your profile…

Memory use by level

Against a 24 GB card.

F16
1342.3 GBmeasured
Too large
Q4_K_M
recommended
404.4 GBmeasured
Too large
Q5_K_M
506.3 GBest
Too large
Q8_0
713.3 GBmeasured
Too large

Check against your own machine → · All 71 devices, with every size →

02

Or rent it from someone else

Cheapest published offer

Cheapest of 6 live listings. Picked at the widest standard context we hold, within one quantisation slice, so the numbers beside it are a price one host actually charges.

per 1M tokens
$0.26 in / $1.03 out
Context served
164K
Throughput
Not measured
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
DeepInfrafp4$0.32 / $0.89164Knot measuredUnknownUnknownUnknown
DeepInfrafp4$0.32 / $0.89164K14 tok/sNoNoConfirmed
Novita AI$0.89 / $0.8964Knot measuredUnknownUnknownUnknown
OpenRouter$0.26 / $1.03164Knot measuredUnknownUnknownUnknown
StreamLake$0.26 / $1.03128K26 tok/sNoYesunknown periodUnknown
Novita AIfp8$0.40 / $1.3064K19 tok/sNoNoConfirmed

Across the 6 listings we hold: 3 say they do not train on prompts, 0 say they do and 3 do not say. 2 appear in the zero-retention registry we check; the rest are unknown to us rather than confirmed either way.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

Supported
Not supported
Not published
host gave no parameter list
API features per host
ProviderTool callingJSON outputStrict schema
DeepInfrafp4
DeepInfrafp4
Novita AI
OpenRouter
StreamLake
Novita AIfp8

Tool calling: 4 of 6 listings say yes, 2 publish no parameter list. JSON output: 3 of 6 listings say yes, 1 says no, 2 publish no parameter list. Strict schema: 3 of 6 listings say yes, 1 says no, 2 publish no parameter list.

03

Models people weigh against DeepSeek V3

04

When we formed this view

Dates behind this page

Aug 3, 2026BenchmarkScored 49.6 on LiveCodeBenchleaderboard
Aug 2, 2026BenchmarkScored 1387.7 on Arena Codingleaderboard
Aug 2, 2026BenchmarkScored 1348.9 on Arena Creative Writingleaderboard
Aug 2, 2026BenchmarkScored 1350.3 on Arena Hard Promptsleaderboard
Aug 2, 2026BenchmarkScored 1343.7 on Arena Instruction Followingleaderboard
Aug 2, 2026BenchmarkScored 1310.7 on Arena Mathsleaderboard
Aug 2, 2026BenchmarkScored 1358.5 on Arena Text (overall)leaderboard
Jul 30, 2026Price changeDeepSeek V3 rose across 2 hosts, by up to 29% at StreamLakeDeepSeek V3 moved on 2 hosts: StreamLake: input +29% ($0.20 → $0.26 per 1M tokens); output +29% ($0.80 → $1.03 per 1M tokens); OpenRouter: input +29% ($0.20 → $0.26 per 1M tokens); output +29% ($0.80 → $1.03 per 1M tokens)
Jul 26, 2026ListedListed on LLMapfirst indexed by our pipeline
Dec 25, 2024AnnouncedDeepSeek V3 announced by DeepSeek

Prices last checked 6h ago

What we do not know about this model yet

  • 2 of 6 listings publish no parameter list, so what their API accepts is unknown to us.
  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 3 of 6 listings do not say whether they train on prompts.
  • We hold no cached-input rate for any of its listings.
05

Licence and identifiers

What the licence allowsWe hold no licence record for this model. Inside are the identifiers you need to pull it — its Hugging Face repo where we have one, our slug and a machine-readable card.

Licence

We hold no licence record for this model yet.

Identifiers

Architecture
Mixture of experts
Modality record
text->text
Catalogue slug
deepseek-deepseek-v3

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us