Models / DeepSeek/ DeepSeek V4 Flash

DeepSeek V4 Flash

DeepSeek · released Apr 22, 2026 · deepseek-ai/DeepSeek-V4-Flash

Input: text. Output: text.InputOutput
Type
Open weightsMIT License
Params
291B
Context
1M

37B active per word · about 786K words of context

Our take

Written Aug 3, 2026

DeepSeek V4 Flash is a large downloadable text model built for long documents, with a permissive MIT licence and a one-million-token request limit. Its measured web-development coding score is the strongest of its seven tracked skills, while creative writing is the weakest.

Who should pick it

Choose this for long-document processing at one-million-token scale where open weights and a permissive licence matter. It suits budget-conscious API inference through multiple providers, and web-development coding tasks where its measured Arena Code score is highest. Skip it if you need image, video or audio input, if creative writing quality is central, or if you need consistently fast throughput across every host you might use.

The case for it

  • One-million-token request limit, among the longest in the downloadable-model market.
  • Permissive MIT licence allows commercial use, modification and redistribution without copyleft requirements.
  • Strong measured web-development coding score, with an Arena Code Elo of 1576.5 and a prior peak of 1586.1.
  • Multiple providers price input aggressively, with the cheapest well under a tenth of a cent per million tokens.

The case against it

  • Creative writing is the weakest measured skill area, with an Arena Creative Writing Elo 168.7 points below its own web-development coding score.
  • Throughput varies wildly by host, from 7 to 71 tokens per second — a more than tenfold speed gap for the same model.
  • Text-to-text only; no image, video or audio input supported.
00

How good is it?

IntelligencePuzzles, maths, exam questions

3 of 5

Arena Text (overall)48th of 143 · 1435.6

Arena Hard Prompts 47th of 143Arena Maths 52nd of 139LiveBench Data Analysis 27th of 35LiveBench Mathematics 30th of 35LiveBench Reasoning 33rd of 35

Also on this board: 1437.6 via Thinking (Jul 30, 2026). Read the pair, not the higher one.

CodingWriting and fixing code on its own

3 of 5

Arena Coding48th of 143 · 1483.5

Arena Code (WebDev) 6th of 74 via HighLiveBench Coding 32nd of 35

Also on this board: 1480.9 via Thinking (Jul 30, 2026). Read the pair, not the higher one.

AgenticPlanning, calling tools, staying on task

2 of 5

Arena Agent (IPS)29th of 36 · −0.03

LiveBench Agentic Coding 34th of 35

WritingWe do not rate this

Scored, not ratedThe placings are on the right.

Two boards come close and neither tests writing: Arena Creative Writing asks people which of two replies they prefer, and LiveBench Language tests whether a model understood a passage. So we show where DeepSeek V4 Flash placed and give it no mark out of five.

Arena Creative Writing 43rd of 143 · 1407.4LiveBench Language 33rd of 35 · 70.1
Also scored, on boards we give no mark for
Arena Instruction Following 46th of 143LiveBench Instruction Following 24th of 35LiveBench 32nd of 35

These tests check whether a model follows instructions — a precondition for all the work above, but not a measure of how well that work is done, which is why they get no rating.

Every published score for this model16 scoresEvery figure we hold, from 16 boards, with who ran it and a link to the source — including the boards no rating above is built on.
LiveBenchreasoning
65.5independentsource ↗
69.2independentsource ↗
68independentsource ↗
70.1independentsource ↗
70.6independentsource ↗
−0.03independentsource ↗
1483.5independentsource ↗
1459.3independentsource ↗
1426.6independentsource ↗
1435.6independentsource ↗
1576.5via Highindependentsource ↗
01

Can you run it yourself?

Fits in memory
weights load entirely on the card
Spills to system RAM
some weights offload; much slower
Too large
will not load even with offload
est
size is calculated; the verdict could change by 10%
A card many people ownToo large

GeForce RTX 4090 · 24 GB

Weights at Q4_K_M183.4 / 24 GBest
Usable contextNot calculated for spilled setups
Decode speedNot estimated for spilled setups

Too large for this card. The weights do not fit even with part of them offloaded to system memory.

One step upToo large

Apple M1 Pro (16-core GPU) · 32 GB

Weights at Q4_K_M183.4 / 32 GBest
Usable contextNot calculated for spilled setups
Decode speedNot estimated for spilled setups

Too large for this card. The weights do not fit even with part of them offloaded to system memory.

Comfortable fit

On a MacFits in memory

Apple M3 Ultra (80-core GPU) · 512 GB

Weights at Q4_K_M183.4 / 512 GBest
Spare memory193.7 GB spare
Usable context262K of 1M
Decode speed22 tok/sest

Room to spare. 193.7 GB spare means a 10% error in the size would not change the answer.

Your hardware
Checking your profile…

Memory use by level

Against a 24 GB card.

Q4_K_M
recommended
183.4 GBest
Too large
Q5_K_M
215.2 GBest
Too large
Q8_0
321.4 GBest
Too large

All 3 sizes here are calculated, not measured. We hold no measured file for this model, so each size comes from the parameter count and every verdict above inherits that uncertainty.

Check against your own machine → · All 71 devices, with every size →

02

Or rent it from someone else

Cheapest published offer

Cheapest of 32 live listings. Picked at the widest standard context we hold, within one quantisation slice, so the numbers beside it are a price one host actually charges.

per 1M tokens
$0.14 in / $0.28 out
Context served
1M
Throughput
Not measured
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
StreamLakefp8$0.088 / $0.181M45 tok/sNoYesunknown periodUnknown
Baidufp8$0.088 / $0.181M79 tok/sNoYesunknown periodUnknown
DeepInfrafp4$0.090 / $0.181M66 tok/sNoNoConfirmed
DeepInfrafp4$0.090 / $0.181Mnot measuredUnknownUnknownUnknown
GMICloudfp8$0.094 / $0.191M45 tok/sNoYesunknown periodUnknown
AkashMLfp8$0.098 / $0.20131K10 tok/sNoNoConfirmed
DigitalOcean Gradient$0.11 / $0.22262K12 tok/sNoNoConfirmed
Alibaba Cloudfp8$0.13 / $0.271M66 tok/sNoYesunknown periodUnknown
Alibaba Cloud$0.13 / $0.271M66 tok/sNoYesunknown periodUnknown
Venice AI$0.14 / $0.281M44 tok/sNoNoConfirmed
Morph$0.14 / $0.281M12 tok/sNoNoConfirmed
OpenInferencefp8$0.14 / $0.281M10 tok/sNoNoConfirmed
Novita AIfp8$0.14 / $0.281M43 tok/sNoNoConfirmed
Novita AI$0.14 / $0.281Mnot measuredUnknownUnknownUnknown
SiliconFlowfp8$0.14 / $0.281M58 tok/sNoNoConfirmed
Cloudflare Workers AI$0.14 / $0.28384K29 tok/sNoYesunknown periodUnknown
Cloudflare Workers AIfp8$0.14 / $0.28384K45 tok/sNoYesunknown periodUnknown
AtlasCloudfp4$0.14 / $0.281M72 tok/sNoYesunknown periodUnknown
AtlasCloudfp8$0.14 / $0.28262K47 tok/sNoYesunknown periodUnknown
Parasailfp8$0.14 / $0.281M55 tok/sNoNoConfirmed
CoreWeavefp8$0.14 / $0.281M17 tok/sNoNoConfirmed
Ambientfp4$0.14 / $0.281M40 tok/sNoYesunknown periodUnknown
DeepSeek$0.14 / $0.281M61 tok/sYesYesunknown periodUnknown
DeepSeekfp8$0.14 / $0.281M74 tok/sYesYesunknown periodUnknown
OpenRouter$0.14 / $0.281Mnot measuredUnknownUnknownUnknown
Ionstreamfp4$0.14 / $0.281M39 tok/sNoNoConfirmed
Ionstreamfp4$0.14 / $0.281M16 tok/sNoNoUnknown
Fireworks AI$0.14 / $0.281M59 tok/sNoNoConfirmed
Io Netfp8$0.17 / $0.33262K5 tok/sNoNoConfirmed
Phala$0.20 / $0.401M22 tok/sNoNoConfirmed
Mancer 2fp4$0.20 / $0.501M27 tok/sNoNoConfirmed
Mancer 2fp8$0.25 / $1.001M56 tok/sNoNoConfirmed

Across the 32 listings we hold: 27 say they do not train on prompts, 2 say they do and 3 do not say. 16 appear in the zero-retention registry we check; the rest are unknown to us rather than confirmed either way.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

Supported
Not supported
Not published
host gave no parameter list
API features per host
ProviderTool callingJSON outputStrict schema
StreamLakefp8
Baidufp8
DeepInfrafp4
DeepInfrafp4
GMICloudfp8
AkashMLfp8
DigitalOcean Gradient
Alibaba Cloudfp8
Alibaba Cloud
Venice AI
Morph
OpenInferencefp8
Novita AIfp8
Novita AI
SiliconFlowfp8
Cloudflare Workers AI
Cloudflare Workers AIfp8
AtlasCloudfp4
AtlasCloudfp8
Parasailfp8
CoreWeavefp8
Ambientfp4
DeepSeek
DeepSeekfp8
OpenRouter
Ionstreamfp4
Ionstreamfp4
Fireworks AI
Io Netfp8
Phala
Mancer 2fp4
Mancer 2fp8

Tool calling: 30 of 32 listings say yes, 2 publish no parameter list. JSON output: 29 of 32 listings say yes, 1 says no, 2 publish no parameter list. Strict schema: 26 of 32 listings say yes, 4 say no, 2 publish no parameter list.

03

Models people weigh against DeepSeek V4 Flash

04

When we formed this view

Dates behind this page

Aug 3, 2026Price changeDeepSeek V4 Flash repriced across 2 hosts, from an 8% rise at SiliconFlow to a 10% cut at Io NetDeepSeek V4 Flash moved on 2 hosts: SiliconFlow: input +8% ($0.13 → $0.14 per 1M tokens); Io Net: input −6% ($0.18 → $0.17 per 1M tokens); output −3% ($0.34 → $0.33 per 1M tokens); cache read −10% ($0.080 → $0.072 per 1M tokens)
Aug 3, 2026Price changeStreamLake cut DeepSeek V4 Flash pricing by 6%input −6% ($0.094 → $0.088 per 1M tokens); output −6% ($0.19 → $0.18 per 1M tokens); cache read −6% ($0.019 → $0.018 per 1M tokens)
Aug 2, 2026Price changeIo Net raised DeepSeek V4 Flash pricing by 29%input +29% ($0.14 → $0.18 per 1M tokens); output +22% ($0.28 → $0.34 per 1M tokens); cache read +15% ($0.070 → $0.080 per 1M tokens)
Aug 2, 2026Price changeDeepSeek V4 Flash cut across 2 hosts, by up to 2% at BaiduDeepSeek V4 Flash moved on 2 hosts: Baidu: input −2% ($0.090 → $0.088 per 1M tokens); output −2% ($0.18 → $0.18 per 1M tokens); cache read −2% ($0.018 → $0.018 per 1M tokens); StreamLake: input −3% ($0.090 → $0.087 per 1M tokens); output −3% ($0.18 → $0.17 per 1M tokens); cache read −3% ($0.018 → $0.017 per 1M tokens)
Aug 2, 2026BenchmarkScored 1483.5 on Arena Codingleaderboard
Aug 2, 2026BenchmarkScored 1407.4 on Arena Creative Writingleaderboard
Aug 2, 2026BenchmarkScored 1459.3 on Arena Hard Promptsleaderboard
Aug 2, 2026BenchmarkScored 1428.3 on Arena Instruction Followingleaderboard
Aug 2, 2026BenchmarkScored 1426.6 on Arena Mathsleaderboard
Aug 2, 2026BenchmarkScored 1435.6 on Arena Text (overall)leaderboard

Prices last checked 6h ago

What we do not know about this model yet

  • We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
  • 2 of 32 listings publish no parameter list, so what their API accepts is unknown to us.
  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 3 of 32 listings do not say whether they train on prompts.
05

Licence and identifiers

What the licence allowsMIT License, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.

Licence

MIT License

permissiveCommercial use allowed

Fully permissive: do anything with attribution. No patent grant, unlike Apache-2.0.

Identifiers

Architecture
Mixture of experts
Modality record
text->text
Catalogue slug
deepseek-deepseek-v4-flash

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us