Models / DeepSeek/ DeepSeek V4.1 Flash

DeepSeek V4.1 Flash

DeepSeek · released Sep 10, 2026 · deepseek-ai/DeepSeek-V4.1-Flash

Input: text and images. Output: text.InputOutput
Type
Open weightsMIT License
Params
763B
Context
1M

about 786K words of context

Our take

Written Sep 30, 2026

DeepSeek V4.1 Flash is a downloadable model with a permissive licence, and it places 11th of 168 on Arena Coding via Max as of 25 Sep 2026. At 763.2 billion parameters it is a data-centre job to run yourself, so hosted use is the practical route for most readers.

Who should pick it

Reach for it on agentic coding work, where the model runs inside a harness and the job is to finish a task rather than write a snippet, and on maths-heavy prompts. Long documents can go in whole, though recall across a very long input is unverified in our data. Skip it if you need to run the model on your own machine, or if reliable tool calling and steerability matter more than raw task completion.

The case for it

  • 1st of 58 on LiveBench Agentic Coding via Max as of 25 Jun 2026, a result measured inside an agent harness rather than on standalone code generation.
  • 6th of 163 on Arena Maths via Max as of 25 Sep 2026, a placing drawn from human pairwise votes on maths prompts.
  • 3rd of 55 on Arena Agent · Task outcome via Max as of 25 Sep 2026, so it is a reasonable first pick when the goal is a finished job rather than a draft.

The case against it

  • 763.2 billion parameters in total, with no figure supplied for how many work on any one token, so the memory needed is far beyond a desktop machine.
  • 31st of 55 on Arena Agent · Tool use via Max as of 25 Sep 2026, so check it against your own tool-calling workload before committing.
  • 36th of 55 on Arena Agent · Steerability via Max as of 25 Sep 2026, so plan on more supervision when a task needs redirecting mid-run.
00

How good is it?

An open-weights text model for everyday questions, drafting, coding and multi-step tasks.

Good at
  • getting answers to everyday questionsArena Text (overall) · 21st of 168
  • drafts, rewrites and editingArena Creative Writing · 37th of 168
  • writing and completing codeArena Coding · 11th of 168
  • multi-step work it carries out for youArena Agent · 13th of 55
  • getting back on track after a step failsArena Agent · Recovery · 8th of 55

EverydayGeneral questions and everyday reasoning

4 of 5

Arena Text (overall)21st of 168 · 1477

Arena Hard Prompts 20th of 168Arena Maths 6th of 163LiveBench Data Analysis 12th of 58LiveBench Mathematics 17th of 58LiveBench Reasoning 27th of 58

CodingWriting and fixing code on its own

4.5 of 5

Arena Coding11th of 168 · 1532

Arena Code (WebDev) 13th of 95LiveBench Coding 18th of 58

AgenticPlanning, calling tools, staying on task

3 of 5

Arena Agent13th of 55 · 0.041

LiveBench Agentic Coding 1st of 58

WritingDrafting and rewriting prose

3 of 5

Arena Creative Writing37th of 168 · 1440

LiveBench Language 23rd of 58
How it behaves in an agent loop
Tool usereaches for the right one, and does not invent one31st of 55
Steerabilitydoes what it was asked, and changes course when told36th of 55
Recoverygets back on track after a command fails8th of 55
Task outcomefinishes what the session set out to do3rd of 55

Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.

Other boards it appears on
Arena Instruction Following 14th of 168LiveBench 6th of 58LiveBench Instruction Following 27th of 58Arena Agent · Task outcome 3rd of 55Arena Agent · Recovery 8th of 55Arena Agent · Tool use 31st of 55Arena Agent · Steerability 36th of 55

Boards this model appears on that none of the ratings above are built on.

Every published score for this model20 scoresEvery figure we hold, from 20 boards, with who ran it and a link to the source — including the boards no rating above is built on.
LiveBenchreasoning
81.11source ↗
77.27source ↗
80.04source ↗
79.25source ↗
70.03source ↗
81.19source ↗
93.29source ↗
86.69source ↗
0.041source ↗
0.074source ↗
−0.035source ↗
0.131source ↗
0.002source ↗
1532source ↗
1440source ↗
1501source ↗
1478source ↗
1509source ↗
1477source ↗
1621source ↗
01

Can you run it yourself?

A card many people ownToo large

GeForce RTX 4090 · 24 GB

Weights at 481.2 / 24 GBest
Usable contextNot calculated for spilled setups
Decode speedNot estimated for spilled setups

Too large for this card. The weights do not fit even with part of them offloaded to system memory.

One step upToo large

Apple M1 Pro (16-core GPU) · 32 GB

Weights at 481.2 / 32 GBest
Usable contextNot calculated for spilled setups
Decode speedNot estimated for spilled setups

Too large for this card. The weights do not fit even with part of them offloaded to system memory.

One step downToo large

Radeon RX 7900 XT · 20 GB

Weights at 481.2 / 20 GBest
Usable contextNot calculated for spilled setups
Decode speedNot estimated for spilled setups

Too large for this card. The weights do not fit even with part of them offloaded to system memory.

Your hardware
Checking your profile…

Memory use by level

Against a 24 GB card.

What is quantisation? →
481.2 GBest
Too large
564.5 GBest
Too large
843.3 GBest
Too large
This model on every device we track71 devicesThe Q4 build most people download, on each device: what the weights come to, how much context the memory leaves, and whether it runs. Smallest device that runs it first.
Apple M3 Ultra (80-core GPU)512 GB481.2 GBestnot calculatedSpills to system RAM
Apple M2 Ultra (76-core GPU)192 GB481.2 GBestnot calculatedToo large
B200 (SXM 192GB)192 GB481.2 GBestnot calculatedToo large
Instinct MI300X192 GB481.2 GBestnot calculatedToo large
H200 141GB SXM141 GB481.2 GBestnot calculatedToo large
Apple M1 Ultra (64-core GPU)128 GB481.2 GBestnot calculatedToo large
Apple M3 Max (40-core GPU)128 GB481.2 GBestnot calculatedToo large
Apple M4 Max (40-core GPU)128 GB481.2 GBestnot calculatedToo large
Apple M5 Max (40-core GPU)128 GB481.2 GBestnot calculatedToo large
NVIDIA DGX Spark (GB10)128 GB481.2 GBestnot calculatedToo large
Ryzen AI Max+ 395 (Radeon 8060S)128 GB481.2 GBestnot calculatedToo large
Apple M2 Max (38-core GPU)96 GB481.2 GBestnot calculatedToo large
RTX PRO 6000 Blackwell96 GB481.2 GBestnot calculatedToo large
A100 80GB SXM80 GB481.2 GBestnot calculatedToo large
H100 80GB SXM80 GB481.2 GBestnot calculatedToo large
Apple M1 Max (32-core GPU)64 GB481.2 GBestnot calculatedToo large
Apple M4 Max (32-core GPU)64 GB481.2 GBestnot calculatedToo large
Apple M4 Pro (20-core GPU)64 GB481.2 GBestnot calculatedToo large
Apple M5 Max (32-core GPU)64 GB481.2 GBestnot calculatedToo large
Apple M5 Pro (20-core GPU)64 GB481.2 GBestnot calculatedToo large
L40S48 GB481.2 GBestnot calculatedToo large
RTX 6000 Ada48 GB481.2 GBestnot calculatedToo large
Apple M3 Pro (18-core GPU)36 GB481.2 GBestnot calculatedToo large
Apple M1 Pro (16-core GPU)32 GB481.2 GBestnot calculatedToo large
Apple M2 Pro (19-core GPU)32 GB481.2 GBestnot calculatedToo large
Apple M4 (10-core GPU)32 GB481.2 GBestnot calculatedToo large
Apple M5 (10-core GPU)32 GB481.2 GBestnot calculatedToo large
GeForce RTX 509032 GB481.2 GBestnot calculatedToo large
Apple M2 (10-core GPU)24 GB481.2 GBestnot calculatedToo large
Apple M3 (10-core GPU)24 GB481.2 GBestnot calculatedToo large
GeForce RTX 309024 GB481.2 GBestnot calculatedToo large
GeForce RTX 3090 Ti24 GB481.2 GBestnot calculatedToo large
GeForce RTX 409024 GB481.2 GBestnot calculatedToo large
Radeon RX 7900 XTX24 GB481.2 GBestnot calculatedToo large
Radeon RX 7900 XT20 GB481.2 GBestnot calculatedToo large
Apple M1 (8-core GPU)16 GB481.2 GBestnot calculatedToo large
GeForce RTX 4060 Ti 16GB16 GB481.2 GBestnot calculatedToo large
GeForce RTX 4070 Ti SUPER16 GB481.2 GBestnot calculatedToo large
GeForce RTX 4080 SUPER16 GB481.2 GBestnot calculatedToo large
GeForce RTX 5060 Ti 16GB16 GB481.2 GBestnot calculatedToo large
GeForce RTX 5070 Ti16 GB481.2 GBestnot calculatedToo large
GeForce RTX 508016 GB481.2 GBestnot calculatedToo large
Radeon RX 907016 GB481.2 GBestnot calculatedToo large
Radeon RX 9070 XT16 GB481.2 GBestnot calculatedToo large
Arc B58012 GB481.2 GBestnot calculatedToo large
GeForce RTX 3060 12GB12 GB481.2 GBestnot calculatedToo large
GeForce RTX 4070 SUPER12 GB481.2 GBestnot calculatedToo large
GeForce RTX 507012 GB481.2 GBestnot calculatedToo large
Arc B57010 GB481.2 GBestnot calculatedToo large
GeForce RTX 3080 10GB10 GB481.2 GBestnot calculatedToo large
Android phone · 16 GB · 2024 or newer8 GB481.2 GBestnot calculatedToo large
Apple M1 (8-core GPU, 8GB unified)8 GB481.2 GBestnot calculatedToo large
Apple M2 (8-core GPU, 8GB unified)8 GB481.2 GBestnot calculatedToo large
GeForce RTX 3060 8GB8 GB481.2 GBestnot calculatedToo large
GeForce RTX 4060 8GB8 GB481.2 GBestnot calculatedToo large
Radeon RX 66008 GB481.2 GBestnot calculatedToo large
iPhone 17 Pro6.6 GB481.2 GBestnot calculatedToo large
Android phone · 12 GB · 2023 or newer6 GB481.2 GBestnot calculatedToo large
GeForce GTX 1660 SUPER6 GB481.2 GBestnot calculatedToo large
iPhone 15 Pro4.4 GB481.2 GBestnot calculatedToo large
iPhone 164.4 GB481.2 GBestnot calculatedToo large
iPhone 16 Pro4.4 GB481.2 GBestnot calculatedToo large
iPhone 174.4 GB481.2 GBestnot calculatedToo large
Android phone · 8 GB · 2020–20224 GB481.2 GBestnot calculatedToo large
Android phone · 8 GB · 2023 or newer4 GB481.2 GBestnot calculatedToo large
iPhone 143.3 GB481.2 GBestnot calculatedToo large
iPhone 153.3 GB481.2 GBestnot calculatedToo large
Android phone · 6 GB3 GB481.2 GBestnot calculatedToo large
iPhone 132.2 GB481.2 GBestnot calculatedToo large
iPhone SE (3rd gen)2.2 GB481.2 GBestnot calculatedToo large
Android phone · 4 GB2 GB481.2 GBestnot calculatedToo large

Check against your own machine → · Where to rent it hosted →

02

Or rent it from someone else

Prices checked between 1 hour and 37 hours ago — each listing carries its own date.

Some hosts sell this model at two prices: on their own price list (“direct”) and on their OpenRouter listing (“through OpenRouter”). Where the two differ, the row shows both, each with the date we last read it.

Cheapest published offer

DekaLLM, through OpenRouter

Cheapest of 33 live listings.

per 1M tokens
$0.12 in / $0.40 out
Context served
1M
Throughput
~43 tok/s
Current provider offers with price, context and prompt-privacy answers
ProviderIn / out per 1M tokensContextThroughputTrains on promptsLogs promptsZero retention
Sail Researchfp4Through OpenRouter$0.080 / $0.40checked 1 hour ago1M384K max reply49 tok/sNoNoConfirmed
DekaLLMThrough OpenRouter$0.12 / $0.40checked 1 hour ago1M944K max reply43 tok/sNoNoConfirmed
Io Netfp8Through OpenRouter$0.12 / $0.41checked 1 hour ago1M131K max reply73 tok/sNoNoConfirmed
DeepInfrafp8Direct and through OpenRouter$0.20 / $0.60directchecked 1 hour ago$0.14 / $0.42through OpenRouterchecked 1 hour ago1M131K max reply through OpenRouter75 tok/sthrough OpenRouterDirectUnknownThrough OpenRouterNoDirectUnknownThrough OpenRouterNoDirectUnknownThrough OpenRouterConfirmed
Morphfp8Through OpenRouter$0.053 / $0.43checked 1 hour ago1M944K max reply60 tok/sNoNoConfirmed
OpenInferencefp4Through OpenRouter$0.030 / $0.50checked 1 hour ago1M944K max reply7 tok/sNoNoConfirmed
AtlasCloudfp8Through OpenRouter$0.14 / $0.56checked 1 hour ago1M393K max reply111 tok/sNoYesunknown periodUnknown
StreamLakefp8Through OpenRouter$0.14 / $0.56checked 1 hour ago1M384K max reply75 tok/sNoYesunknown periodUnknown
RekaThrough OpenRouter$0.14 / $0.56checked 1 hour ago1M944K max reply51 tok/sNoNoConfirmed
WaferThrough OpenRouter$0.050 / $0.60checked 1 hour ago1M944K max reply117 tok/sNoNoConfirmed
InferenceNetThrough OpenRouter$0.040 / $0.60checked 1 hour ago1M384K max reply86 tok/sNoNoConfirmed
DeepSeekThrough OpenRouter$0.15 / $0.60checked 7 hours ago1M393K max reply82 tok/sYesYesunknown periodUnknown
RelaceThrough OpenRouter$0.026 / $0.60checked 1 hour ago1M944K max reply33 tok/sNoNoConfirmed
CoreWeavefp8Through OpenRouter$0.20 / $0.65checked 7 hours ago1M944K max reply145 tok/sNoNoConfirmed
Fireworks AIThrough OpenRouter$0.22 / $0.66checked 1 hour ago1M944K max reply68 tok/sNoNoConfirmed
DigitalOcean GradientThrough OpenRouter$0.18 / $0.72checked 1 hour ago1M944K max reply53 tok/sNoNoConfirmed
GMICloudfp8Through OpenRouter$0.18 / $0.72checked 1 hour ago1M944K max reply98 tok/sNoYesunknown periodUnknown
NextBitfp8Through OpenRouter$0.21 / $0.84checked 1 hour ago1M944K max reply78 tok/sNoNoConfirmed
PhalaThrough OpenRouter$0.21 / $0.84checked 1 hour ago1M944K max reply87 tok/sNoNoConfirmed
IonstreamThrough OpenRouter$0.28 / $1.15checked 37 hours ago1M944K max reply157 tok/sNoNoConfirmed
Baidufp8Through OpenRouter$0.30 / $1.20checked 1 hour ago1M393K max reply162 tok/sNoYesunknown periodUnknown
Parasailfp8Through OpenRouter$0.30 / $1.20checked 1 hour ago1M944K max reply131 tok/sNoNoConfirmed
OpenRouterOpenRouter's own listing$0.30 / $1.20checked 37 hours ago1Mnot measuredUnknownUnknownUnknown
Basetenfp8Through OpenRouter$0.30 / $1.20checked 1 hour ago1M262K max reply79 tok/sNoNoConfirmed
Makorafp8Through OpenRouter$0.30 / $1.20checked 1 hour ago1M393K max reply125 tok/sNoNoConfirmed
Novita AIfp8Direct and through OpenRouter$0.30 / $1.20checked 1 hour ago1M393K max reply through OpenRouter78 tok/sthrough OpenRouterDirectUnknownThrough OpenRouterNoDirectUnknownThrough OpenRouterNoDirectUnknownThrough OpenRouterConfirmed
Alibaba CloudThrough OpenRouter$0.30 / $1.20checked 13 hours ago1M393K max reply65 tok/sNoYesunknown periodUnknown
SiliconFlowfp8Through OpenRouter$0.30 / $1.20checked 1 hour ago1M393K max reply93 tok/sNoNoConfirmed
Together AIThrough OpenRouter$0.30 / $1.20checked 1 hour ago1M944K max reply234 tok/sNoNoConfirmed
ModalThrough OpenRouter$0.30 / $1.20checked 1 hour ago1M944K max reply182 tok/sNoNoConfirmed
Venice AIfp8Through OpenRouter$0.38 / $1.50checked 1 hour ago1M131K max reply101 tok/sNoNoConfirmed
Fireworks AIusThrough OpenRouter$0.45 / $1.80checked 1 hour ago1M944K max reply173 tok/sNoNoConfirmed
Basetenfast tierfp32Through OpenRouter$0.60 / $2.40checked 1 hour ago1M131K max reply97 tok/sNoNoUnknown

Across the 33 listings we hold: 31 say they do not train on prompts (2 of them only through OpenRouter), 1 says it does and 1 does not say. 25 appear in the zero-retention registry we check (2 of them only through OpenRouter); the rest are unknown to us.

What each host's API supports

From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.

API features per host
ProviderTool callingJSON outputStrict schema
Sail Researchfp4Through OpenRouter✓✓✓
DekaLLMThrough OpenRouter✓✓✓
Io Netfp8Through OpenRouter✓✓✓
DeepInfrafp8Direct and through OpenRouter✓✓✓
Morphfp8Through OpenRouter✓✓✓
OpenInferencefp4Through OpenRouter✓✓✓
AtlasCloudfp8Through OpenRouter✓✓✓
StreamLakefp8Through OpenRouter✓✓✗
RekaThrough OpenRouter✓✓✓
WaferThrough OpenRouter✓✓✓
InferenceNetThrough OpenRouter✓✓✓
DeepSeekThrough OpenRouter✓✓✗
RelaceThrough OpenRouter✓✗✗
CoreWeavefp8Through OpenRouter✓✓✓
Fireworks AIThrough OpenRouter✓✓✓
DigitalOcean GradientThrough OpenRouter✓✓✓
GMICloudfp8Through OpenRouter✓✓✗
NextBitfp8Through OpenRouter✓✓✓
PhalaThrough OpenRouter✓✓✗
IonstreamThrough OpenRouter✓✓✓
Baidufp8Through OpenRouter✓✓✓
Parasailfp8Through OpenRouter✓✓✓
OpenRouterOpenRouter's own listing✓✓✓
Basetenfp8Through OpenRouter✓✓✓
Makorafp8Through OpenRouter✓✓✓
Novita AIfp8Direct and through OpenRouter✓✓✗
Alibaba CloudThrough OpenRouter✓✓✗
SiliconFlowfp8Through OpenRouter✓✓✗
Together AIThrough OpenRouter✓✓✓
ModalThrough OpenRouter✓✓✓
Venice AIfp8Through OpenRouter✓✓✓
Fireworks AIusThrough OpenRouter✓✓✓
Basetenfast · fp32Through OpenRouter✓✓✓

Tool calling: 33 of 33 listings say yes. JSON output: 32 of 33 listings say yes, 1 says no. Strict schema: 25 of 33 listings say yes, 8 say no.

03

Models people weigh against DeepSeek V4.1 Flash

04

When we formed this view

Recent changes

Oct 1, 2026Price changeDeepSeek V4.1 Flash repriced across 8 hosts: DigitalOcean down 40% on all rates, Wafer output up 36%
What movedDeepSeek V4.1 Flash moved on 8 hosts: DigitalOcean: input −40% ($0.30 → $0.18 per 1M tokens), output −40% ($1.20 → $0.72 per 1M tokens), cache read −40% ($0.0060 → $0.0036 per 1M tokens); Wafer: input −33% ($0.075 → $0.050 per 1M tokens), output +36% ($0.44 → $0.60 per 1M tokens), cache read −11% ($0.045 → $0.040 per 1M tokens); Io Net: input +33% ($0.090 → $0.120 per 1M tokens), cache read +11% ($0.009 → $0.010 per 1M tokens); Reka: input −30% ($0.20 → $0.14 per 1M tokens), output −33% ($0.84 → $0.56 per 1M tokens), cache read +76% ($0.0080 → $0.0141 per 1M tokens); Morph: input −32% ($0.078 → $0.053 per 1M tokens), output −14% ($0.50 → $0.43 per 1M tokens), cache read −50% ($0.010 → $0.005 per 1M tokens); Relace: input +32% ($0.0200 → $0.0264 per 1M tokens), cache read +32% ($0.0200 → $0.0264 per 1M tokens); AtlasCloud: input +24% ($0.114 → $0.141 per 1M tokens), output +24% ($0.456 → $0.564 per 1M tokens), cache read +24% ($0.0114 → $0.0141 per 1M tokens); OpenInference: input −22% ($0.0198 → $0.0155 per 1M tokens)
Sep 30, 2026Price changeDeepSeek V4.1 Flash repriced across 6 hosts: Io Net input down 64%, Ionstream input up 97%
What movedDeepSeek V4.1 Flash moved on 6 hosts: Ionstream: input +97% ($0.145 → $0.285 per 1M tokens); InferenceNet: input −42% ($0.069 → $0.040 per 1M tokens), output +67% ($0.45 → $0.75 per 1M tokens); Io Net: input −64% ($0.250 → $0.090 per 1M tokens), output −59% ($1.00 → $0.41 per 1M tokens), cache read −70% ($0.030 → $0.009 per 1M tokens); OpenInference: input −34% ($0.0300 → $0.0198 per 1M tokens), output −21% ($0.500 → $0.396 per 1M tokens), cache read −71% ($0.0100 → $0.0029 per 1M tokens); Wafer: input −25% ($0.100 → $0.075 per 1M tokens); Morph: input −3% ($0.0805 → $0.0780 per 1M tokens), output −1% ($0.510 → $0.504 per 1M tokens), cache read +270% ($0.0027 → $0.0100 per 1M tokens)
Sep 29, 2026Price changeDeepSeek V4.1 Flash repriced across 7 hosts: OpenInference input down 70%, InferenceNet input up 97%
What movedDeepSeek V4.1 Flash moved on 7 hosts: InferenceNet: input +97% ($0.035 → $0.069 per 1M tokens), output +55% ($0.29 → $0.45 per 1M tokens), cache read +1900% ($0.001 → $0.020 per 1M tokens); OpenInference: input −70% ($0.100 → $0.030 per 1M tokens); Morph: output +64% ($0.3101 → $0.5100 per 1M tokens), cache read +74% ($0.00155 → $0.00270 per 1M tokens); Relace: input −60% ($0.050 → $0.020 per 1M tokens), cache read +100% ($0.010 → $0.020 per 1M tokens); Ionstream: input −48% ($0.280 → $0.145 per 1M tokens); Wafer: input +1% ($0.099 → $0.100 per 1M tokens), output −37% ($0.70 → $0.44 per 1M tokens); Phala: input −13% ($0.345 → $0.300 per 1M tokens), output −13% ($1.38 → $1.20 per 1M tokens), cache read −13% ($0.0069 → $0.0060 per 1M tokens)
Sep 28, 2026Price changeDeepSeek V4.1 Flash repriced across 3 hosts: AtlasCloud down 62% on all rates, Wafer input up 136%
What movedDeepSeek V4.1 Flash moved on 3 hosts: Wafer: input +136% ($0.042 → $0.099 per 1M tokens), output +17% ($0.60 → $0.70 per 1M tokens), cache read +19% ($0.0378 → $0.0450 per 1M tokens); AtlasCloud: input −62% ($0.300 → $0.114 per 1M tokens), output −62% ($1.20 → $0.46 per 1M tokens), cache read −62% ($0.0300 → $0.0114 per 1M tokens); Morph: input −46% ($0.150 → $0.081 per 1M tokens), output −48% ($0.60 → $0.31 per 1M tokens), cache read −48% ($0.00300 → $0.00155 per 1M tokens)
Sep 27, 2026Price changeDeepSeek V4.1 Flash repriced across 3 hosts: DekaLLM output down 33%, Relace output up 50%
What movedDeepSeek V4.1 Flash moved on 3 hosts: Relace: output +50% ($0.40 → $0.60 per 1M tokens), cache read +100% ($0.005 → $0.010 per 1M tokens); DekaLLM: input −20% ($0.15 → $0.12 per 1M tokens), output −33% ($0.60 → $0.40 per 1M tokens); Wafer: input −14% ($0.049 → $0.042 per 1M tokens), cache read −16% ($0.045 → $0.038 per 1M tokens)
Sep 26, 2026Price changeDeepSeek V4.1 Flash cut across 2 hosts, by up to 47% at Sail Research (output) · machine-readable source ↗
What movedDeepSeek V4.1 Flash moved on 2 hosts: Sail Research: input −38% ($0.130 → $0.080 per 1M tokens), output −47% ($0.75 → $0.40 per 1M tokens); Relace: input −44% ($0.090 → $0.050 per 1M tokens), output −11% ($0.45 → $0.40 per 1M tokens), cache read −44% ($0.009 → $0.005 per 1M tokens)
Sep 25, 2026Price changeDeepSeek V4.1 Flash repriced across 2 hosts: DekaLLM output down 40%, input up 275%
What movedDeepSeek V4.1 Flash moved on 2 hosts: DekaLLM: input +275% ($0.040 → $0.150 per 1M tokens), output −40% ($1.00 → $0.60 per 1M tokens), cache read −50% ($0.010 → $0.005 per 1M tokens); NextBit: input −30% ($0.30 → $0.21 per 1M tokens), output −30% ($1.20 → $0.84 per 1M tokens), cache read −33% ($0.006 → $0.004 per 1M tokens)
Sep 25, 2026BenchmarkScored 0.041 via Max on Arena Agent
What movedleaderboard
Sep 25, 2026BenchmarkScored 0.074 via Max on Arena Agent · Recovery
What movedleaderboard
Sep 25, 2026BenchmarkScored −0.035 via Max on Arena Agent · Steerability
What movedleaderboard

Each date is the day we first saw the change, or the day the maker announced it.

What we do not know about this model yet

  • We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
  • Nothing we hold says whether an endpoint streams, so we do not show it either way.
  • 1 of 33 listings does not say whether it trains on prompts, and 2 answer only through OpenRouter, not for their own listing.
  • We hold no batch or off-peak rate for any of its listings.
  • We hold a decode speed for it, but no prompt-processing (prefill) figure, so how long the input side of a job takes is unknown to us.
05

Licence and identifiers

What the licence allowsMIT License, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.

Licence

MIT License

Open, few conditionsCommercial use allowed

Fully permissive: do anything with attribution. No patent grant, unlike Apache-2.0.

Identifiers

Architecture
Dense
Takes in, gives back
Text and images in, text out
Catalogue slug
deepseek-deepseek-v4-1-flash

Machine-readable model card (omc.json) →

Something wrong on this page? Tell us