DeepSeek V4.1 Flash
DeepSeek · released Sep 10, 2026 · deepseek-ai/DeepSeek-V4.1-Flash
- Type
- Open weightsMIT License
- Params
- 763B
- Context
- 1M
about 786K words of context
Our take
Written Sep 30, 2026DeepSeek V4.1 Flash is a downloadable model with a permissive licence, and it places 11th of 168 on Arena Coding via Max as of 25 Sep 2026. At 763.2 billion parameters it is a data-centre job to run yourself, so hosted use is the practical route for most readers.
Reach for it on agentic coding work, where the model runs inside a harness and the job is to finish a task rather than write a snippet, and on maths-heavy prompts. Long documents can go in whole, though recall across a very long input is unverified in our data. Skip it if you need to run the model on your own machine, or if reliable tool calling and steerability matter more than raw task completion.
The case for it
- 1st of 58 on LiveBench Agentic Coding via Max as of 25 Jun 2026, a result measured inside an agent harness rather than on standalone code generation.
- 6th of 163 on Arena Maths via Max as of 25 Sep 2026, a placing drawn from human pairwise votes on maths prompts.
- 3rd of 55 on Arena Agent · Task outcome via Max as of 25 Sep 2026, so it is a reasonable first pick when the goal is a finished job rather than a draft.
The case against it
- 763.2 billion parameters in total, with no figure supplied for how many work on any one token, so the memory needed is far beyond a desktop machine.
- 31st of 55 on Arena Agent · Tool use via Max as of 25 Sep 2026, so check it against your own tool-calling workload before committing.
- 36th of 55 on Arena Agent · Steerability via Max as of 25 Sep 2026, so plan on more supervision when a task needs redirecting mid-run.
How good is it?
An open-weights text model for everyday questions, drafting, coding and multi-step tasks.
- getting answers to everyday questionsArena Text (overall) · 21st of 168
- drafts, rewrites and editingArena Creative Writing · 37th of 168
- writing and completing codeArena Coding · 11th of 168
- multi-step work it carries out for youArena Agent · 13th of 55
- getting back on track after a step failsArena Agent · Recovery · 8th of 55
EverydayGeneral questions and everyday reasoning
Arena Text (overall)21st of 168 · 1477
CodingWriting and fixing code on its own
Arena Coding11th of 168 · 1532
AgenticPlanning, calling tools, staying on task
Arena Agent13th of 55 · 0.041
WritingDrafting and rewriting prose
Arena Creative Writing37th of 168 · 1440
Placings on Arena's agent boards, from live sessions people ran themselves. A model can lead on one of these and sit mid-field on the others.
Boards this model appears on that none of the ratings above are built on.
Every published score for this model20 scoresEvery figure we hold, from 20 boards, with who ran it and a link to the source — including the boards no rating above is built on.
Can you run it yourself?
GeForce RTX 4090 · 24 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Apple M1 Pro (16-core GPU) · 32 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Radeon RX 7900 XT · 20 GB
Too large for this card. The weights do not fit even with part of them offloaded to system memory.
Memory use by level
Against a 24 GB card.
What is quantisation? →This model on every device we track71 devicesThe Q4 build most people download, on each device: what the weights come to, how much context the memory leaves, and whether it runs. Smallest device that runs it first.
Check against your own machine → · Where to rent it hosted →
Or rent it from someone else
Prices checked between 1 hour and 37 hours ago — each listing carries its own date.
Some hosts sell this model at two prices: on their own price list (“direct”) and on their OpenRouter listing (“through OpenRouter”). Where the two differ, the row shows both, each with the date we last read it.
- per 1M tokens
- $0.12 in / $0.40 out
- Context served
- 1M
- Throughput
- ~43 tok/s
| Provider | In / out per 1M tokens | Context | Throughput | Trains on prompts | Logs prompts | Zero retention |
|---|---|---|---|---|---|---|
| Sail Researchfp4Through OpenRouter | $0.080 / $0.40checked 1 hour ago | 1M384K max reply | 49 tok/s | No | No | Confirmed |
| DekaLLMThrough OpenRouter | $0.12 / $0.40checked 1 hour ago | 1M944K max reply | 43 tok/s | No | No | Confirmed |
| Io Netfp8Through OpenRouter | $0.12 / $0.41checked 1 hour ago | 1M131K max reply | 73 tok/s | No | No | Confirmed |
| DeepInfrafp8Direct and through OpenRouter | $0.20 / $0.60directchecked 1 hour ago$0.14 / $0.42through OpenRouterchecked 1 hour ago | 1M131K max reply through OpenRouter | 75 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| Morphfp8Through OpenRouter | $0.053 / $0.43checked 1 hour ago | 1M944K max reply | 60 tok/s | No | No | Confirmed |
| OpenInferencefp4Through OpenRouter | $0.030 / $0.50checked 1 hour ago | 1M944K max reply | 7 tok/s | No | No | Confirmed |
| AtlasCloudfp8Through OpenRouter | $0.14 / $0.56checked 1 hour ago | 1M393K max reply | 111 tok/s | No | Yesunknown period | Unknown |
| StreamLakefp8Through OpenRouter | $0.14 / $0.56checked 1 hour ago | 1M384K max reply | 75 tok/s | No | Yesunknown period | Unknown |
| RekaThrough OpenRouter | $0.14 / $0.56checked 1 hour ago | 1M944K max reply | 51 tok/s | No | No | Confirmed |
| WaferThrough OpenRouter | $0.050 / $0.60checked 1 hour ago | 1M944K max reply | 117 tok/s | No | No | Confirmed |
| InferenceNetThrough OpenRouter | $0.040 / $0.60checked 1 hour ago | 1M384K max reply | 86 tok/s | No | No | Confirmed |
| DeepSeekThrough OpenRouter | $0.15 / $0.60checked 7 hours ago | 1M393K max reply | 82 tok/s | Yes | Yesunknown period | Unknown |
| RelaceThrough OpenRouter | $0.026 / $0.60checked 1 hour ago | 1M944K max reply | 33 tok/s | No | No | Confirmed |
| CoreWeavefp8Through OpenRouter | $0.20 / $0.65checked 7 hours ago | 1M944K max reply | 145 tok/s | No | No | Confirmed |
| Fireworks AIThrough OpenRouter | $0.22 / $0.66checked 1 hour ago | 1M944K max reply | 68 tok/s | No | No | Confirmed |
| DigitalOcean GradientThrough OpenRouter | $0.18 / $0.72checked 1 hour ago | 1M944K max reply | 53 tok/s | No | No | Confirmed |
| GMICloudfp8Through OpenRouter | $0.18 / $0.72checked 1 hour ago | 1M944K max reply | 98 tok/s | No | Yesunknown period | Unknown |
| NextBitfp8Through OpenRouter | $0.21 / $0.84checked 1 hour ago | 1M944K max reply | 78 tok/s | No | No | Confirmed |
| PhalaThrough OpenRouter | $0.21 / $0.84checked 1 hour ago | 1M944K max reply | 87 tok/s | No | No | Confirmed |
| IonstreamThrough OpenRouter | $0.28 / $1.15checked 37 hours ago | 1M944K max reply | 157 tok/s | No | No | Confirmed |
| Baidufp8Through OpenRouter | $0.30 / $1.20checked 1 hour ago | 1M393K max reply | 162 tok/s | No | Yesunknown period | Unknown |
| Parasailfp8Through OpenRouter | $0.30 / $1.20checked 1 hour ago | 1M944K max reply | 131 tok/s | No | No | Confirmed |
| OpenRouterOpenRouter's own listing | $0.30 / $1.20checked 37 hours ago | 1M | not measured | Unknown | Unknown | Unknown |
| Basetenfp8Through OpenRouter | $0.30 / $1.20checked 1 hour ago | 1M262K max reply | 79 tok/s | No | No | Confirmed |
| Makorafp8Through OpenRouter | $0.30 / $1.20checked 1 hour ago | 1M393K max reply | 125 tok/s | No | No | Confirmed |
| Novita AIfp8Direct and through OpenRouter | $0.30 / $1.20checked 1 hour ago | 1M393K max reply through OpenRouter | 78 tok/sthrough OpenRouter | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterNo | DirectUnknownThrough OpenRouterConfirmed |
| Alibaba CloudThrough OpenRouter | $0.30 / $1.20checked 13 hours ago | 1M393K max reply | 65 tok/s | No | Yesunknown period | Unknown |
| SiliconFlowfp8Through OpenRouter | $0.30 / $1.20checked 1 hour ago | 1M393K max reply | 93 tok/s | No | No | Confirmed |
| Together AIThrough OpenRouter | $0.30 / $1.20checked 1 hour ago | 1M944K max reply | 234 tok/s | No | No | Confirmed |
| ModalThrough OpenRouter | $0.30 / $1.20checked 1 hour ago | 1M944K max reply | 182 tok/s | No | No | Confirmed |
| Venice AIfp8Through OpenRouter | $0.38 / $1.50checked 1 hour ago | 1M131K max reply | 101 tok/s | No | No | Confirmed |
| Fireworks AIusThrough OpenRouter | $0.45 / $1.80checked 1 hour ago | 1M944K max reply | 173 tok/s | No | No | Confirmed |
| Basetenfast tierfp32Through OpenRouter | $0.60 / $2.40checked 1 hour ago | 1M131K max reply | 97 tok/s | No | No | Unknown |
Across the 33 listings we hold: 31 say they do not train on prompts (2 of them only through OpenRouter), 1 says it does and 1 does not say. 25 appear in the zero-retention registry we check (2 of them only through OpenRouter); the rest are unknown to us.
What each host's API supports
From the parameter list each endpoint publishes. Streaming is omitted: nothing we hold reports it, for any model.
| Provider | Tool calling | JSON output | Strict schema |
|---|---|---|---|
| Sail Researchfp4Through OpenRouter | ✓ | ✓ | ✓ |
| DekaLLMThrough OpenRouter | ✓ | ✓ | ✓ |
| Io Netfp8Through OpenRouter | ✓ | ✓ | ✓ |
| DeepInfrafp8Direct and through OpenRouter | ✓ | ✓ | ✓ |
| Morphfp8Through OpenRouter | ✓ | ✓ | ✓ |
| OpenInferencefp4Through OpenRouter | ✓ | ✓ | ✓ |
| AtlasCloudfp8Through OpenRouter | ✓ | ✓ | ✓ |
| StreamLakefp8Through OpenRouter | ✓ | ✓ | ✗ |
| RekaThrough OpenRouter | ✓ | ✓ | ✓ |
| WaferThrough OpenRouter | ✓ | ✓ | ✓ |
| InferenceNetThrough OpenRouter | ✓ | ✓ | ✓ |
| DeepSeekThrough OpenRouter | ✓ | ✓ | ✗ |
| RelaceThrough OpenRouter | ✓ | ✗ | ✗ |
| CoreWeavefp8Through OpenRouter | ✓ | ✓ | ✓ |
| Fireworks AIThrough OpenRouter | ✓ | ✓ | ✓ |
| DigitalOcean GradientThrough OpenRouter | ✓ | ✓ | ✓ |
| GMICloudfp8Through OpenRouter | ✓ | ✓ | ✗ |
| NextBitfp8Through OpenRouter | ✓ | ✓ | ✓ |
| PhalaThrough OpenRouter | ✓ | ✓ | ✗ |
| IonstreamThrough OpenRouter | ✓ | ✓ | ✓ |
| Baidufp8Through OpenRouter | ✓ | ✓ | ✓ |
| Parasailfp8Through OpenRouter | ✓ | ✓ | ✓ |
| OpenRouterOpenRouter's own listing | ✓ | ✓ | ✓ |
| Basetenfp8Through OpenRouter | ✓ | ✓ | ✓ |
| Makorafp8Through OpenRouter | ✓ | ✓ | ✓ |
| Novita AIfp8Direct and through OpenRouter | ✓ | ✓ | ✗ |
| Alibaba CloudThrough OpenRouter | ✓ | ✓ | ✗ |
| SiliconFlowfp8Through OpenRouter | ✓ | ✓ | ✗ |
| Together AIThrough OpenRouter | ✓ | ✓ | ✓ |
| ModalThrough OpenRouter | ✓ | ✓ | ✓ |
| Venice AIfp8Through OpenRouter | ✓ | ✓ | ✓ |
| Fireworks AIusThrough OpenRouter | ✓ | ✓ | ✓ |
| Basetenfast · fp32Through OpenRouter | ✓ | ✓ | ✓ |
Tool calling: 33 of 33 listings say yes. JSON output: 32 of 33 listings say yes, 1 says no. Strict schema: 25 of 33 listings say yes, 8 say no.
Models people weigh against DeepSeek V4.1 Flash
When we formed this view
Recent changes
What moved
DeepSeek V4.1 Flash moved on 8 hosts: DigitalOcean: input −40% ($0.30 → $0.18 per 1M tokens), output −40% ($1.20 → $0.72 per 1M tokens), cache read −40% ($0.0060 → $0.0036 per 1M tokens); Wafer: input −33% ($0.075 → $0.050 per 1M tokens), output +36% ($0.44 → $0.60 per 1M tokens), cache read −11% ($0.045 → $0.040 per 1M tokens); Io Net: input +33% ($0.090 → $0.120 per 1M tokens), cache read +11% ($0.009 → $0.010 per 1M tokens); Reka: input −30% ($0.20 → $0.14 per 1M tokens), output −33% ($0.84 → $0.56 per 1M tokens), cache read +76% ($0.0080 → $0.0141 per 1M tokens); Morph: input −32% ($0.078 → $0.053 per 1M tokens), output −14% ($0.50 → $0.43 per 1M tokens), cache read −50% ($0.010 → $0.005 per 1M tokens); Relace: input +32% ($0.0200 → $0.0264 per 1M tokens), cache read +32% ($0.0200 → $0.0264 per 1M tokens); AtlasCloud: input +24% ($0.114 → $0.141 per 1M tokens), output +24% ($0.456 → $0.564 per 1M tokens), cache read +24% ($0.0114 → $0.0141 per 1M tokens); OpenInference: input −22% ($0.0198 → $0.0155 per 1M tokens)What moved
DeepSeek V4.1 Flash moved on 6 hosts: Ionstream: input +97% ($0.145 → $0.285 per 1M tokens); InferenceNet: input −42% ($0.069 → $0.040 per 1M tokens), output +67% ($0.45 → $0.75 per 1M tokens); Io Net: input −64% ($0.250 → $0.090 per 1M tokens), output −59% ($1.00 → $0.41 per 1M tokens), cache read −70% ($0.030 → $0.009 per 1M tokens); OpenInference: input −34% ($0.0300 → $0.0198 per 1M tokens), output −21% ($0.500 → $0.396 per 1M tokens), cache read −71% ($0.0100 → $0.0029 per 1M tokens); Wafer: input −25% ($0.100 → $0.075 per 1M tokens); Morph: input −3% ($0.0805 → $0.0780 per 1M tokens), output −1% ($0.510 → $0.504 per 1M tokens), cache read +270% ($0.0027 → $0.0100 per 1M tokens)What moved
DeepSeek V4.1 Flash moved on 7 hosts: InferenceNet: input +97% ($0.035 → $0.069 per 1M tokens), output +55% ($0.29 → $0.45 per 1M tokens), cache read +1900% ($0.001 → $0.020 per 1M tokens); OpenInference: input −70% ($0.100 → $0.030 per 1M tokens); Morph: output +64% ($0.3101 → $0.5100 per 1M tokens), cache read +74% ($0.00155 → $0.00270 per 1M tokens); Relace: input −60% ($0.050 → $0.020 per 1M tokens), cache read +100% ($0.010 → $0.020 per 1M tokens); Ionstream: input −48% ($0.280 → $0.145 per 1M tokens); Wafer: input +1% ($0.099 → $0.100 per 1M tokens), output −37% ($0.70 → $0.44 per 1M tokens); Phala: input −13% ($0.345 → $0.300 per 1M tokens), output −13% ($1.38 → $1.20 per 1M tokens), cache read −13% ($0.0069 → $0.0060 per 1M tokens)What moved
DeepSeek V4.1 Flash moved on 3 hosts: Wafer: input +136% ($0.042 → $0.099 per 1M tokens), output +17% ($0.60 → $0.70 per 1M tokens), cache read +19% ($0.0378 → $0.0450 per 1M tokens); AtlasCloud: input −62% ($0.300 → $0.114 per 1M tokens), output −62% ($1.20 → $0.46 per 1M tokens), cache read −62% ($0.0300 → $0.0114 per 1M tokens); Morph: input −46% ($0.150 → $0.081 per 1M tokens), output −48% ($0.60 → $0.31 per 1M tokens), cache read −48% ($0.00300 → $0.00155 per 1M tokens)What moved
DeepSeek V4.1 Flash moved on 3 hosts: Relace: output +50% ($0.40 → $0.60 per 1M tokens), cache read +100% ($0.005 → $0.010 per 1M tokens); DekaLLM: input −20% ($0.15 → $0.12 per 1M tokens), output −33% ($0.60 → $0.40 per 1M tokens); Wafer: input −14% ($0.049 → $0.042 per 1M tokens), cache read −16% ($0.045 → $0.038 per 1M tokens)What moved
DeepSeek V4.1 Flash moved on 2 hosts: Sail Research: input −38% ($0.130 → $0.080 per 1M tokens), output −47% ($0.75 → $0.40 per 1M tokens); Relace: input −44% ($0.090 → $0.050 per 1M tokens), output −11% ($0.45 → $0.40 per 1M tokens), cache read −44% ($0.009 → $0.005 per 1M tokens)What moved
DeepSeek V4.1 Flash moved on 2 hosts: DekaLLM: input +275% ($0.040 → $0.150 per 1M tokens), output −40% ($1.00 → $0.60 per 1M tokens), cache read −50% ($0.010 → $0.005 per 1M tokens); NextBit: input −30% ($0.30 → $0.21 per 1M tokens), output −30% ($1.20 → $0.84 per 1M tokens), cache read −33% ($0.006 → $0.004 per 1M tokens)Each date is the day we first saw the change, or the day the maker announced it.
What we do not know about this model yet
- We hold no measured file for it, so all 3 sizes on this page are calculated from the parameter count.
- Nothing we hold says whether an endpoint streams, so we do not show it either way.
- 1 of 33 listings does not say whether it trains on prompts, and 2 answer only through OpenRouter, not for their own listing.
- We hold no batch or off-peak rate for any of its listings.
- We hold a decode speed for it, but no prompt-processing (prefill) figure, so how long the input side of a job takes is unknown to us.
Licence and identifiers
What the licence allowsMIT License, what it allows commercially, and the identifiers you need to pull this model — its Hugging Face repo, our slug and a machine-readable card.
Licence
MIT License
Fully permissive: do anything with attribution. No patent grant, unlike Apache-2.0.
Identifiers
- Hugging Face
- deepseek-ai/DeepSeek-V4.1-Flash
- Architecture
- Dense
- Takes in, gives back
- Text and images in, text out
- Catalogue slug
- deepseek-deepseek-v4-1-flash