DeepSeek V4.1 Flash repriced across 2 hosts: Relace input down 8%, Wafer input up 35%
Published 23 September 2026
Price move, per 1M tokens
- Wafer · input↑ 35%$0.125 → $0.169
- Wafer · outputunchanged$0.60 → $0.60
- Wafer · cache read↑ 20%$0.050 → $0.060
- Relace · input↓ 8%$0.13 → $0.12
- Relace · outputunchanged$0.52 → $0.52
- Relace · cache readunchanged$0.010 → $0.010
| host | rate | was | now | change |
|---|---|---|---|---|
| Wafer | input | $0.125 | $0.169 | ↑ 35% |
| Wafer | output | $0.60 | $0.60 | — |
| Wafer | cache read | $0.050 | $0.060 | ↑ 20% |
| Relace | input | $0.13 | $0.12 | ↓ 8% |
| Relace | output | $0.52 | $0.52 | — |
| Relace | cache read | $0.010 | $0.010 | — |
Wafer input: ↑ 35%Relace input: ↓ 8%
If you buy
Relace now costs less to run for input-heavy work, so check whether your usage is mostly fresh prompts or cached reads before your next top-up there.
The model
DeepSeek V4.1 Flashopen the model →
Open weights763.2B parameters1049K context35 hosts
intelligence21st of 168 writing37th of 168 coding11th of 168 agents13th of 55
DeepSeek V4.1 Flash is a downloadable model with a permissive licence, and it places 11th of 168 on Arena Coding via Max as of 25 Sep 2026. At 763.2 billion parameters it is a data-centre job to run yourself, so hosted use is the practical route for most readers.
Source
Read from the hosts' own published rates · we published this on 23 September 2026.