DeepSeek V4.1 Flash repriced across 2 hosts, from a 400% rise at Relace (cache read) to a 50% cut at Alibaba (all rates)
Published 15 September 2026
Price move, per 1M tokens
- Relace · inputunchanged$0.15 → $0.15
- Relace · outputunchanged$0.60 → $0.60
- Relace · cache read↑ 400%$0.003 → $0.015
- Alibaba · input↓ 50%$0.30 → $0.15
- Alibaba · output↓ 50%$1.20 → $0.60
- Alibaba · cache read↓ 50%$0.030 → $0.015
| host | rate | was | now | change |
|---|---|---|---|---|
| Relace | input | $0.15 | $0.15 | — |
| Relace | output | $0.60 | $0.60 | — |
| Relace | cache read | $0.003 | $0.015 | ↑ 400% |
| Alibaba | input | $0.30 | $0.15 | ↓ 50% |
| Alibaba | output | $1.20 | $0.60 | ↓ 50% |
| Alibaba | cache read | $0.030 | $0.015 | ↓ 50% |
Relace cache read: ↑ 400%Alibaba all rates: ↓ 50%
If you buy
Re-check your cache-read spend on Relace before the next top-up, since cached prompts now cost more there for this model.
The model
DeepSeek V4.1 Flashopen the model →
Open weights763.2B parameters1049K context35 hosts
intelligence21st of 168 writing37th of 168 coding11th of 168 agents13th of 55
DeepSeek V4.1 Flash is a downloadable model with a permissive licence, and it places 11th of 168 on Arena Coding via Max as of 25 Sep 2026. At 763.2 billion parameters it is a data-centre job to run yourself, so hosted use is the practical route for most readers.
Source
Read on openrouter.ai · we published this on 15 September 2026.
read the original at openrouter.ai ↗