Host Wafer cut DeepSeek V4.1 Flash output pricing by 50%
Published 18 September 2026
Price move, per 1M tokens
- Wafer · input↓ 33%$0.30 → $0.20
- Wafer · output↓ 50%$1.20 → $0.60
- Wafer · cache readunchanged$0.006 → $0.006
| host | rate | was | now | change |
|---|---|---|---|---|
| Wafer | input | $0.30 | $0.20 | ↓ 33% |
| Wafer | output | $1.20 | $0.60 | ↓ 50% |
| Wafer | cache read | $0.006 | $0.006 | — |
output: ↓ 50%
If you buy
output-heavy workloads now stretch further on the same top-up, so re-check your usage split before deciding how much credit to load.
The model
DeepSeek V4.1 Flashopen the model →
Open weights763.2B parameters1049K context35 hosts
intelligence21st of 168 writing37th of 168 coding11th of 168 agents13th of 55
DeepSeek V4.1 Flash is a downloadable model with a permissive licence, and it places 11th of 168 on Arena Coding via Max as of 25 Sep 2026. At 763.2 billion parameters it is a data-centre job to run yourself, so hosted use is the practical route for most readers.
Source
Read on openrouter.ai · we published this on 18 September 2026.
read the original at openrouter.ai ↗