Host OpenInference cut DeepSeek V4 Flash 0423 input pricing by 40%
Published 20 September 2026
Price move, per 1M tokens
- OpenInference · input↓ 40%$0.100 → $0.060
- OpenInference · outputunchanged$0.14 → $0.14
- OpenInference · cache readunchanged$0.014 → $0.014
| host | rate | was | now | change |
|---|---|---|---|---|
| OpenInference | input | $0.100 | $0.060 | ↓ 40% |
| OpenInference | output | $0.14 | $0.14 | — |
| OpenInference | cache read | $0.014 | $0.014 | — |
input: ↓ 40%
If you buy
Input-heavy workloads on this model now cost less to run at this host, so it is worth re-checking your usage split before your next top-up.
The model
DeepSeek V4 Flashopen the model →
Open weights290.9B parameters1049K context33 hosts
intelligence65th of 168 writing57th of 168 coding64th of 168 agents20th of 55
DeepSeek V4 Flash is a downloadable text model with a permissive licence, built for long-document work where the bill matters more than peak quality. Its measured quality is weak on the LiveBench boards and mid-pack on the Arena boards, so treat it as a cheap workhorse rather than a quality leader.
Source
Read on openrouter.ai · we published this on 20 September 2026.
read the original at openrouter.ai ↗