News / PRICE

DeepSeek V4 Flash repriced across 4 hosts: Relace output down 50%, OpenInference input up 250%

Published 24 September 2026
Price move, per 1M tokens
  • OpenInference · input↑ 250%
    $0.040 → $0.140
  • OpenInference · output↑ 40%
    $0.50 → $0.70
  • OpenInference · cache read↑ 114%
    $0.014 → $0.030
  • Relace · input↓ 25%
    $0.040 → $0.030
  • Relace · output↓ 50%
    $0.64 → $0.32
  • Relace · cache readunchanged
    $0.016 → $0.016
  • Inceptron · input↓ 9%
    $0.091 → $0.083
  • Inceptron · output↓ 32%
    $0.613 → $0.414
  • Inceptron · cache readunchanged
    $0.060 → $0.060
  • Wafer · input↓ 8%
    $0.087 → $0.080
  • Wafer · outputunchanged
    $0.35 → $0.35
  • Wafer · cache read↓ 75%
    $0.080 → $0.020
OpenInference input: ↑ 250%Relace output: ↓ 50%Inceptron output: ↓ 32%Wafer input: ↓ 8%
If you buy

OpenInference users now pay more for the same model, so re-check your usage mix and whether a lighter prompt or shorter context keeps the bill where you expected it.

The model
DeepSeek V4 Flashopen the model →
Open weights290.9B parameters1049K context33 hosts
intelligence65th of 168 writing57th of 168 coding64th of 168 agents20th of 55

DeepSeek V4 Flash is a downloadable text model with a permissive licence, built for long-document work where the bill matters more than peak quality. Its measured quality is weak on the LiveBench boards and mid-pack on the Arena boards, so treat it as a cheap workhorse rather than a quality leader.

Source

Read on openrouter.ai · we published this on 24 September 2026.

read the original at openrouter.ai ↗