News / PRICE

GLM 5.3 Flash cut across 3 hosts, by up to 50% at OpenInference (input)

Published 25 September 2026
Price move, per 1M tokens
  • OpenInference · input↓ 50%
    $0.100 → $0.050
  • OpenInference · outputunchanged
    $0.50 → $0.50
  • OpenInference · cache read↓ 40%
    $0.025 → $0.015
  • Inceptron · input↓ 27%
    $0.15 → $0.11
  • Inceptron · output↓ 10%
    $0.50 → $0.45
  • Inceptron · cache readunchanged
    $0.070 → $0.070
  • Wafer · input↓ 22%
    $0.089 → $0.069
  • Wafer · outputunchanged
    $0.35 → $0.35
  • Wafer · cache readunchanged
    $0.030 → $0.030
OpenInference input: ↓ 50%Inceptron input: ↓ 27%Wafer input: ↓ 22%
If you buy

Renting this model at Wafer now costs less to run, so heavier prompt and cache traffic fits a tighter budget; re-check your usage tier and cache settings before your next top-up.

The model
GLM 5.3 Flashopen the model →
Open weights321.3B parameters1311K context35 hosts
intelligence26th of 168 writing43rd of 168 coding22nd of 168 agents27th of 55

GLM 5.3 Flash is a downloadable model you can run yourself, and it is at its best when the job is agentic: picking the right tool and finishing the task. It is at its worst when the job is holding to a format under instruction, and we list no licence for it, so the terms need checking at the source.

Source

Read on openrouter.ai · we published this on 25 September 2026.

read the original at openrouter.ai ↗