News / PRICE

GLM 5.3 Flash rose across 3 hosts, by up to 100% at OpenInference (output), with cache read up 133% at Inceptron

Published 22 September 2026
Price move, per 1M tokens
  • OpenInference · input↑ 33%
    $0.075 → $0.100
  • OpenInference · output↑ 100%
    $0.25 → $0.50
  • OpenInference · cache read↑ 25%
    $0.020 → $0.025
  • Inceptron · input↑ 50%
    $0.10 → $0.15
  • Inceptron · output↑ 25%
    $0.40 → $0.50
  • Inceptron · cache read↑ 133%
    $0.030 → $0.070
  • Relace · input↑ 11%
    $0.099 → $0.110
  • Relace · output↑ 9%
    $0.33 → $0.36
  • Relace · cache read↑ 1%
    $0.0198 → $0.0200
OpenInference output: ↑ 100%Inceptron input: ↑ 50%Relace input: ↑ 11%
If you buy

Budget for heavier cache reads on this host before your next top-up, and re-check whether your workload still fits the plan you sized at the old rates.

The model
GLM 5.3 Flashopen the model →
Open weights321.3B parameters1311K context35 hosts
intelligence26th of 168 writing43rd of 168 coding22nd of 168 agents27th of 55

GLM 5.3 Flash is a downloadable model you can run yourself, and it is at its best when the job is agentic: picking the right tool and finishing the task. It is at its worst when the job is holding to a format under instruction, and we list no licence for it, so the terms need checking at the source.

Source

Read on openrouter.ai · we published this on 22 September 2026.

read the original at openrouter.ai ↗