GLM 5.3 Flash rose across 3 hosts, by up to 100% at OpenInference (output), with cache read up 133% at Inceptron
Published 22 September 2026
Price move, per 1M tokens
- OpenInference · input↑ 33%$0.075 → $0.100
- OpenInference · output↑ 100%$0.25 → $0.50
- OpenInference · cache read↑ 25%$0.020 → $0.025
- Inceptron · input↑ 50%$0.10 → $0.15
- Inceptron · output↑ 25%$0.40 → $0.50
- Inceptron · cache read↑ 133%$0.030 → $0.070
- Relace · input↑ 11%$0.099 → $0.110
- Relace · output↑ 9%$0.33 → $0.36
- Relace · cache read↑ 1%$0.0198 → $0.0200
| host | rate | was | now | change |
|---|---|---|---|---|
| OpenInference | input | $0.075 | $0.100 | ↑ 33% |
| OpenInference | output | $0.25 | $0.50 | ↑ 100% |
| OpenInference | cache read | $0.020 | $0.025 | ↑ 25% |
| Inceptron | input | $0.10 | $0.15 | ↑ 50% |
| Inceptron | output | $0.40 | $0.50 | ↑ 25% |
| Inceptron | cache read | $0.030 | $0.070 | ↑ 133% |
| Relace | input | $0.099 | $0.110 | ↑ 11% |
| Relace | output | $0.33 | $0.36 | ↑ 9% |
| Relace | cache read | $0.0198 | $0.0200 | ↑ 1% |
OpenInference output: ↑ 100%Inceptron input: ↑ 50%Relace input: ↑ 11%
If you buy
Budget for heavier cache reads on this host before your next top-up, and re-check whether your workload still fits the plan you sized at the old rates.
The model
GLM 5.3 Flashopen the model →
Open weights321.3B parameters1311K context35 hosts
intelligence26th of 168 writing43rd of 168 coding22nd of 168 agents27th of 55
GLM 5.3 Flash is a downloadable model you can run yourself, and it is at its best when the job is agentic: picking the right tool and finishing the task. It is at its worst when the job is holding to a format under instruction, and we list no licence for it, so the terms need checking at the source.
Source
Read on openrouter.ai · we published this on 22 September 2026.
read the original at openrouter.ai ↗