GLM 5.3 Flash is a native multimodal model from Z.ai with a 1M-token context window, designed for efficient inference on coding and long-horizon agent tasks.
- Strengths
- Its hybrid sparse and linear attention architecture maintains accuracy across long contexts while enabling faster inference compared to GLM 5.3.
- Best for
- Coding tasks, agentic workflows, and applications where you need to process extended contexts without sacrificing response speed.
- Limitations
- As a Flash variant, it trades some reasoning capability for speed; GLM 5.3 remains the choice for tasks requiring deeper reasoning over similarly large contexts.
Input / 1M
$0.15
Output / 1M
$0.5
Cached input / 1M
$0.03
Context window
1.05M
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in increased 100.0% out increased 100.0%
- 90d
- in increased 100.0% out increased 100.0%
- 1y
- in increased 100.0% out increased 100.0%
- Since launch
- in increased 100.0% out increased 100.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 10 Sep 2026 | $0.15 | $0.5 | $0.03 | Imported from OpenRouter | openrouter.ai |
| 26 Aug 2026 | $0.075 | $0.25 | $0.015 | Imported from OpenRouter | openrouter.ai |