GLM 5.3 Flash is a native multimodal model from Z.ai with a 1M-token context window, designed for efficient inference on coding and long-horizon agent tasks.
- Strengths
- It maintains accurate long-context behavior through a hybrid sparse and linear attention architecture while keeping inference costs low.
- Best for
- Batch processing of extended documents, multi-step coding tasks, and agent workflows that benefit from reduced latency.
- Limitations
- It is optimized for batch processing rather than real-time interactions, and trades some capability for speed compared to the full GLM 5.3.
Input / 1M
$0.075
Output / 1M
$0.25
Cached input / 1M
$0.015
Context window
1.05M
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in decreased 50.0% out decreased 50.0%
- 90d
- in decreased 50.0% out decreased 50.0%
- 1y
- in decreased 50.0% out decreased 50.0%
- Since launch
- in decreased 50.0% out decreased 50.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 8 Sep 2026 | $0.075 | $0.25 | $0.015 | Imported from OpenRouter | openrouter.ai |
| 29 Aug 2026 | $0.15 | $0.5 | $0.03 | Imported from OpenRouter | openrouter.ai |