GLM 5.3 FlashX is Z.ai's high-speed variant of GLM 5.3 Flash, a native multimodal model built on hybrid sparse and linear attention architecture.
- Strengths
- Delivers inference speeds up to 200 tokens/s while maintaining support for text, image, and video inputs across a 1M-token context window.
- Best for
- Workloads requiring fast multimodal inference on coding tasks, long-horizon agent processes, and extended document handling where speed is the priority.
- Limitations
- As an optimized speed variant, it trades some capability depth compared to the full GLM 5.3 foundation model; use GLM 5.3 if reasoning complexity outweighs latency concerns.
Input / 1M
$0.37
Output / 1M
$1.25
Cached input / 1M
$0.075
Context window
1.05M
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 19 Sep 2026 | $0.37 | $1.25 | $0.075 | Imported from OpenRouter | openrouter.ai |