GLM 5.3 Prime is Z.ai's high-throughput variant of GLM 5.3, delivering 1.5–2× faster output generation while retaining full model capabilities and 1M-token context.
- Strengths
- It handles long-context tasks at accelerated inference speed without the capability reduction present in lighter variants like Flash.
- Best for
- Workloads requiring both extended context processing and high output throughput, such as real-time generation from large documents or high-volume batch operations.
- Limitations
- It trades some latency optimization for throughput gain compared to the base GLM 5.3, and lacks the multimodal input support of GLM 5.3 Flash.
Input / 1M
$2.80
Output / 1M
$8.80
Cached input / 1M
$0.56
Context window
1M
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 24 Sep 2026 | $2.80 | $8.80 | $0.56 | Imported from OpenRouter | openrouter.ai |