GLM 5 Turbo is Z.ai's optimized inference variant of the GLM 5 foundation model, designed for faster response times while maintaining broad capability across reasoning and long-context tasks.
- Strengths
- It handles complex reasoning, code, and long documents efficiently within its 262K-token context window, balancing speed and quality across general workloads.
- Best for
- Applications requiring solid general-purpose reasoning and long-context processing where faster inference is preferred over maximum model scale, such as real-time agent systems and document analysis.
- Limitations
- It is positioned between the base GLM 5 and the reasoning-optimized GLM 5.2; workloads demanding maximum reasoning depth or the 1M-token context of GLM 5.2 may benefit from those models instead.
Input / 1M
$1.20
Output / 1M
$4.00
Cached input / 1M
$0.24
Context window
202K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $1.20 | $4.00 | $0.24 | Imported from OpenRouter | openrouter.ai |