GLM 5 Turbo is Z.ai's optimized inference variant of their GLM 5 foundation model, offering faster processing with a 202K-token context window.
- Strengths
- Balances inference speed and capability, making it suitable for applications where latency matters but reasoning complexity remains moderate.
- Best for
- Interactive applications, real-time chat, and production workloads where throughput and response time are constraints.
- Limitations
- Smaller context window and lower reasoning capability than GLM 5.1 and later models; superseded by GLM 5V Turbo for multimodal tasks and by GLM 5.3 Flash for long-context and agent workloads.
Input / 1M
$1.20
Output / 1M
$4.00
Cached input / 1M
$0.24
Context window
202K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $1.20 | $4.00 | $0.24 | Imported from OpenRouter | openrouter.ai |