GLM 5V Turbo is Z.ai's optimized inference variant for multimodal processing with text, image, and video inputs.
- Strengths
- It delivers faster multimodal inference while maintaining a large 202K-token context window for processing extended documents and conversations alongside visual content.
- Best for
- Applications requiring quick multimodal responses over substantial text and visual context, where speed is prioritized over the reasoning depth of GLM 5.3.
- Limitations
- As an optimized inference variant, it trades some reasoning capability for speed and does not support the 1M-token context window available in GLM 5.3 models.
Input / 1M
$1.20
Output / 1M
$4.00
Cached input / 1M
$0.24
Context window
202K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 23 Jun 2026 | $1.20 | $4.00 | $0.24 | Imported from OpenRouter | openrouter.ai |