GLM-5V Turbo is Z.ai's native multimodal model that processes image, video, and text inputs for agent-driven tasks with a 202K token context window.
- Strengths
- Handles complex coding tasks and long-horizon planning across vision and text modalities natively in a single model.
- Best for
- Vision-based coding, agentic workflows, and multi-step tasks that require reasoning over images, video, and code together.
- Limitations
- As a specialized multimodal model, it may not match performance of single-modality specialists for pure text or pure vision tasks at the margins.
Input / 1M
$1.20
Output / 1M
$4.00
Cached input / 1M
$0.24
Context window
202K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 23 Jun 2026 | $1.20 | $4.00 | $0.24 | Imported from OpenRouter | openrouter.ai |