GLM 4.6V is Z.ai's multimodal model supporting text, image, and video inputs with a 131K-token context window.
- Strengths
- Handles multimodal reasoning across text, images, and video in a single request.
- Best for
- Tasks requiring analysis or reasoning over mixed media — documents with embedded images, video transcription with context, or cross-modal question answering.
- Limitations
- Superseded by GLM 5V Turbo, which offers expanded capabilities and is part of Z.ai's current foundation model generation.
Input / 1M
$0.3
Output / 1M
$0.9
Cached input / 1M
$0.055
Context window
131K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.3 | $0.9 | $0.055 | Imported from OpenRouter | openrouter.ai |