GLM 4.6V is Z.ai's vision-language model that processes text, image, and video inputs for multimodal reasoning tasks.
- Strengths
- Handles multimodal inputs across images and video alongside text, enabling vision-grounded analysis and understanding.
- Best for
- Applications requiring simultaneous processing of visual and textual content, such as image understanding, video analysis, and cross-modal reasoning.
- Limitations
- Older than the current generation of Z.ai models; GLM 5V Turbo and GLM 5.2 supersede this model with newer architectures and larger context windows.
Input / 1M
$0.3
Output / 1M
$0.9
Cached input / 1M
$0.055
Context window
131K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.3 | $0.9 | $0.055 | Imported from OpenRouter | openrouter.ai |