GLM 4.5V is Z.ai's vision-language model that processes text, image, and video inputs for multimodal reasoning tasks.
- Strengths
- Handles multimodal inputs natively across text, image, and video to support reasoning across different data types.
- Best for
- Applications requiring visual understanding and reasoning combined with text analysis, such as document analysis with images or video understanding tasks.
- Limitations
- Has a smaller context window at 65536 tokens compared to later foundation models like GLM 5 and GLM 5.1, and is superseded by GLM 5V Turbo which adds agent-driven task optimization.
Input / 1M
$0.6
Output / 1M
$1.80
Cached input / 1M
$0.11
Context window
65K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.6 | $1.80 | $0.11 | Imported from OpenRouter | openrouter.ai |