GLM 4.5V is Z.ai's multimodal model supporting text, image, and video inputs with a 65K-token context window.
- Strengths
- Handles multimodal input across text, images, and video within a single request.
- Best for
- Applications requiring simultaneous processing of text and visual media like video analysis, image-to-text tasks, and multimodal document understanding.
- Limitations
- The 65K-token context window is smaller than most current Z.ai foundation models, and it predates GLM 4.6V and later multimodal variants that offer larger context.
Input / 1M
$0.6
Output / 1M
$1.80
Cached input / 1M
$0.11
Context window
65K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.6 | $1.80 | $0.11 | Imported from OpenRouter | openrouter.ai |