GLM 4.5 Air is Z.ai's lightweight inference variant of the GLM 4.5 foundation model, designed for faster processing with a 131K-token context window.
- Strengths
- Optimized for lower-latency inference compared to the full GLM 4.5, making it suitable for applications where response speed is a priority.
- Best for
- Real-time chat, interactive applications, and workloads where inference speed matters more than maximum capability.
- Limitations
- As a lighter variant of GLM 4.5, it trades some model capacity for speed and is superseded by newer foundation models like GLM 4.6, 4.7, and later releases in the lineup.
Input / 1M
$0.13
Output / 1M
$0.85
Cached input / 1M
$0.025
Context window
131K
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in — out —
- 90d
- in increased 4.0% out decreased 0.0%
- 1y
- in increased 4.0% out decreased 0.0%
- Since launch
- in increased 4.0% out decreased 0.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 17 Jun 2026 | $0.13 | $0.85 | $0.025 | Imported from OpenRouter | openrouter.ai |
| 11 Jun 2026 | $0.125 | $0.85 | $0.06 | Imported from OpenRouter | openrouter.ai |