GLM 4.7 Flash is Z.ai's optimized inference variant of the GLM 4.7 foundation model, designed for faster processing with a 202K-token context window.
- Strengths
- The model delivers faster inference latency compared to the full GLM 4.7 foundation model while maintaining a substantial context window for handling extended documents.
- Best for
- Applications requiring lower latency responses on moderately complex tasks, such as real-time chat, content generation, and document analysis.
- Limitations
- As an optimized variant rather than a foundation release, it may trade reasoning depth or accuracy versus the full GLM 4.7 model for speed gains; newer foundation models like GLM 5 and GLM 5.1 supersede it for tasks requiring state-of-the-art capabilities.
Input / 1M
$0.0605
Output / 1M
$0.4
Cached input / 1M
—
Context window
131K
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in increased 0.8% out decreased 0.0%
- 90d
- in increased 0.8% out decreased 0.0%
- 1y
- in increased 0.8% out decreased 0.0%
- Since launch
- in increased 0.8% out decreased 0.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 8 Sep 2026 | $0.0605 | $0.4 | — | Imported from OpenRouter | openrouter.ai |
| 23 Jul 2026 | $0.06 | $0.4 | $0.01 | Imported from OpenRouter | openrouter.ai |
| 17 Jul 2026 | $0.0605 | $0.4 | — | Imported from OpenRouter | openrouter.ai |
| 16 Jul 2026 | $0.0605 | $0.4 | — | Imported from OpenRouter | openrouter.ai |
| 15 Jul 2026 | $0.0605 | $0.4 | — | Imported from OpenRouter | openrouter.ai |
| 11 Jun 2026 | $0.06 | $0.4 | $0.01 | Imported from OpenRouter | openrouter.ai |