prices synced 2026-07-28
Z

GLM 4.7 Flash

Z.ai · Released Jan 2026

Compare

GLM 4.7 Flash is Z.ai's optimized inference variant of the GLM 4.7 foundation model, designed for faster response times with a 202K-token context window.

Strengths
Delivers reduced latency while retaining the reasoning and code-generation capabilities of the full GLM 4.7 model.
Best for
Latency-sensitive applications that need strong reasoning and code handling without the computational overhead of the full model.
Limitations
As a speed-optimized variant, it may trade some accuracy or capability depth compared to the full GLM 4.7 foundation model.

Input / 1M

$0.06

Output / 1M

$0.4

Cached input / 1M

$0.01

Context window

202K

Price history

Input (solid)Output (dashed)
Price per 1M tokens over timeOutput went from $0.4 to $0.4; Input went from $0.06 to $0.06 between Jun 2026 and Jul 2026.$0$0.1$0.2$0.3$0.4Output on 11 Jun 2026: $0.4Output on 15 Jul 2026: $0.4Output on 16 Jul 2026: $0.4Output on 17 Jul 2026: $0.4Output on 23 Jul 2026: $0.4Out $0.4Input on 11 Jun 2026: $0.06Input on 15 Jul 2026: $0.0605Input on 16 Jul 2026: $0.0605Input on 17 Jul 2026: $0.0605Input on 23 Jul 2026: $0.06In $0.06Jun 2026Jul 2026

Price change

30d
in decreased 0.0% out decreased 0.0%
90d
in decreased 0.0% out decreased 0.0%
1y
in decreased 0.0% out decreased 0.0%
Since launch
in decreased 0.0% out decreased 0.0%

Snapshots

Effective Input Output Cached in Note Source
23 Jul 2026 $0.06 $0.4 $0.01 Imported from OpenRouter openrouter.ai
17 Jul 2026 $0.0605 $0.4 Imported from OpenRouter openrouter.ai
16 Jul 2026 $0.0605 $0.4 Imported from OpenRouter openrouter.ai
15 Jul 2026 $0.0605 $0.4 Imported from OpenRouter openrouter.ai
11 Jun 2026 $0.06 $0.4 $0.01 Imported from OpenRouter openrouter.ai

More from Z.ai

Report a problem