prices synced 2026-09-11
Z

GLM 5.3 Flash

Z.ai · Multimodal · Released Aug 2026

Compare

GLM 5.3 Flash is a native multimodal model from Z.ai with a 1M-token context window, designed for efficient inference on coding and long-horizon agent tasks.

Strengths
Its hybrid sparse and linear attention architecture maintains accuracy across long contexts while enabling faster inference compared to GLM 5.3.
Best for
Coding tasks, agentic workflows, and applications where you need to process extended contexts without sacrificing response speed.
Limitations
As a Flash variant, it trades some reasoning capability for speed; GLM 5.3 remains the choice for tasks requiring deeper reasoning over similarly large contexts.

Input / 1M

$0.15

Output / 1M

$0.5

Cached input / 1M

$0.03

Context window

1.05M

Price history

Input (solid)Output (dashed)
Price per 1M tokens over timeOutput went from $0.25 to $0.5; Input went from $0.075 to $0.15 between Aug 2026 and Sep 2026.$0$0.1$0.2$0.3$0.4$0.5Output on 26 Aug 2026: $0.25Output on 10 Sep 2026: $0.5Out $0.5Input on 26 Aug 2026: $0.075Input on 10 Sep 2026: $0.15In $0.15Aug 2026Sep 2026

Price change

30d
in increased 100.0% out increased 100.0%
90d
in increased 100.0% out increased 100.0%
1y
in increased 100.0% out increased 100.0%
Since launch
in increased 100.0% out increased 100.0%

Snapshots

Effective Input Output Cached in Note Source
10 Sep 2026 $0.15 $0.5 $0.03 Imported from OpenRouter openrouter.ai
26 Aug 2026 $0.075 $0.25 $0.015 Imported from OpenRouter openrouter.ai

More from Z.ai

Report a problem