prices synced 2026-09-11
Z

GLM 5.3 Flash (batch)

Z.ai · Multimodal · Released Aug 2026

Compare

GLM 5.3 Flash is a native multimodal model from Z.ai with a 1M-token context window, designed for efficient inference on coding and long-horizon agent tasks.

Strengths
It maintains accurate long-context behavior through a hybrid sparse and linear attention architecture while keeping inference costs low.
Best for
Batch processing of extended documents, multi-step coding tasks, and agent workflows that benefit from reduced latency.
Limitations
It is optimized for batch processing rather than real-time interactions, and trades some capability for speed compared to the full GLM 5.3.

Input / 1M

$0.075

Output / 1M

$0.25

Cached input / 1M

$0.015

Context window

1.05M

Price history

Input (solid)Output (dashed)
Price per 1M tokens over timeOutput went from $0.5 to $0.25; Input went from $0.15 to $0.075 between Aug 2026 and Sep 2026.$0$0.1$0.2$0.3$0.4$0.5Output on 29 Aug 2026: $0.5Output on 8 Sep 2026: $0.25Out $0.25Input on 29 Aug 2026: $0.15Input on 8 Sep 2026: $0.075In $0.075Aug 2026Sep 2026

Price change

30d
in decreased 50.0% out decreased 50.0%
90d
in decreased 50.0% out decreased 50.0%
1y
in decreased 50.0% out decreased 50.0%
Since launch
in decreased 50.0% out decreased 50.0%

Snapshots

Effective Input Output Cached in Note Source
8 Sep 2026 $0.075 $0.25 $0.015 Imported from OpenRouter openrouter.ai
29 Aug 2026 $0.15 $0.5 $0.03 Imported from OpenRouter openrouter.ai

More from Z.ai

Report a problem