prices synced 2026-09-19
Z

GLM 5.3 FlashX

Z.ai · Multimodal · Released Sep 2026

Compare

GLM 5.3 FlashX is Z.ai's high-speed variant of GLM 5.3 Flash, a native multimodal model built on hybrid sparse and linear attention architecture.

Strengths
Delivers inference speeds up to 200 tokens/s while maintaining support for text, image, and video inputs across a 1M-token context window.
Best for
Workloads requiring fast multimodal inference on coding tasks, long-horizon agent processes, and extended document handling where speed is the priority.
Limitations
As an optimized speed variant, it trades some capability depth compared to the full GLM 5.3 foundation model; use GLM 5.3 if reasoning complexity outweighs latency concerns.

Input / 1M

$0.37

Output / 1M

$1.25

Cached input / 1M

$0.075

Context window

1.05M

Price history

Price per 1M tokens over timeOutput $1.25; Input $0.37 as of Sep 2026.$0$0.5$1.00Output on 19 Sep 2026: $1.25Out $1.25Input on 19 Sep 2026: $0.37In $0.37Sep 2026

Snapshots

Effective Input Output Cached in Note Source
19 Sep 2026 $0.37 $1.25 $0.075 Imported from OpenRouter openrouter.ai

More from Z.ai

Report a problem