prices synced 2026-09-24
Z

GLM 5.3 Prime

Z.ai · Released Sep 2026

Compare

GLM 5.3 Prime is Z.ai's high-throughput variant of GLM 5.3, delivering 1.5–2× faster output generation while retaining full model capabilities and 1M-token context.

Strengths
It handles long-context tasks at accelerated inference speed without the capability reduction present in lighter variants like Flash.
Best for
Workloads requiring both extended context processing and high output throughput, such as real-time generation from large documents or high-volume batch operations.
Limitations
It trades some latency optimization for throughput gain compared to the base GLM 5.3, and lacks the multimodal input support of GLM 5.3 Flash.

Input / 1M

$2.80

Output / 1M

$8.80

Cached input / 1M

$0.56

Context window

1M

Price history

Price per 1M tokens over timeOutput $8.80; Input $2.80 as of Sep 2026.$0$2.00$4.00$6.00$8.00$10.00Output on 24 Sep 2026: $8.80Out $8.80Input on 24 Sep 2026: $2.80In $2.80Sep 2026

Snapshots

Effective Input Output Cached in Note Source
24 Sep 2026 $2.80 $8.80 $0.56 Imported from OpenRouter openrouter.ai

More from Z.ai

Report a problem