Compare two models
Select any two models. Input, output, cached, and context rarely move together — the cheaper model on one line isn't always the cheaper one on the next.
Description
Qwen3 VL 235B A22B Instruct is a 235-billion-parameter multimodal vision-language model from Alibaba that processes text, images, and video within a 131k-token context window.
OpenAI's frontier model for complex professional workloads. 1M token context, text + image input, native computer use.
I/O price /1M
$0.21/$1.90
$5.00/$30.00
Input /1M
$0.21
$5.00
Output /1M
$1.90
$30.00
Cached in /1M
$0.1
$0.5
Context
131K
1M
Δ input since launch
increased 5.0%
—
Δ output since launch
increased 115.9%
—
Released
Sep 2025
Apr 2026
Status
—
—
Provider
Alibaba
OpenAI