Qwen3 VL 30B A3B Instruct
Alibaba · Multimodal · Released Oct 2025
A 30B multimodal vision-language model from Alibaba that processes text, images, and video with a 262k-token context window.
- Strengths
- Handles longer documents and complex multimodal tasks with its extended context window, larger than the 8B and 32B VL models in the same generation.
- Best for
- Long-form document analysis with visual content, where extended context and parameter count matter more than reasoning overhead.
- Limitations
- Lacks the extended reasoning capabilities of the Qwen3 VL 30B A3B Thinking variant, which offers structured problem-solving for more complex tasks.
Input / 1M
$0.15
Output / 1M
$0.6
Cached input / 1M
—
Context window
262K
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in increased 15.4% out increased 15.4%
- 90d
- in increased 15.4% out increased 15.4%
- 1y
- in increased 15.4% out increased 15.4%
- Since launch
- in increased 15.4% out increased 15.4%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 23 Jul 2026 | $0.15 | $0.6 | — | Imported from OpenRouter | openrouter.ai |
| 11 Jun 2026 | $0.13 | $0.52 | — | Imported from OpenRouter | openrouter.ai |