Qwen3 VL 30B A3B Thinking
Alibaba · Multimodal · Released Oct 2025
A 30B multimodal vision-language model from Alibaba that processes text, images, and video with extended reasoning capabilities and a 131k-token context window.
- Strengths
- Combines multimodal understanding across text, images, and video with reasoning chains for structured problem-solving across domain tasks.
- Best for
- Visual reasoning tasks that benefit from step-by-step inference—image analysis requiring explanation, video understanding with interpretation, or multimodal problems needing deliberate reasoning.
- Limitations
- The 30B parameter size sits between smaller and larger models in the Qwen3 VL lineup; reasoning capabilities are specialized and may add latency for tasks needing only fast inference without intermediate reasoning steps.
Input / 1M
$0.13
Output / 1M
$1.56
Cached input / 1M
—
Context window
131K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.13 | $1.56 | — | Imported from OpenRouter | openrouter.ai |