Qwen3 VL 32B Instruct
Alibaba · Multimodal · Released Oct 2025
A 32-billion-parameter multimodal vision-language model from Alibaba that processes text, images, and video with a 131k-token context window.
- Strengths
- Handles vision and text tasks in a single model with moderate parameter count, balancing capability and resource efficiency compared to the 235B variant.
- Best for
- Multimodal applications involving images and text where the 8B variant offers insufficient capacity but the 235B variant is overkill for inference cost or latency constraints.
- Limitations
- Positioned between the smaller 8B and larger 235B VL models; if reasoning over visual input is required, the Thinking variants are better suited.
Input / 1M
$0.104
Output / 1M
$0.416
Cached input / 1M
—
Context window
131K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.104 | $0.416 | — | Imported from OpenRouter | openrouter.ai |