prices synced 2026-09-11

Qwen3 VL 32B Instruct

Alibaba · Multimodal · Released Oct 2025

Compare

A 32-billion-parameter multimodal vision-language model from Alibaba that processes text, images, and video with a 131k-token context window.

Strengths
Handles vision and text tasks in a single model with moderate parameter count, balancing capability and resource efficiency compared to the 235B variant.
Best for
Multimodal applications involving images and text where the 8B variant offers insufficient capacity but the 235B variant is overkill for inference cost or latency constraints.
Limitations
Positioned between the smaller 8B and larger 235B VL models; if reasoning over visual input is required, the Thinking variants are better suited.

Input / 1M

$0.104

Output / 1M

$0.416

Cached input / 1M

Context window

131K

Price history

Price per 1M tokens over timeOutput $0.416; Input $0.104 as of Jun 2026.$0$0.1$0.2$0.3$0.4Output on 11 Jun 2026: $0.416Out $0.416Input on 11 Jun 2026: $0.104In $0.104Jun 2026

Snapshots

Effective Input Output Cached in Note Source
11 Jun 2026 $0.104 $0.416 Imported from OpenRouter openrouter.ai

More from Alibaba

Report a problem