prices synced 2026-07-28

Qwen3 VL 8B Instruct

Alibaba · Multimodal · Released Oct 2025

Compare

An 8B multimodal vision-language model from Alibaba that processes text, images, and video with a 131k-token context window.

Strengths
Delivers fast inference and low resource overhead while handling multimodal tasks across text, images, and video simultaneously.
Best for
Lightweight multimodal applications where speed and efficiency matter more than reasoning depth, such as image captioning, visual Q&A, and video summarization.
Limitations
At 8B parameters, it lacks the reasoning capabilities of larger models like Qwen3 VL 30B A3B or the extended reasoning mode of Qwen3 VL 8B Thinking, making it less suitable for complex problem-solving tasks.

Input / 1M

$0.117

Output / 1M

$0.455

Cached input / 1M

Context window

131K

Price history

Input (solid)Output (dashed)
Price per 1M tokens over timeOutput went from $0.5 to $0.455; Input went from $0.08 to $0.117 between Jun 2026 and Jul 2026.$0$0.1$0.2$0.3$0.4$0.5Output on 11 Jun 2026: $0.5Output on 3 Jul 2026: $0.455Out $0.455Input on 11 Jun 2026: $0.08Input on 3 Jul 2026: $0.117In $0.117Jun 2026Jul 2026

Price change

30d
in increased 46.3% out decreased 9.0%
90d
in increased 46.3% out decreased 9.0%
1y
in increased 46.3% out decreased 9.0%
Since launch
in increased 46.3% out decreased 9.0%

Snapshots

Effective Input Output Cached in Note Source
3 Jul 2026 $0.117 $0.455 Imported from OpenRouter openrouter.ai
11 Jun 2026 $0.08 $0.5 Imported from OpenRouter openrouter.ai

More from Alibaba

Report a problem