prices synced 2026-09-11

Qwen3 VL 8B Instruct

Alibaba · Multimodal · Released Oct 2025

Compare

An 8-billion-parameter multimodal vision-language model from Alibaba that processes text, images, and video with a 131k-token context window.

Strengths
Delivers faster inference and lower resource requirements than larger multimodal models while retaining broad multimodal capabilities across text, image, and video inputs.
Best for
Applications requiring multimodal understanding where inference speed and computational efficiency matter, or where model size constraints exist.
Limitations
Smaller parameter count means reduced reasoning depth and domain expertise compared to larger Qwen3 VL variants; for tasks needing extended reasoning, Qwen3 VL 8B Thinking is the better fit in this size class.

Input / 1M

$0.117

Output / 1M

$0.455

Cached input / 1M

Context window

131K

Price history

Input (solid)Output (dashed)
Price per 1M tokens over timeOutput went from $0.5 to $0.455; Input went from $0.08 to $0.117 between Jun 2026 and Jul 2026.$0$0.1$0.2$0.3$0.4$0.5Output on 11 Jun 2026: $0.5Output on 3 Jul 2026: $0.455Out $0.455Input on 11 Jun 2026: $0.08Input on 3 Jul 2026: $0.117In $0.117Jun 2026Jul 2026

Price change

30d
in out
90d
in increased 46.3% out decreased 9.0%
1y
in increased 46.3% out decreased 9.0%
Since launch
in increased 46.3% out decreased 9.0%

Snapshots

Effective Input Output Cached in Note Source
3 Jul 2026 $0.117 $0.455 Imported from OpenRouter openrouter.ai
11 Jun 2026 $0.08 $0.5 Imported from OpenRouter openrouter.ai

More from Alibaba

Report a problem