Qwen3 VL 8B Instruct
Alibaba · Multimodal · Released Oct 2025
An 8-billion-parameter multimodal vision-language model from Alibaba that processes text, images, and video with a 131k-token context window.
- Strengths
- Delivers faster inference and lower resource requirements than larger multimodal models while retaining broad multimodal capabilities across text, image, and video inputs.
- Best for
- Applications requiring multimodal understanding where inference speed and computational efficiency matter, or where model size constraints exist.
- Limitations
- Smaller parameter count means reduced reasoning depth and domain expertise compared to larger Qwen3 VL variants; for tasks needing extended reasoning, Qwen3 VL 8B Thinking is the better fit in this size class.
Input / 1M
$0.117
Output / 1M
$0.455
Cached input / 1M
—
Context window
131K
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in — out —
- 90d
- in increased 46.3% out decreased 9.0%
- 1y
- in increased 46.3% out decreased 9.0%
- Since launch
- in increased 46.3% out decreased 9.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 3 Jul 2026 | $0.117 | $0.455 | — | Imported from OpenRouter | openrouter.ai |
| 11 Jun 2026 | $0.08 | $0.5 | — | Imported from OpenRouter | openrouter.ai |