Qwen3 VL 8B Instruct
Alibaba · Multimodal · Released Oct 2025
An 8B multimodal vision-language model from Alibaba that processes text, images, and video with a 131k-token context window.
- Strengths
- Delivers fast inference and low resource overhead while handling multimodal tasks across text, images, and video simultaneously.
- Best for
- Lightweight multimodal applications where speed and efficiency matter more than reasoning depth, such as image captioning, visual Q&A, and video summarization.
- Limitations
- At 8B parameters, it lacks the reasoning capabilities of larger models like Qwen3 VL 30B A3B or the extended reasoning mode of Qwen3 VL 8B Thinking, making it less suitable for complex problem-solving tasks.
Input / 1M
$0.117
Output / 1M
$0.455
Cached input / 1M
—
Context window
131K
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in increased 46.3% out decreased 9.0%
- 90d
- in increased 46.3% out decreased 9.0%
- 1y
- in increased 46.3% out decreased 9.0%
- Since launch
- in increased 46.3% out decreased 9.0%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 3 Jul 2026 | $0.117 | $0.455 | — | Imported from OpenRouter | openrouter.ai |
| 11 Jun 2026 | $0.08 | $0.5 | — | Imported from OpenRouter | openrouter.ai |