Qwen2.5 VL 72B Instruct
Alibaba · Multimodal · Released Feb 2025
Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images.
- Strengths
- Handles visual question answering, document analysis, and text-in-image recognition including charts, diagrams, and layout comprehension.
- Best for
- Workflows requiring extraction and interpretation of information from images with complex visual elements like forms, graphs, and structured layouts.
- Limitations
- At 72B parameters it has higher latency and throughput constraints compared to smaller multimodal models; performance on specialized or domain-specific visual tasks may vary.
Input / 1M
$0.8
Output / 1M
$1.00
Cached input / 1M
$0.4
Context window
128K
Price history
Input (solid)Output (dashed)
Price change
- 30d
- in decreased 0.0% out decreased 0.0%
- 90d
- in increased 220.0% out increased 33.3%
- 1y
- in increased 220.0% out increased 33.3%
- Since launch
- in increased 220.0% out increased 33.3%
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 15 Jul 2026 | $0.8 | $1.00 | $0.4 | Imported from OpenRouter | openrouter.ai |
| 12 Jul 2026 | $0.25 | $0.75 | — | Imported from OpenRouter | openrouter.ai |
| 12 Jun 2026 | $0.8 | $1.00 | $0.4 | Imported from OpenRouter | openrouter.ai |
| 11 Jun 2026 | $0.25 | $0.75 | — | Imported from OpenRouter | openrouter.ai |