A 7B multimodal vision-language model from ByteDance designed to understand and interact with graphical user interfaces across desktop, web, mobile, and game environments.
- Strengths
- Excels at visual understanding of UI elements, screen layouts, and GUI-based interaction tasks through a combination of vision encoding and language reasoning.
- Best for
- Automating tasks that involve interpreting screenshots, identifying UI components, and generating actions based on visual interface states.
- Limitations
- As a smaller model, it may struggle with complex reasoning tasks or nuanced context that larger models handle more reliably, and performance depends on clear, well-structured UI visuals.
Input / 1M
$0.1
Output / 1M
$0.2
Cached input / 1M
$0.1
Context window
128K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.1 | $0.2 | $0.1 | Imported from OpenRouter | openrouter.ai |