A 7-billion-parameter multimodal model that processes images, videos, and text to generate text outputs.
- Strengths
- Delivers efficient image and video understanding while maintaining a small model size suitable for edge deployment.
- Best for
- Applications requiring multimodal vision-language capabilities on resource-constrained hardware or with latency constraints.
- Limitations
- As a smaller model, it may lack the depth of reasoning or nuance in complex multimodal tasks compared to larger vision-language models.
Input / 1M
$0.1
Output / 1M
$0.1
Cached input / 1M
—
Context window
16K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.1 | $0.1 | — | Imported from OpenRouter | openrouter.ai |