Amazon's lightweight multimodal model that processes text, images, and video inputs with a 300K token context window.
- Strengths
- Handles image and video understanding while maintaining low latency and computational requirements.
- Best for
- Applications requiring fast multimodal inference on images, video, and text at scale.
- Limitations
- As a smaller model, it may struggle with complex reasoning tasks or long-form document analysis compared to larger alternatives.
Input / 1M
$0.06
Output / 1M
$0.24
Cached input / 1M
—
Context window
300K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.06 | $0.24 | — | Imported from OpenRouter | openrouter.ai |