A multimodal model from Amazon designed to balance accuracy and speed across a broad range of tasks, with a 300K-token context window.
- Strengths
- Handles both text and image inputs with reasonable latency, making it suitable for interactive applications that need multimodal understanding.
- Best for
- General-purpose text and vision tasks where you need acceptable quality without specializing for a single domain.
- Limitations
- Not optimized for specialized domains like advanced reasoning, code generation at scale, or handling very long documents requiring deep context retention.
Input / 1M
$0.8
Output / 1M
$3.20
Cached input / 1M
—
Context window
300K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.8 | $3.20 | — | Imported from OpenRouter | openrouter.ai |