Llama 3.2 11B Vision Instruct
Meta · Multimodal · Released Sep 2024
An 11-billion-parameter instruction-tuned vision model from Meta that processes both text and images.
- Strengths
- Handles multimodal input with reasonable efficiency at 11B parameters, making it lighter than the 70B vision variant while retaining image understanding capabilities.
- Best for
- Applications requiring vision understanding without the computational overhead of larger models — document analysis, image captioning, visual question answering on edge or cost-constrained deployments.
- Limitations
- Smaller scale limits performance on complex reasoning tasks compared to larger models; the 11B size trades capability for speed and resource efficiency in vision tasks.
Input / 1M
$0.345
Output / 1M
$0.345
Cached input / 1M
—
Context window
131K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.345 | $0.345 | — | Imported from OpenRouter | openrouter.ai |