prices synced 2026-07-28

Llama 3.2 11B Vision Instruct

Meta · Multimodal · Released Sep 2024

Compare

An 11-billion-parameter instruction-tuned vision model from Meta that processes both text and images with a 131k token context window.

Strengths
Handles multimodal tasks combining text and image understanding in a single forward pass, with sufficient scale to manage complex visual reasoning.
Best for
Applications requiring both image and text interpretation in a lightweight form factor, such as document analysis or visual question-answering on consumer hardware or edge deployments.
Limitations
Smaller than the 70B instruction-tuned variants, so may struggle with complex reasoning tasks that benefit from additional model capacity; newer agentic models like Muse Spark 1.1 are purpose-built for tool use and multimodal reasoning at scale.

Input / 1M

$0.345

Output / 1M

$0.345

Cached input / 1M

Context window

131K

Price history

Price per 1M tokens over timeOutput $0.345; Input $0.345 as of Jun 2026.$0$0.1$0.2$0.3Output on 11 Jun 2026: $0.345Out $0.345Input on 11 Jun 2026: $0.345In $0.345Jun 2026

Snapshots

Effective Input Output Cached in Note Source
11 Jun 2026 $0.345 $0.345 Imported from OpenRouter openrouter.ai

More from Meta

Report a problem