prices synced 2026-07-28

Llama 3.1 8B Instruct

Meta · Released Jul 2024

Compare

Meta's 8-billion-parameter instruction-tuned language model from the Llama 3.1 family, with a 131k token context window.

Strengths
Offers a good balance of capability and speed for a compact model, with support for long context windows that enable processing of substantial documents or conversation histories.
Best for
Tasks where response latency matters and you don't need the reasoning depth of larger models, such as conversational agents, customer support, and simple content generation.
Limitations
Smaller parameter count limits performance on complex reasoning, code generation, and tasks requiring broad knowledge compared to larger models like Llama 3.1 70B Instruct; a newer release like Llama 3.3 70B Instruct provides more capability in the instruction-tuned family.

Input / 1M

$0.05

Output / 1M

$0.08

Cached input / 1M

$0.025

Context window

131K

Price history

Input (solid)Output (dashed)
Price per 1M tokens over timeOutput went from $0.03 to $0.08; Input went from $0.02 to $0.05 between Jun 2026 and Jul 2026.$0$0.02$0.04$0.06$0.08Output on 11 Jun 2026: $0.03Output on 15 Jul 2026: $0.08Out $0.08Input on 11 Jun 2026: $0.02Input on 15 Jul 2026: $0.05In $0.05Jun 2026Jul 2026

Price change

30d
in increased 150.0% out increased 166.7%
90d
in increased 150.0% out increased 166.7%
1y
in increased 150.0% out increased 166.7%
Since launch
in increased 150.0% out increased 166.7%

Snapshots

Effective Input Output Cached in Note Source
15 Jul 2026 $0.05 $0.08 $0.025 Imported from OpenRouter openrouter.ai
11 Jun 2026 $0.02 $0.03 Imported from OpenRouter openrouter.ai

More from Meta

Report a problem