prices synced 2026-09-11

Llama 3.1 8B Instruct

Meta · Released Jul 2024

Compare

An 8-billion-parameter instruction-tuned language model from Meta released in July 2024 with a 131,072-token context window.

Strengths
Delivers solid performance on general language tasks and reasoning while maintaining a small parameter footprint, making it efficient to run and cost-effective to deploy.
Best for
Latency-sensitive applications, edge deployment, and workloads where model size and inference speed matter more than maximum reasoning capability.
Limitations
The 8B scale means it underperforms compared to larger models like Llama 3.1 70B on complex reasoning, coding, and specialized tasks; for demanding use cases, a larger model will produce noticeably better results.

Input / 1M

$0.05

Output / 1M

$0.08

Cached input / 1M

$0.025

Context window

131K

Price history

Input (solid)Output (dashed)
Price per 1M tokens over timeOutput went from $0.03 to $0.08; Input went from $0.02 to $0.05 between Jun 2026 and Jul 2026.$0$0.02$0.04$0.06$0.08Output on 11 Jun 2026: $0.03Output on 15 Jul 2026: $0.08Out $0.08Input on 11 Jun 2026: $0.02Input on 15 Jul 2026: $0.05In $0.05Jun 2026Jul 2026

Price change

30d
in out
90d
in increased 150.0% out increased 166.7%
1y
in increased 150.0% out increased 166.7%
Since launch
in increased 150.0% out increased 166.7%

Snapshots

Effective Input Output Cached in Note Source
15 Jul 2026 $0.05 $0.08 $0.025 Imported from OpenRouter openrouter.ai
11 Jun 2026 $0.02 $0.03 Imported from OpenRouter openrouter.ai

More from Meta

Report a problem