R1 Distill Llama 70B
DeepSeek · Released Jan 2025
A 70B parameter model distilled from DeepSeek R1 using Llama 3.3 as a base, designed to replicate advanced reasoning capabilities at reduced computational cost.
- Strengths
- Delivers reasoning and problem-solving performance close to the larger R1 model while running with lower latency and memory requirements than the full R1.
- Best for
- Workloads requiring strong reasoning, code generation, or complex analysis where full R1 performance would be overkill or resource constraints matter.
- Limitations
- The 8K token context window is tight for longer documents or multi-turn conversations; reasoning capability may not match the original R1 across all task types.
Input / 1M
$0.8
Output / 1M
$0.8
Cached input / 1M
—
Context window
8K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.8 | $0.8 | — | Imported from OpenRouter | openrouter.ai |