R1 Distill Llama 70B
DeepSeek · Released Jan 2025
A 70-billion-parameter reasoning model distilled from DeepSeek's R1, designed to run efficiently on consumer hardware while retaining chain-of-thought capabilities.
- Strengths
- Delivers reasoning performance on much smaller hardware than the full R1, making it practical for deployment on GPUs with limited memory.
- Best for
- Reasoning and complex problem-solving tasks where you need chain-of-thought inference but want faster inference and lower resource requirements than the full R1 model.
- Limitations
- Smaller than the full R1 and likely trades some reasoning depth for efficiency; superseded by newer distilled variants like R1 Distill Qwen 32B for even more compact deployment.
Input / 1M
$0.8
Output / 1M
$0.8
Cached input / 1M
—
Context window
8K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.8 | $0.8 | — | Imported from OpenRouter | openrouter.ai |