prices synced 2026-09-11

R1 Distill Llama 70B

DeepSeek · Released Jan 2025

Compare

A 70-billion-parameter reasoning model distilled from DeepSeek's R1, designed to run efficiently on consumer hardware while retaining chain-of-thought capabilities.

Strengths
Delivers reasoning performance on much smaller hardware than the full R1, making it practical for deployment on GPUs with limited memory.
Best for
Reasoning and complex problem-solving tasks where you need chain-of-thought inference but want faster inference and lower resource requirements than the full R1 model.
Limitations
Smaller than the full R1 and likely trades some reasoning depth for efficiency; superseded by newer distilled variants like R1 Distill Qwen 32B for even more compact deployment.

Input / 1M

$0.8

Output / 1M

$0.8

Cached input / 1M

Context window

8K

Price history

Price per 1M tokens over timeOutput $0.8; Input $0.8 as of Jun 2026.$0$0.2$0.4$0.6$0.8Output on 11 Jun 2026: $0.8Out $0.8Input on 11 Jun 2026: $0.8In $0.8Jun 2026

Snapshots

Effective Input Output Cached in Note Source
11 Jun 2026 $0.8 $0.8 Imported from OpenRouter openrouter.ai

More from DeepSeek

Report a problem