Mercury 2.5 is a diffusion-based language model from Inception that generates and refines multiple tokens in parallel instead of sequentially.
- Strengths
- Parallel token generation and refinement enables faster inference latency compared to autoregressive approaches.
- Best for
- Workloads where inference speed is critical and a 260k token context window is sufficient.
- Limitations
- As a newer diffusion-based approach, it may have different output characteristics or quality trade-offs relative to established autoregressive models; Mercury 2 may still be preferred for tasks requiring the most proven reasoning patterns.
Input / 1M
$0.04
Output / 1M
$0.15
Cached input / 1M
$0.004
Context window
260K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 9 Sep 2026 | $0.04 | $0.15 | $0.004 | Imported from OpenRouter | openrouter.ai |