A Mixture-of-Experts language model from Upstage with 102B parameters total and 12B active per forward pass, operating on a 128K token context window.
- Strengths
- MoE architecture allows it to process requests with lower compute requirements than comparably-sized dense models while maintaining performance across diverse tasks.
- Best for
- Tasks requiring broad capability and efficiency trade-offs, or applications where reduced latency and throughput per request matters.
- Limitations
- Mixture-of-Experts models can exhibit uneven performance across specialized domains and may require careful prompt engineering to fully activate relevant expert pathways.
Input / 1M
$0.15
Output / 1M
$0.6
Cached input / 1M
$0.015
Context window
128K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 11 Jun 2026 | $0.15 | $0.6 | $0.015 | Imported from OpenRouter | openrouter.ai |