I
Compare
Ling-3.0-flash
inclusionAI · Released Jul 2026
Ling-3.0-flash is a 124B-parameter Mixture-of-Experts model from inclusionAI with approximately 5.1B parameters activated per token.
- Strengths
- The model prioritizes token efficiency and low per-token activation, making it resource-efficient for long-context reasoning and high-throughput inference scenarios.
- Best for
- Production-scale agentic inference and multi-turn reasoning where token efficiency and operational cost per request matter.
- Limitations
- As a newer release in the Ling series, it follows Ling-2.6-flash (104B, 7.4B active), so you should verify it offers meaningful improvements in your specific workload before migrating.
Input / 1M
$0.075
Output / 1M
$0.22
Cached input / 1M
$0.015
Context window
131K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 6 Aug 2026 | $0.075 | $0.22 | $0.015 | Imported from OpenRouter | openrouter.ai |