prices synced 2026-08-06
I

Ling-3.0-flash

inclusionAI · Released Jul 2026

Compare

Ling-3.0-flash is a 124B-parameter Mixture-of-Experts model from inclusionAI with approximately 5.1B parameters activated per token.

Strengths
The model prioritizes token efficiency and low per-token activation, making it resource-efficient for long-context reasoning and high-throughput inference scenarios.
Best for
Production-scale agentic inference and multi-turn reasoning where token efficiency and operational cost per request matter.
Limitations
As a newer release in the Ling series, it follows Ling-2.6-flash (104B, 7.4B active), so you should verify it offers meaningful improvements in your specific workload before migrating.

Input / 1M

$0.075

Output / 1M

$0.22

Cached input / 1M

$0.015

Context window

131K

Price history

Price per 1M tokens over timeOutput $0.22; Input $0.075 as of Aug 2026.$0$0.05$0.1$0.15$0.2$0.25Output on 6 Aug 2026: $0.22Out $0.22Input on 6 Aug 2026: $0.075In $0.075Aug 2026

Snapshots

Effective Input Output Cached in Note Source
6 Aug 2026 $0.075 $0.22 $0.015 Imported from OpenRouter openrouter.ai

More from inclusionAI

Report a problem