prices synced 2026-08-07
T

Inkling (batch)

Thinking Machines · Multimodal · Released Jul 2026

Compare

Open-weight multimodal mixture-of-experts model with 41B active parameters from 975B total and a 524K token context window.

Strengths
Handles multimodal inputs, supports long-context reasoning with its 524K token window, and uses mixture-of-experts routing to activate only a portion of its parameters per request.
Best for
General-purpose reasoning tasks, coding, agentic workflows, and tool-use systems that benefit from larger capacity than Inkling Small.
Limitations
As an open-weight model, it requires self-hosting or managed deployment; the larger parameter count and context window increase computational requirements compared to Inkling Small.

Input / 1M

$0.5

Output / 1M

$2.02

Cached input / 1M

$0.085

Context window

524K

Price history

Price per 1M tokens over timeOutput $2.02; Input $0.5 as of Aug 2026.$0$0.5$1.00$1.50$2.00Output on 6 Aug 2026: $2.02Out $2.02Input on 6 Aug 2026: $0.5In $0.5Aug 2026

Snapshots

Effective Input Output Cached in Note Source
6 Aug 2026 $0.5 $2.02 $0.085 Imported from OpenRouter openrouter.ai

More from Thinking Machines

Report a problem