prices synced 2026-07-31
T

Inkling Small

Thinking Machines · Multimodal · Released Jul 2026

Compare

Open-weight multimodal mixture-of-experts model with 12B active parameters from 276B total and 524K token context.

Strengths
Efficient inference from a sparse parameter architecture while maintaining multimodal capabilities across long contexts.
Best for
Applications where reduced computational requirements and faster latency are prioritized over the scale of the full Inkling model.
Limitations
Smaller active parameter count compared to Inkling means reduced reasoning depth and knowledge capacity for complex tasks.

Input / 1M

$0.58

Output / 1M

$1.44

Cached input / 1M

$0.116

Context window

524K

Price history

Price per 1M tokens over timeOutput $1.44; Input $0.58 as of Jul 2026.$0$0.5$1.00$1.50Output on 31 Jul 2026: $1.44Out $1.44Input on 31 Jul 2026: $0.58In $0.58Jul 2026

Snapshots

Effective Input Output Cached in Note Source
31 Jul 2026 $0.58 $1.44 $0.116 Imported from OpenRouter openrouter.ai

More from Thinking Machines

Report a problem