prices synced 2026-09-11
T

Inkling Small (batch)

Thinking Machines · Multimodal · Released Jul 2026

Compare

Open-weight multimodal mixture-of-experts model with 12B active parameters out of 276B total and a 524K token context window, optimized for batch processing.

Strengths
Routes computation efficiently across sparse parameters, reducing inference overhead while maintaining multimodal capabilities across text and images.
Best for
Batch processing jobs where latency is not a constraint and you want to process large volumes of multimodal content with lower compute requirements than the larger Inkling variant.
Limitations
The smaller active parameter count makes it less capable than Inkling on complex reasoning tasks; the batch-only interface means real-time or streaming use cases are not supported.

Input / 1M

$0.5

Output / 1M

$1.20

Cached input / 1M

$0.1

Context window

524K

Price history

Price per 1M tokens over timeOutput $1.20; Input $0.5 as of Aug 2026.$0$0.5$1.00Output on 29 Aug 2026: $1.20Out $1.20Input on 29 Aug 2026: $0.5In $0.5Aug 2026

Snapshots

Effective Input Output Cached in Note Source
29 Aug 2026 $0.5 $1.20 $0.1 Imported from OpenRouter openrouter.ai

More from Thinking Machines

Report a problem