T
Compare
Inkling Small (batch)
Thinking Machines · Multimodal · Released Jul 2026
Open-weight multimodal mixture-of-experts model with 12B active parameters out of 276B total and a 524K token context window, optimized for batch processing.
- Strengths
- Routes computation efficiently across sparse parameters, reducing inference overhead while maintaining multimodal capabilities across text and images.
- Best for
- Batch processing jobs where latency is not a constraint and you want to process large volumes of multimodal content with lower compute requirements than the larger Inkling variant.
- Limitations
- The smaller active parameter count makes it less capable than Inkling on complex reasoning tasks; the batch-only interface means real-time or streaming use cases are not supported.
Input / 1M
$0.5
Output / 1M
$1.20
Cached input / 1M
$0.1
Context window
524K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 29 Aug 2026 | $0.5 | $1.20 | $0.1 | Imported from OpenRouter | openrouter.ai |