P
Compare
Perceptron Mk1.5
Perceptron · Multimodal · Released Sep 2026
Perceptron Mk1.5 is an embodied reasoning model that processes text, image, video, and audio to produce text responses with structured spatial annotations like points, boxes, polygons, and tracks.
- Strengths
- It handles multimodal input across video and audio in addition to text and images, and can ground responses with precise spatial or temporal annotations.
- Best for
- Tasks involving physical agents, scene understanding, object tracking, and spatial reasoning where structured output about locations or movements is needed.
- Limitations
- Specialized for embodied reasoning and spatial tasks; Mk1, the earlier general-purpose model, may be more suitable for text-centric workloads without the need for video, audio, or spatial annotations.
Input / 1M
$0.15
Output / 1M
$1.50
Cached input / 1M
—
Context window
36K
Price history
Snapshots
| Effective | Input | Output | Cached in | Note | Source |
|---|---|---|---|---|---|
| 25 Sep 2026 | $0.15 | $1.50 | — | Imported from OpenRouter | openrouter.ai |