LLM API pricing, tracked from launch.
424 language models across 88 providers — input, output, and cached rates per 1M tokens, updated daily, with full price history.
Price table
| Model | Input /1M | Output /1M | Cached in /1M | Context | Released | ||
|---|---|---|---|---|---|---|---|
| $150.00 | $600.00 | — | 200K | Mar 2025 | |||
| $75.00 | $300.00 | — | 200K | Mar 2025 | |||
| $30.00 | $180.00 | — | 1.05M | Mar 2026 | |||
| $30.00 | $180.00 | — | 1M | Apr 2026 | |||
| $21.00 | $168.00 | — | 400K | Dec 2025 | |||
| $15.00 | $120.00 | — | 400K | Oct 2025 | |||
| $15.00 | $90.00 | — | 1.05M | Mar 2026 | |||
| $15.00 | $90.00 | — | 1.05M | Apr 2026 | |||
| $10.50 | $84.00 | — | 400K | Dec 2025 | |||
| $20.00 | $80.00 | — | 200K | Jun 2025 | |||
| $20.00 | $80.00 | — | 200K | Jun 2025 | |||
| $7.50 | $60.00 | — | 400K | Oct 2025 | |||
| $10.00 | $50.00 | $1.00 | 1.05M | Sep 2026 | |||
| $10.00 | $50.00 | $0.25 | 1M | Sep 2026 | |||
| $10.00 | $50.00 | $1.00 | 1M | Jun 2026 | |||
| $10.00 | $40.00 | — | 200K | Jun 2025 | |||
| $10.00 | $40.00 | $2.50 | 200K | Oct 2025 | |||
| $7.50 | $37.50 | $0.75 | 200K | Aug 2025 | |||
|
S
Fugu Ultra v2 is a multi-agent orchestration model that routes requests across specialized language models rather than relying on a single monolithic architecture.
$5.00/$30.00I/O
1M ctx
Sep 2026
|
$5.00 | $30.00 | $0.5 | 1M | Sep 2026 | ||
| $7.50 | $30.00 | $3.75 | 200K | Dec 2024 | |||
|
S
A large-context model from Sakana with a 1-million-token window for processing extended documents and codebases.
$5.00/$30.00I/O
1M ctx
Jun 2026
|
$5.00 | $30.00 | $0.5 | 1M | Jun 2026 | ||
| $10.00 | $30.00 | — | 128K | Jan 2024 | |||
| $5.00 | $30.00 | $0.5 | 1M | Apr 2026 | |||
| $5.00 | $25.00 | $0.5 | 1.05M | Sep 2026 | |||
| $5.00 | $25.00 | $0.5 | 1.05M | Sep 2026 | |||
| $5.00 | $25.00 | $0.125 | 1M | Sep 2026 | |||
| $5.00 | $25.00 | $0.5 | 1M | Jun 2026 | |||
| $5.00 | $25.00 | $0.5 | 1M | Jul 2026 | |||
| $5.00 | $25.00 | $0.5 | 200K | Nov 2025 | |||
| $5.00 | $25.00 | $0.5 | 1M | Feb 2026 | |||
| $5.00 | $25.00 | $0.5 | 1M | Apr 2026 | |||
| $5.00 | $25.00 | $0.5 | 1M | May 2026 | |||
| $3.00 | $15.00 | $0.3 | 1.05M | Jul 2026 | |||
| $5.00 | $15.00 | — | 128K | Apr 2024 | |||
| $2.50 | $15.00 | $0.25 | 1.05M | Apr 2026 | |||
| $5.00 | $15.00 | — | 128K | May 2024 | |||
|
P
Sonar Pro is a large language model from Perplexity with a 200k token context window that integrates web search capabilities.
$3.00/$15.00I/O
200K ctx
Mar 2025
|
$3.00 | $15.00 | — | 200K | Mar 2025 | ||
| $2.50 | $15.00 | $0.25 | 1.05M | Mar 2026 | |||
| $3.00 | $15.00 | $0.3 | 1M | Sep 2025 | |||
| $3.00 | $15.00 | $0.3 | 1M | Feb 2026 | |||
| $1.75 | $14.00 | $0.175 | 400K | Dec 2025 | |||
| $1.75 | $14.00 | $0.175 | 400K | Jan 2026 | |||
| $1.75 | $14.00 | $0.175 | 400K | Feb 2026 | |||
| $1.75 | $14.00 | $0.175 | 128K | Mar 2026 | |||
| $2.50 | $12.50 | $0.25 | 200K | Nov 2025 | |||
| $2.50 | $12.50 | $0.25 | 1M | Feb 2026 | |||
| $2.50 | $12.50 | $0.25 | 1M | Apr 2026 | |||
| $2.50 | $12.50 | $0.25 | 1M | May 2026 | |||
| $2.50 | $12.50 | $0.25 | 1M | Jul 2026 | |||
|
A
Nova Premier 1.0 is Amazon's flagship general-purpose language model with a 1-million-token context window.
$2.50/$12.50I/O
1M ctx
Oct 2025
|
$2.50 | $12.50 | $0.625 | 1M | Oct 2025 | ||
| $2.00 | $12.00 | $0.2 | 1.05M | Jul 2026 | |||
| $2.00 | $12.00 | $0.2 | 1M | Nov 2025 | |||
| $2.00 | $12.00 | $0.2 | 1M | Feb 2026 | |||
| $2.34 | $11.70 | $0.261 | 1.05M | Jul 2026 | |||
| $2.00 | $10.00 | $0.2 | 1.05M | Jul 2026 | |||
| $2.00 | $10.00 | $0.2 | 1M | Jun 2026 | |||
| $2.50 | $10.00 | — | 128K | Apr 2024 | |||
| $2.50 | $10.00 | $1.25 | 128K | Aug 2024 | |||
| $2.50 | $10.00 | — | 128K | Aug 2024 | |||
|
I
Inflection 3 Pi is a conversational model that powers the Pi chatbot, designed for dialogue with built-in knowledge of recent events.
$2.50/$10.00I/O
8K ctx
Oct 2024
|
$2.50 | $10.00 | — | 8K | Oct 2024 | ||
|
I
Inflection 3 Productivity is a compact instruction-following model optimized for structured outputs and rule-based tasks with access to recent information.
$2.50/$10.00I/O
8K ctx
Oct 2024
|
$2.50 | $10.00 | — | 8K | Oct 2024 | ||
| $2.50 | $10.00 | $1.25 | 128K | Nov 2024 | |||
| $2.50 | $10.00 | — | 128K | Mar 2025 | |||
| $2.50 | $10.00 | — | 256K | Mar 2025 | |||
| $1.25 | $10.00 | $0.125 | 128K | Aug 2025 | |||
| $1.25 | $10.00 | $0.125 | 400K | Sep 2025 | |||
| $1.25 | $10.00 | $0.13 | 400K | Nov 2025 | |||
| $1.25 | $10.00 | $0.125 | 400K | Nov 2025 | |||
| $1.25 | $10.00 | $0.125 | 400K | Dec 2025 | |||
| $2.50 | $10.00 | — | 128K | Jan 2026 | |||
| $1.25 | $10.00 | $0.125 | 1M | Jun 2025 | |||
| $1.25 | $10.00 | $0.125 | 1M | Aug 2025 | |||
| $1.50 | $9.00 | $0.15 | 1M | May 2026 | |||
|
A
Aion-1.0 is a large language model from AionLabs with a 131k-token context window, released before the company's shift toward multi-model systems.
$4.00/$8.00I/O
131K ctx
Feb 2025
|
$4.00 | $8.00 | — | 131K | Feb 2025 | ||
|
P
A research-focused model from Perplexity that autonomously searches, retrieves, and synthesizes information across complex topics with a 128k token context window.
$2.00/$8.00I/O
128K ctx
Mar 2025
|
$2.00 | $8.00 | — | 128K | Mar 2025 | ||
|
P
A reasoning model based on DeepSeek R1 that uses chain-of-thought processing to work through complex problems step-by-step, with access to Perplexity's search integration.
$2.00/$8.00I/O
128K ctx
Mar 2025
|
$2.00 | $8.00 | — | 128K | Mar 2025 | ||
|
A
Jamba Large 1.7 is a hybrid architecture model from AI21 that combines Transformers with Mamba state-space layers to handle extended context efficiently.
$2.00/$8.00I/O
256K ctx
Aug 2025
|
$2.00 | $8.00 | — | 256K | Aug 2025 | ||
| $2.00 | $8.00 | $0.5 | 200K | Oct 2025 | |||
| $2.00 | $8.00 | $0.5 | 1M | Apr 2025 | |||
| $2.00 | $8.00 | $0.5 | 200K | Apr 2025 | |||
| $1.50 | $7.50 | $0.15 | 1M | Feb 2026 | |||
| $1.50 | $7.50 | $0.15 | 1M | Sep 2025 | |||
| $1.25 | $7.50 | $0.125 | 1.05M | Mar 2026 | |||
| $2.50 | $7.50 | — | 1M | May 2026 | |||
| $1.50 | $7.50 | — | 128K | Apr 2026 | |||
| $0.875 | $7.00 | $0.0875 | 400K | Dec 2025 | |||
| $1.03 | $6.16 | — | 262K | Apr 2026 | |||
|
S
Fugu Max is a multi-agent orchestration model that routes requests across specialized language models to balance performance and efficiency.
$2.00/$6.00I/O
1M ctx
Sep 2026
|
$2.00 | $6.00 | $0.25 | 1M | Sep 2026 | ||
| $2.00 | $6.00 | $0.25 | 1.01M | Aug 2026 | |||
| $1.00 | $6.00 | $0.1 | 1.05M | Jul 2026 | |||
| $2.00 | $6.00 | $0.5 | 500K | Aug 2026 | |||
| $2.00 | $6.00 | $0.25 | 1M | Aug 2026 | |||
| $1.00 | $6.00 | $0.1 | 1.05M | Jul 2026 | |||
| $2.00 | $6.00 | $0.25 | 1M | Aug 2026 | |||
| $1.00 | $6.00 | — | 1.05M | Feb 2026 | |||
| $2.00 | $6.00 | $0.3 | 500K | Jul 2026 | |||
|
A
A multi-model system from AionLabs that combines specialized models for collaborative generation, designed around roleplaying and storytelling tasks.
$3.00/$6.00I/O
131K ctx
Jul 2026
|
$3.00 | $6.00 | $0.75 | 131K | Jul 2026 | ||
| $2.00 | $6.00 | $0.2 | 65K | Apr 2024 | |||
| $2.00 | $6.00 | $0.2 | 131K | Nov 2024 | |||
|
W
Palmyra X5 is Writer's large language model with a 1 million token context window.
$0.6/$6.00I/O
1.04M ctx
Jan 2026
|
$0.6 | $6.00 | — | 1.04M | Jan 2026 | ||
| $1.25 | $5.00 | $0.625 | 128K | May 2024 | |||
| $0.625 | $5.00 | $0.0625 | 400K | Sep 2025 | |||
| $1.00 | $5.00 | $0.1 | 1.05M | Jul 2026 | |||
| $1.00 | $5.00 | $0.1 | 1.05M | Jul 2026 | |||
| $0.625 | $5.00 | $0.125 | 1.05M | Jun 2025 | |||
| $0.625 | $5.00 | $0.0625 | 400K | Aug 2025 | |||
| $0.625 | $5.00 | $0.0625 | 400K | Nov 2025 | |||
| $1.00 | $5.00 | $0.1 | 1M | Jun 2026 | |||
|
A
This is a series of models designed to replicate the prose quality of the Claude 3 models, specifically Sonnet(https://openrouter.ai/anthropic/claude-3.5-sonnet) and Opus(https://openrouter.ai/anthropic/claude-3-opus). The model is fine-tuned on top of [Qwen2.5 72B](https://openrouter.ai/qwen/qwen-2.5-72b-instruct).
$2.50/$5.00I/O
32K ctx
Oct 2024
|
$2.50 | $5.00 | — | 32K | Oct 2024 | ||
| $1.00 | $5.00 | $0.1 | 200K | Oct 2025 | |||
| $0.75 | $4.50 | $0.075 | 1.05M | May 2026 | |||
| $0.75 | $4.50 | $0.075 | 400K | Mar 2026 | |||
|
Z
GLM 5.3 is Z.ai's latest foundation model with a 1M-token context window for processing extended documents and complex reasoning tasks.
$1.40/$4.40I/O
1.05M ctx
Aug 2026
|
$1.40 | $4.40 | $0.26 | 1.05M | Aug 2026 | ||
| $1.10 | $4.40 | $0.55 | 200K | Feb 2025 | |||
| $1.10 | $4.40 | $0.275 | 200K | Apr 2025 | |||
| $1.10 | $4.40 | $0.275 | 200K | Apr 2025 | |||
| $1.25 | $4.25 | $0.15 | 1.05M | Sep 2026 | |||
| $1.25 | $4.25 | $0.15 | 1.05M | Aug 2026 | |||
| $1.25 | $4.25 | — | 1M | Jul 2026 | |||
|
T
Open-weight multimodal mixture-of-experts model with 41B active parameters from 975B total and a 524K token context window.
$1.00/$4.05I/O
524K ctx
Jul 2026
|
$1.00 | $4.05 | $0.17 | 524K | Jul 2026 | ||
|
T
An open-weight multimodal mixture-of-experts model with 12B active parameters from 276B total parameters and a 524K token context window.
$1.00/$4.05I/O
1.05M ctx
Jul 2026
|
$1.00 | $4.05 | $0.17 | 1.05M | Jul 2026 | ||
|
S
A Japanese-specialized reasoning model from Sakana AI, built on Kimi K2.6 with additional training for Japanese language and business contexts.
$0.95/$4.00I/O
262K ctx
Aug 2026
|
$0.95 | $4.00 | $0.15 | 262K | Aug 2026 | ||
| $1.00 | $4.00 | $0.25 | 1.05M | Apr 2025 | |||
| $1.00 | $4.00 | $0.25 | 200K | Apr 2025 | |||
| $0.95 | $4.00 | $0.19 | 262K | Jun 2026 | |||
|
Z
GLM 5V Turbo is Z.ai's optimized multimodal inference variant with support for text, image, and video inputs alongside a 202K-token context window.
$1.20/$4.00I/O
202K ctx
Apr 2026
|
$1.20 | $4.00 | $0.24 | 202K | Apr 2026 | ||
| $3.00 | $4.00 | — | 16K | Aug 2023 | |||
| $0.4 | $4.00 | — | 131K | Sep 2025 | |||
|
Z
GLM 5 Turbo is Z.ai's optimized inference variant of their GLM 5 foundation model, offering faster processing with a 202K-token context window.
$1.20/$4.00I/O
202K ctx
Mar 2026
|
$1.20 | $4.00 | $0.24 | 202K | Mar 2026 | ||
| $0.95 | $4.00 | $0.16 | 256K | Apr 2026 | |||
| $0.78 | $3.90 | — | 262K | Feb 2026 | |||
| $0.75 | $3.75 | $0.075 | 1.05M | Sep 2026 | |||
| $0.75 | $3.75 | — | 262K | Apr 2026 | |||
| $0.75 | $3.75 | $0.075 | 1.05M | Aug 2026 | |||
| $0.75 | $3.75 | $0.075 | 1.05M | Jul 2026 | |||
|
N
A 55-billion-parameter mixture-of-experts model from NVIDIA with a hybrid Transformer-Mamba architecture and 512K token context window.
$0.6/$3.60I/O
512K ctx
Jun 2026
|
$0.6 | $3.60 | $0.2 | 512K | Jun 2026 | ||
| $0.71 | $3.50 | $0.15 | 262K | Jun 2026 | |||
| $0.55 | $3.50 | $0.225 | 262K | Feb 2026 | |||
|
S
A routing model that directs requests to other LLMs based on task characteristics and cost-performance tradeoffs.
$0.85/$3.40I/O
131K ctx
Jul 2025
|
$0.85 | $3.40 | — | 131K | Jul 2025 | ||
| $0.65 | $3.25 | $0.13 | 1M | Sep 2025 | |||
|
A
A multimodal model from Amazon designed to balance accuracy and speed across a broad range of tasks, with a 300K-token context window.
$0.8/$3.20I/O
300K ctx
Dec 2024
|
$0.8 | $3.20 | — | 300K | Dec 2024 | ||
|
N
A 55-billion-parameter mixture-of-experts model from NVIDIA with a hybrid Transformer-Mamba architecture and 512K token context window.
$0.625/$3.12I/O
256K ctx
Jun 2026
|
$0.625 | $3.12 | $0.1875 | 256K | Jun 2026 | ||
|
Z
GLM 5.1 is Z.ai's foundation model with a 202K-token context window for processing extended documents and conversations.
$0.966/$3.04I/O
200K ctx
Apr 2026
|
$0.966 | $3.04 | $0.1794 | 200K | Apr 2026 | ||
| $0.42 | $3.00 | $0.085 | 1M | Aug 2026 | |||
|
B
Seed-2.0-Code is a specialized coding model from ByteDance Seed with a 262k token context window, designed for software development tasks.
$0.5/$3.00I/O
262K ctx
Jul 2026
|
$0.5 | $3.00 | — | 262K | Jul 2026 | ||
|
S
This is [Sao10K](/sao10k)'s experiment over [Euryale v2.2](/sao10k/l3.1-euryale-70b).
$3.00/$3.00I/O
16K ctx
Jan 2025
|
$3.00 | $3.00 | — | 16K | Jan 2025 | ||
|
N
Hermes 4 405B is a 405-billion parameter instruction-tuned language model from Nous with a 131K token context window.
$1.00/$3.00I/O
131K ctx
Aug 2025
|
$1.00 | $3.00 | — | 131K | Aug 2025 | ||
|
R
Relace Search is a specialized model designed to locate and extract relevant code segments and documentation from large codebases and repositories.
$1.00/$3.00I/O
256K ctx
Dec 2025
|
$1.00 | $3.00 | — | 256K | Dec 2025 | ||
| $0.6 | $3.00 | $0.1 | 256K | Jan 2026 | |||
| $0.5 | $3.00 | $0.05 | 1M | Dec 2025 | |||
|
K
A coding-focused language model from Kwaipilot with a 256K token context window.
$0.74/$2.96I/O
262K ctx
Jul 2026
|
$0.74 | $2.96 | $0.15 | 262K | Jul 2026 | ||
|
T
Hy4 preview is a mixture-of-experts model from Tencent with 49B active parameters and a 1M-token context window, designed for coding, tool use, and multi-step reasoning tasks.
$0.834/$2.50I/O
1.05M ctx
Aug 2026
|
$0.834 | $2.50 | $0.042 | 1.05M | Aug 2026 | ||
|
B
Seed 2.1 Turbo is a multimodal model from ByteDance Seed designed for coding tasks and agentic workflows, with a 262k token context window.
$0.5/$2.50I/O
262K ctx
Aug 2026
|
$0.5 | $2.50 | — | 262K | Aug 2026 | ||
| $0.5 | $2.50 | $0.05 | 200K | Oct 2025 | |||
| $0.3 | $2.50 | $0.03 | 1.05M | Jul 2026 | |||
| $0.7 | $2.50 | — | 64K | Jan 2025 | |||
| $0.6 | $2.50 | — | 262K | Sep 2025 | |||
| $0.6 | $2.50 | $0.15 | 262K | Nov 2025 | |||
|
A
Nova 2 Lite is Amazon's lightweight multimodal model that processes text, images, and video with a 1-million-token context window.
$0.3/$2.50I/O
1M ctx
Dec 2025
|
$0.3 | $2.50 | — | 1M | Dec 2025 | ||
| $1.25 | $2.50 | $0.2 | 2M | Mar 2026 | |||
| $1.25 | $2.50 | $0.2 | 2M | Mar 2026 | |||
| $1.25 | $2.50 | $0.2 | 1M | Apr 2026 | |||
| $0.3 | $2.50 | $0.03 | 1M | Jun 2025 | |||
| $0.2 | $2.40 | — | 81K | Aug 2025 | |||
| $0.2 | $2.40 | — | 131K | Oct 2025 | |||
| $0.6 | $2.40 | — | 128K | Jan 2026 | |||
| $0.57 | $2.30 | — | 131K | Jul 2025 | |||
| $0.23 | $2.30 | — | 131K | Jul 2025 | |||
| $0.375 | $2.25 | $0.0375 | 400K | Mar 2026 | |||
|
Z
GLM 5.3 (batch) is Z.ai's reasoning model optimized for asynchronous batch processing, with a 1M-token context window for handling extended documents and complex multi-step tasks.
$0.7/$2.20I/O
1.05M ctx
Aug 2026
|
$0.7 | $2.20 | $0.13 | 1.05M | Aug 2026 | ||
| $0.55 | $2.20 | $0.275 | 200K | Jan 2025 | |||
| $0.55 | $2.20 | $0.275 | 200K | Feb 2025 | |||
| $0.55 | $2.20 | $0.1375 | 200K | Apr 2025 | |||
| $0.55 | $2.20 | $0.1375 | 200K | Apr 2025 | |||
|
Z
GLM 5.2 is Z.ai's foundation model with a 1M-token context window for processing extended documents and complex reasoning tasks.
$0.7/$2.20I/O
1.05M ctx
Jun 2026
|
$0.7 | $2.20 | $0.07 | 1.05M | Jun 2026 | ||
|
M
MiniMax M1 is a large language model with a 1 million token context window from MiniMax's earlier model lineups.
$0.55/$2.20I/O
1M ctx
Jun 2025
|
$0.55 | $2.20 | — | 1M | Jun 2025 | ||
|
Z
GLM 4.5 is Z.ai's foundation model released in mid-2025 with a 131K-token context window for processing documents and multi-turn conversations.
$0.6/$2.20I/O
131K ctx
Jul 2025
|
$0.6 | $2.20 | $0.11 | 131K | Jul 2025 | ||
| $0.5 | $2.15 | $0.35 | 163K | May 2025 | |||
| $0.18 | $2.10 | — | 131K | Oct 2025 | |||
| $0.26 | $2.08 | — | 262K | Feb 2026 | |||
| $1.00 | $2.00 | $0.16 | 1M | Apr 2026 | |||
|
Z
GLM 5.2 is Z.ai's foundation model with a 1M-token context window for processing extended documents and complex reasoning tasks.
$0.6/$2.00I/O
202K ctx
Jun 2026
|
$0.6 | $2.00 | $0.15 | 202K | Jun 2026 | ||
| $1.50 | $2.00 | — | 4K | Sep 2023 | |||
| $1.00 | $2.00 | — | 4K | Jan 2024 | |||
| $0.4 | $2.00 | $0.04 | 131K | May 2025 | |||
| $0.25 | $2.00 | $0.025 | 400K | Aug 2025 | |||
| $0.4 | $2.00 | $0.04 | 131K | Aug 2025 | |||
| $0.25 | $2.00 | $0.03 | 400K | Nov 2025 | |||
| $0.4 | $2.00 | $0.04 | 262K | Dec 2025 | |||
|
B
A general-purpose model from ByteDance Seed with 262k token context window, superseded by the Seed-2.0 family released in 2026.
$0.25/$2.00I/O
262K ctx
Dec 2025
|
$0.25 | $2.00 | — | 262K | Dec 2025 | ||
|
B
A lightweight general-purpose model from ByteDance Seed with a 262k token context window, positioned as a faster inference alternative within the Seed-2.0 family.
$0.25/$2.00I/O
262K ctx
Mar 2026
|
$0.25 | $2.00 | — | 262K | Mar 2026 | ||
| $0.3 | $2.00 | $0.03 | 262K | Apr 2026 | |||
| $1.00 | $2.00 | $0.2 | 256K | May 2026 | |||
| $0.66 | $1.98 | $0.022 | 1.05M | Aug 2026 | |||
| $0.325 | $1.95 | — | 1M | Apr 2026 | |||
|
Z
GLM 5 is Z.ai's foundation model with a 198K-token context window for processing extended documents and multi-turn conversations.
$0.6/$1.92I/O
198K ctx
Feb 2026
|
$0.6 | $1.92 | $0.12 | 198K | Feb 2026 | ||
|
M
Morph V3 Large is a code-focused model with a 262k token context window from Morph.
$0.9/$1.90I/O
262K ctx
Jul 2025
|
$0.9 | $1.90 | — | 262K | Jul 2025 | ||
| $0.21 | $1.90 | $0.1 | 131K | Sep 2025 | |||
| $0.9478 | $1.90 | $0.079 | 1.02M | Apr 2026 | |||
| $0.375 | $1.88 | $0.0375 | 1.05M | Sep 2026 | |||
| $0.375 | $1.88 | $0.0375 | 1.05M | Aug 2026 | |||
| $0.375 | $1.88 | $0.0375 | 1.05M | Jul 2026 | |||
| $0.455 | $1.82 | — | 131K | Apr 2025 | |||
|
Z
GLM 4.5V is Z.ai's multimodal model supporting text, image, and video inputs with a 65K-token context window.
$0.6/$1.80I/O
65K ctx
Aug 2025
|
$0.6 | $1.80 | $0.11 | 65K | Aug 2025 | ||
| $0.3 | $1.80 | — | 1M | Apr 2026 | |||
|
Z
GLM 4.6 is Z.ai's foundation model with a 198K-token context window released in late 2025.
$0.43/$1.75I/O
198K ctx
Sep 2025
|
$0.43 | $1.75 | $0.08 | 198K | Sep 2025 | ||
|
Z
GLM 4.7 is Z.ai's foundation model with a 202K-token context window for processing extended documents and conversations.
$0.4/$1.75I/O
202K ctx
Dec 2025
|
$0.4 | $1.75 | $0.08 | 202K | Dec 2025 | ||
|
A
Aion-RP 1.0 is an 8-billion-parameter language model from AionLabs designed for roleplaying and narrative generation tasks.
$0.8/$1.60I/O
32K ctx
Feb 2025
|
$0.8 | $1.60 | — | 32K | Feb 2025 | ||
|
A
Aion-2.0 is a multi-model system from AionLabs that builds on the reasoning and coding strengths of Aion-1.0 with broader capability across language tasks.
$0.8/$1.60I/O
131K ctx
Feb 2026
|
$0.8 | $1.60 | $0.2 | 131K | Feb 2026 | ||
| $0.4 | $1.60 | $0.1 | 1M | Apr 2025 | |||
| $0.26 | $1.56 | — | 1M | Feb 2026 | |||
| $0.195 | $1.56 | — | 262K | Feb 2026 | |||
| $0.25 | $1.50 | — | 1.05M | Dec 2025 | |||
| $0.5 | $1.50 | — | 16K | May 2023 | |||
| $0.25 | $1.50 | $0.025 | 1.05M | May 2026 | |||
|
P
A 32k-token context model from Perceptron designed for general-purpose language tasks.
$0.15/$1.50I/O
32K ctx
May 2026
|
$0.15 | $1.50 | — | 32K | May 2026 | ||
| $0.5 | $1.50 | — | 262K | Dec 2025 | |||
| $0.36 | $1.43 | — | 256K | Sep 2025 | |||
|
A
Aion-3.0-Mini is a multi-model roleplaying and storytelling system from AionLabs built on the DeepSeek family, using collaborative generation from specialized models.
$0.7/$1.40I/O
131K ctx
Jul 2026
|
$0.7 | $1.40 | $0.18 | 131K | Jul 2026 | ||
|
A
Aion-1.0-Mini is a smaller parameter variant of Aion-1.0 from AionLabs, sharing the same 131k-token context window as its full-size sibling.
$0.7/$1.40I/O
131K ctx
Feb 2025
|
$0.7 | $1.40 | — | 131K | Feb 2025 | ||
| $0.32 | $1.28 | $0.064 | 1M | Jun 2026 | |||
| $0.15 | $1.25 | $0.03 | 1.05M | Jun 2025 | |||
| $0.15 | $1.25 | $0.015 | 1.05M | Jul 2026 | |||
|
B
ERNIE 4.5 VL 424B A47B is Baidu's multimodal model that processes both text and images with a 123,000-token context window.
$0.42/$1.25I/O
123K ctx
Jun 2025
|
$0.42 | $1.25 | — | 123K | Jun 2025 | ||
|
R
Relace Apply 3 is a general-purpose LLM with a 256k token context window for processing large documents and code repositories.
$0.85/$1.25I/O
256K ctx
Sep 2025
|
$0.85 | $1.25 | — | 256K | Sep 2025 | ||
|
D
Cogito v2.1 is a 671-billion parameter language model with a 128k token context window.
$1.25/$1.25I/O
128K ctx
Nov 2025
|
$1.25 | $1.25 | — | 128K | Nov 2025 | ||
| $0.3125 | $1.25 | $0.1562 | 256K | Feb 2026 | |||
| $0.2 | $1.25 | $0.02 | 400K | Mar 2026 | |||
|
T
Open-weight multimodal mixture-of-experts model with 12B active parameters out of 276B total and a 524K token context window, optimized for batch processing.
$0.5/$1.20I/O
524K ctx
Jul 2026
|
$0.5 | $1.20 | $0.1 | 524K | Jul 2026 | ||
|
T
Open-weight multimodal mixture-of-experts model with 12B active parameters from 276B total and 524K token context.
$0.45/$1.20I/O
524K ctx
Jul 2026
|
$0.45 | $1.20 | $0.1 | 524K | Jul 2026 | ||
|
M
MiniMax M3 is a multimodal foundation model that processes text, image, and video inputs to produce text output, with a 524K token context window.
$0.3/$1.20I/O
524K ctx
May 2026
|
$0.3 | $1.20 | $0.06 | 524K | May 2026 | ||
|
M
A sparse mixture-of-experts model from Meituan with 48B active parameters designed to handle very long contexts up to 1M tokens.
$0.3/$1.20I/O
1.05M ctx
Jul 2026
|
$0.3 | $1.20 | $0.006 | 1.05M | Jul 2026 | ||
| $0.2 | $1.20 | $0.02 | 1.05M | Jul 2026 | |||
|
A
Virtuoso Large is a general-purpose language model from Arcee AI with a 131k token context window.
$0.75/$1.20I/O
131K ctx
May 2025
|
$0.75 | $1.20 | — | 131K | May 2025 | ||
|
M
Morph V3 Fast is a smaller, faster variant of the Morph V3 line with an 81k token context window.
$0.8/$1.20I/O
81K ctx
Jul 2025
|
$0.8 | $1.20 | — | 81K | Jul 2025 | ||
| $0.15 | $1.20 | — | 262K | Sep 2025 | |||
|
M
MiniMax M2.1 is a large language model with a 204K token context window, released as part of MiniMax's M2 generation lineup.
$0.3/$1.20I/O
204K ctx
Dec 2025
|
$0.3 | $1.20 | $0.03 | 204K | Dec 2025 | ||
|
M
MiniMax M2-her is a large language model with a 65K token context window, positioned as a lightweight tier within MiniMax's M2 generation.
$0.3/$1.20I/O
65K ctx
Jan 2026
|
$0.3 | $1.20 | $0.03 | 65K | Jan 2026 | ||
|
M
MiniMax M2.7 is a large language model positioned between the M2.5 productivity model and the M3 multimodal generation, with a 204K token context window.
$0.3/$1.20I/O
204K ctx
Mar 2026
|
$0.3 | $1.20 | $0.06 | 204K | Mar 2026 | ||
|
K
A code-focused language model from Kwaipilot with a 256K token context window.
$0.3/$1.20I/O
262K ctx
Mar 2026
|
$0.3 | $1.20 | $0.06 | 262K | Mar 2026 | ||
|
M
MiniMax M3 is a multimodal foundation model that processes text, image, and video inputs to produce text output, with a 524K token context window.
$0.3/$1.20I/O
524K ctx
May 2026
|
$0.3 | $1.20 | $0.06 | 524K | May 2026 | ||
|
S
Step 3.7 Flash is a sparse Mixture of Experts model from StepFun that activates a subset of its parameters per token, designed for fast inference across a 256K token context.
$0.2/$1.15I/O
256K ctx
May 2026
|
$0.2 | $1.15 | $0.04 | 256K | May 2026 | ||
| $0.1875 | $1.12 | — | 1M | Apr 2026 | |||
| $0.3 | $1.10 | $0.04 | 131K | Aug 2026 | |||
|
M
MiniMax-01 is a multimodal model combining text generation and image understanding with 456 billion parameters and a 1 million token context window.
$0.2/$1.10I/O
1M ctx
Jan 2025
|
$0.2 | $1.10 | — | 1M | Jan 2025 | ||
| $0.09 | $1.10 | — | 262K | Sep 2025 | |||
|
P
INTELLECT-3 is a large language model from Prime Intellect with a 131k token context window.
$0.2/$1.10I/O
131K ctx
Nov 2025
|
$0.2 | $1.10 | — | 131K | Nov 2025 | ||
|
M
MiniMax M2.5 is a large language model with a 204K token context window, positioned as a general-purpose productivity model within MiniMax's lineup.
$0.27/$1.08I/O
200K ctx
Feb 2026
|
$0.27 | $1.08 | $0.027 | 200K | Feb 2026 | ||
|
M
MiniMax M2 is a large language model with a 204K token context window from MiniMax's M2 generation.
$0.255/$1.02I/O
204K ctx
Oct 2025
|
$0.255 | $1.02 | — | 204K | Oct 2025 | ||
| $0.2 | $1.00 | $0.02 | 131K | Aug 2025 | |||
| $0.125 | $1.00 | $0.0125 | 400K | Aug 2025 | |||
|
N
Nex-N2-Pro is a mixture-of-experts model from Nex AGI that processes text and image inputs with a 262K token context window.
$0.25/$1.00I/O
262K ctx
Jun 2026
|
$0.25 | $1.00 | $0.025 | 262K | Jun 2026 | ||
|
N
Hermes 3 405B Instruct is a large generalist language model from Nous with a 131k token context window, designed for multi-turn conversations and agentic tasks.
$1.00/$1.00I/O
131K ctx
Aug 2024
|
$1.00 | $1.00 | — | 131K | Aug 2024 | ||
| $0.66 | $1.00 | — | 32K | Nov 2024 | |||
|
P
Perplexity's Sonar is a lightweight model designed for question-answering tasks with citation support and customizable source integration.
$1.00/$1.00I/O
127K ctx
Jan 2025
|
$1.00 | $1.00 | — | 127K | Jan 2025 | ||
| $0.8 | $1.00 | $0.4 | 128K | Feb 2025 | |||
| $0.25 | $1.00 | — | 163K | Mar 2025 | |||
| $0.3 | $1.00 | $0.1 | 262K | Jul 2025 | |||
| $0.27 | $1.00 | $0.135 | 131K | Sep 2025 | |||
| $0.195 | $0.975 | $0.039 | 1M | Sep 2025 | |||
| $0.39 | $0.97 | — | 262K | Apr 2026 | |||
| $0.25 | $0.95 | $0.13 | 163K | Aug 2025 | |||
| $0.2275 | $0.91 | — | 131K | Apr 2025 | |||
|
V
An instruction-tuned 24B model based on Mistral that minimizes refusals on sensitive or controversial topics.
$0.2/$0.9I/O
128K ctx
Jul 2025
|
$0.2 | $0.9 | — | 128K | Jul 2025 | ||
| $0.3 | $0.9 | $0.03 | 256K | Aug 2025 | |||
|
Z
GLM 4.6V is Z.ai's multimodal model supporting text, image, and video inputs with a 131K-token context window.
$0.3/$0.9I/O
131K ctx
Dec 2025
|
$0.3 | $0.9 | $0.055 | 131K | Dec 2025 | ||
| $0.1 | $0.9 | $0.05 | 262K | Apr 2026 | |||
| $0.22 | $0.88 | — | 262K | Jul 2025 | |||
|
X
Xiaomi's omnimodal model that processes text, images, and video in a single forward pass, with a 1M token context window.
$0.435/$0.87I/O
1.05M ctx
Apr 2026
|
$0.435 | $0.87 | $0.0036 | 1.05M | Apr 2026 | ||
| $0.435 | $0.87 | $0.0036 | 1M | Apr 2026 | |||
|
S
Euryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.1](/models/sao10k/l3-euryale-70b).
$0.85/$0.85I/O
131K ctx
Aug 2024
|
$0.85 | $0.85 | — | 131K | Aug 2024 | ||
|
Z
GLM 4.5 Air is Z.ai's lightweight inference variant of the GLM 4.5 foundation model, designed for faster processing with a 131K-token context window.
$0.13/$0.85I/O
131K ctx
Jul 2025
|
$0.13 | $0.85 | $0.025 | 131K | Jul 2025 | ||
| $0.2 | $0.8 | $0.05 | 1.05M | Apr 2025 | |||
| $0.8 | $0.8 | — | 8K | Jan 2025 | |||
|
T
Skyfall 36B v2 is an enhanced iteration of Mistral Small 2501, specifically fine-tuned for improved creativity, nuanced writing, role-playing, and coherent storytelling.
$0.55/$0.8I/O
32K ctx
Mar 2025
|
$0.55 | $0.8 | $0.25 | 32K | Mar 2025 | ||
|
A
Coder Large is a code-focused language model from Arcee AI with a 32k token context window.
$0.5/$0.8I/O
32K ctx
May 2025
|
$0.5 | $0.8 | — | 32K | May 2025 | ||
| $0.12 | $0.8 | $0.07 | 262K | Feb 2026 | |||
|
A
A large reasoning model from Arcee AI with a 262k token context window, designed to handle extended reasoning tasks and complex problem-solving.
$0.25/$0.8I/O
262K ctx
Apr 2026
|
$0.25 | $0.8 | $0.06 | 262K | Apr 2026 | ||
| $0.26 | $0.78 | $0.052 | 1M | Feb 2025 | |||
| $0.26 | $0.78 | — | 1M | Sep 2025 | |||
|
I
Mercury 2.5 Preview is an early version of Inception's diffusion-based language model that generates and refines multiple tokens in parallel.
$0.2/$0.75I/O
260K ctx
Aug 2026
|
$0.2 | $0.75 | $0.02 | 260K | Aug 2026 | ||
| $0.175 | $0.75 | $0.02 | 131K | Aug 2026 | |||
| $0.25 | $0.75 | $0.025 | 262K | Dec 2025 | |||
| $0.25 | $0.75 | — | 16K | May 2023 | |||
| $0.125 | $0.75 | $0.0125 | 1.05M | May 2026 | |||
|
M
An attempt to recreate Claude-style verbosity, but don't expect the same level of coherence or memory. Meant for use in roleplay/narrative situations.
$0.4/$0.75I/O
8K ctx
Aug 2023
|
$0.4 | $0.75 | — | 8K | Aug 2023 | ||
|
S
Euryale L3.3 70B is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.2](/models/sao10k/l3-euryale-70b).
$0.65/$0.75I/O
131K ctx
Dec 2024
|
$0.65 | $0.75 | — | 131K | Dec 2024 | ||
|
I
Mercury 2 is a large language model from Inception with a 128k token context window.
$0.25/$0.75I/O
128K ctx
Mar 2026
|
$0.25 | $0.75 | $0.025 | 128K | Mar 2026 | ||
| $0.51 | $0.74 | — | 8K | Apr 2024 | |||
|
N
Hermes 3 70B Instruct is a 70-billion parameter generalist model with a 131K token context window, trained to follow instructions and support agentic workflows.
$0.7/$0.7I/O
131K ctx
Aug 2024
|
$0.7 | $0.7 | — | 131K | Aug 2024 | ||
| $0.22 | $0.66 | $0.007 | 1.05M | Aug 2026 | |||
|
U
A recreation trial of the original MythoMax-L2-B13 but with updated models. #merge
$0.35/$0.65I/O
6K ctx
Jul 2023
|
$0.35 | $0.65 | — | 6K | Jul 2023 | ||
| $0.65 | $0.65 | — | 8K | Jul 2024 | |||
| $0.1 | $0.625 | $0.01 | 400K | Mar 2026 | |||
|
I
Ling-2.6-1T is a trillion-parameter instruct model from inclusionAI with a 262K-token context window.
$0.075/$0.625I/O
262K ctx
Apr 2026
|
$0.075 | $0.625 | $0.015 | 262K | Apr 2026 | ||
|
I
A trillion-parameter instruct model from inclusionAI with a 262K-token context window.
$0.075/$0.625I/O
262K ctx
May 2026
|
$0.075 | $0.625 | $0.015 | 262K | May 2026 | ||
|
M
A 176-billion-parameter mixture-of-experts model from Microsoft that combines eight 22-billion experts with a large context window for handling extended documents and conversations.
$0.62/$0.62I/O
65K ctx
Apr 2024
|
$0.62 | $0.62 | — | 65K | Apr 2024 | ||
| $0.15 | $0.6 | $0.003 | 1.05M | Sep 2026 | |||
| $0.15 | $0.6 | — | 131K | Aug 2025 | |||
| $0.1 | $0.6 | $0.01 | 1.05M | Jul 2026 | |||
| $0.1 | $0.6 | $0.01 | 1.05M | Jul 2026 | |||
|
K
A code-focused language model from Kwaipilot with a 256k token context window.
$0.15/$0.6I/O
256K ctx
Jul 2026
|
$0.15 | $0.6 | $0.03 | 256K | Jul 2026 | ||
| $0.15 | $0.6 | — | 128K | Mar 2024 | |||
| $0.15 | $0.6 | $0.075 | 128K | Jul 2024 | |||
| $0.2 | $0.6 | $0.02 | 32K | Feb 2025 | |||
| $0.15 | $0.6 | — | 128K | Mar 2025 | |||
| $0.15 | $0.6 | — | 262K | Oct 2025 | |||
|
U
Solar Pro 3 is a large language model from Upstage with a 131K token context window.
$0.15/$0.6I/O
131K ctx
Jan 2026
|
$0.15 | $0.6 | $0.015 | 131K | Jan 2026 | ||
|
T
Hy3 preview is Tencent's mixture-of-experts reasoning model with a 262k-token context window.
$0.18/$0.6I/O
262K ctx
Apr 2026
|
$0.18 | $0.6 | $0.06 | 262K | Apr 2026 | ||
| $0.15 | $0.6 | — | 1M | Apr 2025 | |||
|
T
Hunyuan A13B Instruct is a 13-billion-parameter instruction-tuned language model from Tencent with a 131k-token context window.
$0.14/$0.57I/O
131K ctx
Jul 2025
|
$0.14 | $0.57 | — | 131K | Jul 2025 | ||
| $0.351 | $0.555 | — | 128K | Mar 2025 | |||
|
Z
GLM 5.3 Flash is a native multimodal model from Z.ai with a 1M-token context window, designed for efficient inference on coding and long-horizon agent tasks.
$0.15/$0.5I/O
1.05M ctx
Aug 2026
|
$0.15 | $0.5 | $0.03 | 1.05M | Aug 2026 | ||
|
T
Rocinante 12B is a 12-billion-parameter model from TheDrummer designed for narrative generation and creative writing with a 32K token context window.
$0.25/$0.5I/O
65K ctx
Sep 2024
|
$0.25 | $0.5 | — | 65K | Sep 2024 | ||
| $0.12 | $0.5 | — | 40K | Apr 2025 | |||
|
T
Cydonia 24B V4.1 is a 24-billion-parameter model from TheDrummer with a 131K token context window.
$0.3/$0.5I/O
131K ctx
Sep 2025
|
$0.3 | $0.5 | $0.15 | 131K | Sep 2025 | ||
|
A
Olmo 3 32B Think is a 32-billion-parameter language model from Allen Institute for AI with an extended reasoning capability and 65k token context window.
$0.15/$0.5I/O
65K ctx
Nov 2025
|
$0.15 | $0.5 | — | 65K | Nov 2025 | ||
| $0.15 | $0.47 | $0.016 | 1M | Aug 2026 | |||
| $0.117 | $0.455 | — | 131K | Apr 2025 | |||
| $0.117 | $0.455 | — | 131K | Oct 2025 | |||
| $0.15 | $0.45 | $0.015 | 256K | Aug 2025 | |||
| $0.08 | $0.45 | $0.04 | 131K | Mar 2025 | |||
| $0.28 | $0.42 | $0.028 | 128K | Dec 2024 | |||
| $0.28 | $0.42 | $0.028 | 128K | Jan 2025 | |||
| $0.104 | $0.416 | — | 131K | Oct 2025 | |||
| $0.27 | $0.41 | — | 163K | Sep 2025 | |||
|
P
Laguna M.1 is a mid-sized coding model from Poolside with a 262k token context window.
$0.2/$0.4I/O
262K ctx
Apr 2026
|
$0.2 | $0.4 | $0.1 | 262K | Apr 2026 | ||
| $0.4 | $0.4 | — | 131K | Jul 2024 | |||
| $0.36 | $0.4 | — | 32K | Sep 2024 | |||
|
T
UnslopNemo v4.1 is the latest addition from the creator of Rocinante, designed for adventure writing and role-play scenarios.
$0.4/$0.4I/O
1.02M ctx
Nov 2024
|
$0.4 | $0.4 | — | 1.02M | Nov 2024 | ||
| $0.1 | $0.4 | $0.01 | 1.05M | Jul 2025 | |||
| $0.05 | $0.4 | $0.005 | 400K | Aug 2025 | |||
|
N
Hermes 4 70B is a 70-billion parameter instruction-tuned model from Nous with a 131K token context window.
$0.13/$0.4I/O
131K ctx
Aug 2025
|
$0.13 | $0.4 | — | 131K | Aug 2025 | ||
|
N
A 49-billion-parameter model from NVIDIA with a 131K token context window, positioned between the earlier Nemotron 3 Super and the newer mixture-of-experts models in NVIDIA's lineup.
$0.4/$0.4I/O
131K ctx
Oct 2025
|
$0.4 | $0.4 | — | 131K | Oct 2025 | ||
| $0.269 | $0.4 | $0.1345 | 163K | Dec 2025 | |||
|
Z
GLM 4.7 Flash is Z.ai's optimized inference variant of the GLM 4.7 foundation model, designed for faster processing with a 202K-token context window.
$0.0605/$0.4I/O
131K ctx
Jan 2026
|
$0.0605 | $0.4 | — | 131K | Jan 2026 | ||
|
B
A lightweight general-purpose model from ByteDance Seed with a 262k token context window.
$0.1/$0.4I/O
262K ctx
Feb 2026
|
$0.1 | $0.4 | — | 262K | Feb 2026 | ||
|
N
A 49-billion-parameter model from NVIDIA optimized for reasoning and agentic workflows, with a 262K token context window.
$0.085/$0.4I/O
262K ctx
Mar 2026
|
$0.085 | $0.4 | — | 262K | Mar 2026 | ||
| $0.1 | $0.4 | $0.025 | 1M | Apr 2025 | |||
|
U
Solar Pro 4 is a large language model from Upstage with a 524K token context window designed for agentic workflows, document processing, office tasks, and code generation.
$0.09/$0.36I/O
524K ctx
Aug 2026
|
$0.09 | $0.36 | $0.018 | 524K | Aug 2026 | ||
|
M
A smaller, instruction-tuned variant of Phi 4 designed to run efficiently on resource-constrained hardware while handling a 128K token context.
$0.08/$0.35I/O
128K ctx
Oct 2025
|
$0.08 | $0.35 | $0.08 | 128K | Oct 2025 | ||
| $0.345 | $0.345 | — | 131K | Sep 2024 | |||
| $0.09 | $0.34 | $0.05 | 262K | Apr 2026 | |||
| $0.11 | $0.33 | $0.0035 | 1.05M | Aug 2026 | |||
| $0.11 | $0.33 | $0.0035 | 1.05M | Jul 2026 | |||
|
T
Hy3 is Tencent's mixture-of-experts reasoning model with a 262k-token context window.
$0.0825/$0.33I/O
262K ctx
Jul 2026
|
$0.0825 | $0.33 | $0.0206 | 262K | Jul 2026 | ||
| $0.05 | $0.33 | — | 131K | Sep 2024 | |||
| $0.1 | $0.32 | — | 131K | Dec 2024 | |||
| $0.075 | $0.3 | $0.0075 | 262K | Mar 2026 | |||
| $0.075 | $0.3 | $0.0375 | 128K | Jul 2024 | |||
| $0.09 | $0.3 | — | 262K | Jul 2025 | |||
| $0.075 | $0.3 | $0.0375 | 131K | Oct 2025 | |||
| $0.1 | $0.3 | $0.01 | 32K | Oct 2025 | |||
|
X
A text-based language model from Xiaomi with a 262K token context window.
$0.1/$0.3I/O
262K ctx
Dec 2025
|
$0.1 | $0.3 | $0.01 | 262K | Dec 2025 | ||
|
B
Seed 1.6 Flash is a lightweight variant of the Seed 1.6 general-purpose model from ByteDance Seed, optimized for faster inference with a 262k token context window.
$0.075/$0.3I/O
262K ctx
Dec 2025
|
$0.075 | $0.3 | — | 262K | Dec 2025 | ||
|
S
Step 3.5 Flash is a lightweight language model from StepFun with a 262K token context window, designed for fast inference on tasks that don't require the full capacity of larger models.
$0.1/$0.3I/O
262K ctx
Jan 2026
|
$0.1 | $0.3 | — | 262K | Jan 2026 | ||
| $0.1 | $0.3 | — | 262K | Mar 2026 | |||
| $0.08 | $0.3 | — | 1M | Apr 2025 | |||
|
T
Hy-MT2-7B is a 7-billion-parameter translation model from Tencent supporting multilingual translation across dozens of language pairs.
$0.074/$0.295I/O
8K ctx
Aug 2026
|
$0.074 | $0.295 | — | 8K | Aug 2026 | ||
|
T
Hy-MT2-30B-A3B is a 30-billion-parameter translation model from Tencent.
$0.074/$0.295I/O
8K ctx
Aug 2026
|
$0.074 | $0.295 | — | 8K | Aug 2026 | ||
| $0.29 | $0.29 | — | 32K | Jan 2025 | |||
| $0.08 | $0.28 | — | 40K | Apr 2025 | |||
| $0.07 | $0.28 | — | 262K | Jul 2025 | |||
|
X
A language model from Xiaomi with a 1M token context window for processing extended text inputs.
$0.14/$0.28I/O
1.05M ctx
Apr 2026
|
$0.14 | $0.28 | $0.0028 | 1.05M | Apr 2026 | ||
| $0.14 | $0.28 | $0.0028 | 1M | Apr 2026 | |||
| $0.065 | $0.26 | — | 1M | Feb 2026 | |||
|
I
Granite 4.2 8B is an 8-billion-parameter dense reasoning model from IBM with a 131K-token context window, designed for tasks requiring multi-step reasoning.
$0.06/$0.25I/O
131K ctx
Aug 2026
|
$0.06 | $0.25 | $0.015 | 131K | Aug 2026 | ||
| $0.17 | $0.25 | — | 262K | Mar 2026 | |||
|
Z
GLM 5.3 Flash is a native multimodal model from Z.ai with a 1M-token context window, designed for efficient inference on coding and long-horizon agent tasks.
$0.075/$0.25I/O
1.05M ctx
Aug 2026
|
$0.075 | $0.25 | $0.015 | 1.05M | Aug 2026 | ||
|
A
Amazon's lightweight multimodal model that processes text, images, and video inputs with a 300K token context window.
$0.06/$0.24I/O
300K ctx
Dec 2024
|
$0.06 | $0.24 | — | 300K | Dec 2024 | ||
| $0.042 | $0.22 | — | 131K | Apr 2026 | |||
| $0.027 | $0.201 | — | 60K | Sep 2024 | |||
|
N
A 4-billion-parameter multimodal guardrail model from NVIDIA designed to moderate both inputs to and outputs from language and vision models.
$0.2/$0.2I/O
131K ctx
Jun 2026
|
$0.2 | $0.2 | — | 131K | Jun 2026 | ||
| $0.1 | $0.2 | $0.002 | 1.05M | Sep 2026 | |||
| $0.05 | $0.2 | — | 131K | Aug 2025 | |||
| $0.1 | $0.2 | $0.002 | 1.05M | Aug 2026 | |||
|
N
A 30-billion-parameter mixture-of-experts model from NVIDIA with 3B active parameters per token and a 262K token context window.
$0.08/$0.2I/O
262K ctx
Aug 2026
|
$0.08 | $0.2 | $0.04 | 262K | Aug 2026 | ||
| $0.05 | $0.2 | $0.0125 | 1.05M | Apr 2025 | |||
| $0.05 | $0.2 | $0.01 | 1.05M | Jul 2025 | |||
| $0.025 | $0.2 | $0.0025 | 400K | Aug 2025 | |||
|
P
Laguna XS.2 is a lightweight coding model from Poolside with a 262k token context window.
$0.1/$0.2I/O
262K ctx
Apr 2026
|
$0.1 | $0.2 | $0.05 | 262K | Apr 2026 | ||
| $0.1 | $0.2 | — | 32K | Oct 2024 | |||
|
R
Reka Flash 3 is a multimodal model from Rekaai with a 65K token context window, positioned as a mid-tier offering between the smaller Edge variant and larger models in the Reka lineup.
$0.1/$0.2I/O
65K ctx
Mar 2025
|
$0.1 | $0.2 | — | 65K | Mar 2025 | ||
| $0.075 | $0.2 | — | 128K | Jun 2025 | |||
|
B
UI-TARS 7B is a 7-billion-parameter language model from ByteDance designed for understanding and reasoning about user interface interactions.
$0.1/$0.2I/O
128K ctx
Jul 2025
|
$0.1 | $0.2 | $0.1 | 128K | Jul 2025 | ||
| $0.2 | $0.2 | $0.02 | 262K | Dec 2025 | |||
|
N
A 30-billion-parameter mixture-of-experts model from NVIDIA with 3 billion active parameters per token and 262K token context window.
$0.05/$0.2I/O
262K ctx
Dec 2025
|
$0.05 | $0.2 | $0.03 | 262K | Dec 2025 | ||
|
I
Ling 3.0 Flash VL is a multimodal instruct model from InclusionAI that combines efficient language processing with native visual perception capabilities.
$0.06/$0.18I/O
131K ctx
Sep 2026
|
$0.06 | $0.18 | $0.012 | 131K | Sep 2026 | ||
|
I
Ling 3.0 Flash Fin is a finance-specialized mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters and a 262K-token context window.
$0.06/$0.18I/O
262K ctx
Aug 2026
|
$0.06 | $0.18 | $0.012 | 262K | Aug 2026 | ||
|
P
Laguna S 2.1 is a 118B parameter coding agent model from Poolside with a 1M token context window, designed for extended code understanding and generation tasks.
$0.09/$0.18I/O
1.05M ctx
Jul 2026
|
$0.09 | $0.18 | $0.009 | 1.05M | Jul 2026 | ||
| $0.18 | $0.18 | — | 163K | Apr 2025 | |||
|
T
Hy-MT2-1.8B is a 1.8-billion-parameter translation model from Tencent supporting 33 language pairs and Chinese dialects and minority languages.
$0.044/$0.177I/O
8K ctx
Aug 2026
|
$0.044 | $0.177 | — | 8K | Aug 2026 | ||
| $0.037 | $0.17 | — | 131K | Aug 2025 | |||
|
I
Mercury 2.5 is a diffusion-based language model from Inception that generates and refines multiple tokens in parallel instead of sequentially.
$0.04/$0.15I/O
260K ctx
Sep 2026
|
$0.04 | $0.15 | $0.004 | 260K | Sep 2026 | ||
| $0.0375 | $0.15 | — | 128K | Dec 2024 | |||
| $0.05 | $0.15 | — | 131K | Mar 2025 | |||
|
A
Trinity Mini is a smaller reasoning model from Arcee AI with a 131k token context window for handling intermediate problem-solving and reasoning tasks.
$0.045/$0.15I/O
131K ctx
Dec 2025
|
$0.045 | $0.15 | — | 131K | Dec 2025 | ||
| $0.15 | $0.15 | $0.015 | 262K | Dec 2025 | |||
|
E
An instruction-tuned language model from EssentialAI with a 32K token context window.
$0.15/$0.15I/O
32K ctx
Dec 2025
|
$0.15 | $0.15 | — | 32K | Dec 2025 | ||
| $0.1 | $0.15 | — | 262K | Mar 2026 | |||
| $0.14 | $0.14 | — | 8K | Apr 2024 | |||
|
A
A lightweight text-only model from Amazon's Nova family designed for low-latency, cost-efficient inference.
$0.035/$0.14I/O
128K ctx
Dec 2024
|
$0.035 | $0.14 | — | 128K | Dec 2024 | ||
|
M
A 14-billion-parameter model from Microsoft Research designed for complex reasoning with efficient performance on limited hardware.
$0.07/$0.14I/O
16K ctx
Jan 2025
|
$0.07 | $0.14 | — | 16K | Jan 2025 | ||
| $0.03 | $0.13 | $0.006 | 1M | Jul 2026 | |||
| $0.03 | $0.13 | $0.03 | 131K | Aug 2025 | |||
|
P
Laguna XS 2.1 is a lightweight coding model from Poolside with a 262k token context window.
$0.06/$0.12I/O
262K ctx
Jul 2026
|
$0.06 | $0.12 | $0.03 | 262K | Jul 2026 | ||
| $0.06 | $0.12 | — | 32K | May 2025 | |||
|
L
A 24-billion-parameter open-weight model from LiquidAI with a 32K token context window.
$0.03/$0.12I/O
32K ctx
Feb 2026
|
$0.03 | $0.12 | — | 32K | Feb 2026 | ||
|
I
A lightweight language model from IBM's Granite family with a 131K-token context window, designed for efficient inference on resource-constrained environments.
$0.017/$0.112I/O
131K ctx
Oct 2025
|
$0.017 | $0.112 | — | 131K | Oct 2025 | ||
| $0.11 | $0.11 | — | 128K | Oct 2024 | |||
|
N
Nex-N2-Mini is an open-source mixture-of-experts model from Nex AGI designed for agentic tasks, accepting text and image inputs with a 262K token context window.
$0.025/$0.1I/O
262K ctx
Jun 2026
|
$0.025 | $0.1 | $0.0025 | 262K | Jun 2026 | ||
| $0.05 | $0.1 | — | 131K | Mar 2025 | |||
| $0.1 | $0.1 | $0.01 | 131K | Dec 2025 | |||
|
R
Reka Edge is a smaller instruction-tuned model from Reka with a 16K token context window.
$0.1/$0.1I/O
16K ctx
Mar 2026
|
$0.1 | $0.1 | — | 16K | Mar 2026 | ||
|
I
An 8-billion-parameter language model from IBM's Granite family with a 131K-token context window.
$0.05/$0.1I/O
131K ctx
Apr 2026
|
$0.05 | $0.1 | $0.05 | 131K | Apr 2026 | ||
| $0.05 | $0.08 | $0.025 | 131K | Jul 2024 | |||
| $0.05 | $0.08 | — | 32K | Jan 2025 | |||
| $0.075 | $0.075 | $0.0075 | 262K | Dec 2025 | |||
|
I
Ling-3.0-flash is a fast, efficient instruct model from inclusionAI with a 262K-token context window.
$0.021/$0.063I/O
262K ctx
Jul 2026
|
$0.021 | $0.063 | $0.0042 | 262K | Jul 2026 | ||
|
G
One of the highest performing and most popular fine-tunes of Llama 2 13B, with rich descriptions and roleplay. #merge
$0.06/$0.06I/O
4K ctx
Jul 2023
|
$0.06 | $0.06 | — | 4K | Jul 2023 | ||
|
S
An 8B parameter model based on Llama 3, created through merging multiple models to balance creative and logical capabilities.
$0.04/$0.05I/O
8K ctx
Aug 2024
|
$0.04 | $0.05 | — | 8K | Aug 2024 | ||
| $0.019 | $0.03 | — | 131K | Jul 2024 | |||
| $0.484 | $0.03 | — | 131K | Feb 2025 | |||
|
I
Ling-2.6-flash is a fast, efficient instruct model from inclusionAI with a 262K-token context window.
$0.01/$0.03I/O
262K ctx
Apr 2026
|
$0.01 | $0.03 | $0.002 | 262K | Apr 2026 |
Sorted by output price by default — output tokens are typically 3–5× more expensive than input and dominate the bill for generation-heavy workloads. Sort by input or cached input when your mix differs.