LLM API pricing, tracked from launch.
300 language models across 88 providers — input, output, and cached rates per 1M tokens, updated daily, with full price history.
Price table
| Model | Input /1M | Output /1M | Cached in /1M | Context | Released | ||
|---|---|---|---|---|---|---|---|
|
I
Ling-2.6-flash is a 104B parameter model from inclusionAI with 7.4B active parameters, designed for fast inference and efficient token usage in agent applications.
$0.01/$0.03I/O
262K ctx
Apr 2026
|
$0.01 | $0.03 | $0.002 | 262K | Apr 2026 | ||
| $0.484 | $0.03 | — | 131K | Feb 2025 | |||
| $0.019 | $0.03 | — | 131K | Jul 2024 | |||
|
S
An 8B parameter model based on Llama 3, created through merging multiple models to balance creative and logical capabilities.
$0.04/$0.05I/O
8K ctx
Aug 2024
|
$0.04 | $0.05 | — | 8K | Aug 2024 | ||
|
G
One of the highest performing and most popular fine-tunes of Llama 2 13B, with rich descriptions and roleplay. #merge
$0.06/$0.06I/O
4K ctx
Jul 2023
|
$0.06 | $0.06 | — | 4K | Jul 2023 | ||
| $0.05 | $0.08 | — | 32K | Jan 2025 | |||
| $0.05 | $0.08 | $0.025 | 131K | Jul 2024 | |||
|
I
Granite 4.1 8B is an 8-billion-parameter decoder-only language model from IBM with a 131K-token context window, designed for enterprise applications.
$0.05/$0.1I/O
131K ctx
Apr 2026
|
$0.05 | $0.1 | $0.05 | 131K | Apr 2026 | ||
|
R
A 7-billion-parameter multimodal model that processes images, videos, and text to generate text outputs.
$0.1/$0.1I/O
16K ctx
Mar 2026
|
$0.1 | $0.1 | — | 16K | Mar 2026 | ||
| $0.1 | $0.1 | $0.01 | 131K | Dec 2025 | |||
| $0.05 | $0.1 | — | 131K | Mar 2025 | |||
| $0.04 | $0.1 | — | 32K | Oct 2024 | |||
|
N
Nex-N2-Mini is an open-source mixture-of-experts model from Nex AGI designed for agentic tasks, accepting text and image inputs with a 262K token context window.
$0.025/$0.1I/O
262K ctx
Jun 2026
|
$0.025 | $0.1 | $0.0025 | 262K | Jun 2026 | ||
|
I
A 3B parameter model from IBM's Granite 4 family, designed to handle long context windows efficiently.
$0.017/$0.112I/O
131K ctx
Oct 2025
|
$0.017 | $0.112 | — | 131K | Oct 2025 | ||
|
L
A 24-billion parameter mixture-of-experts model from LiquidAI that activates only 2 billion parameters per token for efficient inference.
$0.03/$0.12I/O
32K ctx
Feb 2026
|
$0.03 | $0.12 | — | 32K | Feb 2026 | ||
| $0.06 | $0.12 | — | 32K | May 2025 | |||
|
P
Laguna XS 2.1 is a lightweight coding model from Poolside with a 262k token context window.
$0.06/$0.12I/O
262K ctx
Jul 2026
|
$0.06 | $0.12 | $0.03 | 262K | Jul 2026 | ||
| $0.03 | $0.13 | $0.006 | 1M | Jul 2026 | |||
| $0.03 | $0.14 | — | 131K | Aug 2025 | |||
|
M
A 14-billion-parameter model from Microsoft Research designed for complex reasoning with efficient performance on limited hardware.
$0.07/$0.14I/O
16K ctx
Jan 2025
|
$0.07 | $0.14 | — | 16K | Jan 2025 | ||
|
A
A lightweight text-only model from Amazon's Nova family designed for low-latency, cost-efficient inference.
$0.035/$0.14I/O
128K ctx
Dec 2024
|
$0.035 | $0.14 | — | 128K | Dec 2024 | ||
| $0.14 | $0.14 | — | 8K | Apr 2024 | |||
| $0.1 | $0.15 | — | 262K | Mar 2026 | |||
|
E
An 8-billion-parameter dense model trained by EssentialAI for programming, math, and scientific reasoning tasks.
$0.15/$0.15I/O
32K ctx
Dec 2025
|
$0.15 | $0.15 | — | 32K | Dec 2025 | ||
| $0.15 | $0.15 | $0.015 | 262K | Dec 2025 | |||
|
A
A 26B sparse mixture-of-experts language model with 3B parameters active per token and 131k token context window, designed for efficient reasoning over extended documents.
$0.045/$0.15I/O
131K ctx
Dec 2025
|
$0.045 | $0.15 | — | 131K | Dec 2025 | ||
| $0.05 | $0.15 | — | 131K | Mar 2025 | |||
| $0.0375 | $0.15 | — | 128K | Dec 2024 | |||
| $0.037 | $0.17 | — | 131K | Aug 2025 | |||
| $0.18 | $0.18 | — | 163K | Apr 2025 | |||
| $0.0482 | $0.193 | — | 128K | Jul 2025 | |||
|
N
A 30-billion-parameter mixture-of-experts model from NVIDIA designed for efficient inference and specialized task performance.
$0.05/$0.2I/O
262K ctx
Dec 2025
|
$0.05 | $0.2 | — | 262K | Dec 2025 | ||
| $0.2 | $0.2 | $0.02 | 262K | Dec 2025 | |||
|
B
A 7B multimodal vision-language model from ByteDance designed to understand and interact with graphical user interfaces across desktop, web, mobile, and game environments.
$0.1/$0.2I/O
128K ctx
Jul 2025
|
$0.1 | $0.2 | $0.1 | 128K | Jul 2025 | ||
|
R
Reka Flash 3 is a 21-billion-parameter instruction-tuned model from Reka designed for general-purpose tasks including chat, coding, and function calling.
$0.1/$0.2I/O
65K ctx
Mar 2025
|
$0.1 | $0.2 | — | 65K | Mar 2025 | ||
|
P
Laguna XS.2 is a lightweight coding model from Poolside with a 262k token context window.
$0.1/$0.2I/O
262K ctx
Apr 2026
|
$0.1 | $0.2 | $0.05 | 262K | Apr 2026 | ||
|
P
Laguna S 2.1 is a 118B parameter coding agent model from Poolside with a 1M token context window, designed for extended code understanding and generation tasks.
$0.1/$0.2I/O
1.05M ctx
Jul 2026
|
$0.1 | $0.2 | $0.01 | 1.05M | Jul 2026 | ||
| $0.027 | $0.201 | — | 60K | Sep 2024 | |||
|
T
Hy3 preview is an early-release version of Tencent's mixture-of-experts reasoning model, giving early access to features and capability development before the full Hy3 release.
$0.063/$0.21I/O
262K ctx
Apr 2026
|
$0.063 | $0.21 | $0.021 | 262K | Apr 2026 | ||
|
A
Amazon's lightweight multimodal model that processes text, images, and video inputs with a 300K token context window.
$0.06/$0.24I/O
300K ctx
Dec 2024
|
$0.06 | $0.24 | — | 300K | Dec 2024 | ||
| $0.065 | $0.26 | — | 1M | Feb 2026 | |||
| $0.07 | $0.27 | — | 160K | Jul 2025 | |||
| $0.14 | $0.28 | $0.0028 | 1M | Apr 2026 | |||
|
X
Xiaomi's native omnimodal model that processes text, images, and video in a single forward pass.
$0.14/$0.28I/O
1.05M ctx
Apr 2026
|
$0.14 | $0.28 | $0.0028 | 1.05M | Apr 2026 | ||
| $0.08 | $0.28 | — | 40K | Apr 2025 | |||
| $0.29 | $0.29 | — | 32K | Jan 2025 | |||
| $0.08 | $0.3 | — | 1M | Apr 2025 | |||
| $0.1 | $0.3 | — | 262K | Mar 2026 | |||
|
S
StepFun's sparse Mixture of Experts model that activates 11B of 196B parameters per token, providing a balance between capability and efficiency.
$0.1/$0.3I/O
262K ctx
Jan 2026
|
$0.1 | $0.3 | — | 262K | Jan 2026 | ||
|
B
A fast multimodal model from ByteDance Seed that handles text and image inputs with a 256k token context window.
$0.075/$0.3I/O
262K ctx
Dec 2025
|
$0.075 | $0.3 | — | 262K | Dec 2025 | ||
|
X
MiMo-V2-Flash is a Mixture-of-Experts language model from Xiaomi with 309B total parameters and 15B active parameters, supporting a 262K token context window.
$0.1/$0.3I/O
262K ctx
Dec 2025
|
$0.1 | $0.3 | $0.01 | 262K | Dec 2025 | ||
| $0.1 | $0.3 | $0.01 | 32K | Oct 2025 | |||
| $0.075 | $0.3 | $0.0375 | 131K | Oct 2025 | |||
| $0.1 | $0.3 | $0.01 | 131K | Jun 2025 | |||
| $0.05 | $0.33 | — | 131K | Sep 2024 | |||
| $0.345 | $0.345 | — | 131K | Sep 2024 | |||
|
M
Phi-4-mini-instruct is a lightweight open model trained on synthetic and curated public data, designed to handle reasoning-heavy tasks efficiently.
$0.08/$0.35I/O
128K ctx
Oct 2025
|
$0.08 | $0.35 | $0.08 | 128K | Oct 2025 | ||
| $0.1 | $0.4 | $0.025 | 1M | Apr 2025 | |||
| $0.14 | $0.4 | — | 262K | Apr 2026 | |||
|
N
A 120B-parameter open hybrid MoE model from NVIDIA that activates 12B parameters per token, combining Mamba and Transformer architectures for inference efficiency.
$0.085/$0.4I/O
262K ctx
Mar 2026
|
$0.085 | $0.4 | — | 262K | Mar 2026 | ||
|
B
A lightweight model from ByteDance designed for low-latency inference at high concurrency, with 256k context window and multimodal understanding.
$0.1/$0.4I/O
262K ctx
Feb 2026
|
$0.1 | $0.4 | — | 262K | Feb 2026 | ||
|
Z
GLM 4.7 Flash is Z.ai's optimized inference variant of the GLM 4.7 foundation model, designed for faster response times with a 202K-token context window.
$0.06/$0.4I/O
202K ctx
Jan 2026
|
$0.06 | $0.4 | $0.01 | 202K | Jan 2026 | ||
| $0.269 | $0.4 | $0.1345 | 163K | Dec 2025 | |||
|
N
A 49B-parameter model optimized for reasoning and agentic workflows, derived from Llama 3.3 and supporting 131K token context.
$0.4/$0.4I/O
131K ctx
Oct 2025
|
$0.4 | $0.4 | — | 131K | Oct 2025 | ||
|
N
Hermes 4 70B is a hybrid reasoning model from Nous Research that can switch between fast and deliberate reasoning modes to balance speed and accuracy on a per-query basis.
$0.13/$0.4I/O
131K ctx
Aug 2025
|
$0.13 | $0.4 | — | 131K | Aug 2025 | ||
| $0.05 | $0.4 | $0.005 | 400K | Aug 2025 | |||
| $0.1 | $0.4 | $0.01 | 1.05M | Jul 2025 | |||
| $0.13 | $0.4 | — | 131K | Dec 2024 | |||
|
T
UnslopNemo v4.1 is the latest addition from the creator of Rocinante, designed for adventure writing and role-play scenarios.
$0.4/$0.4I/O
32K ctx
Nov 2024
|
$0.4 | $0.4 | — | 32K | Nov 2024 | ||
| $0.36 | $0.4 | — | 32K | Sep 2024 | |||
| $0.4 | $0.4 | — | 131K | Jul 2024 | |||
|
P
Laguna M.1 is a mid-sized coding model from Poolside with a 262k token context window.
$0.2/$0.4I/O
262K ctx
Apr 2026
|
$0.2 | $0.4 | $0.1 | 262K | Apr 2026 | ||
| $0.27 | $0.41 | — | 163K | Sep 2025 | |||
| $0.104 | $0.416 | — | 131K | Oct 2025 | |||
| $0.28 | $0.42 | $0.028 | 128K | Jan 2025 | |||
| $0.28 | $0.42 | $0.028 | 128K | Dec 2024 | |||
| $0.14 | $0.42 | $0.05 | 262K | Apr 2026 | |||
| $0.08 | $0.45 | $0.04 | 131K | Mar 2025 | |||
| $0.117 | $0.455 | — | 131K | Oct 2025 | |||
| $0.117 | $0.455 | — | 131K | Apr 2025 | |||
|
A
Olmo 3 32B Think is a 32-billion-parameter reasoning model from AllenAI designed for complex logical inference and extended problem-solving.
$0.15/$0.5I/O
65K ctx
Nov 2025
|
$0.15 | $0.5 | — | 65K | Nov 2025 | ||
|
T
Uncensored and creative writing model based on Mistral Small 3.2 24B with good recall, prompt adherence, and intelligence.
$0.3/$0.5I/O
131K ctx
Sep 2025
|
$0.3 | $0.5 | $0.15 | 131K | Sep 2025 | ||
| $0.12 | $0.5 | — | 40K | Apr 2025 | |||
|
T
Rocinante 12B is a 12-billion-parameter model from TheDrummer designed for narrative generation and creative writing with a 32K token context window.
$0.25/$0.5I/O
65K ctx
Sep 2024
|
$0.25 | $0.5 | — | 65K | Sep 2024 | ||
|
T
Hy3 is a 295B-parameter mixture-of-experts model from Tencent with 21B active parameters that supports configurable reasoning effort for complex problem-solving.
$0.132/$0.528I/O
262K ctx
Jul 2026
|
$0.132 | $0.528 | $0.033 | 262K | Jul 2026 | ||
| $0.09 | $0.55 | — | 262K | Jul 2025 | |||
| $0.351 | $0.555 | — | 128K | Mar 2025 | |||
|
T
A 13-billion-parameter instruction-tuned language model from Tencent with a 131k-token context window.
$0.14/$0.57I/O
131K ctx
Jul 2025
|
$0.14 | $0.57 | — | 131K | Jul 2025 | ||
| $0.15 | $0.6 | — | 1M | Apr 2025 | |||
|
U
A Mixture-of-Experts language model from Upstage with 102B parameters total and 12B active per forward pass, operating on a 128K token context window.
$0.15/$0.6I/O
128K ctx
Jan 2026
|
$0.15 | $0.6 | $0.015 | 128K | Jan 2026 | ||
| $0.15 | $0.6 | — | 262K | Oct 2025 | |||
| $0.15 | $0.6 | — | 128K | Mar 2025 | |||
| $0.2 | $0.6 | $0.02 | 32K | Feb 2025 | |||
| $0.15 | $0.6 | $0.075 | 128K | Jul 2024 | |||
| $0.15 | $0.6 | — | 128K | Mar 2024 | |||
|
K
A code-focused language model from Kwaipilot with a 256k token context window.
$0.15/$0.6I/O
256K ctx
Jul 2026
|
$0.15 | $0.6 | $0.03 | 256K | Jul 2026 | ||
|
M
A 176-billion-parameter mixture-of-experts model from Microsoft that combines eight 22-billion experts with a large context window for handling extended documents and conversations.
$0.62/$0.62I/O
65K ctx
Apr 2024
|
$0.62 | $0.62 | — | 65K | Apr 2024 | ||
|
I
Ring-2.6-1T is a 1 trillion parameter mixture-of-experts model from inclusionAI with 63B active parameters, featuring a 262K token context window optimized for agentic workflows.
$0.075/$0.625I/O
262K ctx
May 2026
|
$0.075 | $0.625 | $0.015 | 262K | May 2026 | ||
|
I
Ling-2.6-1T is a trillion-parameter instruct model from inclusionAI designed for fast execution and efficient operation at scale.
$0.075/$0.625I/O
262K ctx
Apr 2026
|
$0.075 | $0.625 | $0.015 | 262K | Apr 2026 | ||
| $0.65 | $0.65 | — | 8K | Jul 2024 | |||
|
U
A recreation trial of the original MythoMax-L2-B13 but with updated models. #merge
$0.45/$0.65I/O
6K ctx
Jul 2023
|
$0.45 | $0.65 | — | 6K | Jul 2023 | ||
|
N
Hermes 3 70B Instruct is a 70-billion parameter generalist model with a 131K token context window, trained to follow instructions and support agentic workflows.
$0.7/$0.7I/O
131K ctx
Aug 2024
|
$0.7 | $0.7 | — | 131K | Aug 2024 | ||
| $0.51 | $0.74 | — | 8K | Apr 2024 | |||
|
I
Mercury 2 is a reasoning model from Inception that uses diffusion-based token generation to produce and refine multiple tokens in parallel rather than sequentially.
$0.25/$0.75I/O
128K ctx
Mar 2026
|
$0.25 | $0.75 | $0.025 | 128K | Mar 2026 | ||
|
S
Euryale L3.3 70B is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.2](/models/sao10k/l3-euryale-70b).
$0.65/$0.75I/O
131K ctx
Dec 2024
|
$0.65 | $0.75 | — | 131K | Dec 2024 | ||
|
M
An attempt to recreate Claude-style verbosity, but don't expect the same level of coherence or memory. Meant for use in roleplay/narrative situations.
$0.5/$0.75I/O
8K ctx
Aug 2023
|
$0.5 | $0.75 | — | 8K | Aug 2023 | ||
| $0.0975 | $0.78 | — | 131K | Sep 2025 | |||
| $0.26 | $0.78 | — | 1M | Sep 2025 | |||
| $0.26 | $0.78 | $0.052 | 1M | Feb 2025 | |||
|
A
Coder Large is a 32B parameter model from Arcee AI built on Qwen 2.5-Instruct and trained on open-source code repositories and synthetic bug-fix examples to specialize in code generation and analysis.
$0.5/$0.8I/O
32K ctx
May 2025
|
$0.5 | $0.8 | — | 32K | May 2025 | ||
|
T
Skyfall 36B v2 is an enhanced iteration of Mistral Small 2501, specifically fine-tuned for improved creativity, nuanced writing, role-playing, and coherent storytelling.
$0.55/$0.8I/O
32K ctx
Mar 2025
|
$0.55 | $0.8 | $0.25 | 32K | Mar 2025 | ||
| $0.8 | $0.8 | — | 8K | Jan 2025 | |||
|
A
An open-source reasoning model from Arcee AI designed for complex problem-solving and agentic tasks, with extended thinking capabilities.
$0.22/$0.85I/O
262K ctx
Apr 2026
|
$0.22 | $0.85 | $0.06 | 262K | Apr 2026 | ||
|
Z
GLM 4.5 Air is Z.ai's lightweight inference variant of the GLM 4.5 foundation model, maintaining a 131K-token context window with optimizations for faster response times.
$0.13/$0.85I/O
131K ctx
Jul 2025
|
$0.13 | $0.85 | $0.025 | 131K | Jul 2025 | ||
|
S
Euryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.1](/models/sao10k/l3-euryale-70b).
$0.85/$0.85I/O
131K ctx
Aug 2024
|
$0.85 | $0.85 | — | 131K | Aug 2024 | ||
| $0.435 | $0.87 | $0.0036 | 1M | Apr 2026 | |||
|
X
MiMo-V2.5-Pro is Xiaomi's flagship model with a 1M-token context window designed for complex reasoning and extended reasoning over long inputs.
$0.435/$0.87I/O
1.05M ctx
Apr 2026
|
$0.435 | $0.87 | $0.0036 | 1.05M | Apr 2026 | ||
|
M
MiniMax M2.5 is a large language model designed for productivity tasks, trained on complex real-world digital working environments with particular strength in code understanding.
$0.15/$0.9I/O
196K ctx
Feb 2026
|
$0.15 | $0.9 | $0.05 | 196K | Feb 2026 | ||
| $0.18 | $0.9 | $0.036 | 256K | Feb 2026 | |||
|
Z
GLM 4.6V is Z.ai's vision-language model that processes text, image, and video inputs for multimodal reasoning tasks.
$0.3/$0.9I/O
131K ctx
Dec 2025
|
$0.3 | $0.9 | $0.055 | 131K | Dec 2025 | ||
| $0.3 | $0.9 | $0.03 | 256K | Aug 2025 | |||
|
V
An instruction-tuned 24B model based on Mistral that minimizes refusals on sensitive or controversial topics.
$0.2/$0.9I/O
128K ctx
Jul 2025
|
$0.2 | $0.9 | — | 128K | Jul 2025 | ||
| $0.2275 | $0.91 | — | 131K | Apr 2025 | |||
| $0.25 | $0.95 | $0.13 | 163K | Aug 2025 | |||
| $0.195 | $0.975 | $0.039 | 1M | Sep 2025 | |||
| $0.14 | $1.00 | — | 262K | Apr 2026 | |||
|
M
MiniMax M2.7 is a compact language model designed for agentic tasks and real-world automation with a 196k token context window.
$0.25/$1.00I/O
196K ctx
Mar 2026
|
$0.25 | $1.00 | $0.05 | 196K | Mar 2026 | ||
| $0.14 | $1.00 | — | 262K | Feb 2026 | |||
| $0.27 | $1.00 | $0.135 | 131K | Sep 2025 | |||
| $0.3 | $1.00 | $0.1 | 262K | Jul 2025 | |||
| $0.8 | $1.00 | $0.4 | 128K | Feb 2025 | |||
|
P
Perplexity's Sonar is a lightweight model designed for question-answering tasks with citation support and customizable source integration.
$1.00/$1.00I/O
127K ctx
Jan 2025
|
$1.00 | $1.00 | — | 127K | Jan 2025 | ||
| $0.66 | $1.00 | — | 32K | Nov 2024 | |||
|
N
Hermes 3 405B Instruct is a large generalist language model from Nous with a 131k token context window, designed for multi-turn conversations and agentic tasks.
$1.00/$1.00I/O
131K ctx
Aug 2024
|
$1.00 | $1.00 | — | 131K | Aug 2024 | ||
|
N
Nex-N2-Pro is a mixture-of-experts model from Nex AGI that processes text and image inputs with a 262K token context window.
$0.25/$1.00I/O
262K ctx
Jun 2026
|
$0.25 | $1.00 | $0.025 | 262K | Jun 2026 | ||
|
M
MiniMax M2 is a 10B activated parameter model with a 196K token context window, designed for coding and agentic workflows.
$0.255/$1.02I/O
204K ctx
Oct 2025
|
$0.255 | $1.02 | — | 204K | Oct 2025 | ||
|
P
A 106B-parameter mixture-of-experts model from Prime Intellect with 12B active parameters, post-trained with supervised fine-tuning and reinforcement learning on a 131k-token context window.
$0.2/$1.10I/O
131K ctx
Nov 2025
|
$0.2 | $1.10 | — | 131K | Nov 2025 | ||
| $0.1 | $1.10 | $0.07 | 262K | Sep 2025 | |||
|
M
MiniMax-01 is a multimodal model combining text generation and image understanding with 456 billion parameters and a 1 million token context window.
$0.2/$1.10I/O
1M ctx
Jan 2025
|
$0.2 | $1.10 | — | 1M | Jan 2025 | ||
| $0.27 | $1.12 | $0.135 | 163K | Mar 2025 | |||
| $0.1875 | $1.12 | — | 1M | Apr 2026 | |||
|
S
StepFun's multimodal Mixture-of-Experts model that processes text, images, and video with a 196B parameter backbone while activating approximately 11B parameters per inference.
$0.2/$1.15I/O
256K ctx
May 2026
|
$0.2 | $1.15 | $0.04 | 256K | May 2026 | ||
|
M
MiniMax M3 is a multimodal foundation model that processes text, image, and video inputs to produce text output, with a 524K token context window.
$0.3/$1.20I/O
524K ctx
May 2026
|
$0.3 | $1.20 | $0.06 | 524K | May 2026 | ||
|
K
A code-focused language model from Kwaipilot with a 256K token context window.
$0.3/$1.20I/O
256K ctx
Mar 2026
|
$0.3 | $1.20 | $0.06 | 256K | Mar 2026 | ||
|
M
MiniMax M2-her is a dialogue-focused model designed for character-driven conversations and roleplay scenarios with a 65K token context window.
$0.3/$1.20I/O
65K ctx
Jan 2026
|
$0.3 | $1.20 | $0.03 | 65K | Jan 2026 | ||
|
M
MiniMax M2.1 is a 10-billion-parameter lightweight language model designed for coding, agentic workflows, and application development.
$0.3/$1.20I/O
204K ctx
Dec 2025
|
$0.3 | $1.20 | $0.03 | 204K | Dec 2025 | ||
|
M
Morph's fastest model for applying code edits and transformations, optimized to process code modifications at high throughput.
$0.8/$1.20I/O
81K ctx
Jul 2025
|
$0.8 | $1.20 | — | 81K | Jul 2025 | ||
|
A
Virtuoso Large is a 72B general-purpose language model from Arcee AI with a 131k token context window.
$0.75/$1.20I/O
131K ctx
May 2025
|
$0.75 | $1.20 | — | 131K | May 2025 | ||
|
M
A sparse mixture-of-experts model from Meituan with 48B active parameters designed to handle very long contexts up to 1M tokens.
$0.3/$1.20I/O
1.05M ctx
Jul 2026
|
$0.3 | $1.20 | $0.006 | 1.05M | Jul 2026 | ||
| $0.2 | $1.25 | $0.02 | 400K | Mar 2026 | |||
|
D
Cogito v2.1 671B is a 671-billion parameter mixture-of-experts model trained with self-play reinforcement learning, offering a 128k token context window.
$1.25/$1.25I/O
128K ctx
Nov 2025
|
$1.25 | $1.25 | — | 128K | Nov 2025 | ||
|
R
A specialized code-patching model that applies AI-suggested edits directly into source files from multiple LLM providers.
$0.85/$1.25I/O
256K ctx
Sep 2025
|
$0.85 | $1.25 | — | 256K | Sep 2025 | ||
|
B
A multimodal mixture-of-experts model from Baidu with 424B total parameters and 47B active per token, trained on both text and image data.
$0.42/$1.25I/O
123K ctx
Jun 2025
|
$0.42 | $1.25 | — | 123K | Jun 2025 | ||
| $0.32 | $1.28 | $0.064 | 1M | Jun 2026 | |||
| $0.117 | $1.36 | — | 131K | Oct 2025 | |||
|
A
Aion-1.0-Mini is a smaller parameter variant of Aion-1.0 from AionLabs, sharing the same 131k-token context window as its full-size sibling.
$0.7/$1.40I/O
131K ctx
Feb 2025
|
$0.7 | $1.40 | — | 131K | Feb 2025 | ||
|
A
Aion-3.0-Mini is a multi-model roleplaying and storytelling system from AionLabs built on the DeepSeek family, using collaborative generation from specialized models.
$0.7/$1.40I/O
131K ctx
Jul 2026
|
$0.7 | $1.40 | $0.18 | 131K | Jul 2026 | ||
| $0.36 | $1.43 | — | 256K | Sep 2025 | |||
| $0.5 | $1.50 | — | 262K | Dec 2025 | |||
|
P
A vision-language model from Perceptron designed to process images and videos alongside natural language queries for visual understanding tasks.
$0.15/$1.50I/O
32K ctx
May 2026
|
$0.15 | $1.50 | — | 32K | May 2026 | ||
| $0.25 | $1.50 | $0.025 | 1.05M | May 2026 | |||
| $0.5 | $1.50 | — | 16K | May 2023 | |||
| $0.195 | $1.56 | — | 262K | Feb 2026 | |||
| $0.26 | $1.56 | — | 1M | Feb 2026 | |||
| $0.13 | $1.56 | — | 131K | Oct 2025 | |||
| $0.13 | $1.56 | — | 81K | Aug 2025 | |||
| $0.4 | $1.60 | $0.1 | 1M | Apr 2025 | |||
|
A
Aion-2.0 is a multi-model system from AionLabs that builds on the reasoning and coding strengths of Aion-1.0 with broader capability across language tasks.
$0.8/$1.60I/O
131K ctx
Feb 2026
|
$0.8 | $1.60 | $0.2 | 131K | Feb 2026 | ||
|
A
Aion-RP 1.0 is an 8-billion-parameter language model from AionLabs designed for roleplaying and narrative generation tasks.
$0.8/$1.60I/O
32K ctx
Feb 2025
|
$0.8 | $1.60 | — | 32K | Feb 2025 | ||
|
Z
GLM 4.7 is Z.ai's foundation model with a 202K-token context window for processing extended documents and conversations.
$0.4/$1.75I/O
202K ctx
Dec 2025
|
$0.4 | $1.75 | $0.08 | 202K | Dec 2025 | ||
| $0.3 | $1.80 | — | 1M | Apr 2026 | |||
|
Z
GLM 4.5V is Z.ai's vision-language model that processes text, image, and video inputs for multimodal reasoning tasks.
$0.6/$1.80I/O
65K ctx
Aug 2025
|
$0.6 | $1.80 | $0.11 | 65K | Aug 2025 | ||
| $0.455 | $1.82 | — | 131K | Apr 2025 | |||
| $0.21 | $1.90 | $0.1 | 131K | Sep 2025 | |||
|
M
Morph V3 Large is a code-specialized model designed for precise code transformations and edits at scale with a 262K token context window.
$0.9/$1.90I/O
262K ctx
Jul 2025
|
$0.9 | $1.90 | — | 262K | Jul 2025 | ||
| $0.325 | $1.95 | — | 1M | Apr 2026 | |||
| $1.00 | $2.00 | $0.2 | 256K | May 2026 | |||
| $0.3 | $2.00 | $0.15 | 262K | Apr 2026 | |||
|
B
Seed-2.0-Lite is a multimodal model from ByteDance with a 262k token context window, designed for enterprise production workloads with emphasis on lower latency.
$0.25/$2.00I/O
262K ctx
Mar 2026
|
$0.25 | $2.00 | — | 262K | Mar 2026 | ||
|
B
Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256K context window.
$0.25/$2.00I/O
262K ctx
Dec 2025
|
$0.25 | $2.00 | — | 262K | Dec 2025 | ||
| $0.4 | $2.00 | $0.04 | 262K | Dec 2025 | |||
| $0.25 | $2.00 | $0.03 | 400K | Nov 2025 | |||
|
Z
GLM 4.6 is Z.ai's foundation model with a 202K-token context window for processing extended documents and conversations.
$0.5/$2.00I/O
202K ctx
Sep 2025
|
$0.5 | $2.00 | $0.1 | 202K | Sep 2025 | ||
| $0.4 | $2.00 | $0.04 | 131K | Aug 2025 | |||
| $0.25 | $2.00 | $0.025 | 400K | Aug 2025 | |||
| $0.4 | $2.00 | $0.04 | 131K | May 2025 | |||
| $1.00 | $2.00 | — | 4K | Jan 2024 | |||
| $1.50 | $2.00 | — | 4K | Sep 2023 | |||
| $0.26 | $2.08 | — | 262K | Feb 2026 | |||
| $0.5 | $2.15 | $0.35 | 163K | May 2025 | |||
|
N
A 55B-parameter mixture-of-experts model from NVIDIA with a 262K token context window, using a hybrid Transformer-Mamba architecture for reasoning and task orchestration.
$0.5/$2.20I/O
262K ctx
Jun 2026
|
$0.5 | $2.20 | $0.1 | 262K | Jun 2026 | ||
|
Z
GLM 4.5 is Z.ai's foundation model with a 131K-token context window for processing documents and conversations.
$0.6/$2.20I/O
131K ctx
Jul 2025
|
$0.6 | $2.20 | $0.11 | 131K | Jul 2025 | ||
|
M
MiniMax M1 is an open-weight reasoning model with a 1M-token context window and hybrid mixture-of-experts architecture designed for efficient inference at scale.
$0.55/$2.20I/O
1M ctx
Jun 2025
|
$0.55 | $2.20 | — | 1M | Jun 2025 | ||
| $0.57 | $2.30 | — | 131K | Jul 2025 | |||
| $0.39 | $2.34 | — | 262K | Feb 2026 | |||
| $0.6 | $2.40 | — | 128K | Jan 2026 | |||
|
Z
GLM 5.2 is a large-scale reasoning model from Z.ai with a 1M-token context window for processing long documents and complex tasks.
$0.7686/$2.42I/O
1.05M ctx
Jun 2026
|
$0.7686 | $2.42 | $0.1427 | 1.05M | Jun 2026 | ||
| $0.3 | $2.50 | $0.03 | 1M | Jun 2025 | |||
| $1.25 | $2.50 | $0.2 | 1M | Apr 2026 | |||
| $1.25 | $2.50 | $0.2 | 2M | Mar 2026 | |||
| $1.25 | $2.50 | $0.2 | 2M | Mar 2026 | |||
|
A
Nova 2 Lite is Amazon's lightweight multimodal model that handles text, images, and video input with a 1M token context window.
$0.3/$2.50I/O
1M ctx
Dec 2025
|
$0.3 | $2.50 | — | 1M | Dec 2025 | ||
| $0.6 | $2.50 | $0.15 | 262K | Nov 2025 | |||
| $0.6 | $2.50 | — | 262K | Sep 2025 | |||
| $0.7 | $2.50 | — | 64K | Jan 2025 | |||
| $0.3 | $2.50 | $0.03 | 1.05M | Jul 2026 | |||
|
Z
GLM 5 is Z.ai's foundation model with a 198K-token context window for processing extended documents and conversations.
$0.95/$2.55I/O
204K ctx
Feb 2026
|
$0.95 | $2.55 | $0.2 | 204K | Feb 2026 | ||
| $0.26 | $2.60 | — | 131K | Sep 2025 | |||
|
K
A coding-focused language model from Kwaipilot with a 256K token context window.
$0.74/$2.96I/O
256K ctx
Jul 2026
|
$0.74 | $2.96 | $0.15 | 256K | Jul 2026 | ||
| $0.5 | $3.00 | $0.05 | 1M | Dec 2025 | |||
| $0.6 | $3.00 | $0.1 | 256K | Jan 2026 | |||
|
R
Relace Search is an agentic code search tool that uses parallel file and grep operations to explore codebases and locate relevant files.
$1.00/$3.00I/O
256K ctx
Dec 2025
|
$1.00 | $3.00 | — | 256K | Dec 2025 | ||
|
N
A 405-billion parameter reasoning model based on Llama 3.1 that can choose between direct answers and extended internal deliberation modes.
$1.00/$3.00I/O
131K ctx
Aug 2025
|
$1.00 | $3.00 | — | 131K | Aug 2025 | ||
| $0.3 | $3.00 | — | 131K | Jul 2025 | |||
|
S
This is [Sao10K](/sao10k)'s experiment over [Euryale v2.2](/sao10k/l3.1-euryale-70b).
$3.00/$3.00I/O
16K ctx
Jan 2025
|
$3.00 | $3.00 | — | 16K | Jan 2025 | ||
| $0.5 | $3.00 | $0.05 | 1.05M | Jul 2026 | |||
|
Z
GLM 5.1 is Z.ai's foundation model with a 200K-token context window, positioned between the earlier GLM 5 and the newer reasoning-focused GLM 5.2.
$0.966/$3.04I/O
200K ctx
Apr 2026
|
$0.966 | $3.04 | $0.1794 | 200K | Apr 2026 | ||
|
A
A multimodal model from Amazon designed to balance accuracy and speed across a broad range of tasks, with a 300K-token context window.
$0.8/$3.20I/O
300K ctx
Dec 2024
|
$0.8 | $3.20 | — | 300K | Dec 2024 | ||
| $0.65 | $3.25 | $0.13 | 1M | Sep 2025 | |||
|
S
A routing system that analyzes requests and directs them to specialized LLMs from a managed library based on task requirements.
$0.85/$3.40I/O
131K ctx
Jul 2025
|
$0.85 | $3.40 | — | 131K | Jul 2025 | ||
| $0.73 | $3.50 | $0.15 | 262K | Jun 2026 | |||
| $0.78 | $3.90 | — | 262K | Feb 2026 | |||
| $0.95 | $4.00 | $0.16 | 256K | Apr 2026 | |||
|
Z
GLM 5 Turbo is Z.ai's optimized inference variant of the GLM 5 foundation model, designed for faster response times while maintaining broad capability across reasoning and long-context tasks.
$1.20/$4.00I/O
202K ctx
Mar 2026
|
$1.20 | $4.00 | $0.24 | 202K | Mar 2026 | ||
| $3.00 | $4.00 | — | 16K | Aug 2023 | |||
|
Z
GLM-5V Turbo is Z.ai's native multimodal model that processes image, video, and text inputs for agent-driven tasks with a 202K token context window.
$1.20/$4.00I/O
202K ctx
Apr 2026
|
$1.20 | $4.00 | $0.24 | 202K | Apr 2026 | ||
|
T
Inkling is an open-weight mixture-of-experts model with 41B active parameters out of 975B total, supporting multimodal inputs and a 524K token context window.
$1.00/$4.05I/O
524K ctx
Jul 2026
|
$1.00 | $4.05 | $0.17 | 524K | Jul 2026 | ||
| $1.25 | $4.25 | — | 1M | Jul 2026 | |||
| $1.10 | $4.40 | $0.275 | 200K | Apr 2025 | |||
| $1.10 | $4.40 | $0.275 | 200K | Apr 2025 | |||
| $1.10 | $4.40 | $0.55 | 200K | Feb 2025 | |||
| $0.75 | $4.50 | $0.075 | 400K | Mar 2026 | |||
| $1.00 | $5.00 | $0.1 | 200K | Oct 2025 | |||
|
A
This is a series of models designed to replicate the prose quality of the Claude 3 models, specifically Sonnet(https://openrouter.ai/anthropic/claude-3.5-sonnet) and Opus(https://openrouter.ai/anthropic/claude-3-opus). The model is fine-tuned on top of [Qwen2.5 72B](https://openrouter.ai/qwen/qwen-2.5-72b-instruct).
$3.00/$5.00I/O
16K ctx
Oct 2024
|
$3.00 | $5.00 | — | 16K | Oct 2024 | ||
|
W
Writer's most advanced model designed for building and scaling AI agents in enterprise environments, with a context window supporting up to 1 million tokens.
$0.6/$6.00I/O
1.04M ctx
Jan 2026
|
$0.6 | $6.00 | — | 1.04M | Jan 2026 | ||
| $2.00 | $6.00 | $0.2 | 131K | Nov 2024 | |||
| $2.00 | $6.00 | $0.2 | 65K | Apr 2024 | |||
|
A
A multi-model system from AionLabs that combines specialized models for collaborative generation, designed around roleplaying and storytelling tasks.
$3.00/$6.00I/O
131K ctx
Jul 2026
|
$3.00 | $6.00 | $0.75 | 131K | Jul 2026 | ||
| $2.00 | $6.00 | $0.3 | 500K | Jul 2026 | |||
| $1.04 | $6.24 | — | 262K | Apr 2026 | |||
| $1.50 | $7.50 | — | 128K | Apr 2026 | |||
| $2.50 | $7.50 | — | 1M | May 2026 | |||
| $1.25 | $7.50 | $0.125 | 1.05M | Jul 2026 | |||
| $1.50 | $7.50 | $0.15 | 1.05M | Jul 2026 | |||
| $2.00 | $8.00 | $0.5 | 200K | Apr 2025 | |||
| $2.00 | $8.00 | $0.5 | 1M | Apr 2025 | |||
| $2.00 | $8.00 | $0.5 | 200K | Oct 2025 | |||
|
A
Jamba Large 1.7 is a hybrid SSM-Transformer model from AI21 with a 256K token context window, designed for efficient processing of long documents and complex tasks.
$2.00/$8.00I/O
256K ctx
Aug 2025
|
$2.00 | $8.00 | — | 256K | Aug 2025 | ||
|
P
A reasoning model based on DeepSeek R1 that uses chain-of-thought processing to work through complex problems step-by-step, with access to Perplexity's search integration.
$2.00/$8.00I/O
128K ctx
Mar 2025
|
$2.00 | $8.00 | — | 128K | Mar 2025 | ||
|
P
A research-focused model from Perplexity that autonomously searches, retrieves, and synthesizes information across complex topics with a 128k token context window.
$2.00/$8.00I/O
128K ctx
Mar 2025
|
$2.00 | $8.00 | — | 128K | Mar 2025 | ||
|
A
Aion-1.0 is a large language model from AionLabs with a 131k-token context window, released before the company's shift toward multi-model systems.
$4.00/$8.00I/O
131K ctx
Feb 2025
|
$4.00 | $8.00 | — | 131K | Feb 2025 | ||
| $1.50 | $9.00 | $0.15 | 1M | May 2026 | |||
| $1.25 | $10.00 | $0.125 | 1M | Aug 2025 | |||
| $1.25 | $10.00 | $0.125 | 1M | Jun 2025 | |||
| $2.50 | $10.00 | — | 128K | Jan 2026 | |||
| $1.25 | $10.00 | $0.125 | 400K | Dec 2025 | |||
| $1.25 | $10.00 | $0.125 | 400K | Nov 2025 | |||
| $1.25 | $10.00 | $0.13 | 400K | Nov 2025 | |||
| $1.25 | $10.00 | $0.125 | 400K | Sep 2025 | |||
| $1.25 | $10.00 | $0.125 | 128K | Aug 2025 | |||
| $2.50 | $10.00 | — | 256K | Mar 2025 | |||
| $2.50 | $10.00 | — | 128K | Mar 2025 | |||
| $2.50 | $10.00 | $1.25 | 128K | Nov 2024 | |||
|
I
Inflection 3 Productivity is a compact instruction-following model optimized for structured outputs and rule-based tasks with access to recent information.
$2.50/$10.00I/O
8K ctx
Oct 2024
|
$2.50 | $10.00 | — | 8K | Oct 2024 | ||
|
I
Inflection 3 Pi is a conversational model that powers the Pi chatbot, designed for dialogue with built-in knowledge of recent events.
$2.50/$10.00I/O
8K ctx
Oct 2024
|
$2.50 | $10.00 | — | 8K | Oct 2024 | ||
| $2.50 | $10.00 | — | 128K | Aug 2024 | |||
| $2.50 | $10.00 | $1.25 | 128K | Aug 2024 | |||
| $2.50 | $10.00 | — | 128K | Apr 2024 | |||
| $2.00 | $10.00 | $0.2 | 1M | Jun 2026 | |||
| $2.00 | $12.00 | $0.2 | 1M | Feb 2026 | |||
| $2.00 | $12.00 | $0.2 | 1M | Nov 2025 | |||
|
A
Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling custom models.
$2.50/$12.50I/O
1M ctx
Oct 2025
|
$2.50 | $12.50 | $0.625 | 1M | Oct 2025 | ||
| $1.75 | $14.00 | $0.175 | 128K | Mar 2026 | |||
| $1.75 | $14.00 | $0.175 | 400K | Feb 2026 | |||
| $1.75 | $14.00 | $0.175 | 400K | Jan 2026 | |||
| $1.75 | $14.00 | $0.175 | 400K | Dec 2025 | |||
| $3.00 | $15.00 | $0.3 | 1M | Feb 2026 | |||
| $3.00 | $15.00 | $0.3 | 1M | Sep 2025 | |||
| $2.50 | $15.00 | $0.25 | 1.05M | Mar 2026 | |||
|
P
Sonar Pro is a large language model from Perplexity with a 200k token context window that integrates web search capabilities.
$3.00/$15.00I/O
200K ctx
Mar 2025
|
$3.00 | $15.00 | — | 200K | Mar 2025 | ||
| $5.00 | $15.00 | — | 128K | May 2024 | |||
| $3.00 | $15.00 | $0.3 | 1.05M | Jul 2026 | |||
| $5.00 | $25.00 | $0.5 | 1M | May 2026 | |||
| $5.00 | $25.00 | $0.5 | 1M | Apr 2026 | |||
| $5.00 | $25.00 | $0.5 | 1M | Feb 2026 | |||
| $5.00 | $25.00 | $0.5 | 200K | Nov 2025 | |||
| $5.00 | $25.00 | $0.5 | 1M | Jul 2026 | |||
| $5.00 | $30.00 | $0.5 | 1M | Apr 2026 | |||
| $10.00 | $30.00 | — | 128K | Jan 2024 | |||
|
S
Fugu Ultra is Sakana's high-performance multi-agent orchestration system that routes requests across specialized models rather than using a single monolithic architecture.
$5.00/$30.00I/O
1M ctx
Jun 2026
|
$5.00 | $30.00 | $0.5 | 1M | Jun 2026 | ||
| $5.00 | $30.00 | $0.5 | 1.05M | Jul 2026 | |||
| $10.00 | $40.00 | $2.50 | 200K | Oct 2025 | |||
| $10.00 | $50.00 | $1.00 | 1M | Jun 2026 | |||
| $20.00 | $80.00 | — | 200K | Jun 2025 | |||
| $20.00 | $80.00 | — | 200K | Jun 2025 | |||
| $15.00 | $120.00 | — | 400K | Oct 2025 | |||
| $21.00 | $168.00 | — | 400K | Dec 2025 | |||
| $30.00 | $180.00 | — | 1M | Apr 2026 | |||
| $30.00 | $180.00 | — | 1.05M | Mar 2026 | |||
| $150.00 | $600.00 | — | 200K | Mar 2025 |
Sorted by output price by default — output tokens are typically 3–5× more expensive than input and dominate the bill for generation-heavy workloads. Sort by input or cached input when your mix differs.