Market events
Every pricing and market event we've tracked — 34 in all, going back to 2023. Curated from the headlines we scan each day.
2026
-
GPT-5.6 Terra cut 20%
OpenAI cut GPT-5.6 Terra API pricing by 20%, alongside an 80% cut to Luna.
Terra's blended rate now works out to $14 per million tokens ($2 in / $12 out), which puts it below Anthropic's mid-tier Claude Sonnet 4.6, priced at $3 per million input tokens and $15 per million output tokens — so any routing logic picking models on cost-per-token alone now favors Terra over Sonnet by a real marg...
-
Claude Opus 5: Fable perf, half price
Anthropic launches Claude Opus 5 (ECI 159, SWE-ECI 161) matching Fable-level performance at Opus pricing, roughly half of Fable's cost.
I have enough grounding to write the note. Based on the sources gathered: The new model costs $5 per million input tokens and $25 per million output tokens, matching the price of its predecessor, Opus 4.8, and Anthropic released Claude Opus 5, a model the company says delivers nearly all the intelligence of its top...
-
Kimi K3: Opus-class at $3/$15
Moonshot's Kimi K3 (2.8T-A50B), the largest open-weights model yet, launches at $3/$15 per MTok—Opus 4.8-class quality at Sonnet 5 pricing.
-
Inkling opens Apache 2.0 frontier
Thinking Machines' Inkling (975B/41B active multimodal) launches under Apache 2.0 with API pricing of $1.87–$3.74 per 1M input tokens on Tinker depending on context window.
With thinking tokens billed like any other output tokens, longer chains of thought directly raise the cost of each response — which makes Inkling's claim of comparable Terminal Bench 2.1 results at roughly one-third the thinking tokens a pricing event, not just a launch. Open weights under Apache 2.0 mean the model is served by a field of competing inference providers rather than a single lab, pushing margins toward compute cost; the token efficiency drops the effective price of reasoning work on top of that. Together they set a commodity price floor under a workload tier — agentic reasoning — that closed frontier APIs have so far been able to price at a premium.
-
DeepSeek V4 adds peak-valley pricing
DeepSeek V4 introduces time-based peak/off-peak pricing, varying API token costs by demand.
During peak hours—9:00–12:00 and 14:00–18:00 Beijing time—API pricing will be doubled compared to regular rates, so for deepseek-v4-pro output jumps from ¥6.00 to ¥12.00 per million tokens during those windows. Anyone batching large jobs or building cost estimates now needs to schedule non-urgent inference for off-p...
Source -
DeepSeek V4 Pro 75% cut
DeepSeek makes a 75% promotional discount permanent, pricing V4 Pro at $0.435/$0.87.
Making the discount permanent removes the promo-expiry risk that kept teams from architecting around it, and locks in output pricing at roughly $0.87 per million tokens versus OpenAI's GPT-5 charges $2.50 per million input tokens and $10 per million output tokens and Anthropic's Claude Opus 4.7 is priced at $5 input...
-
Cheap Flash era ends
Gemini 3.5 Flash ships at $1.50/$9 per 1M — triple its predecessor — as Google reprices Flash from budget tier toward Pro territory.
High-volume workloads that routed to Flash specifically because it was the cheap tier now absorb a 3x cost jump — Google's Gemini 3.5 Flash has shattered that assumption; released on May 19, it costs $1.50 per million input tokens and $9 per million output tokens, and the model it effectively replaces, Gemini 3 Flas...
-
GPT-5.5 raises frontier prices
GPT-5.5 launches at $5/$30 per 1M, a sharp step up from GPT-5's commodity pricing — frontier labs begin testing price tolerance.
A workload that cost $8.81 to summarize a 10-page contract on GPT-5 (at $1.25 per 1e6 for input and $10 per 1e6 for output) now costs roughly 4x on input and 3x on output at $5.00 input, $30.00 output per 1M tokens for GPT-5.5, and that gap widens further since prompts with >272K input tokens are priced at 2x input ...
2025
-
Mistral Large 3: 75% cheaper
Mistral's open-weight frontier flagship lands at $0.50/$1.50 per 1M — 75% below Large 2 — keeping open-model pressure on closed pricing.
Because Mistral released Large 3 under Apache 2.0 as a 675B-parameter MoE with a sparse mixture-of-experts trained with 41B active and 675B total parameters, the $0.50/$1.50 API rate is only Mistral's own hosted price, not a floor — any inference provider can pull the weights and compete on cost, so expect further u...
-
Opus gets 67% cheaper
Anthropic drops Opus pricing from $15/$75 to $5/$25 with the Opus 4.5 release.
A million output tokens now runs $25 instead of $75, so any workload that was previously avoiding Opus for output-heavy tasks (long completions, agentic loops with lots of generated code) sees the same three-quarters cost cut on the more expensive side of the ledger. It also puts Opus's output rate below GPT-5.5's $...
-
Qwen3 Max halved in price war
Alibaba cuts Qwen3 Max roughly 50% as China's AI price war reignites, pressuring domestic and global rivals alike.
A trillion-parameter frontier model now runs at $0.459/M input and $1.836/M output tokens after the cut, Alibaba Group Holding has slashed charges for its biggest artificial intelligence model by as much as half, triggering speculation of another price war in China's highly competitive AI market, so anyone benchmark...
-
GPT-5 sparks price war
GPT-5 launches at commodity pricing ($1.25/$10) — TechCrunch calls it a price-war trigger.
Anyone paying Claude Opus 4.1 rates for frontier-tier work now has a flagship alternative at roughly a sixth the input cost and a seventh the output cost — GPT-5's API landing at $1.25 per million input tokens and $10 per million output tokens undercuts Anthropic's pricing, which could pressure rivals to respond, as...
-
o3 price cut 80%
OpenAI slashes o3 by 80% ($10→$2/1M input), making frontier reasoning mainstream-affordable.
Output tokens dropped the same 80%, from $40 to $8 per million "Now, OpenAI has dropped prices to $2 per million input tokens and $8 per million output tokens." That's cheap enough that resellers repriced instantly — Cursor now counts one o3 request the same as a GPT-4o call, and Windsurf lowered the "o3-reasoning" ...
-
Long-context goes cheap
GPT-4.1 lands a 1M-token window at mid-tier pricing, matching Gemini on context economics.
Feeding a full 1M-token prompt into GPT‑4.1 still bills at the flat $2/$8 per million rate, with no long-context surcharge, whereas Gemini 2.5 Pro charges $1.25 per million input tokens for contexts under 200K, but this jumps to $2.50 per million for longer contexts — so anyone routing large-document or long-chat-hi...
-
GPT-4.5: ultra-premium experiment
OpenAI tests $75/$150 pricing with GPT-4.5 — the most expensive API model ever offered. Quickly superseded by cheaper, better models.
At $75 per million input tokens and $150 per million output tokens, GPT-4.5 priced out at roughly 15x GPT-4o's rate, and OpenAI announced it would remove GPT-4.5 Preview from the API on July 14, 2025, just months after launch — a reminder that top-of-market pricing on a new flagship model tends to be a temporary tol...
-
The DeepSeek moment
DeepSeek R1 ships near-frontier reasoning at ~1/20th the price. Markets jolt; pricing pressure spikes industry-wide.
A workload that would cost $100 on OpenAI's o1 API runs about $3.60 on R1 for the same token volume, since input tokens cost $0.55 per million on DeepSeek R1 versus $15 per million on o1, and output tokens cost $2.19 per million on DeepSeek R1 versus $60 per million on o1, and on most reasoning benchmarks, DeepSeek ...
2024
-
DeepSeek V3: frontier for cents
DeepSeek releases a GPT-4o-class open model priced around $0.27/$1.10 per 1M, previewing the shock R1 would deliver a month later.
A workload benchmarked at GPT-4o quality that costs $2.50/$10 per million tokens on OpenAI drops to $0.27 per 1M input tokens and $1.10 per 1M output tokens on DeepSeek's API, roughly a 9x cut on both sides of the ledger — meaning teams already budgeting against 4o-class pricing get room to either slash spend or run...
-
OpenAI caching: auto 50% off
DevDay brings automatic prompt caching — a no-code 50% discount on recently seen input tokens across GPT-4o and o1 models.
Because the discount applies automatically to any request reusing the same prefix — no cache-write step or special API parameter required, unlike Anthropic's opt-in caching — Starting today, we will give you a 50% discount for input tokens that the model has seen recently, no action required. Starts at 1024 cached t...
-
Gemini 1.5 Pro cut 64%
Google cuts Gemini 1.5 Pro to $1.25/$5.00 per 1M on prompts under 128K — the flagship tier joins the price war.
At $1.25/$5.00 per 1M tokens, Gemini 1.5 Pro now undercuts GPT-4o's $2.50/$10.00 pricing by half on both input and output, while still offering the 2M-token context window that OpenAI's flagship doesn't match — a real option if you're routing large-context workloads purely on cost. Starting October 1, 2024, Gemini 1...
-
o1 creates reasoning price tier
OpenAI's o1-preview launches at $15/$60 — a new pricing category where chain-of-thought tokens make effective costs 2–5× the list price.
The $60/M list price is only the floor: the hidden reasoning tokens make it harder to manage costs compared to previous GPT models, and o1-preview charges $15 per million input tokens and $60 per million output tokens versus $5 and $15 for GPT-4o, and because reasoning tokens count as part of context window usage an...