prices synced 2026-08-10

Market events

Every pricing and market event we've tracked — 34 in all, going back to 2023. Curated from the headlines we scan each day.

2026

  1. GPT-5.6 Terra cut 20%

    OpenAI cut GPT-5.6 Terra API pricing by 20%, alongside an 80% cut to Luna.

    Terra's blended rate now works out to $14 per million tokens ($2 in / $12 out), which puts it below Anthropic's mid-tier Claude Sonnet 4.6, priced at $3 per million input tokens and $15 per million output tokens — so any routing logic picking models on cost-per-token alone now favors Terra over Sonnet by a real marg...

  2. Claude Opus 5: Fable perf, half price

    Anthropic launches Claude Opus 5 (ECI 159, SWE-ECI 161) matching Fable-level performance at Opus pricing, roughly half of Fable's cost.

    I have enough grounding to write the note. Based on the sources gathered: The new model costs $5 per million input tokens and $25 per million output tokens, matching the price of its predecessor, Opus 4.8, and Anthropic released Claude Opus 5, a model the company says delivers nearly all the intelligence of its top...

  3. Kimi K3: Opus-class at $3/$15

    Moonshot's Kimi K3 (2.8T-A50B), the largest open-weights model yet, launches at $3/$15 per MTok—Opus 4.8-class quality at Sonnet 5 pricing.

  4. Inkling opens Apache 2.0 frontier

    Thinking Machines' Inkling (975B/41B active multimodal) launches under Apache 2.0 with API pricing of $1.87–$3.74 per 1M input tokens on Tinker depending on context window.

    With thinking tokens billed like any other output tokens, longer chains of thought directly raise the cost of each response — which makes Inkling's claim of comparable Terminal Bench 2.1 results at roughly one-third the thinking tokens a pricing event, not just a launch. Open weights under Apache 2.0 mean the model is served by a field of competing inference providers rather than a single lab, pushing margins toward compute cost; the token efficiency drops the effective price of reasoning work on top of that. Together they set a commodity price floor under a workload tier — agentic reasoning — that closed frontier APIs have so far been able to price at a premium.

  5. DeepSeek V4 adds peak-valley pricing

    DeepSeek V4 introduces time-based peak/off-peak pricing, varying API token costs by demand.

    During peak hours—9:00–12:00 and 14:00–18:00 Beijing time—API pricing will be doubled compared to regular rates, so for deepseek-v4-pro output jumps from ¥6.00 to ¥12.00 per million tokens during those windows. Anyone batching large jobs or building cost estimates now needs to schedule non-urgent inference for off-p...

    Source
  6. DeepSeek V4 Pro 75% cut

    DeepSeek makes a 75% promotional discount permanent, pricing V4 Pro at $0.435/$0.87.

    Making the discount permanent removes the promo-expiry risk that kept teams from architecting around it, and locks in output pricing at roughly $0.87 per million tokens versus OpenAI's GPT-5 charges $2.50 per million input tokens and $10 per million output tokens and Anthropic's Claude Opus 4.7 is priced at $5 input...

  7. Cheap Flash era ends

    Gemini 3.5 Flash ships at $1.50/$9 per 1M — triple its predecessor — as Google reprices Flash from budget tier toward Pro territory.

    High-volume workloads that routed to Flash specifically because it was the cheap tier now absorb a 3x cost jump — Google's Gemini 3.5 Flash has shattered that assumption; released on May 19, it costs $1.50 per million input tokens and $9 per million output tokens, and the model it effectively replaces, Gemini 3 Flas...

  8. GPT-5.5 raises frontier prices

    GPT-5.5 launches at $5/$30 per 1M, a sharp step up from GPT-5's commodity pricing — frontier labs begin testing price tolerance.

    A workload that cost $8.81 to summarize a 10-page contract on GPT-5 (at $1.25 per 1e6 for input and $10 per 1e6 for output) now costs roughly 4x on input and 3x on output at $5.00 input, $30.00 output per 1M tokens for GPT-5.5, and that gap widens further since prompts with >272K input tokens are priced at 2x input ...

2025

  1. Mistral Large 3: 75% cheaper

    Mistral's open-weight frontier flagship lands at $0.50/$1.50 per 1M — 75% below Large 2 — keeping open-model pressure on closed pricing.

    Because Mistral released Large 3 under Apache 2.0 as a 675B-parameter MoE with a sparse mixture-of-experts trained with 41B active and 675B total parameters, the $0.50/$1.50 API rate is only Mistral's own hosted price, not a floor — any inference provider can pull the weights and compete on cost, so expect further u...

  2. Opus gets 67% cheaper

    Anthropic drops Opus pricing from $15/$75 to $5/$25 with the Opus 4.5 release.

    A million output tokens now runs $25 instead of $75, so any workload that was previously avoiding Opus for output-heavy tasks (long completions, agentic loops with lots of generated code) sees the same three-quarters cost cut on the more expensive side of the ledger. It also puts Opus's output rate below GPT-5.5's $...

  3. Qwen3 Max halved in price war

    Alibaba cuts Qwen3 Max roughly 50% as China's AI price war reignites, pressuring domestic and global rivals alike.

    A trillion-parameter frontier model now runs at $0.459/M input and $1.836/M output tokens after the cut, Alibaba Group Holding has slashed charges for its biggest artificial intelligence model by as much as half, triggering speculation of another price war in China's highly competitive AI market, so anyone benchmark...

  4. GPT-5 sparks price war

    GPT-5 launches at commodity pricing ($1.25/$10) — TechCrunch calls it a price-war trigger.

    Anyone paying Claude Opus 4.1 rates for frontier-tier work now has a flagship alternative at roughly a sixth the input cost and a seventh the output cost — GPT-5's API landing at $1.25 per million input tokens and $10 per million output tokens undercuts Anthropic's pricing, which could pressure rivals to respond, as...

  5. o3 price cut 80%

    OpenAI slashes o3 by 80% ($10→$2/1M input), making frontier reasoning mainstream-affordable.

    Output tokens dropped the same 80%, from $40 to $8 per million "Now, OpenAI has dropped prices to $2 per million input tokens and $8 per million output tokens." That's cheap enough that resellers repriced instantly — Cursor now counts one o3 request the same as a GPT-4o call, and Windsurf lowered the "o3-reasoning" ...

  6. Long-context goes cheap

    GPT-4.1 lands a 1M-token window at mid-tier pricing, matching Gemini on context economics.

    Feeding a full 1M-token prompt into GPT‑4.1 still bills at the flat $2/$8 per million rate, with no long-context surcharge, whereas Gemini 2.5 Pro charges $1.25 per million input tokens for contexts under 200K, but this jumps to $2.50 per million for longer contexts — so anyone routing large-document or long-chat-hi...

  7. GPT-4.5: ultra-premium experiment

    OpenAI tests $75/$150 pricing with GPT-4.5 — the most expensive API model ever offered. Quickly superseded by cheaper, better models.

    At $75 per million input tokens and $150 per million output tokens, GPT-4.5 priced out at roughly 15x GPT-4o's rate, and OpenAI announced it would remove GPT-4.5 Preview from the API on July 14, 2025, just months after launch — a reminder that top-of-market pricing on a new flagship model tends to be a temporary tol...

  8. The DeepSeek moment

    DeepSeek R1 ships near-frontier reasoning at ~1/20th the price. Markets jolt; pricing pressure spikes industry-wide.

    A workload that would cost $100 on OpenAI's o1 API runs about $3.60 on R1 for the same token volume, since input tokens cost $0.55 per million on DeepSeek R1 versus $15 per million on o1, and output tokens cost $2.19 per million on DeepSeek R1 versus $60 per million on o1, and on most reasoning benchmarks, DeepSeek ...

2024

  1. DeepSeek V3: frontier for cents

    DeepSeek releases a GPT-4o-class open model priced around $0.27/$1.10 per 1M, previewing the shock R1 would deliver a month later.

    A workload benchmarked at GPT-4o quality that costs $2.50/$10 per million tokens on OpenAI drops to $0.27 per 1M input tokens and $1.10 per 1M output tokens on DeepSeek's API, roughly a 9x cut on both sides of the ledger — meaning teams already budgeting against 4o-class pricing get room to either slash spend or run...

  2. OpenAI caching: auto 50% off

    DevDay brings automatic prompt caching — a no-code 50% discount on recently seen input tokens across GPT-4o and o1 models.

    Because the discount applies automatically to any request reusing the same prefix — no cache-write step or special API parameter required, unlike Anthropic's opt-in caching — Starting today, we will give you a 50% discount for input tokens that the model has seen recently, no action required. Starts at 1024 cached t...

  3. Gemini 1.5 Pro cut 64%

    Google cuts Gemini 1.5 Pro to $1.25/$5.00 per 1M on prompts under 128K — the flagship tier joins the price war.

    At $1.25/$5.00 per 1M tokens, Gemini 1.5 Pro now undercuts GPT-4o's $2.50/$10.00 pricing by half on both input and output, while still offering the 2M-token context window that OpenAI's flagship doesn't match — a real option if you're routing large-context workloads purely on cost. Starting October 1, 2024, Gemini 1...

  4. o1 creates reasoning price tier

    OpenAI's o1-preview launches at $15/$60 — a new pricing category where chain-of-thought tokens make effective costs 2–5× the list price.

    The $60/M list price is only the floor: the hidden reasoning tokens make it harder to manage costs compared to previous GPT models, and o1-preview charges $15 per million input tokens and $60 per million output tokens versus $5 and $15 for GPT-4o, and because reasoning tokens count as part of context window usage an...

  5. Anthropic ships prompt caching

    Cached input tokens cost 90% less on Claude, making long-system-prompt and RAG workloads dramatically cheaper.

    Because Anthropic introduced prompt caching as a beta feature in August 2024 with a clean economic structure: cache reads at 0.1× the standard input rate (a 90% discount), cache writes at 1.25× the standard rate for a 5-minute TTL, 2× for a 1-hour TTL, any workload that repeatedly sends the same system prompt, tool ...

  6. Google slashes Flash 78%

    Gemini 1.5 Flash drops 78% on input / 71% on output to $0.075/$0.30 per 1M, undercutting GPT-4o mini by half.

    At $0.075/$0.30 per 1M tokens, running high-volume workloads like summarization or classification on Flash costs roughly half what GPT-4o mini charges for the same job, so throughput-heavy pipelines see input costs cut by more than 4x overnight the input price by 78% to $0.075/1 million tokens and the output price b...

  7. OpenAI cuts GPT-4o 50%

    GPT-4o input drops $5→$2.50/1M and prompt caching arrives — the first big frontier price war shot.

    A repeated system prompt or long context now costs $2.50/1M on input instead of $5, and once prompt caching landed that same reused content dropped a further 50% for any prefix over 1024 tokens"we will give you a 50% discount for input tokens that the model has seen recently, no action required. Starts at 1024 cache...

  8. Llama 3.1 405B opens frontier

    Meta releases a GPT-4-class model as open weights; hosted 405B undercuts proprietary frontier pricing and anchors expectations lower.

    This gives me exactly the comparison I need. Hosted 405B launched at roughly comparable-to-cheaper rates than the incumbent frontier model: the 405B model pricing on APIs is very similar to GPT-4o, ranging from $3-9 per million input tokens and $3-15 per million output tokens where GPT-4o is $5 per million input, $...

  9. GPT-4o mini at $0.15

    OpenAI replaces GPT-3.5 Turbo with GPT-4o mini at $0.15/$0.60 per 1M — over 60% cheaper and far more capable.

    This makes running default/fallback workloads on GPT-3.5 Turbo hard to justify: for a typical 1K-in/500-out request, cost drops from roughly $2.50 to about $0.45 per million tokens blended, while GPT-4o mini scores 82% on MMLU and currently outperforms GPT-4 on chat preferences, so anyone still budgeting around 3.5 ...

  10. Claude 3.5 Sonnet resets value

    Anthropic ships a model beating Opus at one-fifth Opus's price ($3/$15), collapsing the gap between mid-tier price and frontier quality.

    Anyone still paying Opus's $15/$75 rate for frontier quality had no more reason to, since a model at $3/$15 matched or beat it on outperforming competitor models and Claude 3 Opus on a wide range of evaluations, with the speed and cost of our mid-tier model — an 80% discount that made picking the pricier tier for ca...

  11. China's LLM price war erupts

    After DeepSeek-V2 launched at ~$0.14/1M input on May 6, Alibaba cuts Qwen prices up to 97% and Baidu makes Ernie Speed/Lite free hours later.

    At $0.14 per million input tokens, DeepSeek-V2 undercut GPT-4 Turbo pricing by roughly one-seventieth, and the response was immediate: Alibaba matched with cuts of up to 97% on a range of models, with Baidu and others following within hours. For anyone budgeting API spend, this means the floor for a capable model's ...

  12. GPT-4o: frontier at half price

    GPT-4o launches at $5/$15 per 1M — half of GPT-4 Turbo's price with better performance, resetting the frontier price bar.

    Anyone still budgeting against GPT-4 Turbo's rates gets an immediate 50% cut on the same workload: GPT-4o was advertised as 50% cheaper than GPT-4 Turbo, which implied roughly $15/million output if GPT-4 Turbo was $30, so switching models alone — with no drop in quality — halves the API bill on every request.

  13. Claude 3 brings $0.25 Haiku

    The Claude 3 family ships with Haiku at $0.25/$1.25 per 1M, staking out the fast-and-cheap tier against GPT-3.5 Turbo.

    Undercuts GPT-3.5 Turbo's $0.50/$1.50 per 1M rate by half on input and a third on output, so any workload already tuned for 3.5 Turbo's cost profile gets a straight price cut just by swapping endpoints, before even weighing the quality difference.

  14. Gemini 1.5 Pro: 1M context

    Google announces a 1M-token context window, an order of magnitude beyond rivals — long context becomes a $/token battleground.

    A 1M-token window means a full codebase or hundreds of PDFs fit in one call, but Google splits pricing at the 128K mark—for prompts up to 128,000 tokens in size, the price is $1.25 per 1 million tokens, going up to $2.50 per 1 million tokens for prompts longer than 128,000 tokens, so pushing a request past that thre...

  15. OpenAI cuts GPT-3.5 Turbo 50%

    Third GPT-3.5 Turbo cut in a year: input drops 50% to $0.50/1M, output 25% to $1.50. New embedding models arrive 5x cheaper too.

    At $0.50/1M input and $1.50/1M output, GPT-3.5 Turbo now undercuts the per-token cost of running Claude 2.0 or 2.1, since GPT-3.5 Turbo input prices have now been reduced by 50% to $0.0005 per 1k tokens and output prices are reduced by 25% to $0.0015/1K tokens, putting it at a cost per 1k tokens less than Anthropic'...

2023

  1. Mixtral 8x7B goes open-weight

    Mistral releases Mixtral 8x7B under Apache 2.0 — GPT-3.5-class quality anyone can host, setting a price floor under proprietary small models.

    Because Mixtral 8x7B ships as open weights under Apache 2.0 and matches or outperforms GPT3.5 on most standard benchmarks, anyone can self-host GPT-3.5-class inference instead of paying OpenAI's rate — and third-party providers moved fast on this, with Fireworks pricing it at $0.4/million for prompt and $1.6/million...

  2. GPT-4 Turbo: 3× cheaper

    GPT-4 Turbo launches at $10/$30 with 128K context, cutting the frontier price by two-thirds and kicking off a year of rapid cuts.

    A workload that previously required chunking an 8K-context prompt into multiple GPT-4 calls now fits in a single 128K call, while the per-token bill itself drops: When GPT-4 launched in March 2023, it cost $30 per million input tokens and $60 per million output tokens, versus GPT-4 Turbo pricing starting at $10.00 p...

  3. GPT-4 sets the frontier price

    GPT-4 launches at $30/$60 per MTok — 10× the cost of GPT-3.5 Turbo, establishing the first frontier pricing baseline.

    Running a chat product on GPT-4 at launch meant paying $30/$60 per MTok versus GPT-3.5 Turbo's $0.002 per 1K tokens, so any workload swapped over to the frontier model saw its per-token bill jump roughly 15x on input alone — a gap large enough that most teams had to route only the hardest prompts to GPT-4 and keep e...

No more events.