prices synced 2026-09-24

Market events

Every pricing and market event we've tracked — 58 in all, going back to 2023. Curated from the headlines we scan each day.

2026

  1. Mercury 2.5 launches at $0.04/$0.15

    Inception's Mercury 2.5 launches at $0.04/$0.15 per MTok, an aggressively cheap entry to the LLM API market.

    At $0.04/$0.15 per MTok, Mercury 2.5 prices roughly 6x below Mercury 2's Output cost ($0.75/1M) dramatically undercuts Claude Haiku ($4.00) and Gemini Flash ($2.50) baseline, while still matching quality on cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.

  2. GPT-6 Astra launches at $50/MTok

    OpenAI debuts GPT-6 Astra, an automated AI engineer priced at $50/million tokens, roughly under $6/hour of agent work.

    Astra prices at $10 input / $50 output per million tokens, matching the same rate as Anthropic's Fable 5.1, while running 2.5 times the price of GPT-5.6 Sol, closing the window to shop between the two frontier labs on price alone. Requests above 272K input tokens bill at 2x input and 1.5x output for the full request, a real cliff for anyone running long-context agent work.

    Source
  3. Muse Spark 1.3: 90% training discount

    Meta's Muse Spark 1.3 matches GPT-5.6-Sol performance and launches with a >90% discount for training workloads, establishing Meta Superintelligence as a new frontier lab.

    Muse Spark 1.3 costs $1.25/M input tokens and $4.25/M output tokens, but the discounted "Contributor" tier drops to $0.10 per million input tokens, $0.20 per million output tokens because prompts and outputs may be used to improve Meta's products — so the 90%+ savings come from handing over your data for training, not from a cheaper base rate.

  4. Claude Fable/Mythos 5.1: 75% cache cut

    Anthropic's Claude Fable 5.1 and Mythos 5.1 launch with a 75% cache-read price cut, though 1.7x higher output token usage raises per-task costs ~20%.

    Anthropic has cut a Fable 5.1 cache hit to just $0.25 on input, down from $1.00 for Fable 5. That helps cache-heavy agent loops, cutting typical workload costs by about 25% and agentic ones by up to 45%, but base $10/$50 per-million input/output pricing is unchanged and the model's ~1.7x heavier output usage means low-cache-reuse tasks can net roughly 20% higher spend despite the cache discount.

    Source
  5. Meta Muse Image: $0.01 per image

    Meta's Muse Image launches with pricing at $0.01 per generated image, undercutting rival image APIs.

    At $0.01 per image, Muse Image runs roughly 2-6x cheaper than Google's Imagen 4 tier ($0.02-$0.06) Per-image costs: Google's Imagen 4 ranges from $0.02–$0.06 and over 16x cheaper than OpenAI's top GPT Image 1 tier OpenAI (GPT Image 1 High, $0.167/image), making high-volume jobs like ad variants or catalog imagery viable at scale when images cost $0.01 each.

  6. GLM-5.3-Flash launches at $0.15/$0.50

    Z.ai's GLM-5.3-Flash launches at $0.15/$0.50 per MTok, positioning as a price-competitive Flash-tier alternative.

    GLM-5.3-Flash costs $0.15 per 1M input tokens (very competitive, median: $0.53) and $0.50 per 1M output tokens (very competitive, median: $2.20), based on Z AI's API. At 49 tokens per second, GLM-5.3-Flash is notably slow (67). That undercut comes with a throughput tax, so the price wins on cost-per-token but loses ground in latency-sensitive or agentic pipelines where output speed compounds.

  7. GPT-5.6 Sol cut to $4/$20

    OpenAI cut GPT-5.6 Sol API pricing to $4/M input and $20/M output tokens.

    GPT‑5.6 Sol's new rate undercuts Anthropic's Claude Opus 5, listed at $5 per 1 million input tokens and $25 per 1 million output tokens, but the discount only covers requests using no more than 272,000 input tokens and is guaranteed at least through November 21, so budgets built on $4/$20 need a fallback rate once the promo lapses.

  8. Anthropic Opus 5 undercuts rivals

    Anthropic cut Opus 5 pricing, overtaking competitors on cost per token, per tldr_ai coverage.

    Opus 5 holds at $5/$25 per million tokens while undercutting OpenAI's flagship on output cost — Opus 5 has lower output pricing ($25 against $30 per million tokens) than GPT-5.5, so frontier-tier reasoning now costs less than the "cheaper" competitor's top model, flipping the usual assumption about which lab is the budget pick at the high end.

  9. GPT-5.6 Sol: developer price cut 20%

    OpenAI cuts developer pricing for its frontier GPT-5.6 Sol model by more than 20%.

    The cut is temporary — three months only — and lands entirely on output tokens ($30→$20/M), the line item that scales fastest for coding agents and chatbots, while input drops less ($5→$4/M).

  10. Stripe buys OpenRouter for $7B

    Stripe acquired LLM API aggregator OpenRouter for $7B, reshaping the model-routing and pricing brokerage layer.

    Stripe finalizes a $7B+ acquisition of OpenRouter, the multi-model AI gateway used by 8M developers. is the routing algorithm genuinely agnostic to Stripe's billing interests, or will cost optimization toward Chinese models — which is often the economically rational developer choice — be quietly deprioritized as Stripe manages competing commercial relationships?

    Source
  11. GPT-5.6 Sol: pricing cut 50%

    OpenAI cut GPT-5.6 Sol API pricing by 50%, extending the frontier price war.

    Sol's cut brings its flagship rate to $2.50 per million input tokens and $15.00 per million output tokens, down from $5 input / $30 output per million tokens — the same price OpenAI was charging for the mid-tier Terra model just weeks earlier, so top-tier reasoning now costs what the step-down tier used to.

    Source
  12. Gemini 3.7 Flash cut 50%

    Gemini 3.7 Flash launches with a 50% introductory cut to $0.75/$3.75 per 1M tokens.

    Lock in current pricing before it reverts: the $0.75/$3.75 rate is temporary, jumping to permanent $1.50/$7.50 per 1M tokens on January 1, 2027 On January 1, 2027, the price reverts to permanent rates of $1.50 per million input tokens and $7.50 per million output tokens. Until then it undercuts rivals sharply — roughly a third the blended cost of Claude Sonnet 5 or GPT-5.6 Terra.

  13. Grok 4.6 launches at $2/$6

    xAI's Grok 4.6 launches at $2/$6 per MTok input/output, positioned as a frontier model undercutting competitors on price.

    Grok 4.6 pricing is $2 input and $6 output per 1M tokens, but the measured effective input price is $0.74, since nine in ten input tokens across live Grok 4.6 traffic are served from cache.

  14. DeepSeek V4 Pro raises prices

    DeepSeek-V4-Pro (0813) implements a pricing increase that reduces its prior cost advantage.

    Off-peak V4-Pro rates now run $0.66/M input and $1.98/M output, up from $0.435/$0.87 during off-peak hours, V4-Pro input goes from $0.435 to $0.66 per million tokens, and output jumps from $0.87 to $1.98, with those rates double at peak hours — so cost-modeling now needs a peak/off-peak multiplier, not a flat rate, and the gap to Anthropic/OpenAI pricing narrows accordingly.

  15. Claude Sonnet 5: permanent $2/$10

    Anthropic made its reduced Claude Sonnet 5 pricing permanent at $2/M input and $10/M output.

    Introductory pricing of $2/$10 per million input/output tokens is in effect through August 31, 2026, after which the standard pricing of $3/$15 per million input/output tokens will take effect was the plan before Anthropic made $2/$10 permanent — so anyone who modeled budgets on $2/$10 and braced for a scheduled hike now keeps output costs locked at roughly a third below that planned $15/M rate,…

  16. DeepSeek plans price increase

    DeepSeek announced plans to significantly raise its LLM API prices, reversing its aggressive discount strategy.

    DeepSeek said it plans to significantly increase prices for its API services, though it did not disclose the timing or size of the increase, urging customers to plan usage accordingly.

    Source
  17. Muse Spark 1.2 launches at $1.25/$4.25

    Meta's Muse Spark 1.2 launches at $1.25/$4.25 per 1M tokens, setting a new mid-tier price benchmark.

    At $4.25/M output versus a market median near $10, this undercuts most mid-tier reasoning models on the output side specifically, where usage-heavy agentic workloads rack up cost fastest. Muse Spark 1.2 (xhigh) costs $1.25 per 1M input tokens (better than average, median: $1.75) and $4.25 per 1M output tokens (very competitive, median: $10.00).

  18. DeepSeek V4-Flash: 105× cheaper

    DeepSeek V4-Flash is reported as the cheapest model to run, at 105× lower cost than Claude, undercutting the market.

    Cheap per-token pricing at $0.14/$0.28 per million doesn't guarantee cheap per-task cost: the model is unusually verbose, generating roughly twice the median token volume during evaluation, which means real-world costs depend heavily on the task, so a low headline rate can still lose to a pricier model if it needs far more tokens to finish the job.

  19. GPT-5.6 Terra cut 20%

    OpenAI cut GPT-5.6 Terra API pricing by 20%, alongside an 80% cut to Luna.

    Terra's blended rate now works out to $14 per million tokens ($2 in / $12 out), which puts it below Anthropic's mid-tier Claude Sonnet 4.6, priced at $3 per million input tokens and $15 per million output tokens — so any routing logic picking models on cost-per-token alone now favors Terra over Sonnet by a real marg...

  20. DeepSeek V4 Flash launches

    DeepSeek released V4-Flash on 2026-07-31 with intelligence, performance and price analysis, but concrete per-MTok pricing figures are not yet confirmed from these items.

    At $0.14 per million input tokens and $0.28 per million output on a cache miss, V4 Flash undercuts the median frontier model by roughly 3x on input and over 4x on output pricing for DeepSeek V4 Flash is $0.14 per 1M input tokens (competitively priced, median: $0.43) and $0.28 per 1M output tokens (competitively pric...

  21. Grok Voice launches at $0.09/min

    SpaceX AI's Grok Voice debuts with per-minute pricing at $0.09/min, setting a new benchmark for voice API costs.

    The new Think Fast 2.0 rate is a 60% jump from the original $0.05/min Grok Voice Think Fast was upgraded to version 2.0, priced at $0.08 per minute, with Grok-Voice-Latest transitioning from version 1.0 to 2.0 on August 5, so anyone pinned to the `grok-voice-latest` alias for budget predictability gets auto-migrated...

    Source
  22. GPT-5.6 cuts output tokens 6×

    OpenAI's GPT-5.6 advances the price-performance frontier with efficiency gains reducing output tokens by roughly 6×, lowering effective inference cost.

    Source
  23. GPT-5.6 Luna cut 80%

    OpenAI cut GPT-5.6 Luna prices by 80% and Terra by 20%, and introduced a new Sol tier at 2.5× lower latency for 2× standard price.

    Luna now clears at $1.40 combined per million tokens, undercutting Google's Gemini 3.5 Flash-Lite ($2.80) and Gemini 3.6 Flash ($9), so high-volume routine calls that were marginal on Luna at the old $7 rate become a clear default choice, while the widened gap to Terra (now $2/$12 combined) makes correct tier-routin...

  24. OpenAI Transcribe cut 25%

    OpenAI cut GPT Transcribe pricing from $6 to $4.50 per 1,000 minutes, a 25% reduction.

    Flagship transcription now runs $0.27/hour instead of $0.36, while GPT-4o Mini Transcribe remains available at $3.00 per 1,000 minutes — so the mini tier's cost edge over the top model shrinks from 2x to about 1.5x, and it comes with a small accuracy bump too, since GPT Transcribe scores 3.31% on AA-WER, improving 0...

  25. Claude Opus 5: Fable perf, half price

    Anthropic launches Claude Opus 5 (ECI 159, SWE-ECI 161) matching Fable-level performance at Opus pricing, roughly half of Fable's cost.

    I have enough grounding to write the note. Based on the sources gathered: The new model costs $5 per million input tokens and $25 per million output tokens, matching the price of its predecessor, Opus 4.8, and Anthropic released Claude Opus 5, a model the company says delivers nearly all the intelligence of its top...

  26. Laguna S 2.1 undercuts DeepSeek V4

    Poolside's open-weight Laguna S 2.1 launches claiming cheaper pricing than DeepSeek V4 Flash while outperforming V4 Pro, though exact per-MTok figures are unconfirmed.

    At $0.10/$0.20 per million input/output tokens, Laguna S 2.1 is priced at $0.10 per million input tokens and $0.20 per million output tokens, versus DeepSeek V4 Pro at $0.435 per million input tokens, $0.87 per million output tokens — roughly 4x cheaper — while on Terminal-Bench 2.1 and SWE-Bench Pro, Laguna S 2.1 m...

    Source
  27. Kimi K3: Opus-class at $3/$15

    Moonshot's Kimi K3 (2.8T-A50B), the largest open-weights model yet, launches at $3/$15 per MTok—Opus 4.8-class quality at Sonnet 5 pricing.

  28. Inkling opens Apache 2.0 frontier

    Thinking Machines' Inkling (975B/41B active multimodal) launches under Apache 2.0 with API pricing of $1.87–$3.74 per 1M input tokens on Tinker depending on context window.

    With thinking tokens billed like any other output tokens, longer chains of thought directly raise the cost of each response — which makes Inkling's claim of comparable Terminal Bench 2.1 results at roughly one-third the thinking tokens a pricing event, not just a launch. Open weights under Apache 2.0 mean the model is served by a field of competing inference providers rather than a single lab, pushing margins toward compute cost; the token efficiency drops the effective price of reasoning work on top of that. Together they set a commodity price floor under a workload tier — agentic reasoning — that closed frontier APIs have so far been able to price at a premium.

  29. DeepSeek V4 adds peak-valley pricing

    DeepSeek V4 introduces time-based peak/off-peak pricing, varying API token costs by demand.

    During peak hours—9:00–12:00 and 14:00–18:00 Beijing time—API pricing will be doubled compared to regular rates, so for deepseek-v4-pro output jumps from ¥6.00 to ¥12.00 per million tokens during those windows. Anyone batching large jobs or building cost estimates now needs to schedule non-urgent inference for off-p...

    Source
  30. DeepSeek V4 Pro 75% cut

    DeepSeek makes a 75% promotional discount permanent, pricing V4 Pro at $0.435/$0.87.

    Making the discount permanent removes the promo-expiry risk that kept teams from architecting around it, and locks in output pricing at roughly $0.87 per million tokens versus OpenAI's GPT-5 charges $2.50 per million input tokens and $10 per million output tokens and Anthropic's Claude Opus 4.7 is priced at $5 input...

  31. Cheap Flash era ends

    Gemini 3.5 Flash ships at $1.50/$9 per 1M — triple its predecessor — as Google reprices Flash from budget tier toward Pro territory.

    High-volume workloads that routed to Flash specifically because it was the cheap tier now absorb a 3x cost jump — Google's Gemini 3.5 Flash has shattered that assumption; released on May 19, it costs $1.50 per million input tokens and $9 per million output tokens, and the model it effectively replaces, Gemini 3 Flas...

  32. GPT-5.5 raises frontier prices

    GPT-5.5 launches at $5/$30 per 1M, a sharp step up from GPT-5's commodity pricing — frontier labs begin testing price tolerance.

    A workload that cost $8.81 to summarize a 10-page contract on GPT-5 (at $1.25 per 1e6 for input and $10 per 1e6 for output) now costs roughly 4x on input and 3x on output at $5.00 input, $30.00 output per 1M tokens for GPT-5.5, and that gap widens further since prompts with >272K input tokens are priced at 2x input ...

2025

  1. Mistral Large 3: 75% cheaper

    Mistral's open-weight frontier flagship lands at $0.50/$1.50 per 1M — 75% below Large 2 — keeping open-model pressure on closed pricing.

    Because Mistral released Large 3 under Apache 2.0 as a 675B-parameter MoE with a sparse mixture-of-experts trained with 41B active and 675B total parameters, the $0.50/$1.50 API rate is only Mistral's own hosted price, not a floor — any inference provider can pull the weights and compete on cost, so expect further u...

  2. Opus gets 67% cheaper

    Anthropic drops Opus pricing from $15/$75 to $5/$25 with the Opus 4.5 release.

    A million output tokens now runs $25 instead of $75, so any workload that was previously avoiding Opus for output-heavy tasks (long completions, agentic loops with lots of generated code) sees the same three-quarters cost cut on the more expensive side of the ledger. It also puts Opus's output rate below GPT-5.5's $...

  3. Qwen3 Max halved in price war

    Alibaba cuts Qwen3 Max roughly 50% as China's AI price war reignites, pressuring domestic and global rivals alike.

    A trillion-parameter frontier model now runs at $0.459/M input and $1.836/M output tokens after the cut, Alibaba Group Holding has slashed charges for its biggest artificial intelligence model by as much as half, triggering speculation of another price war in China's highly competitive AI market, so anyone benchmark...

  4. GPT-5 sparks price war

    GPT-5 launches at commodity pricing ($1.25/$10) — TechCrunch calls it a price-war trigger.

    Anyone paying Claude Opus 4.1 rates for frontier-tier work now has a flagship alternative at roughly a sixth the input cost and a seventh the output cost — GPT-5's API landing at $1.25 per million input tokens and $10 per million output tokens undercuts Anthropic's pricing, which could pressure rivals to respond, as...

  5. o3 price cut 80%

    OpenAI slashes o3 by 80% ($10→$2/1M input), making frontier reasoning mainstream-affordable.

    Output tokens dropped the same 80%, from $40 to $8 per million "Now, OpenAI has dropped prices to $2 per million input tokens and $8 per million output tokens." That's cheap enough that resellers repriced instantly — Cursor now counts one o3 request the same as a GPT-4o call, and Windsurf lowered the "o3-reasoning" ...

  6. Long-context goes cheap

    GPT-4.1 lands a 1M-token window at mid-tier pricing, matching Gemini on context economics.

    Feeding a full 1M-token prompt into GPT‑4.1 still bills at the flat $2/$8 per million rate, with no long-context surcharge, whereas Gemini 2.5 Pro charges $1.25 per million input tokens for contexts under 200K, but this jumps to $2.50 per million for longer contexts — so anyone routing large-document or long-chat-hi...

  7. GPT-4.5: ultra-premium experiment

    OpenAI tests $75/$150 pricing with GPT-4.5 — the most expensive API model ever offered. Quickly superseded by cheaper, better models.

    At $75 per million input tokens and $150 per million output tokens, GPT-4.5 priced out at roughly 15x GPT-4o's rate, and OpenAI announced it would remove GPT-4.5 Preview from the API on July 14, 2025, just months after launch — a reminder that top-of-market pricing on a new flagship model tends to be a temporary tol...

  8. The DeepSeek moment

    DeepSeek R1 ships near-frontier reasoning at ~1/20th the price. Markets jolt; pricing pressure spikes industry-wide.

    A workload that would cost $100 on OpenAI's o1 API runs about $3.60 on R1 for the same token volume, since input tokens cost $0.55 per million on DeepSeek R1 versus $15 per million on o1, and output tokens cost $2.19 per million on DeepSeek R1 versus $60 per million on o1, and on most reasoning benchmarks, DeepSeek ...