prices synced 2026-09-24

Market events

Every pricing and market event we've tracked — 58 in all, going back to 2023. Curated from the headlines we scan each day.

2026

  1. Mercury 2.5 launches at $0.04/$0.15

    Inception's Mercury 2.5 launches at $0.04/$0.15 per MTok, an aggressively cheap entry to the LLM API market.

    At $0.04/$0.15 per MTok, Mercury 2.5 prices roughly 6x below Mercury 2's Output cost ($0.75/1M) dramatically undercuts Claude Haiku ($4.00) and Gemini Flash ($2.50) baseline, while still matching quality on cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.

  2. GPT-6 Astra launches at $50/MTok

    OpenAI debuts GPT-6 Astra, an automated AI engineer priced at $50/million tokens, roughly under $6/hour of agent work.

    Astra prices at $10 input / $50 output per million tokens, matching the same rate as Anthropic's Fable 5.1, while running 2.5 times the price of GPT-5.6 Sol, closing the window to shop between the two frontier labs on price alone. Requests above 272K input tokens bill at 2x input and 1.5x output for the full request, a real cliff for anyone running long-context agent work.

    Source
  3. Muse Spark 1.3: 90% training discount

    Meta's Muse Spark 1.3 matches GPT-5.6-Sol performance and launches with a >90% discount for training workloads, establishing Meta Superintelligence as a new frontier lab.

    Muse Spark 1.3 costs $1.25/M input tokens and $4.25/M output tokens, but the discounted "Contributor" tier drops to $0.10 per million input tokens, $0.20 per million output tokens because prompts and outputs may be used to improve Meta's products — so the 90%+ savings come from handing over your data for training, not from a cheaper base rate.

  4. Claude Fable/Mythos 5.1: 75% cache cut

    Anthropic's Claude Fable 5.1 and Mythos 5.1 launch with a 75% cache-read price cut, though 1.7x higher output token usage raises per-task costs ~20%.

    Anthropic has cut a Fable 5.1 cache hit to just $0.25 on input, down from $1.00 for Fable 5. That helps cache-heavy agent loops, cutting typical workload costs by about 25% and agentic ones by up to 45%, but base $10/$50 per-million input/output pricing is unchanged and the model's ~1.7x heavier output usage means low-cache-reuse tasks can net roughly 20% higher spend despite the cache discount.

    Source
  5. Meta Muse Image: $0.01 per image

    Meta's Muse Image launches with pricing at $0.01 per generated image, undercutting rival image APIs.

    At $0.01 per image, Muse Image runs roughly 2-6x cheaper than Google's Imagen 4 tier ($0.02-$0.06) Per-image costs: Google's Imagen 4 ranges from $0.02–$0.06 and over 16x cheaper than OpenAI's top GPT Image 1 tier OpenAI (GPT Image 1 High, $0.167/image), making high-volume jobs like ad variants or catalog imagery viable at scale when images cost $0.01 each.

  6. GLM-5.3-Flash launches at $0.15/$0.50

    Z.ai's GLM-5.3-Flash launches at $0.15/$0.50 per MTok, positioning as a price-competitive Flash-tier alternative.

    GLM-5.3-Flash costs $0.15 per 1M input tokens (very competitive, median: $0.53) and $0.50 per 1M output tokens (very competitive, median: $2.20), based on Z AI's API. At 49 tokens per second, GLM-5.3-Flash is notably slow (67). That undercut comes with a throughput tax, so the price wins on cost-per-token but loses ground in latency-sensitive or agentic pipelines where output speed compounds.

  7. GPT-5.6 Sol cut to $4/$20

    OpenAI cut GPT-5.6 Sol API pricing to $4/M input and $20/M output tokens.

    GPT‑5.6 Sol's new rate undercuts Anthropic's Claude Opus 5, listed at $5 per 1 million input tokens and $25 per 1 million output tokens, but the discount only covers requests using no more than 272,000 input tokens and is guaranteed at least through November 21, so budgets built on $4/$20 need a fallback rate once the promo lapses.

  8. Anthropic Opus 5 undercuts rivals

    Anthropic cut Opus 5 pricing, overtaking competitors on cost per token, per tldr_ai coverage.

    Opus 5 holds at $5/$25 per million tokens while undercutting OpenAI's flagship on output cost — Opus 5 has lower output pricing ($25 against $30 per million tokens) than GPT-5.5, so frontier-tier reasoning now costs less than the "cheaper" competitor's top model, flipping the usual assumption about which lab is the budget pick at the high end.

  9. GPT-5.6 Sol: developer price cut 20%

    OpenAI cuts developer pricing for its frontier GPT-5.6 Sol model by more than 20%.

    The cut is temporary — three months only — and lands entirely on output tokens ($30→$20/M), the line item that scales fastest for coding agents and chatbots, while input drops less ($5→$4/M).

  10. Stripe buys OpenRouter for $7B

    Stripe acquired LLM API aggregator OpenRouter for $7B, reshaping the model-routing and pricing brokerage layer.

    Stripe finalizes a $7B+ acquisition of OpenRouter, the multi-model AI gateway used by 8M developers. is the routing algorithm genuinely agnostic to Stripe's billing interests, or will cost optimization toward Chinese models — which is often the economically rational developer choice — be quietly deprioritized as Stripe manages competing commercial relationships?

    Source
  11. GPT-5.6 Sol: pricing cut 50%

    OpenAI cut GPT-5.6 Sol API pricing by 50%, extending the frontier price war.

    Sol's cut brings its flagship rate to $2.50 per million input tokens and $15.00 per million output tokens, down from $5 input / $30 output per million tokens — the same price OpenAI was charging for the mid-tier Terra model just weeks earlier, so top-tier reasoning now costs what the step-down tier used to.

    Source
  12. Gemini 3.7 Flash cut 50%

    Gemini 3.7 Flash launches with a 50% introductory cut to $0.75/$3.75 per 1M tokens.

    Lock in current pricing before it reverts: the $0.75/$3.75 rate is temporary, jumping to permanent $1.50/$7.50 per 1M tokens on January 1, 2027 On January 1, 2027, the price reverts to permanent rates of $1.50 per million input tokens and $7.50 per million output tokens. Until then it undercuts rivals sharply — roughly a third the blended cost of Claude Sonnet 5 or GPT-5.6 Terra.

  13. Grok 4.6 launches at $2/$6

    xAI's Grok 4.6 launches at $2/$6 per MTok input/output, positioned as a frontier model undercutting competitors on price.

    Grok 4.6 pricing is $2 input and $6 output per 1M tokens, but the measured effective input price is $0.74, since nine in ten input tokens across live Grok 4.6 traffic are served from cache.

  14. DeepSeek V4 Pro raises prices

    DeepSeek-V4-Pro (0813) implements a pricing increase that reduces its prior cost advantage.

    Off-peak V4-Pro rates now run $0.66/M input and $1.98/M output, up from $0.435/$0.87 during off-peak hours, V4-Pro input goes from $0.435 to $0.66 per million tokens, and output jumps from $0.87 to $1.98, with those rates double at peak hours — so cost-modeling now needs a peak/off-peak multiplier, not a flat rate, and the gap to Anthropic/OpenAI pricing narrows accordingly.

  15. Claude Sonnet 5: permanent $2/$10

    Anthropic made its reduced Claude Sonnet 5 pricing permanent at $2/M input and $10/M output.

    Introductory pricing of $2/$10 per million input/output tokens is in effect through August 31, 2026, after which the standard pricing of $3/$15 per million input/output tokens will take effect was the plan before Anthropic made $2/$10 permanent — so anyone who modeled budgets on $2/$10 and braced for a scheduled hike now keeps output costs locked at roughly a third below that planned $15/M rate,…

  16. DeepSeek plans price increase

    DeepSeek announced plans to significantly raise its LLM API prices, reversing its aggressive discount strategy.

    DeepSeek said it plans to significantly increase prices for its API services, though it did not disclose the timing or size of the increase, urging customers to plan usage accordingly.

    Source
  17. Muse Spark 1.2 launches at $1.25/$4.25

    Meta's Muse Spark 1.2 launches at $1.25/$4.25 per 1M tokens, setting a new mid-tier price benchmark.

    At $4.25/M output versus a market median near $10, this undercuts most mid-tier reasoning models on the output side specifically, where usage-heavy agentic workloads rack up cost fastest. Muse Spark 1.2 (xhigh) costs $1.25 per 1M input tokens (better than average, median: $1.75) and $4.25 per 1M output tokens (very competitive, median: $10.00).

  18. DeepSeek V4-Flash: 105× cheaper

    DeepSeek V4-Flash is reported as the cheapest model to run, at 105× lower cost than Claude, undercutting the market.

    Cheap per-token pricing at $0.14/$0.28 per million doesn't guarantee cheap per-task cost: the model is unusually verbose, generating roughly twice the median token volume during evaluation, which means real-world costs depend heavily on the task, so a low headline rate can still lose to a pricier model if it needs far more tokens to finish the job.

  19. GPT-5.6 Terra cut 20%

    OpenAI cut GPT-5.6 Terra API pricing by 20%, alongside an 80% cut to Luna.

    Terra's blended rate now works out to $14 per million tokens ($2 in / $12 out), which puts it below Anthropic's mid-tier Claude Sonnet 4.6, priced at $3 per million input tokens and $15 per million output tokens — so any routing logic picking models on cost-per-token alone now favors Terra over Sonnet by a real marg...

  20. DeepSeek V4 Flash launches

    DeepSeek released V4-Flash on 2026-07-31 with intelligence, performance and price analysis, but concrete per-MTok pricing figures are not yet confirmed from these items.

    At $0.14 per million input tokens and $0.28 per million output on a cache miss, V4 Flash undercuts the median frontier model by roughly 3x on input and over 4x on output pricing for DeepSeek V4 Flash is $0.14 per 1M input tokens (competitively priced, median: $0.43) and $0.28 per 1M output tokens (competitively pric...