Market events
Every pricing and market event we've tracked — 58 in all, going back to 2023. Curated from the headlines we scan each day.
2026
-
Mercury 2.5 launches at $0.04/$0.15
Inception's Mercury 2.5 launches at $0.04/$0.15 per MTok, an aggressively cheap entry to the LLM API market.
At $0.04/$0.15 per MTok, Mercury 2.5 prices roughly 6x below Mercury 2's Output cost ($0.75/1M) dramatically undercuts Claude Haiku ($4.00) and Gemini Flash ($2.50) baseline, while still matching quality on cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.
-
GPT-6 Astra launches at $50/MTok
OpenAI debuts GPT-6 Astra, an automated AI engineer priced at $50/million tokens, roughly under $6/hour of agent work.
Astra prices at $10 input / $50 output per million tokens, matching the same rate as Anthropic's Fable 5.1, while running 2.5 times the price of GPT-5.6 Sol, closing the window to shop between the two frontier labs on price alone. Requests above 272K input tokens bill at 2x input and 1.5x output for the full request, a real cliff for anyone running long-context agent work.
Source -
Muse Spark 1.3: 90% training discount
Meta's Muse Spark 1.3 matches GPT-5.6-Sol performance and launches with a >90% discount for training workloads, establishing Meta Superintelligence as a new frontier lab.
Muse Spark 1.3 costs $1.25/M input tokens and $4.25/M output tokens, but the discounted "Contributor" tier drops to $0.10 per million input tokens, $0.20 per million output tokens because prompts and outputs may be used to improve Meta's products — so the 90%+ savings come from handing over your data for training, not from a cheaper base rate.
-
Claude Fable/Mythos 5.1: 75% cache cut
Anthropic's Claude Fable 5.1 and Mythos 5.1 launch with a 75% cache-read price cut, though 1.7x higher output token usage raises per-task costs ~20%.
Anthropic has cut a Fable 5.1 cache hit to just $0.25 on input, down from $1.00 for Fable 5. That helps cache-heavy agent loops, cutting typical workload costs by about 25% and agentic ones by up to 45%, but base $10/$50 per-million input/output pricing is unchanged and the model's ~1.7x heavier output usage means low-cache-reuse tasks can net roughly 20% higher spend despite the cache discount.
Source -
Meta Muse Image: $0.01 per image
Meta's Muse Image launches with pricing at $0.01 per generated image, undercutting rival image APIs.
At $0.01 per image, Muse Image runs roughly 2-6x cheaper than Google's Imagen 4 tier ($0.02-$0.06) Per-image costs: Google's Imagen 4 ranges from $0.02–$0.06 and over 16x cheaper than OpenAI's top GPT Image 1 tier OpenAI (GPT Image 1 High, $0.167/image), making high-volume jobs like ad variants or catalog imagery viable at scale when images cost $0.01 each.
-
GLM-5.3-Flash launches at $0.15/$0.50
Z.ai's GLM-5.3-Flash launches at $0.15/$0.50 per MTok, positioning as a price-competitive Flash-tier alternative.
GLM-5.3-Flash costs $0.15 per 1M input tokens (very competitive, median: $0.53) and $0.50 per 1M output tokens (very competitive, median: $2.20), based on Z AI's API. At 49 tokens per second, GLM-5.3-Flash is notably slow (67). That undercut comes with a throughput tax, so the price wins on cost-per-token but loses ground in latency-sensitive or agentic pipelines where output speed compounds.
-
GPT-5.6 Sol cut to $4/$20
OpenAI cut GPT-5.6 Sol API pricing to $4/M input and $20/M output tokens.
GPT‑5.6 Sol's new rate undercuts Anthropic's Claude Opus 5, listed at $5 per 1 million input tokens and $25 per 1 million output tokens, but the discount only covers requests using no more than 272,000 input tokens and is guaranteed at least through November 21, so budgets built on $4/$20 need a fallback rate once the promo lapses.
-
Anthropic Opus 5 undercuts rivals
Anthropic cut Opus 5 pricing, overtaking competitors on cost per token, per tldr_ai coverage.
Opus 5 holds at $5/$25 per million tokens while undercutting OpenAI's flagship on output cost — Opus 5 has lower output pricing ($25 against $30 per million tokens) than GPT-5.5, so frontier-tier reasoning now costs less than the "cheaper" competitor's top model, flipping the usual assumption about which lab is the budget pick at the high end.
-
GPT-5.6 Sol: developer price cut 20%
OpenAI cuts developer pricing for its frontier GPT-5.6 Sol model by more than 20%.
The cut is temporary — three months only — and lands entirely on output tokens ($30→$20/M), the line item that scales fastest for coding agents and chatbots, while input drops less ($5→$4/M).
-
Stripe buys OpenRouter for $7B
Stripe acquired LLM API aggregator OpenRouter for $7B, reshaping the model-routing and pricing brokerage layer.
Stripe finalizes a $7B+ acquisition of OpenRouter, the multi-model AI gateway used by 8M developers. is the routing algorithm genuinely agnostic to Stripe's billing interests, or will cost optimization toward Chinese models — which is often the economically rational developer choice — be quietly deprioritized as Stripe manages competing commercial relationships?
Source -
GPT-5.6 Sol: pricing cut 50%
OpenAI cut GPT-5.6 Sol API pricing by 50%, extending the frontier price war.
Sol's cut brings its flagship rate to $2.50 per million input tokens and $15.00 per million output tokens, down from $5 input / $30 output per million tokens — the same price OpenAI was charging for the mid-tier Terra model just weeks earlier, so top-tier reasoning now costs what the step-down tier used to.
Source -
Gemini 3.7 Flash cut 50%
Gemini 3.7 Flash launches with a 50% introductory cut to $0.75/$3.75 per 1M tokens.
Lock in current pricing before it reverts: the $0.75/$3.75 rate is temporary, jumping to permanent $1.50/$7.50 per 1M tokens on January 1, 2027 On January 1, 2027, the price reverts to permanent rates of $1.50 per million input tokens and $7.50 per million output tokens. Until then it undercuts rivals sharply — roughly a third the blended cost of Claude Sonnet 5 or GPT-5.6 Terra.
-
Grok 4.6 launches at $2/$6
xAI's Grok 4.6 launches at $2/$6 per MTok input/output, positioned as a frontier model undercutting competitors on price.
Grok 4.6 pricing is $2 input and $6 output per 1M tokens, but the measured effective input price is $0.74, since nine in ten input tokens across live Grok 4.6 traffic are served from cache.
-
DeepSeek V4 Pro raises prices
DeepSeek-V4-Pro (0813) implements a pricing increase that reduces its prior cost advantage.
Off-peak V4-Pro rates now run $0.66/M input and $1.98/M output, up from $0.435/$0.87 during off-peak hours, V4-Pro input goes from $0.435 to $0.66 per million tokens, and output jumps from $0.87 to $1.98, with those rates double at peak hours — so cost-modeling now needs a peak/off-peak multiplier, not a flat rate, and the gap to Anthropic/OpenAI pricing narrows accordingly.
-
Claude Sonnet 5: permanent $2/$10
Anthropic made its reduced Claude Sonnet 5 pricing permanent at $2/M input and $10/M output.
Introductory pricing of $2/$10 per million input/output tokens is in effect through August 31, 2026, after which the standard pricing of $3/$15 per million input/output tokens will take effect was the plan before Anthropic made $2/$10 permanent — so anyone who modeled budgets on $2/$10 and braced for a scheduled hike now keeps output costs locked at roughly a third below that planned $15/M rate,…
-
DeepSeek plans price increase
DeepSeek announced plans to significantly raise its LLM API prices, reversing its aggressive discount strategy.
DeepSeek said it plans to significantly increase prices for its API services, though it did not disclose the timing or size of the increase, urging customers to plan usage accordingly.
Source -
Muse Spark 1.2 launches at $1.25/$4.25
Meta's Muse Spark 1.2 launches at $1.25/$4.25 per 1M tokens, setting a new mid-tier price benchmark.
At $4.25/M output versus a market median near $10, this undercuts most mid-tier reasoning models on the output side specifically, where usage-heavy agentic workloads rack up cost fastest. Muse Spark 1.2 (xhigh) costs $1.25 per 1M input tokens (better than average, median: $1.75) and $4.25 per 1M output tokens (very competitive, median: $10.00).
-
DeepSeek V4-Flash: 105× cheaper
DeepSeek V4-Flash is reported as the cheapest model to run, at 105× lower cost than Claude, undercutting the market.
Cheap per-token pricing at $0.14/$0.28 per million doesn't guarantee cheap per-task cost: the model is unusually verbose, generating roughly twice the median token volume during evaluation, which means real-world costs depend heavily on the task, so a low headline rate can still lose to a pricier model if it needs far more tokens to finish the job.
-
GPT-5.6 Terra cut 20%
OpenAI cut GPT-5.6 Terra API pricing by 20%, alongside an 80% cut to Luna.
Terra's blended rate now works out to $14 per million tokens ($2 in / $12 out), which puts it below Anthropic's mid-tier Claude Sonnet 4.6, priced at $3 per million input tokens and $15 per million output tokens — so any routing logic picking models on cost-per-token alone now favors Terra over Sonnet by a real marg...
-
DeepSeek V4 Flash launches
DeepSeek released V4-Flash on 2026-07-31 with intelligence, performance and price analysis, but concrete per-MTok pricing figures are not yet confirmed from these items.
At $0.14 per million input tokens and $0.28 per million output on a cache miss, V4 Flash undercuts the median frontier model by roughly 3x on input and over 4x on output pricing for DeepSeek V4 Flash is $0.14 per 1M input tokens (competitively priced, median: $0.43) and $0.28 per 1M output tokens (competitively pric...
-
Grok Voice launches at $0.09/min
SpaceX AI's Grok Voice debuts with per-minute pricing at $0.09/min, setting a new benchmark for voice API costs.
The new Think Fast 2.0 rate is a 60% jump from the original $0.05/min Grok Voice Think Fast was upgraded to version 2.0, priced at $0.08 per minute, with Grok-Voice-Latest transitioning from version 1.0 to 2.0 on August 5, so anyone pinned to the `grok-voice-latest` alias for budget predictability gets auto-migrated...
Source -
GPT-5.6 cuts output tokens 6×
OpenAI's GPT-5.6 advances the price-performance frontier with efficiency gains reducing output tokens by roughly 6×, lowering effective inference cost.
Source -
GPT-5.6 Luna cut 80%
OpenAI cut GPT-5.6 Luna prices by 80% and Terra by 20%, and introduced a new Sol tier at 2.5× lower latency for 2× standard price.
Luna now clears at $1.40 combined per million tokens, undercutting Google's Gemini 3.5 Flash-Lite ($2.80) and Gemini 3.6 Flash ($9), so high-volume routine calls that were marginal on Luna at the old $7 rate become a clear default choice, while the widened gap to Terra (now $2/$12 combined) makes correct tier-routin...
-
OpenAI Transcribe cut 25%
OpenAI cut GPT Transcribe pricing from $6 to $4.50 per 1,000 minutes, a 25% reduction.
Flagship transcription now runs $0.27/hour instead of $0.36, while GPT-4o Mini Transcribe remains available at $3.00 per 1,000 minutes — so the mini tier's cost edge over the top model shrinks from 2x to about 1.5x, and it comes with a small accuracy bump too, since GPT Transcribe scores 3.31% on AA-WER, improving 0...
-
Claude Opus 5: Fable perf, half price
Anthropic launches Claude Opus 5 (ECI 159, SWE-ECI 161) matching Fable-level performance at Opus pricing, roughly half of Fable's cost.
I have enough grounding to write the note. Based on the sources gathered: The new model costs $5 per million input tokens and $25 per million output tokens, matching the price of its predecessor, Opus 4.8, and Anthropic released Claude Opus 5, a model the company says delivers nearly all the intelligence of its top...
-
Laguna S 2.1 undercuts DeepSeek V4
Poolside's open-weight Laguna S 2.1 launches claiming cheaper pricing than DeepSeek V4 Flash while outperforming V4 Pro, though exact per-MTok figures are unconfirmed.
At $0.10/$0.20 per million input/output tokens, Laguna S 2.1 is priced at $0.10 per million input tokens and $0.20 per million output tokens, versus DeepSeek V4 Pro at $0.435 per million input tokens, $0.87 per million output tokens — roughly 4x cheaper — while on Terminal-Bench 2.1 and SWE-Bench Pro, Laguna S 2.1 m...
Source -
Kimi K3: Opus-class at $3/$15
Moonshot's Kimi K3 (2.8T-A50B), the largest open-weights model yet, launches at $3/$15 per MTok—Opus 4.8-class quality at Sonnet 5 pricing.
-
Inkling opens Apache 2.0 frontier
Thinking Machines' Inkling (975B/41B active multimodal) launches under Apache 2.0 with API pricing of $1.87–$3.74 per 1M input tokens on Tinker depending on context window.
With thinking tokens billed like any other output tokens, longer chains of thought directly raise the cost of each response — which makes Inkling's claim of comparable Terminal Bench 2.1 results at roughly one-third the thinking tokens a pricing event, not just a launch. Open weights under Apache 2.0 mean the model is served by a field of competing inference providers rather than a single lab, pushing margins toward compute cost; the token efficiency drops the effective price of reasoning work on top of that. Together they set a commodity price floor under a workload tier — agentic reasoning — that closed frontier APIs have so far been able to price at a premium.
-
DeepSeek V4 adds peak-valley pricing
DeepSeek V4 introduces time-based peak/off-peak pricing, varying API token costs by demand.
During peak hours—9:00–12:00 and 14:00–18:00 Beijing time—API pricing will be doubled compared to regular rates, so for deepseek-v4-pro output jumps from ¥6.00 to ¥12.00 per million tokens during those windows. Anyone batching large jobs or building cost estimates now needs to schedule non-urgent inference for off-p...
Source -
DeepSeek V4 Pro 75% cut
DeepSeek makes a 75% promotional discount permanent, pricing V4 Pro at $0.435/$0.87.
Making the discount permanent removes the promo-expiry risk that kept teams from architecting around it, and locks in output pricing at roughly $0.87 per million tokens versus OpenAI's GPT-5 charges $2.50 per million input tokens and $10 per million output tokens and Anthropic's Claude Opus 4.7 is priced at $5 input...
-
Cheap Flash era ends
Gemini 3.5 Flash ships at $1.50/$9 per 1M — triple its predecessor — as Google reprices Flash from budget tier toward Pro territory.
High-volume workloads that routed to Flash specifically because it was the cheap tier now absorb a 3x cost jump — Google's Gemini 3.5 Flash has shattered that assumption; released on May 19, it costs $1.50 per million input tokens and $9 per million output tokens, and the model it effectively replaces, Gemini 3 Flas...
-
GPT-5.5 raises frontier prices
GPT-5.5 launches at $5/$30 per 1M, a sharp step up from GPT-5's commodity pricing — frontier labs begin testing price tolerance.
A workload that cost $8.81 to summarize a 10-page contract on GPT-5 (at $1.25 per 1e6 for input and $10 per 1e6 for output) now costs roughly 4x on input and 3x on output at $5.00 input, $30.00 output per 1M tokens for GPT-5.5, and that gap widens further since prompts with >272K input tokens are priced at 2x input ...
2025
-
Mistral Large 3: 75% cheaper
Mistral's open-weight frontier flagship lands at $0.50/$1.50 per 1M — 75% below Large 2 — keeping open-model pressure on closed pricing.
Because Mistral released Large 3 under Apache 2.0 as a 675B-parameter MoE with a sparse mixture-of-experts trained with 41B active and 675B total parameters, the $0.50/$1.50 API rate is only Mistral's own hosted price, not a floor — any inference provider can pull the weights and compete on cost, so expect further u...
-
Opus gets 67% cheaper
Anthropic drops Opus pricing from $15/$75 to $5/$25 with the Opus 4.5 release.
A million output tokens now runs $25 instead of $75, so any workload that was previously avoiding Opus for output-heavy tasks (long completions, agentic loops with lots of generated code) sees the same three-quarters cost cut on the more expensive side of the ledger. It also puts Opus's output rate below GPT-5.5's $...
-
Qwen3 Max halved in price war
Alibaba cuts Qwen3 Max roughly 50% as China's AI price war reignites, pressuring domestic and global rivals alike.
A trillion-parameter frontier model now runs at $0.459/M input and $1.836/M output tokens after the cut, Alibaba Group Holding has slashed charges for its biggest artificial intelligence model by as much as half, triggering speculation of another price war in China's highly competitive AI market, so anyone benchmark...
-
GPT-5 sparks price war
GPT-5 launches at commodity pricing ($1.25/$10) — TechCrunch calls it a price-war trigger.
Anyone paying Claude Opus 4.1 rates for frontier-tier work now has a flagship alternative at roughly a sixth the input cost and a seventh the output cost — GPT-5's API landing at $1.25 per million input tokens and $10 per million output tokens undercuts Anthropic's pricing, which could pressure rivals to respond, as...
-
o3 price cut 80%
OpenAI slashes o3 by 80% ($10→$2/1M input), making frontier reasoning mainstream-affordable.
Output tokens dropped the same 80%, from $40 to $8 per million "Now, OpenAI has dropped prices to $2 per million input tokens and $8 per million output tokens." That's cheap enough that resellers repriced instantly — Cursor now counts one o3 request the same as a GPT-4o call, and Windsurf lowered the "o3-reasoning" ...
-
Long-context goes cheap
GPT-4.1 lands a 1M-token window at mid-tier pricing, matching Gemini on context economics.
Feeding a full 1M-token prompt into GPT‑4.1 still bills at the flat $2/$8 per million rate, with no long-context surcharge, whereas Gemini 2.5 Pro charges $1.25 per million input tokens for contexts under 200K, but this jumps to $2.50 per million for longer contexts — so anyone routing large-document or long-chat-hi...
-
GPT-4.5: ultra-premium experiment
OpenAI tests $75/$150 pricing with GPT-4.5 — the most expensive API model ever offered. Quickly superseded by cheaper, better models.
At $75 per million input tokens and $150 per million output tokens, GPT-4.5 priced out at roughly 15x GPT-4o's rate, and OpenAI announced it would remove GPT-4.5 Preview from the API on July 14, 2025, just months after launch — a reminder that top-of-market pricing on a new flagship model tends to be a temporary tol...
-
The DeepSeek moment
DeepSeek R1 ships near-frontier reasoning at ~1/20th the price. Markets jolt; pricing pressure spikes industry-wide.
A workload that would cost $100 on OpenAI's o1 API runs about $3.60 on R1 for the same token volume, since input tokens cost $0.55 per million on DeepSeek R1 versus $15 per million on o1, and output tokens cost $2.19 per million on DeepSeek R1 versus $60 per million on o1, and on most reasoning benchmarks, DeepSeek ...
2024
-
DeepSeek V3: frontier for cents
DeepSeek releases a GPT-4o-class open model priced around $0.27/$1.10 per 1M, previewing the shock R1 would deliver a month later.
A workload benchmarked at GPT-4o quality that costs $2.50/$10 per million tokens on OpenAI drops to $0.27 per 1M input tokens and $1.10 per 1M output tokens on DeepSeek's API, roughly a 9x cut on both sides of the ledger — meaning teams already budgeting against 4o-class pricing get room to either slash spend or run...
-
OpenAI caching: auto 50% off
DevDay brings automatic prompt caching — a no-code 50% discount on recently seen input tokens across GPT-4o and o1 models.
Because the discount applies automatically to any request reusing the same prefix — no cache-write step or special API parameter required, unlike Anthropic's opt-in caching — Starting today, we will give you a 50% discount for input tokens that the model has seen recently, no action required. Starts at 1024 cached t...
-
Gemini 1.5 Pro cut 64%
Google cuts Gemini 1.5 Pro to $1.25/$5.00 per 1M on prompts under 128K — the flagship tier joins the price war.
At $1.25/$5.00 per 1M tokens, Gemini 1.5 Pro now undercuts GPT-4o's $2.50/$10.00 pricing by half on both input and output, while still offering the 2M-token context window that OpenAI's flagship doesn't match — a real option if you're routing large-context workloads purely on cost. Starting October 1, 2024, Gemini 1...
-
o1 creates reasoning price tier
OpenAI's o1-preview launches at $15/$60 — a new pricing category where chain-of-thought tokens make effective costs 2–5× the list price.
The $60/M list price is only the floor: the hidden reasoning tokens make it harder to manage costs compared to previous GPT models, and o1-preview charges $15 per million input tokens and $60 per million output tokens versus $5 and $15 for GPT-4o, and because reasoning tokens count as part of context window usage an...
-
Anthropic ships prompt caching
Cached input tokens cost 90% less on Claude, making long-system-prompt and RAG workloads dramatically cheaper.
Because Anthropic introduced prompt caching as a beta feature in August 2024 with a clean economic structure: cache reads at 0.1× the standard input rate (a 90% discount), cache writes at 1.25× the standard rate for a 5-minute TTL, 2× for a 1-hour TTL, any workload that repeatedly sends the same system prompt, tool ...
-
Google slashes Flash 78%
Gemini 1.5 Flash drops 78% on input / 71% on output to $0.075/$0.30 per 1M, undercutting GPT-4o mini by half.
At $0.075/$0.30 per 1M tokens, running high-volume workloads like summarization or classification on Flash costs roughly half what GPT-4o mini charges for the same job, so throughput-heavy pipelines see input costs cut by more than 4x overnight the input price by 78% to $0.075/1 million tokens and the output price b...
-
OpenAI cuts GPT-4o 50%
GPT-4o input drops $5→$2.50/1M and prompt caching arrives — the first big frontier price war shot.
A repeated system prompt or long context now costs $2.50/1M on input instead of $5, and once prompt caching landed that same reused content dropped a further 50% for any prefix over 1024 tokens"we will give you a 50% discount for input tokens that the model has seen recently, no action required. Starts at 1024 cache...
-
Llama 3.1 405B opens frontier
Meta releases a GPT-4-class model as open weights; hosted 405B undercuts proprietary frontier pricing and anchors expectations lower.
This gives me exactly the comparison I need. Hosted 405B launched at roughly comparable-to-cheaper rates than the incumbent frontier model: the 405B model pricing on APIs is very similar to GPT-4o, ranging from $3-9 per million input tokens and $3-15 per million output tokens where GPT-4o is $5 per million input, $...
-
GPT-4o mini at $0.15
OpenAI replaces GPT-3.5 Turbo with GPT-4o mini at $0.15/$0.60 per 1M — over 60% cheaper and far more capable.
This makes running default/fallback workloads on GPT-3.5 Turbo hard to justify: for a typical 1K-in/500-out request, cost drops from roughly $2.50 to about $0.45 per million tokens blended, while GPT-4o mini scores 82% on MMLU and currently outperforms GPT-4 on chat preferences, so anyone still budgeting around 3.5 ...
-
Claude 3.5 Sonnet resets value
Anthropic ships a model beating Opus at one-fifth Opus's price ($3/$15), collapsing the gap between mid-tier price and frontier quality.
Anyone still paying Opus's $15/$75 rate for frontier quality had no more reason to, since a model at $3/$15 matched or beat it on outperforming competitor models and Claude 3 Opus on a wide range of evaluations, with the speed and cost of our mid-tier model — an 80% discount that made picking the pricier tier for ca...
-
China's LLM price war erupts
After DeepSeek-V2 launched at ~$0.14/1M input on May 6, Alibaba cuts Qwen prices up to 97% and Baidu makes Ernie Speed/Lite free hours later.
At $0.14 per million input tokens, DeepSeek-V2 undercut GPT-4 Turbo pricing by roughly one-seventieth, and the response was immediate: Alibaba matched with cuts of up to 97% on a range of models, with Baidu and others following within hours. For anyone budgeting API spend, this means the floor for a capable model's ...
-
GPT-4o: frontier at half price
GPT-4o launches at $5/$15 per 1M — half of GPT-4 Turbo's price with better performance, resetting the frontier price bar.
Anyone still budgeting against GPT-4 Turbo's rates gets an immediate 50% cut on the same workload: GPT-4o was advertised as 50% cheaper than GPT-4 Turbo, which implied roughly $15/million output if GPT-4 Turbo was $30, so switching models alone — with no drop in quality — halves the API bill on every request.
-
Claude 3 brings $0.25 Haiku
The Claude 3 family ships with Haiku at $0.25/$1.25 per 1M, staking out the fast-and-cheap tier against GPT-3.5 Turbo.
Undercuts GPT-3.5 Turbo's $0.50/$1.50 per 1M rate by half on input and a third on output, so any workload already tuned for 3.5 Turbo's cost profile gets a straight price cut just by swapping endpoints, before even weighing the quality difference.
-
Gemini 1.5 Pro: 1M context
Google announces a 1M-token context window, an order of magnitude beyond rivals — long context becomes a $/token battleground.
A 1M-token window means a full codebase or hundreds of PDFs fit in one call, but Google splits pricing at the 128K mark—for prompts up to 128,000 tokens in size, the price is $1.25 per 1 million tokens, going up to $2.50 per 1 million tokens for prompts longer than 128,000 tokens, so pushing a request past that thre...
-
OpenAI cuts GPT-3.5 Turbo 50%
Third GPT-3.5 Turbo cut in a year: input drops 50% to $0.50/1M, output 25% to $1.50. New embedding models arrive 5x cheaper too.
At $0.50/1M input and $1.50/1M output, GPT-3.5 Turbo now undercuts the per-token cost of running Claude 2.0 or 2.1, since GPT-3.5 Turbo input prices have now been reduced by 50% to $0.0005 per 1k tokens and output prices are reduced by 25% to $0.0015/1K tokens, putting it at a cost per 1k tokens less than Anthropic'...
2023
-
Mixtral 8x7B goes open-weight
Mistral releases Mixtral 8x7B under Apache 2.0 — GPT-3.5-class quality anyone can host, setting a price floor under proprietary small models.
Because Mixtral 8x7B ships as open weights under Apache 2.0 and matches or outperforms GPT3.5 on most standard benchmarks, anyone can self-host GPT-3.5-class inference instead of paying OpenAI's rate — and third-party providers moved fast on this, with Fireworks pricing it at $0.4/million for prompt and $1.6/million...
-
GPT-4 Turbo: 3× cheaper
GPT-4 Turbo launches at $10/$30 with 128K context, cutting the frontier price by two-thirds and kicking off a year of rapid cuts.
A workload that previously required chunking an 8K-context prompt into multiple GPT-4 calls now fits in a single 128K call, while the per-token bill itself drops: When GPT-4 launched in March 2023, it cost $30 per million input tokens and $60 per million output tokens, versus GPT-4 Turbo pricing starting at $10.00 p...
-
GPT-4 sets the frontier price
GPT-4 launches at $30/$60 per MTok — 10× the cost of GPT-3.5 Turbo, establishing the first frontier pricing baseline.
Running a chat product on GPT-4 at launch meant paying $30/$60 per MTok versus GPT-3.5 Turbo's $0.002 per 1K tokens, so any workload swapped over to the frontier model saw its per-token bill jump roughly 15x on input alone — a gap large enough that most teams had to route only the hardest prompts to GPT-4 and keep e...
No more events.