Inference prices fell 9x a year. Also 900x.
The most cited number in AI economics is a single rate. The study behind it reports a range spanning two orders of magnitude, and a caveat that undercuts its own fastest figures.
TL;DR. The price of running a model has collapsed, and this is real. GPT-4 launched in March 2023 at $30 per million input tokens and $60 per million output; equivalent capability is now available for well under a dollar. The figure everyone quotes is roughly 10x per year. The careful study behind that framing reports something different. Epoch AI measured price declines against six benchmarks and found rates ranging from 9x to 900x per year depending on which performance milestone you pick, with GPT-4-level performance on PhD-level science questions falling 40x per year. And the same study notes two things that sit awkwardly together: the fastest declines all begin after January 2024, and benchmark contamination seems to have become more common since 2024. Epoch says both. The decline is real; its rate is a choice of milestone.
---
Status: established, with the source's own caveats carried. Primary source: Epoch AI, LLM inference prices have fallen rapidly but unequally across tasks, 12 March 2025. API prices are from published vendor pricing. This article does not dispute that prices fell steeply. It disputes that a single rate describes it.
---
The collapse is real
Start with the part nobody contests.
GPT-4's API launched in March 2023 at $30 per million input tokens and $60 per million output tokens. By 2026, models matching or exceeding that capability on the benchmarks that existed at launch are available at a fraction of it, and economy-tier models sit at around $0.10 per million input tokens.
The drivers are well understood and independent of one another. Smaller models reaching the performance of older larger ones. Quantisation, from 16-bit to 4-bit arithmetic. Serving optimisations including speculative decoding, continuous batching and paged attention. Hardware generations. And open-weight competition compressing margins.
Cloud GPU rental prices fell too, with H100 hourly rates stabilising around $2.85 to $3.50 after declines of 64 to 75% from their peaks.
This is one of the steepest cost curves in the history of any technology, and the point of this article is not to deny it.
The number that gets quoted
"Roughly 10x per year" is the figure in general circulation, popularised as LLMflation and attributed to a 2024 analysis tracking a fixed quality bar.
It is a reasonable summary. It is also a single number standing in for something that is not a single number.
What the study actually found
Epoch AI took six benchmarks, identified the price required to reach specific performance milestones on each, and tracked how that price moved over three years.
The rates ranged from 9x to 900x per year.
Two orders of magnitude, within one study, using one method.
Specific results: the price to achieve GPT-4's performance on GPQA Diamond, a set of PhD-level science questions, fell 40x per year. For a later model's performance on the same benchmark, the study measured 200x per year for evaluation cost against 400x per year for price per token, which is the largest divergence it found between those two ways of counting.
So the answer to "how fast did AI get cheaper" is: pick a capability, pick a benchmark, pick whether you mean price per token or cost to run an evaluation, and the answer moves by a factor of a hundred.
Any single rate is a choice about all three, and the choices are almost never stated alongside the number.
The caveat Epoch put in its own study
This is the part worth sitting with.
Epoch notes that the fastest trends, up to 900x per year, all begin after January 2024. It flags this itself and says it is therefore less clear those rates persist.
And in its limitations, Epoch notes that AI developers may train on benchmark data accidentally, or use knowledge of a benchmark to inform training, and that this seems to have become more common since 2024.
Both statements are in the same analysis. The fastest measured price declines start after January 2024. Benchmark contamination became more common after 2024.
Epoch does not claim these are connected, and neither does this article. What can be said is that the measure used to define "equivalent capability" became less reliable over exactly the window in which the measured declines became fastest, and a price-per-capability figure is only as solid as the capability measurement underneath it.
That is construct validity applied to a cost curve, and it is the reason the 900x figure deserves less confidence than the 9x one, quite apart from which is more impressive.
Why the framing matters
Two consequences follow, in opposite directions.
It understates the case for building. A team that priced a product against 2023 tokens is working with an assumption that has moved by orders of magnitude, and use cases that were uneconomic then may be trivially economic now. The correct response to the collapse is to re-run the arithmetic, not to quote the rate.
And it overstates the case for extrapolation. A rate derived from the last eighteen months, on benchmarks whose contamination increased over the same period, projected forward, is three assumptions stacked. Analysis suggesting the pace moderates to 3-5x annually through 2027 and then 1.5-2x may or may not be right, and is at least reasoning about a rate rather than assuming one.
Neither reading survives quoting a single number without its basis, which is the same problem as with revenue figures and for the same reason.
And it does not reduce total spend
Worth connecting explicitly, because the deflation story is often used to answer the consumption story.
Falling cost per token has not reduced total expenditure or total energy. One analysis of the physics puts it directly: efficiency gains of this kind expand the token budget rather than shrinking the bill, and do not imply reduced total energy consumption.
That is the same pattern as electricity: per-unit cost falls, usage grows faster, total rises. A 1,000x price decline over three years coexisting with hundreds of billions in capital expenditure is not a contradiction. It is what makes the expenditure rational.
Three things this establishes
A rate quoted without its milestone is uninterpretable. 9x and 900x came from the same study on the same day. Which capability, which benchmark, and price-per-token or cost-per-evaluation are all required for the number to mean anything.
Benchmark-denominated cost curves inherit benchmark problems. If "GPT-4-equivalent" is defined by benchmark scores, and benchmark contamination rose, then the denominator drifted while the numerator was being measured precisely.
And deflation and rising total spend are consistent. Anyone using the price collapse to argue that AI's resource footprint is solved has substituted a per-unit measure for a total, which is the error the electricity article documented from the other direction.
What it does not establish
That prices did not fall steeply. They did, by any measure, across every method, on every benchmark. This is a dispute about the rate, not the direction.
That Epoch's analysis is flawed. It is unusually careful: it states its method, publishes its range rather than a headline, and volunteers the contamination caveat that weakens its own most striking numbers. The problem is what happens to the finding downstream, not the finding.
That contamination explains the fast declines. No causal claim is made and none is available. The two facts are adjacent and neither Epoch nor this article connects them.
And nothing about future rates. Extrapolation is exactly what the range makes unsafe.
What is unresolved
Whether contaminated benchmarks materially distorted the measured curve. Answering it requires uncontaminated held-out evaluations across the same period, which do not exist retrospectively.
Which milestone is the right one. A rate for frontier capability and a rate for commodity capability are different economic facts, and there is no principled basis for treating either as "the" rate.
How much of the decline is margin compression rather than cost reduction. Competitive pricing and genuine efficiency both lower prices, and only the second is durable. Published prices cannot distinguish them.
And whether the pace holds. The low-hanging optimisation has been taken, and reasoning models consume far more tokens per task, which pushes cost per useful output in the other direction.
The counter-argument
Insisting on a range where a rate would do is unhelpful. Practitioners need a planning number. "Somewhere between 9x and 900x" is not usable, and 10x per year has been a serviceable approximation that got most people to roughly the right decisions.
The contamination point may be overstated. Epoch flags it as a general limitation of benchmark-based measurement, not as a specific defect in this analysis, and treating a standard limitations paragraph as undermining the headline reads more into it than the authors did.
The range is partly definitional rather than a finding. Of course the price to reach an easy milestone falls faster than the price to reach a hard one, because commodity models reach easy milestones and only frontier models reach hard ones. The 9x to 900x spread may be describing benchmark difficulty rather than anything about cost.
And margin compression is a real cost reduction to a buyer. Distinguishing durable efficiency from competitive pricing matters for forecasting and not for a team deciding whether to build something this quarter.
The short version
GPT-4 launched in March 2023 at $30 per million input tokens and $60 per million output. Equivalent capability now costs a small fraction of that, economy models sit near $0.10 per million, and H100 rental fell 64 to 75% from peak. The collapse is real and is one of the steepest cost curves in any technology.
The quoted rate is about 10x per year. The study behind that framing reports 9x to 900x per year, depending on which performance milestone is chosen, with GPT-4-level performance on PhD science questions falling 40x per year and one later milestone measured at 200x for evaluation cost against 400x for price per token.
Two orders of magnitude, one study, one method. A single rate is a choice of capability, benchmark and counting basis, and those choices almost never travel with the number.
Epoch also notes, in the same analysis, that the fastest declines all begin after January 2024, and separately that benchmark contamination seems to have become more common since 2024. It draws no connection and neither does this. What follows is narrower: a price-per-capability figure is only as solid as the capability measurement underneath it, and that measurement became less reliable over the window where the curve steepened.
And falling per-token cost has not reduced total spend or total energy, which is the same shape as the electricity finding from the other side. Per-unit deflation alongside rising totals is not a contradiction. It is the reason the spending is rational.
Common questions
How much have inference prices actually fallen? Steeply, by every measure. GPT-4's API launched in March 2023 at $30 per million input tokens and $60 per million output tokens; models matching or exceeding that capability on the benchmarks that existed at launch are now available for a small fraction of that, with economy-tier models around $0.10 per million input tokens. Cloud H100 rental prices also fell 64 to 75% from their peaks.
Is the decline really 10x per year? That is one summary of it. Epoch AI's analysis, which measured the price to reach specific performance milestones across six benchmarks over three years, found rates ranging from 9x to 900x per year depending on the milestone chosen. GPT-4's performance on PhD-level science questions fell 40x per year. For one later milestone the study measured 200x per year for evaluation cost against 400x per year for price per token. Any single rate is a choice about which capability, which benchmark, and which counting basis.
Why does the range span two orders of magnitude? Partly because easy milestones and hard milestones behave differently: commodity models reach easy ones and only frontier models reach hard ones, so the price to reach an easy milestone falls faster. That is the strongest objection to treating the range as a finding rather than a definitional artefact. It also means no rate is "the" rate without saying which capability level it describes.
What is the contamination issue? Epoch notes in its limitations that developers may train on benchmark data accidentally, or use knowledge of a benchmark to inform training, and that this seems to have become more common since 2024. Separately, it notes that the fastest measured declines all begin after January 2024. Epoch draws no connection between these and neither does this article. The narrow point is that a price-per-capability figure depends on the capability measurement, and that measurement became less reliable over the same window in which the curve steepened.
Does the price collapse mean AI's energy footprint is solved? No, and using it that way substitutes a per-unit measure for a total. Falling cost per token has not reduced total expenditure or total energy: efficiency gains expand the token budget rather than shrinking the bill. A thousandfold price decline coexisting with hundreds of billions in capital expenditure is not a contradiction; it is what makes the expenditure rational.
How much of the decline is real efficiency and how much is competition? Both are present and published prices cannot separate them. Competitive pricing pressure, notably from open-weight and low-cost entrants, lowers prices without lowering costs, while quantisation, smaller models, serving optimisations and new hardware lower costs genuinely. Only the second is durable, which matters for forecasting and not for a team deciding whether something is affordable this quarter.
Will the pace continue? Unknown, and the range is exactly what makes extrapolation unsafe. Analysis suggesting moderation to 3-5x annually through 2027 and then 1.5-2x is at least reasoning about the rate rather than assuming it. Two pressures point the other way: the easiest optimisations have been taken, and reasoning models consume far more tokens per task, which raises cost per useful output even as cost per token falls.
What should a practitioner take from this? Re-run the arithmetic rather than quoting the rate. A product priced against 2023 token costs is working from an assumption that has moved by orders of magnitude, and use cases that were uneconomic then may be trivially economic now. That is the actionable consequence, and it does not require knowing whether the true rate was 9x or 900x.
Sources
Primary documents only. Where a claim rests on a single report, the entry says so.
- LLM inference prices have fallen rapidly but unequally across tasks Epoch AI Data Insights, 12 March 2025 The 9x to 900x range across six benchmarks, the 40x per year figure for GPT-4-level GPQA Diamond performance, the note that the fastest trends begin after January 2024, and the limitations paragraph on benchmark contamination becoming more common since 2024.
- Published API pricing, March 2023 to 2026 Vendor pricing pages The $30 and $60 per million token launch pricing for GPT-4 and subsequent economy-tier pricing. Vendor-published and subject to change, which is why this article gives ratios rather than treating current prices as fixed.
Further reading
The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.
- Raji et al. (2021), AI and the Everything in the Whole Wide World Benchmark — why general benchmarks cannot carry general claims. :: https://arxiv.org/abs/2111.15366 Construct Validity
- Bowman & Dahl (2021), What Will it Take to Fix Benchmarking in Natural Language Understanding? — what a benchmark must satisfy to support inference. :: https://arxiv.org/abs/2104.02145 Construct Validity
Related articles
- Same company, same year, revenue figures 63% apartThe most quoted numbers in AI are run-rates from private companies, reported by outlets triangulating from leaks, and sources covering the same period differ by multiples.
- The water bottle was per 10 to 50 responsesThe most repeated environmental claim about AI dropped a qualifier from the study it cites. The concern it created is still justified, for different reasons than the number suggested.
- Three years in: no disruption, and one 20% holeEvery aggregate measure of AI's labour effect shows continuity. One within-firm comparison shows a fifth of a cohort gone. Both are well evidenced.
- The 13% traveled. The authors' caveat did not.A careful study found entry-level employment falling in AI-exposed jobs. Its own authors later narrowed when that becomes significant, and a serious alternative explanation predicts the same pattern.