Unit economics
Inference is COGS, and COGS is what decides whether you are a software company.
A traditional SaaS business books almost nothing against revenue. An AI-native business books a variable, usage-linked cost that grows with every successful user. That single structural difference is why the AI cohort trades at a discount, and it is the thing an inference provider can actually move. This page sets out how far, and where it does not help at all.
01The problem
The margin gap is real and it is measured
This is not a thesis, it is a set of published survey results. Software companies building AI products report gross margins in the forties and fifties. The band traditional SaaS uses as its own comparator is the eighties. The gap is closing, but it is closing slowly and from a long way back.
Reported AI product gross margin against the traditional SaaS reference band
AI products is ICONIQ's all-respondent series across roughly 300 executives at software companies building AI products. App layer is the application-only subset of the same survey, which runs 7 to 8 points below it. SaaS band is the traditional-SaaS reference figure ICONIQ uses as its own comparator, held flat for reference.
- AI products
- App layer
- SaaS band
| Period | AI products | App layer | SaaS band |
|---|---|---|---|
| 2024 | 41% | 33% | 80% |
| 2025 | 45% | 38% | 80% |
| 2026 | 53% | 45% | 80% |
| 2027 | 59% | 80% |
The 2026 and 2027 points are ICONIQ projections, not actuals, and the January edition projected 52 percent for 2026 where the July edition says 53. ICONIQ attributes the expansion to model routing, falling token prices and scale, not to pricing power.
ICONIQ, State of AI Bi-Annual Snapshot (Jan 2026) and State of AI 2026: The Builder's Economy (July 2026)
The cohort average hides a 35-point spread
Bessemer splits its AI companies into two cohorts and the gap between them is wider than the gap between AI and SaaS overall. Any single industry-average AI gross margin number is close to meaningless next to this.
Bessemer's own description is 'often negative'. Explosively scaling companies at roughly $40M year-one ARR and $125M year-two ARR, with $1.13M ARR per employee and retention Bessemer calls fragile. Bessemer's framing: they trade distribution for profit in the short term.
Bessemer Venture Partners, The State of AI 2025 (Aug 2025)reported
Fast-growing, capital-efficient, strong product-market fit. Roughly $3M year-one to $103M year-four ARR, $164K ARR per employee. The 35-point gap between the two Bessemer cohorts is why a single industry-average AI gross margin number is close to meaningless.
Bessemer Venture Partners, The State of AI 2025 (Aug 2025)reported
What individual companies have reported
Read the confidence column before reading the numbers. Roughly half of these are third-party bottom-up models rather than company disclosures, and several of those models omit forward-deployed engineering, implementation and compliance costs that are material in enterprise deployments. We have marked them rather than laundering them into the same table as audited figures.
| Company | Year | Gross margin |
|---|---|---|
Cursor (Anysphere)reportedAI-native | 2026 | -23% |
Perplexity (as reported)reportedAI-native | 2024 | 60% |
Perplexity (free-tier compute booked as COGS)reportedAI-native | 2024 | -64% |
Lovableest.AI-native | 2025 | 35% |
StackBlitzest.AI-native | 2025 | 40% |
Replitest.AI-native | 2025 | 23% |
Harveyest.AI-native | 2026 | 73% |
Sierraest.AI-native | 2026 | 60% |
Decagonest.AI-native | 2026 | 68% |
Abridgeest.AI-native | 2026 | 85% |
OpenAIreportedModel provider | 2024 | 40% |
OpenAIreportedModel provider | 2025 | 33% |
AnthropicreportedModel provider | 2024 | -94% |
AnthropicreportedModel provider | 2025 | 40% |
Rows marked est. are third-party bottom-up models, not company disclosures. Harvey, Sierra, Decagon and Abridge have not published gross margins; the figures shown are external reconstructions from assumed pricing and assumed token cost, and at least two of them omit forward-deployed engineering, implementation and compliance costs. Lovable, StackBlitz and Replit rest on a single weak secondary source. Companies whose margin could not be sourced at all, including Clay, are absent rather than estimated. Negative figures are not data errors.
The same company, two reported margins, 124% apart
This figure is an accounting artifact and should never be shown without its counterpart below. The Information found Perplexity classified about $33M of roughly $57M in compute and web costs, the portion serving non-paying users, as R&D rather than cost of revenue.
On a raw basis 2024 compute was 164 percent of revenue, which implies roughly -64 percent gross margin. Identical operations, two reported margins 124 points apart. Perplexity is the clearest demonstration that AI gross margin is partly an accounting policy choice, and that any benchmark comparison which does not control for where free-tier inference is booked is noise.
The lesson is not that anyone lied. It is that gross margin in this cohort is partly an accounting policy choice, and any benchmark comparison that does not control for where free-tier inference is booked is noise. Ask where a competitor books it before you compare yourself to them.
02Four shapes
Token volume is a property of the product, not the model
These four archetypes cover most of what production AI software looks like. Each one is modelled from a published token shape, and every volume below is reproducible from the assumptions attached to it. The single most important empirical point: real workloads run at input to output ratios of 40:1 to 187:1, not the 3:1 used in most cross-vendor price comparisons.
Monthly inference cost by archetype, frontier tier against frontier tier
Both columns price the identical token volume, cache hit rate and cache write share. Only the rate card changes.
- Claude Opus 4.8
- Seldon Prime 1
| Category | Workload | Claude Opus 4.8 | Seldon Prime 1 |
|---|---|---|---|
| Coding agent | 419.3B in / 2.2B out | $390K | $106K |
| Support agent | 12.1B in / 150M out | $31K | $8.3K |
| Document / RAG | 50.7B in / 1.3B out | $90K | $21K |
| Classification | 80.0B in / 1.5B out | $171K | $48K |
Cache writes are billed at 1.25x base input on the comparator, per its published 5-minute TTL, and at standard input rates on the Seldon side. Ignoring that premium would understate the coding agent's comparator bill by roughly 7 percent, so it is applied.
Token shapes from research/econ-production-app-cost.md. Comparator rates from vendor rate cards, July 2026. Seldon rates are our published list prices.
AI coding agent
1,000 developer seatsA terminal or IDE agent that reads a repository, plans, edits files and runs tools across long autonomous sessions. Sold per developer seat.
- Input tokens / month
- 419.3B
- Output tokens / month
- 2.2B
- Input to output
- 187:1
- Cache hit rate
- 94.8%
- Monthly revenue
- $200,000
- Non-inference COGS
- $16,000
- Inference on Claude Opus 4.8
- $389,586
- Inference on Seldon Prime 1
- $106,343
- Gross margin, before and after
- -102.8% to 38.8%
Assumptions behind this archetype
- Scale unit is 1,000 developer seats, not monthly active users. Seats are how agentic coding is sold and how Anthropic reports its own consumption guidance.
- Session token shape is taken verbatim from Anthropic's Claude Code /usage documentation example: 1,200 fresh input, 940,000 cache read, 50,000 cache write, 5,300 output per session. That example reconciles arithmetically to its stated $0.55 on Sonnet 4.6 rates, which is why it is trusted as a statement of token shape.
- 23.5 session-equivalents per active day across 18 active days per month gives 423 sessions per seat per month.
- Per seat per month: 0.51 M fresh input, 397.6 M cache read, 21.2 M cache write, 2.24 M output. At 1,000 seats that is 419.278 B total input and 2.242 B output.
- Input to output ratio is 187:1. Cache hit rate is 94.8 percent, consistent with the University of Washington TraceLab study of roughly 4,300 real Claude Code and Codex sessions, which measured a 95.7 percent aggregate prefix cache hit rate over 55 B tokens.
- Volume cross-check: the model implies 23.4 M tokens per developer per active day. Anthropic's separately published per-user rate limit guidance implies roughly 24 M. Two unrelated first-party figures agree within 3 percent.
- Cost cross-check: priced on Sonnet 4.6 this model produces $234 per seat per month, inside Anthropic's published $150 to $250 per developer per month band and close to its $215 typical enterprise starting point.
- Revenue is MODELLED at $200 per seat per month. It is not a disclosed price for any named company. It is anchored to Anthropic's published enterprise consumption starting point of $215 per developer per month for Claude Code, which is evidence of demonstrated enterprise willingness to pay at that level.
- Non-inference COGS is modelled at 8 percent of revenue, covering hosting, eval engineering, observability and support. SFAI Labs puts the new AI-specific COGS lines (eval team 1 to 2.5 percent, observability about 1.1 percent, customer success uplift 0.8 to 1.5 percent) at roughly 3 to 5 points of revenue on top of conventional hosting and support.
- At $200 per seat this archetype has an Inference Efficiency Ratio of about 1.3:1 on frontier pricing, far below Ben Murray's 3:1 warning threshold for AI-native products. That is the structural reason coding agents are the hardest margin case in AI application software.
Customer support agent
100,000 conversations per monthA retrieval-grounded support agent that resolves customer conversations end to end, priced per resolved outcome rather than per seat.
- Input tokens / month
- 12.1B
- Output tokens / month
- 150M
- Input to output
- 80:1
- Cache hit rate
- 66.5%
- Monthly revenue
- $99,000
- Non-inference COGS
- $14,850
- Inference on Claude Opus 4.8
- $30,656
- Inference on Seldon Prime 1
- $8,280
- Gross margin, before and after
- 54.0% to 76.6%
Assumptions behind this archetype
- Scale unit is 100,000 conversations per month, roughly a one million user product at a 10 percent monthly contact rate.
- Single-turn token shape generalises Fireworks' published chatbot workload archetype of 5,000 input / 500 output / 1,500 cached.
- Six turns per conversation. A 5,000-token cached prefix holds system prompt, tool definitions and policy. Each turn adds about 3,000 tokens of retrieved knowledge base chunks (roughly four chunks of 750 tokens), a 100-token user message and a 250-token reply. Each increment is written to cache and re-read on subsequent turns.
- Per conversation: 18,600 fresh input, 80,250 cache read, 21,750 cache write, 1,500 output. At 100,000 conversations that is 12.060 B total input and 0.150 B output.
- Input to output ratio is 80:1. Cache hit rate is 66.5 percent, materially lower than the coding agent because the retrieved chunk set changes every turn.
- The six-turn structure is ESTIMATED. It is the author's construction generalising a verified single-turn shape. Turn count is the most sensitive free parameter: cost is roughly linear in turns above turn two, because cache reads grow with the conversation prefix.
- Revenue uses Intercom's published Fin price of $0.99 per resolution, verified across every plan tier, modelled at one billable resolution per conversation. A real deflection rate below 100 percent scales revenue down proportionally and is the dominant lever on this archetype's P&L.
- Non-inference COGS is modelled at 15 percent of revenue, the highest of the four. Enterprise support deployments carry implementation, knowledge base curation and customer success load that a token calculator does not see.
- The strategic read from the research: the spread between the most and least expensive model here is 26x, but in absolute terms only $0.118 per conversation. Against a $0.99 price point every credible model clears 87 percent gross margin on inference alone. Model choice barely moves this P&L; resolution rate does.
Document and RAG analysis product
10,000 monthly active users processing 200,000 documents per monthA product that ingests long documents, summarises them, then answers follow-up questions against a cached full-document context.
- Input tokens / month
- 50.7B
- Output tokens / month
- 1.3B
- Input to output
- 40:1
- Cache hit rate
- 88.3%
- Monthly revenue
- $300,000
- Non-inference COGS
- $30,000
- Inference on Claude Opus 4.8
- $90,500
- Inference on Seldon Prime 1
- $21,176
- Gross margin, before and after
- 59.8% to 82.9%
Assumptions behind this archetype
- Scale unit is 10,000 monthly active users at 20 documents each, giving 200,000 documents per month, 40 pages per document.
- Tokens per page is taken as 650, roughly 500 words at 1.3 tokens per word. This is an industry rule of thumb and is ESTIMATED. No primary source was located for it.
- Per document: one full ingest pass writes 26,000 document tokens plus a 2,000-token system prefix to cache and produces a 1,500-token summary. Eight follow-up questions each re-read the cached 28,000-token prefix with a 200-token question and a 600-token answer.
- Per document: 1,600 fresh input, 224,000 cache read, 28,000 cache write, 6,300 output. At 200,000 documents that is 50.720 B total input and 1.260 B output.
- Input to output ratio is 40:1. Cache hit rate is 88.3 percent.
- The structure is ESTIMATED. 650 tokens per page, 20 documents per user, 40 pages per document and 8 questions per document are all modelling assumptions. The token SHAPE (large cached prefix, tiny fresh input, modest output) is the durable part; absolute volume scales linearly in those four parameters and should be user-editable in any calculator.
- Embeddings are a rounding error and the research says so plainly. Indexing 200,000 documents at 26,000 tokens each is 5.20 B tokens per month, which at Together's $0.02 per million costs about $104 per month, or 0.3 percent of the frontier inference bill. Vector storage and re-embedding on document churn are the real RAG infrastructure costs, not embedding inference.
- Long-context pricing cliff: a 40-page document is about 26,000 tokens and sits under every model's short-context tier. Past OpenAI's 272,000-token threshold, gpt-5.6-sol input doubles and output rises 1.5x. Anthropic is the exception and bills the full 1M window at standard rates. A document product that lets users upload arbitrarily large files crosses that cliff silently.
- Revenue is MODELLED at $30 per monthly active user. It is not a disclosed price for any named company. Cross-check: at that price, frontier inference is 12.1 percent of revenue, inside SFAI Labs' 8 to 12 percent band for chat-shaped AI-native products and close to the roughly 11 percent implied by ICONIQ's own cost structure data.
- Non-inference COGS is modelled at 10 percent of revenue, covering vector storage, re-embedding on document churn, eval engineering and observability.
High-volume classification and extraction pipeline
50,000,000 items per monthA batch or streaming pipeline that classifies or extracts structured fields from a very large number of small items: transaction categorisation, content moderation, ticket routing, product tagging.
- Input tokens / month
- 80.0B
- Output tokens / month
- 1.5B
- Input to output
- 53:1
- Cache hit rate
- 74.3%
- Monthly revenue
- $250,000
- Non-inference COGS
- $12,500
- Inference on Claude Opus 4.8
- $170,950
- Inference on Seldon Prime 1
- $47,864
- Gross margin, before and after
- 26.6% to 75.9%
Assumptions behind this archetype
- Scale unit is 50,000,000 items per month.
- Per item: a 1,200-token cached prefix holding instructions, few-shot examples and the output schema; a 400-token payload; a 30-token structured output.
- At high queries per second the prefix is effectively always warm, so the per-item cache hit on the prefix is about 99 percent and cache writes amortise to near zero: 12 written tokens per item against 1,188 read.
- Per item: 400 fresh input, 1,188 cache read, 12 cache write, 30 output. At 50 M items that is 80.000 B total input and 1.500 B output.
- Input to output ratio is 53:1. Cache hit rate is 74.2 percent, which is lower than the coding agent despite the near-perfect prefix warmth, because the per-item payload is fresh and is a third of every request.
- The structure is ESTIMATED. It is the author's construction.
- GOTCHA, and it is a real one: Claude Haiku 4.5's minimum cacheable prefix is 4,096 tokens and this prefix is 1,200. On Haiku the entire prefix bills as fresh input on every call, costing about $87,500 per month rather than the $34,190 the cached model implies, a 2.6x difference. The Haiku figure quoted anywhere against this archetype assumes the prompt has been padded past 4,096 tokens, which is a real if slightly absurd optimisation. Sonnet 5 and Opus 4.8 have a 1,024-token floor and do not have this problem.
- This is the archetype where model choice dominates. The research measures a 13.7x spread between the most and least expensive credible option, and 27x once the 50 percent batch discount is applied. Unlike the support agent, that spread is real money.
- Revenue is MODELLED at $0.005 per item, or $5 per 1,000 items. It is not a disclosed price for any named company. Cross-check: that gives an Inference Efficiency Ratio of about 7.3:1 against frontier inference cost, inside Ben Murray's 5:1-and-above healthy band for AI-native products.
- Non-inference COGS is modelled at 5 percent of revenue, the lowest of the four. This is a machine-to-machine pipeline with negligible support and customer success load.
03Live model
Move the inputs and watch the margin
Same arithmetic as the sections above, same pure functions, exposed. Pick your shape, pick what you are currently paying, scale it to your volume. The assumptions behind every archetype are listed under the result rather than hidden behind the number.
A terminal or IDE agent that reads a repository, plans, edits files and runs tools across long autonomous sessions. Sold per developer seat.
Baseline is 1,000 developer seats. Token volume, revenue, and other COGS scale together.
Monthly volume
- Input tokens
- 419KM
- Output tokens
- 2.2KM
- Cache hit rate
- 94.8%
- Revenue
- $200,000
Claude Opus 4.8
Seldon Prime 1
Gross margin
75% = healthy SaaS
Margin improvement
+141.6%
percentage points of gross margin
Retained annually
$3,398,921
inference spend that stays in the business
Assumptions behind this model
- Scale unit is 1,000 developer seats, not monthly active users. Seats are how agentic coding is sold and how Anthropic reports its own consumption guidance.
- Session token shape is taken verbatim from Anthropic's Claude Code /usage documentation example: 1,200 fresh input, 940,000 cache read, 50,000 cache write, 5,300 output per session. That example reconciles arithmetically to its stated $0.55 on Sonnet 4.6 rates, which is why it is trusted as a statement of token shape.
- 23.5 session-equivalents per active day across 18 active days per month gives 423 sessions per seat per month.
- Per seat per month: 0.51 M fresh input, 397.6 M cache read, 21.2 M cache write, 2.24 M output. At 1,000 seats that is 419.278 B total input and 2.242 B output.
- Input to output ratio is 187:1. Cache hit rate is 94.8 percent, consistent with the University of Washington TraceLab study of roughly 4,300 real Claude Code and Codex sessions, which measured a 95.7 percent aggregate prefix cache hit rate over 55 B tokens.
- Volume cross-check: the model implies 23.4 M tokens per developer per active day. Anthropic's separately published per-user rate limit guidance implies roughly 24 M. Two unrelated first-party figures agree within 3 percent.
- Cost cross-check: priced on Sonnet 4.6 this model produces $234 per seat per month, inside Anthropic's published $150 to $250 per developer per month band and close to its $215 typical enterprise starting point.
- Revenue is MODELLED at $200 per seat per month. It is not a disclosed price for any named company. It is anchored to Anthropic's published enterprise consumption starting point of $215 per developer per month for Claude Code, which is evidence of demonstrated enterprise willingness to pay at that level.
- Non-inference COGS is modelled at 8 percent of revenue, covering hosting, eval engineering, observability and support. SFAI Labs puts the new AI-specific COGS lines (eval team 1 to 2.5 percent, observability about 1.1 percent, customer success uplift 0.8 to 1.5 percent) at roughly 3 to 5 points of revenue on top of conventional hosting and support.
- At $200 per seat this archetype has an Inference Efficiency Ratio of about 1.3:1 on frontier pricing, far below Ben Murray's 3:1 warning threshold for AI-native products. That is the structural reason coding agents are the hardest margin case in AI application software.
- Seldon rates are our published list prices. Comparator rates are published vendor list prices as of July 2026 and exclude negotiated enterprise discounts, which materially change the picture at large volume. If you already have a negotiated rate, use it here instead, and the honest answer may be that switching is not worth it.
The comparator rates in this calculator are published list prices. They exclude negotiated enterprise discounts and the 50 percent batch discount that three of the four closed vendors offer, both of which narrow the gap materially at volume. If you have a negotiated rate, use it instead of the list price and the honest answer may be that the switch is not worth making.
04The counterweight
When changing inference provider will not fix your margins
This is the section a vendor is not supposed to write. It is here because the model above produces a result that undercuts our own pitch in one of the four archetypes, and hiding it would make everything else on this page worth less.
The support agent’s inference bill falls 9.6x on the same switch that adds 35.9 points of gross margin to the coding agent. It adds 5.7. Both numbers are correct.
The reason is arithmetic rather than a defect in the product. Inference was 6.3% of revenue before the switch, so 6.3%was the entire prize, and no price on earth could have won more than that. Inference was never the binding constraint on that P&L. The binding constraints were the 15% of revenue spent on implementation, knowledge base curation and customer success, and the resolution rate that determines whether a conversation is billable at all. Neither of those is something we sell.
| Archetype | Inference as % of revenue | COGS reduction | Margin points gained |
|---|---|---|---|
AI coding agent1,000 developer seats | 40.1% | 9.5x | 35.9 |
Customer support agent100,000 conversations per month | 6.3% | 9.6x | 5.7 |
Document and RAG analysis product10,000 monthly active users processing 200,000 documents per month | 6.5% | 10.4x | 5.8 |
High-volume classification and extraction pipeline50,000,000 items per month | 14.3% | 9.1x | 12.7 |
The COGS reduction column barely moves across the four rows, because it is a ratio of two rate cards and the token shape mostly cancels. The margin column moves by a factor of six. The difference between those two columns is the entire point of this section: what a provider switch is worth to you is set by the second column, not the third, and the second column is a fact about your product rather than about any vendor. Seldon rates are our published list prices.
Do this division before you talk to any inference vendor
Take last month’s inference invoice and divide it by last month’s revenue. That percentage is the absolute ceiling on what any provider change can add to your gross margin, and you can compute it in ten seconds without a sales call. If the answer is under about five percent, the highest-value work in your business is somewhere else and you should go do that instead.
Four situations where the answer is that we cannot help you.
- 1
Inference is a small share of revenue
The ceiling argument above. A vertical SaaS product with AI in the workflow typically runs 3 to 8 percent of revenue on inference. Cutting that by an order of magnitude moves gross margin by less than the rounding on a good sales quarter.
- 2
The real COGS line is human
Forward-deployed engineering, implementation, knowledge base curation, eval labelling and human-in-the-loop review. The research is emphatic that human-in-the-loop is the most commonly mis-booked line: put in operating expense it makes the AI layer look profitable, booked honestly as cost of revenue it often does not. Token prices are irrelevant to that line.
- 3
The problem is where the cost is booked, not what it costs
Free-tier compute classified as research and development rather than cost of revenue can swing a reported gross margin by more than a hundred points, as the pair of figures in section 01 shows. If your margin problem is a free-tier subsidy, the fix is product and pricing, not procurement.
- 4
The binding constraint is product effectiveness
On outcome-priced products, revenue moves with the resolution or acceptance rate, not with the token bill. Doubling the share of conversations your agent actually resolves is worth more than a 26x model price spread, because that spread is only about eleven cents per conversation in absolute terms.
Inference as a share of revenue, by product shape
Find your shape, then read the ceiling. These bands are third-party estimates rather than audited figures and are marked accordingly.
| Product shape | Inference / revenue |
|---|---|
| Traditional SaaS, no meaningful AI | 0% to 2% |
| Vertical SaaS with AI core to the workflow | 3% to 8% |
| Bessemer vertical AI portfolio | 10% |
| Derived from ICONIQ 2026 using the correct denominator | 11% |
| AI-native, chat-shaped | 8% to 12% |
| AI-native, agentic or reasoning-heavy | 14% to 22% |
A widely repeated error is worth flagging here: ICONIQ reports inference at 20 to 23 percent of total AI product cost, and many secondary sources restate that as 23 percent of revenue. The denominator is wrong. Corrected against ICONIQ's own cost structure it lands near 11 percent of revenue, which is the last row.
Only the bottom two rows describe products where inference cost is the dominant lever. If you are in one of those, the rest of this site is worth your time. If you are not, we would rather you knew that now.
How the cost reduction works05Deflation
Waiting for prices to fall is a strategy with two known failure modes
Cost per unit of intelligence falls roughly 5x to 10x per year at benchmark level, per MIT FutureTech (arXiv:2511.23455), with Epoch AI measuring 9x to 900x per year depending on the capability threshold and the accounting basis.
Price of a fixed capability level: MMLU 42, the level GPT-3 reached in 2021
GPT-3 at MMLU 42 cost $60 per million tokens in November 2021. By November 2024 the cheapest model clearing MMLU 42 was Llama 3.2 3B on Together.ai at $0.06 per million tokens. That is a 1,000x decline in three years, or about 10x per year.
- Cheapest model at MMLU 42
| Period | Cheapest model at MMLU 42 |
|---|---|
| Nov 2021 | $60.00 |
| Nov 2024 | $0.06 |
Plotted separately and never merged, because the two series hold different capability thresholds constant and are not two samples of one curve. Only the published endpoints are shown; intermediate points were not invented. Note what the axis means: a model's own price does not fall, a cheaper model arrives that clears the same bar.
a16z, Welcome to LLMflation (Guido Appenzeller, Nov 2024). Price is the average of input and output; historical prices reconstructed from the Internet Archive.
Price of a fixed capability level: MMLU 64.8, GPT-3.5 level
A GPT-3.5-equivalent query cost $20.00 per million tokens in November 2022 and $0.07 per million tokens by October 2024, served by Gemini 1.5 Flash 8B. The AI Index states this as a 280x decline over roughly 18 months.
- Cheapest model at MMLU 64.8
| Period | Cheapest model at MMLU 64.8 |
|---|---|
| Nov 2022 | $20.00 |
| Oct 2024 | $0.07 |
Plotted separately and never merged, because the two series hold different capability thresholds constant and are not two samples of one curve. Only the published endpoints are shown; intermediate points were not invented. Note what the axis means: a model's own price does not fall, a cheaper model arrives that clears the same bar.
Stanford HAI, 2025 AI Index
Why the published rates disagree by two orders of magnitude
Published estimates of the decline run from 1.7x to 900x per year. That spread is not one estimate being wrong. It is three methodological choices, compounded by the finding that deflation is distributed very unevenly across performance tiers. The most careful study located measures 1.7x per year at the bottom of the market, where intelligence is already near its price floor, against 32x per year at the top.
- Benchmark-level versus token-level accounting. Reasoning models have lower per-token prices but emit far more tokens. Measuring dollars per token flatters the trend; measuring dollars per benchmark run does not. Epoch mitigates this by excluding reasoning models entirely, which is a different fix for the same problem.
- Window selection. MIT uses April 2024 to November 2025. a16z and Epoch include 2022 and 2023, the period when the fastest declines occurred. Early-stage price collapse is partly margin compression and land-grab pricing, not durable efficiency.
- Sample density. MIT includes roughly 10x more models per benchmark in each fit.
The counter-trends, which is the part usually left out
These are the reason cost control is an engineering discipline rather than something that arrives on its own. Deflationism as a permanent planning assumption is an error.
While per-token prices have generally declined, the cost of running frontier-level models has risen approximately exponentially, at about 3x to 18x per year.
Not a contradiction. If inference prices fall 10x but improving performance by one percent requires 100x more inference, the total cost of running the state of the art at that new level increases 10x. MIT also finds that for GPQA Diamond, roughly half of measured benchmark progress is attributable to increased inference spending rather than to price-independent technical advance.
MIT FutureTech, arXiv:2511.23455
Demand for inference is growing roughly 10x per year while supply grows roughly 3.4x per year, which mechanically implies rising prices for large-model access.
Corroborating evidence: SemiAnalysis H100 on-demand capacity showed as sold out in February, March and April 2026, and one-year contract prices rose from about $1.70 per GPU-hour in late 2025 to about $2.65 in June 2026, a 56 percent increase, even as spot prices fell. Stanford independently measures global AI compute capacity growing 3.3x per year since 2022. Deflationism as a permanent planning assumption is an error.
Epoch AI compute crunch analysis (May 2026); Stanford AI Index 2026
Falling token prices do not automatically become margin. Companies absorbed the roughly 100x inference cost decline of the preceding 18 months by moving to more complex, longer-running tasks and to newer, more expensive models.
Bessemer's observation is that margin expansion is happening at scale but more slowly than the pace of model cost improvements would theoretically allow. This is the single best argument that cost control is an engineering discipline rather than something that arrives on its own.
Bessemer Venture Partners (Talia Goldberg, Aug 2025)reported
Between the most and least expensive credible models as of 2026-07. In 2023 the equivalent spread was roughly one order of magnitude.
Derived from the July 2026 price table: Anthropic Fable 5 at $10.00 input and $50.00 output against DeepSeek V4-Flash on DeepInfra at $0.09 input and $0.18 output.est.
Dispersion is widening, not narrowing, and that is the practical consequence of everything above. When the market spread across credible models is 110x on input, the routing decision inside your own application is worth more than the provider decision outside it. We would rather sell you the cheapest tier that clears your quality bar than the most expensive one you can afford.
Pick a tierModel this against your own P&L.
Send a month of invoice shape: input and output token volume, cache hit rate, revenue and non-inference COGS. We will return the same arithmetic run on your numbers, with the assumptions written out, and we will say so plainly if inference is not your binding constraint.