← back to blog

ai, engineering, thoughts

the cost of intelligence is not going down

everyone's favorite chart is the most expensive lie in tech.

fake series a pitch deck slide titled why now: a chart of cost per token going down and to the right, bullets promising 10x cheaper inference every year, commodity prices by 2027, margins improving automatically, TAM everyone. footnote: a16z 2024, numbers not re-checked since

every model launch, same post. the cost of intelligence is collapsing, down and to the right, 10x a year, look at the chart. sam altman wrote it like a law of physics: the cost to use a given level of ai falls about 10x every 12 months. a16z gave it a name, llmflation. linkedin turned it into a carousel. every ai pitch deck with negative margins has it on the slide right before "path to profitability."

one problem. it's a lie.

not the chart. the chart is real. the lie is what everyone decides it means: that the intelligence you're actually buying is getting cheaper. it isn't. it's getting more expensive, and the receipts have been public for two years.

the chart they post
$ per 1M output tokens, gpt-4-level intelligence. log scale.
$60gpt-4$15gpt-4o$1.5mini tier$0.50flash tier
down 120x. real. also not what anyone ships.
the chart nobody posts
cost per task, artificial analysis intelligence index, max effort.
sonnet 4.6sonnet 5+91%
$1.20
$2.29
opus 4.8opus 5+13%
$1.80
$2.03
$2.75
fable 5. the new ceiling, same suite.
same benchmark, newer generation, higher cost per finished task. every time. data: artificialanalysis.ai, june and july 2026.

gpt-4 is free now and nobody wants it

here's the true part, because it's genuinely true. a fixed level of intelligence gets dirt cheap, fast. what needed gpt-4 at $60 per million tokens in 2023, a flash model does today for fifty cents. stanford measured the gpt-3.5 tier getting 280x cheaper in two years. extraction, classification, reading the boring stuff, i route that kind of work to flash models every day and the prices feel like rounding errors. that part of the chart is real.

and for some companies it stays true. if llms unlocked your task years ago and the task hasn't moved since, if it needs a fixed level of intelligence and not the growing kind, then congratulations, your costs genuinely collapse every year. classify the ticket, extract the fields, tag the photo. same job in 2027 as in 2024, cheaper every quarter. if that's your startup, the chart is yours. enjoy it.

the vast majority of ai startups are not that startup. their entire pitch is the opposite: the product gets better as the models get better. that's not a deflation story. that's a promise to live on the frontier forever, and the frontier is the one place the chart doesn't apply.

because the second a smarter model exists, everyone stops trusting the old one. you've felt it yourself. the new model drops and yesterday's model instantly feels like asking the intern when the senior engineer is sitting right there. the old model didn't get worse. the standards moved.

fake linkedin post from Thought Leader, founder and ai visionary: inference costs have dropped 100x in three years. if your ai startup is not profitable yet, it is not the models. it is you. agree? repost to help a founder in your network. 4,271 reactions

so demand migrates to the frontier the week it ships. every product runs on the best available model because the competition does. ethan ding called this a year ago, nobody wants yesterday's newspaper. the 10x chart tracks the price of intelligence nobody buys for anything that matters.

$1.25 was eleven months ago

so what does the frontier cost? gpt-5 launched in august 2025 at $1.25 in, $10 out. gpt-5.6 sol, july 2026: $5 in, $30 out. that's 4x on input and 3x on output, eleven months apart, same company, while the deflation chart was getting reposted daily.

anthropic walked opus down to $5/$25 and then shipped fable 5 above it at $10/$50. the ceiling bounced right back up. google tripled gemini flash prices in may. and there's a whole premium shelf now that didn't exist in 2023. o1-pro was $600 per million output. gpt-5.5 pro sits at $180.

what the best model costs, at launch
$ per 1M output tokens, official list price. log scale.
$10$25$60$150$6002023202420252026gpt-4 $60turbo $304o $15o1 $60o3 $40o3 cut $8gpt-5 $105.6 sol $30opus 3 $75opus 4.5 $25fable 5 $50opus 5 $25o1-pro $600gpt-4.5 $150gpt-5 pro $1205.5 pro $180
openai flagshipanthropic flagshipthe premium shelf (did not exist in 2023)
a sawtooth, not a collapse. gpt-5 at $10 lasted eleven months before the frontier moved to $30. fable 5 put anthropic's ceiling back at $50. swipe sideways on a phone. data: openai and anthropic pricing pages, artificialanalysis.ai.

my favorite detail in all of this: the a16z article that coined "llmflation," the one everyone cites for the 10x number, admits halfway down that o1 cost exactly what gpt-3 cost per output token at launch. sixty dollars. 2020's price on 2025's model. the deflation article debunks itself and nobody read past the chart.

yes, there were real cuts. o3 dropped 80% in a day. opus 4.5 was a genuine cut. it's a sawtooth, not a straight line up. but the teeth keep ending higher, and per-token price is the smallest number in this story anyway. the real damage comes from the two multipliers nobody puts on a slide.

agents bill by the hour now

multiplier one: the size of the job. metr measures the longest task the best model can finish on its own. october 2024: a thirty minute task. today: twelve hours for the released stuff, seventeen for the previews. and nobody buys a twelve hour model to do thirty minute jobs. the second agents could carry bigger work, everyone started handing them bigger work. the task didn't get slower. the ambition got bigger.

how long a task the best model can finish on its own
metr 50%-success time horizon. log scale. dashed line = doubling every 7 months.
30m1h2h4h8h16h20252026claude 3.5 sonnet31 min3.7 sonnet55 minopus 4.54.9 hgpt-5.25.9 hopus 4.612 hmythos preview17 h
thirty minutes to seventeen hours in eighteen months. every one of those hours is tokens. swipe sideways on a phone. data: metr.org time horizons, updated may 2026.

and agents don't pay per token the way a chat does. every step re-sends the whole conversation, so a session's cost grows quadratically with its length. i wrote about that in february when it was a curiosity about one laptop. it's not a curiosity anymore, it's the unit economics of every agent product on the market. an agent fixing one real github issue reads one to eight MILLION tokens. one issue. multiply by every seat at every company that put "agentic" in the deck.

the job you hand the model grew 24x in eighteen months, at flat-to-higher token prices. that's the multiplication the deflation chart has never met.

paying for tokens you're not allowed to read

multiplier two is dumber. reasoning models bill their thinking as output tokens. nobody sees those tokens. nobody reads them. everybody pays for them. reasoning models burn around 18x the tokens of non-reasoning models on the same work, and one xhigh call can think 20,000 tokens before it says a single word. sixty cents of private thoughts per api call.

fake api usage statement: tokens you read $12.40, tokens it thought about privately $98.60, steps you did not ask for $61.20, re-reading everything it already read $174.75, total: you do not want to know

and every frontier model now ships with an effort dial the pricing page doesn't mention. same model, same rate card, around 8x the token burn depending on how hard it thinks. artificial analysis measured exactly that spread on opus 5 across its five effort settings. the pricing page has two numbers. the model has five gears.

sonnet 5, one coding task, five gears
drag the effort. watch the bill.
rate card: $3 / $15 per 1M
does not move. ever.
lowmediumhighxhighmaxeffort
$7.43per task
48.2% of tasks solved
opus 4.8 one gear down: $3.44 for 48.7%. cheaper AND better.
$2.19$26.40
the bill scales with steps, not with tasks. at max effort sonnet 5 grinds 260 steps and reads 72 million tokens for one task. data: deepswe budget-matched runs at standard $3/$15 rates, july 2026.

drag it yourself. the rate card never moves. the bill moves 12x.

the cheap model is the expensive model

this is the one that should end the argument forever. artificial analysis runs every model through the same benchmark suite and publishes what the run cost. sonnet 5, the cheap model, two dollars per million input on promo pricing: the run cost $4,010. opus 5, the expensive model, five dollars per million: $3,835.

THE CHEAP MODEL COSTS MORE TO RUN.

cost to run the artificial analysis intelligence index
full suite, max effort. rows sorted by per-token price, cheapest first.
sonnet 5 (max)$2 / $10 promoscores 53
$4,010.12
cheapest per token. most expensive to run.
opus 4.8 (max)$5 / $25scores 56
$3,752.55
opus 5 (max)$5 / $25scores 61
$3,835.51
gpt-5.6 sol (max)$5 / $30scores 59
$3,442.81
priciest per token. cheapest to run.
the rate card and the bill are not the same axis. sonnet's run is on promotional pricing and it still costs the most. data: artificialanalysis.ai model pages, retrieved july 31, 2026.

because per-token price is the price of a brick, not the price of the house. sonnet grinds through 260 steps where opus takes 116. at max effort it reads 72 million tokens per task where opus reads 17. same rate card as sonnet 4.6, double the cost per task. one german outlet called it "hiding price increases behind unchanged token rates," which is the politest available way of saying the pricing page is a decoy.

two buttons, sweating guy: use sonnet, it is cheaper / use opus, it is cheaper

opus at high effort beats sonnet at max effort on cost AND quality. "which model is cheaper" stopped being a question that means anything. the only number that's real is cost per finished task, and that number has gone up every generation for two years straight.

cursor apologized for doing math in public

now connect it to the people who believed the chart.

cursor built a flat $20 plan on the assumption that model costs fall. then agents got long and users got hungry, and the ceo had to write an actual apology containing the sentence "new models can spend more tokens per request on longer-horizon tasks." that's this entire post in corporate.

and the subscriptions everyone actually uses, the subsidized $20 plans where someone else eats the token bill? chatgpt grew a $200 tier. claude grew a $100 tier, then a $200 tier, then weekly caps on top of the tiers, and someone is now suing anthropic over what "20x usage" was supposed to mean. anthropic's own announcement said people on $200 plans were burning tens of thousands of dollars of inference. openai floated agent tiers up to twenty thousand a month. the subsidy keeps shrinking because the thing being subsidized keeps getting hungrier.

the startups have it worse. ai apps run 20 to 60% gross margins because 40 to 80% of revenue passes straight through to the model providers. enterprise ai spend tripled to $37 billion last year. every one of those companies has the 10x-cheaper slide somewhere in an old deck.

and the decks, man. i keep hearing pitches with the same slide: token costs drop 10x a year, so the margins fix themselves, so the burn is temporary. if the product runs on frontier intelligence, that slide is fiction. the input those companies actually buy has gotten more expensive every generation for two years. a founder pointing at the deflation chart is pointing at the price of the one model they would never ship.

nobody's costs went down. NOBODY'S.

gru plan: ship an ai product / tokens get 10x cheaper every year / agents use 100x more tokens / agents use 100x more tokens
2023 · $20
one month of gpt-4
the best model on earth, all you could chat
2026 · $20
0.86 of one task
a single max-effort sonnet 5 run, median $23.28. you don't get to finish it.
what $20 buys. the flat subscription never stood a chance. data: deepswe median cost per task at max effort, july 2026.

jevons would like a word

the lie survives because both halves are true separately. tokens got cheaper. spend went up. people quote the first half, build companies on it, then act surprised by the second half.

jevons figured this out in 1865 with coal: make the engine more efficient and the world burns more coal, not less. ai is the most jevons thing ever built. the models got cheaper per token, better per task, and longer per run, so everyone runs them more, harder, and deeper. useful things don't get cheaper in total. they get bought more.

none of this is a complaint. the models keep earning the spend, that's the entire reason the spend keeps growing. but if a plan assumes inference costs fall 10x next year, that plan assumes running last year's model against competitors who won't be. budget for the bill going up. it goes up because the thing works.

cheapest tokens in history. biggest bills the industry has ever paid. both true.

only one of them makes the slide.