Start with the number that reorganizes everything else. In the first half of 2026, the price of frontier-grade artificial intelligence — the good stuff, the models that top the independent leaderboards — fell by roughly five times against where the leading edge sat twelve months earlier. Not five percent. Five times. A unit of top-tier machine reasoning that cost a dollar last summer costs something closer to twenty cents today, and the curve is still pointing down. There is no precedent for this speed in the history of computing. Moore's Law, the benchmark everyone reaches for, roughly halved the price of a transistor every two years. Intelligence is repricing that fast every few months.
We spent a week across three desks — compute, markets, and the frontier labs — trying to answer a question that sounds simple and isn't: what actually happens to an industry, and to everyone standing near it, when its core product gets an order of magnitude cheaper while getting better at the same time? The short version is that a price collapse this violent doesn't just lower a bill. It rewrites who has power. It is, quietly, the most important story of the year — bigger than any single model release, because it's the force underneath all of them.
The numbers, laid flat
Take the current board at face value. OpenAI's GPT-5.6 Sol tops the independent aggregate at a composite strength around 83, and lists near $5 per million input tokens and $25 per million output — priced, deliberately, at roughly half of Anthropic's Fable 5, which still leads the very hardest agentic-coding evals but now charges $10 and $50 for the privilege. A year ago, a model of Sol's caliber would have been the single most expensive thing on the menu. Today it's the *value* play among the flagships. That inversion — where the smartest broadly-available model is also one of the cheaper ones — did not exist in 2025.
Below the flagships, the floor fell out entirely. SpaceXAI's Grok 4.5 ships competitive-but-not-leading intelligence at $2 / $6 — and, per independent accounting, burns a fraction of the tokens per task that the premium models do, so its *effective* cost gap is wider than the sticker. China's open-weight tier went further still: GLM-class models deliver frontier-adjacent strength at something like $0.60 / $2, which is how Chinese models clawed 30-plus percent of US enterprise trials this half. When a model that scores in the mid-70s costs a fifteenth of one that scores 80, the marginal buyer stops asking 'which is smartest' and starts asking 'which is enough.'
The marginal buyer stopped asking which model is smartest. They started asking which one is enough — and 'enough' just got very, very cheap.
Why it collapsed — three forces at once
Price collapses this steep usually have one cause. This one has three, stacked. The first is competition reaching the top: for the first time, four or five labs can all credibly claim a seat at the frontier, and when products commoditize, price is the only lever left. The second is efficiency compounding — better training recipes, mixture-of-experts routing that lights up only the parameters a task needs, and inference tricks that mean each answer costs the lab less silicon-time to serve. A model that's cheaper to *run* can be cheaper to *sell* without bleeding margin. The third, and least understood outside the compute desk, is the open-weight undertow: once a genuinely good model can be downloaded and self-hosted, every closed lab has to price against 'free-if-you-run-it-yourself,' and that gravity pulls the whole market down.
None of this would matter if the demand side were soft. It is the opposite of soft. Hyperscaler capital spending is running up roughly 80% year over year, a $650 billion build-out aimed squarely at making inference cheaper per unit at ever-larger scale. That is the paradox at the heart of the repricing: the industry is spending the GDP of a mid-sized country precisely so it can charge you less. The capex is not a bet that intelligence will be scarce and dear. It's a bet that it will be abundant and cheap, and that whoever owns the cheapest abundant supply wins the next decade.
Who it breaks
Cheap is not free of victims. The most exposed party is the one that looks strongest: the frontier lab whose entire business was *selling tokens of intelligence at a premium*. When your flagship's price halves every few months, you are running up a down escalator — you have to ship a materially better model on a brutal cadence just to hold the same revenue per customer, and each better model costs more to train than the last. That math is why, this half, every major lab quietly pivoted the same direction on the same timeline: into deployment and services — sending human engineers into enterprises to make the models actually land — because the margin was visibly draining out of the model itself and had to be recaptured somewhere. Microsoft put $2.5 billion behind it; OpenAI and Anthropic stood up rival ventures within weeks. When the people who make a thing all rush to sell *help using the thing* instead, that tells you where the profit went.
The second casualty is subtler: the thin startup whose only product was a wrapper around a model's price arbitrage. If your business was 'we buy intelligence wholesale and sell it retail with a nice interface,' the wholesale price just fell through your retail price and kept going. The moat was never the model; it was the margin, and the margin evaporated. What survives at the application layer now is what always survives a commodity price war — proprietary data, real distribution, a workflow so embedded that switching costs more than the software saves.
Who it makes
For almost everyone who *uses* AI rather than sells it, the repricing is the best thing that has ever happened to them, and most of them don't know it yet. A capability that was a boardroom line-item in 2025 is now something a solo founder can put in a product without a second thought. Whole categories that were 'too expensive to run at scale' — always-on agents, per-user models, exhaustive rather than sampled analysis — cross the line from prototype to production the moment the cost of a token stops mattering. The historical rhyme is exact: bandwidth, storage, and cloud compute each got cheap enough to become *plumbing*, and the fortunes made after that point were not made selling the plumbing. They were made by whoever built the most valuable thing *on top* of it, once it was cheap enough to stop thinking about.
This is the part the market keeps mispricing. The value doesn't disappear when intelligence gets cheap — it *migrates*, from the layer that makes the model to the layer that deploys it into something a real business or person needs. The repricing is not the end of the AI trade. It's the starting gun for the second, larger one: the application era, where 'enough' intelligence at near-zero cost gets woven into everything, and the winners are decided by product and distribution, not benchmark scores.
The demand paradox: cheaper doesn't mean less spent
Here is the counterintuitive engine that makes the $650 billion build-out rational instead of insane. When a resource people want gets dramatically cheaper, they do not spend less on it — they spend more, because the lower price unlocks a flood of uses that were never viable before. Economists have a name for it, the Jevons paradox, and it was first observed in coal: more efficient steam engines were supposed to reduce coal consumption, and instead they detonated it, because cheap steam power suddenly made sense for a thousand things it hadn't before. Cheap intelligence is coal-and-steam all over again. Drop the price of a token far enough and you don't get a smaller AI bill — you get always-on agents where you used to run a single query, exhaustive analysis where you used to sample, a model call inside every loop of every product instead of once at the end. Total consumption explodes faster than unit price falls.
That is why the labs and hyperscalers can look at a collapsing per-token price and rationally pour a nation's worth of capital into serving even more of it. They are not betting the price stops falling; they are betting that demand, unleashed by the falling price, grows faster than the price drops — so the total market swells even as each unit gets cheaper. It is the same bet the fiber layers made and lost on timing, and the cloud providers made and won. Whether this cycle's builders have the balance sheet to outlast the gap between spending now and earning later is the multi-hundred-billion-dollar question sitting under every earnings call this year. The paradox guarantees the demand is coming. It guarantees nothing about who is still standing when it arrives.
The geopolitics hiding inside the price tag
There is a second story folded into the first, and it is the one that outlasts any single product cycle. The cheapest capable models on the board are, increasingly, Chinese and open-weight — GLM-class systems delivering frontier-adjacent strength at a fifteenth of the premium price. That is not a coincidence of engineering; it is a strategy. When you cannot reliably win the race for the single smartest model — because the best chips are export-controlled and the training runs that need them are throttled — the rational move is to win the race for the cheapest *good-enough* one, give it away or nearly so, and let the world build on your foundation instead of a rival's. Abundance becomes a weapon precisely when scarcity is being used against you.
So the collapse is not only an economic event; it is a geopolitical one. Every enterprise that adopts a cheap open model because the math is irresistible is also, quietly, making an infrastructure choice with a flag attached to it. The United States spent two years trying to keep the frontier scarce and controlled; the counter-move was to make the near-frontier abundant and nearly free. Whether that works is the open question of the decade — but the mechanism is already visible in the enterprise numbers, where the cheapest option keeps winning trials it has no business winning on raw capability alone. Cheap is not neutral. Cheap has a direction, and right now it points away from the labs that assumed the frontier would stay theirs to price.
What the pattern actually rhymes with
To see where this goes, look at what happened the last three times a foundational input got suddenly, violently cheap. When long-distance bandwidth collapsed after the fiber glut of the early 2000s, most of the companies that laid the fiber went bankrupt — and the companies that built *on top of* the now-nearly-free bandwidth became some of the most valuable enterprises in history. When cloud compute crossed from a capital expense to a metered utility, the winners weren't the people who owned the data centers; they were the people who used them to build products that were unthinkable when servers cost real money. When electricity got cheap and reliable, the dynamo-makers were a footnote inside a decade, and the fortune was in every factory, appliance, and lit-up city that cheap power made possible. The rhyme is exact enough to be uncomfortable: the input becomes plumbing, the plumbing-makers' margins compress, and the value migrates — permanently — one layer up.
The uncomfortable corollary for the labs is that being the smartest maker of the input has never, in any prior cycle, been where the durable money ended up. The frontier labs are, right now, the most celebrated companies in technology — and if the pattern holds, they are also standing in the exact spot the last three cycles' value *left* from. That is not a forecast of their doom; several will thrive, especially the ones already sprinting into deployment and owning real distribution. It's a forecast about where to look for the next decade's winners, and the pattern's answer is blunt: not at the layer that makes intelligence, but at the layer that does something irreplaceable with it now that it's almost free. The labs know this — it's why, to a company, they spent this half pivoting toward services. The question is whether the pivot outruns the compression.
The one thing that could stop it
Every clean narrative deserves its counter-force, and this one has a real one, sitting in an unglamorous place: memory. The binding constraint on serving cheap intelligence at scale is no longer the GPUs themselves; it's the high-bandwidth memory stacked beside them. HBM is the reason a memory-maker's IPO drew a frenzy this half, and the reason a genuine supply shortage now sets the ceiling on how fast the whole industry can grow its cheap-inference supply. If memory stays scarce, the price of intelligence could stop falling — not because anyone wants it to, but because the physical supply of the thing that serves it cheaply can't keep pace with demand. Watch HBM output the way you'd watch an interest rate. It is the most important number nobody outside the compute desk is watching, and it's the single variable that could put a floor under a market that has, so far, refused to find one.
The steelman: why the labs might be fine anyway
We owe the other side its strongest form, because the 'labs are doomed' story is too clean to trust. The bull case for the frontier labs is not that the price stops falling — it's that price was never the whole product. Three things could keep the model-makers rich even as tokens approach free. The first is that the frontier keeps moving: if genuinely new capabilities — long-horizon autonomy, agents that don't need babysitting, whatever comes after — keep arriving only at the top, then there is always a premium tier worth paying for, and the cheap models are forever chasing last year's frontier, never this year's. The second is integration: a lab that owns the model, the tooling, the deployment arm, and the enterprise relationship can capture value across the whole stack even if any single layer commoditizes. The third is trust as a moat — for the workloads where a wrong answer is expensive, 'cheapest that's good enough' quietly becomes 'the one we can actually stake the business on,' and that is a different, stickier purchase.
None of this is guaranteed, and the bear case has the momentum right now. But a reader who walks away certain the labs are finished has learned the wrong lesson. The right one is narrower and harder: the labs' *default* business — selling raw tokens at a premium — is dying, and their survival depends on outrunning that death into something the price collapse can't reach. Some will. The ones who mistake this half's celebration for safety won't.
The honest verdict
Here is what a week across three desks convinced us of. The AI story of 2025 was capability — who could build the smartest model. The AI story of 2026 is price — and price is the more consequential story, because capability changes what's *possible* while price changes what's *done*. Frontier intelligence becoming abundant and cheap is a bigger deal than any one model becoming smart, in the same way that electricity becoming cheap mattered more than any single better dynamo. The labs feel this in their margins, which is why they're all pivoting to services. The market feels it in its confusion about who's actually going to make the money. And the rest of the world will feel it, soon, as the thing that was expensive and rationed a year ago quietly becomes the thing that's just *there*, in everything, too cheap to meter. The repricing is the ballgame. Everything else this year is a footnote to it.
- Frontier AI got ~5x cheaper in a year — the fastest price collapse in computing history.
- Three causes: real competition at the top, compounding efficiency, relentless demand.
- $650B of capex is a Jevons bet: cheaper intelligence means more spent, not less.
- Token-premium sellers and wrapper startups break; AI users win; cheap Chinese open-weights surge.
- Caveat: scarce high-bandwidth memory could stall the whole collapse.
