.avif)
The AI price collapse has only just begun
The economics of AI are about to flip. If your business depends on selling access to AI compute, read this carefully.
For the past three years, the dominant story in AI has been capability: which model is smarter, which benchmark it tops, which company is winning the arms race.
That framing is about to become irrelevant.
The real story in 2026 is economics, and the numbers are brutal for anyone who built a business on the assumption that AI compute would stay expensive.
Token costs are falling off a cliff
The price of AI inference has collapsed faster than almost any technology in history. GPT-4-level intelligence costs roughly $30 per million tokens in 2023. By early 2025, that figure was under $1.50. Today, equivalent capability is available for fractions of a cent. Epoch AI's data shows inference costs declining at a median rate of 50x per year, accelerating to 200x per year after January 2024. Andreessen Horowitz coined the term "LLMflation" to describe the phenomenon, noting that the decline is outpacing even PC compute and dotcom-era bandwidth.
Furthermore, in August 2026, Sam Altman announced further price cuts to OpenAI's flagship models: an 80% price drop for GPT-5.6 Luna and a 20% price drop for GPT-5.6 Terra. GPT-5.6 Sol also gained a new fast mode in the API, running up to 2.5x faster at 2x the price without a change in intelligence.
Gartner has now put a formal stake in the ground: inference costs for a 1-trillion-parameter model will fall more than 90% by 2030 compared to 2025 levels. That's not a long-range forecast. It's four years away, and the trajectory is already well underway.
This is a structural shift driven by hardware, software, and competition simultaneously compressing margins at every layer of the stack.
Anthropic and OpenAI are fast approaching peak revenue
OpenAI and Anthropic have achieved something remarkable: they have convinced the market that frontier AI capabilities justify premium API pricing. And for a window of time, that was true. But that window is closing.
As open-weight models close the capability gap and next-generation silicon makes inference dramatically cheaper, many users will no longer need frontier model capabilities.
Revenue growth driven by rising token volumes is already being offset by falling per-token prices. Enterprise AI token costs fell 67% year-over-year from Q1 2025 to Q1 2026. If you are a compute provider charging on token volume, that is a significant headwind. Both companies have high-margin API businesses today. Whether those margins hold through 2027 is a much harder question.
That said, "peak token revenue" does not necessarily mean "peak revenue."
The more likely trajectory for OpenAI and Anthropic is not decline but a shift in where they capture value. As the token layer commoditizes, both companies have a clear incentive to move up the stack: from selling raw model access to selling full applications, agents, and vertical products built on top of their own models, capturing the same economics that VARs currently extract from them. Consumer products, coding agents, enterprise search, and industry-specific tools are all plays in this direction.
The pure token business may shrink as a share of revenue even as total company revenue keeps growing, just from a different layer of the stack.
Open source has crossed the threshold
The capability gap between open-weight and proprietary frontier models has effectively closed for the vast majority of real-world use cases. Open-weight models captured 38% of enterprise token volume in Q1 2026, up from just 11% a year earlier. By mid-2026, open-weight models route approximately half of all production inference tokens.
DeepSeek was the turning point. When a Chinese lab released a model matching frontier-level performance at a fraction of the training cost, and made the weights freely available, it shattered the assumption that leading capability required paying premium prices for the newest closed models. DeepSeek-V3 benchmarks match GPT-5.1-class performance at roughly one-tenth the inference cost.
The leading open-weight models today (e.g., Meta's Llama 4, Mistral Large, Qwen, DeepSeek) are not catching-up models. In coding, reasoning, and agentic tasks, they are competitive with or superior to closed alternatives.
For the vast majority of business use cases, open-weight models are at or above the performance threshold that matters. You do not need the world's smartest model to answer a customer support question. You need one that is good enough, fast, and cheap. Open source now wins that trade-off decisively.
You do not need the world's smartest model to answer a customer support question. You need one that is good enough, fast, and cheap. Open source now wins that trade-off decisively.
Cost isn't the only factor pushing enterprises toward open weights. A separate and growing concern among enterprise buyers is more structural: every query sent to a closed frontier model is, in principle, a potential input into that same vendor's next model, which raises questions for companies wary of their proprietary workflows, prompts, or data effectively training a competitor's product roadmap.
Vendors publish data-use policies that address this to varying degrees, and enterprise contracts often include opt-outs, but the perception itself is proving to be a real headwind for closed models, regardless of the specifics of any one vendor's terms. Open-weight alternatives don't carry the same overhang: the weights are already public, so there is no proprietary model on the other end that customer data could be feeding.
The enterprise migration is already happening, and the names on the list are not small companies experimenting at the margins. Lindy has moved to DeepSeek V4. Cursor to Kimi K2.5. Coinbase to GLM-5.2 and Kimi 2.7. Shopify and Airbnb to Qwen. Uber Eats to Qwen2. Even Microsoft is testing DeepSeek V4.
One of the last meaningful points of resistance to faster adoption of open-weight formats isn't technical or economic but geopolitical.
A meaningful share of the strongest open-weight models, including DeepSeek and Qwen, originate from Chinese labs, and that has been enough to make some risk and security teams hesitate, particularly in regulated industries and government-adjacent sectors. But this looks more like a temporary friction point than a durable barrier.
Enterprises can, and increasingly do, run these open weights entirely within their own infrastructure or region, so no data has to leave their control, regardless of where the model was originally trained. And Western labs, like Meta, Mistral, and others, are shipping open-weight models that are directly competitive on capability, giving cautious buyers a path to the same economics without the geopolitical question at all. Expect this objection to matter less with each passing quarter, not more.
Forbes called it plainly in April 2026: open-source AI has moved "from sideshow to strategy."
Next-generation silicon will accelerate this further
The hardware layer is about to apply another round of deflationary pressure that most enterprise AI budgets have not yet priced in.
Companies like Cerebras, SambaNova, and Groq are building inference hardware specifically optimized for large-model serving, rather than adapting GPU architectures designed for a different era. Cerebras' wafer-scale chips deliver inference at speeds that were unimaginable on traditional GPU clusters a year ago. SambaNova's RDU architecture achieves dramatically better efficiency on large transformer workloads. Groq's LPU delivers deterministic, ultra-low-latency inference at a fraction of the per-token energy cost.
These represent architectural rethinks of how inference is done. As this silicon reaches scale and competition drives down hardware margins, the cost per token will fall again, regardless of any software-side efficiency gains already baked into the forecast. The 90% cost reduction Gartner projects is likely conservative once next-generation inference silicon is fully in production.
A day of reckoning for AI resellers
The question every AI product company needs to answer honestly is: what is the part of my product that isn't a commodity? Is it proprietary data? Workflow integration that would be genuinely painful to replicate? A network effect? Domain-specific fine-tuning? A customer relationship with real switching costs?
If the honest answer is "we wrap a model with a nice interface," the business model has an expiry date. Customers will eventually either go direct to the underlying model provider, switch to a cheaper open-weight alternative, or find a competitor who built something harder to replicate.
The companies that will thrive are those that treat AI compute as a commodity input, like cloud storage or bandwidth, and build durable value on top of it. The companies that will struggle are those whose margin is primarily a toll on token throughput.
The planning window is now, not later. The mistake most companies in this position will make is treating the 90% cost decline as a future event to react to when it arrives. It isn't. It's already priced into the trajectory, and the businesses that wait until margins compress to start diversifying revenue will be doing so from a position of weakness.
The practical implication is straightforward: any company whose revenue currently depends on the spread between what it pays for compute and what it charges customers needs a second revenue engine that has nothing to do with token throughput. That could mean owning proprietary data no one else has access to, building workflow lock-in that survives a customer switching the underlying model, or pricing on outcomes rather than usage.
Whatever the answer, it needs to be identified, funded, and growing before the margin compression shows up in the P&L. The same logic applies to investors: portfolio companies should be asked now, explicitly, what percentage of their revenue is compute-price-dependent, and what the plan is to replace it. Waiting for the 2027 or 2028 numbers to reveal the problem is waiting too long.
A note to investors: The multiples don't reflect this yet
If, as the evidence suggests, token costs are going to fall 90%, then investors need to more closely examine the AI companies they are valuing today.
There is a second assumption baked into frontier lab valuations that deserves scrutiny: that AI is a winner-take-all market.
The multiples commanded by OpenAI, Anthropic, and others implicitly reflect the kind of structural dominance we saw with Meta in social, Google in search, or Microsoft in enterprise software. These were markets where network effects, data moats, and switching costs compounded into near-unassailable positions. That logic justified extraordinary valuations because the winner's economics only improved over time.
But AI is demonstrably not playing out that way.
There is no single dominant model, nor is there a meaningful switching cost between foundation model providers. Enterprises are already routing workloads across multiple models simultaneously, selecting by price and performance on a task-by-task basis. DeepSeek, Qwen, Mistral, and Llama are genuine competitors, not distant challengers, and new capable models arrive every few months.
The market structure looks far more like a commodity with differentiated tiers than a winner-take-all platform. And network effects simply are not there in the way they were for social.
A significant portion of the current market capitalization of AI software businesses is implicitly pricing in durable revenue streams tied to AI compute consumption. Those streams are about to shrink. Because revenue per unit of AI work will fall dramatically, and many companies have not built enough non-compute value to compensate.
For companies whose revenue is substantially derived from AI token throughput, investors should be asking: What does this business look like when compute is 90% cheaper? If the honest answer is "smaller," then the multiple must reflect that.
This is not a niche concern. A large number of AI companies that attracted capital in 2023 and 2024 did so on the basis of revenue growth that was partly a function of high compute prices. As those prices fall, revenue will face structural pressure even as usage grows.
Investors pricing AI companies on revenue multiples without adjusting for compute deflation are, in effect, capitalizing a headwind as if it were a tailwind. For many companies in this category, significant value should be removed from multiples to reflect the probability that a meaningful share of today's revenue is compute-price-dependent and will not survive the next price cycle intact.
Conclusion
The cost collapse in AI is not a threat to the technology. In fact, it represents a massive expansion of what becomes economically viable to build. Inference that costs one-tenth as much means use cases that were previously marginal become obvious.
But it does require an honest reckoning for anyone operating in the AI value chain.
The winners in this next phase will not be the companies with the most expensive models. They will be the companies that accumulated real, non-commoditisable assets during the expensive era, and are ready to deploy them at scale in the cheap one.
The price collapse is not coming. It is already here. The question is whether your business model and your portfolio are ready for it.


