Why AI Tokens Could Become Significantly Cheaper If the Bubble Bursts — Even Though Providers Are Already Selling “Below Cost”

The apparent paradox is straightforward: AI companies like OpenAI are widely reported to be selling inference (token generation) at a loss, subsidized by investor capital. Yet many analysts argue that a bursting AI investment bubble would ultimately drive token prices lower over time. How can both be true?
The resolution lies in distinguishing between accounting costs (heavily influenced by inflated hardware prices and aggressive depreciation) and true economic marginal costs (what it actually costs to run one more query once the hardware exists). Current “losses” are real in a financial reporting sense, but they are inflated by monopoly-like pricing power upstream. When that power erodes, the floor for sustainable pricing drops sharply.
Nvidia’s Dominant Position and Massive Markups
NVIDIA currently enjoys extraordinary pricing power in AI accelerators. Independent analyses estimate the manufacturing cost of an H100 GPU at roughly $3,300 (including logic die, high-bandwidth memory/HBM, advanced packaging like CoWoS, and assembly/test).
Yet these GPUs sell for $25,000–$40,000 each (depending on configuration and volume). This implies a gross margin in the range of 80–88% or higher on the chip itself for NVIDIA.
This is not a normal competitive market dynamic. NVIDIA’s near-monopoly position in high-end AI GPUs (combined with supply constraints, especially around HBM memory) has allowed it to capture the vast majority of the value created by the AI boom. The “true” cost of the silicon and packaging is a small fraction of the sticker price that downstream providers pay (or implicitly pay through cloud contracts).
Inference Providers Are Burning Cash — But Relative to What?
OpenAI and similar frontier labs report heavy losses. Recent leaked financials for 2025 show revenue around $13 billion against costs and expenses of roughly $34 billion, resulting in an operating loss of about $21 billion (with even larger net loss figures depending on accounting treatments).
This aligns with the pattern of spending more than $1 per dollar of revenue in recent periods. Much of this burn goes toward compute — training new models and serving inference at scale — often through long-term cloud contracts or direct hardware purchases at premium prices.
However, a large portion of the “cost” that providers are failing to cover is the depreciation or rental cost of hardware purchased or contracted at NVIDIA’s inflated prices. When providers say they are operating below cost, they are typically including these high capitalized or contracted expenses. Electricity and other variable operating costs are real but represent a much smaller share of the total.
Marginal Cost vs. Average Cost: The Key Distinction
Once a GPU cluster is built and paid for (or under a sunk-cost contract), the marginal cost of running additional inference is dominated by:
- Electricity (and cooling/PUE overhead).
- Minor incremental maintenance and networking.
Estimates put electricity at roughly $0.20–$0.30 per hour per H100-class GPU under typical data center conditions — a tiny fraction of the $2–$4+ per hour effective rental or depreciation rates seen in the market.
In a post-bubble scenario, several forces align to push effective costs (and thus competitive prices) downward:
1. NVIDIA’s premium compresses. Increased competition from AMD, custom ASICs from hyperscalers (Google TPU, Amazon Trainium/Inferentia, Microsoft Maia, etc.), Chinese alternatives, and potentially new entrants would erode NVIDIA’s ability to maintain 8–10x markups. Gross margins on chips would fall toward more normal semiconductor levels.
2. Existing hardware gets run harder at marginal cost. Billions of dollars worth of GPUs are already deployed or on order. In a downturn or oversupply situation, owners (or distressed sellers) will run this hardware to recover any contribution margin above pure variable costs (mainly power). This is classic economics: price toward marginal cost when there is excess capacity and fixed costs are sunk.
3. Distress and secondary markets emerge. Used or surplus GPUs would trade at steep discounts to original purchase prices, lowering the effective capital cost for new or expanding providers. Cloud spot and secondary-market pricing would reflect this.
4. Efficiency gains continue independently of the bubble. Algorithmic improvements, better quantization, mixture-of-experts architectures, speculative decoding, and hardware-software co-optimization (e.g., Blackwell and beyond) keep reducing tokens per watt and tokens per dollar. These trends predate and will outlast any investment cycle.
What “Below Cost” Really Means Today
The statement that providers are selling tokens below cost is accurate under current accounting — they are not yet covering their full loaded costs (including high hardware acquisition prices and heavy R&D/training spend). However, a substantial slice of that reported cost structure traces back to the upstream monopoly rent captured by NVIDIA.
Remove or sharply reduce that 8–10x chip markup from the equation, and the gap between revenue and true variable + sustainable fixed costs narrows dramatically. Providers would still need to cover ongoing power, operations, model improvement, and a reasonable return on new capital — but the bar for profitability would be much lower than today’s inflated baseline.
In short: today’s pricing subsidizes both market-share battles and NVIDIA’s extraordinary margins. A bubble burst would likely discipline the latter while increasing competitive pressure on the former.
Likely Outcomes for Token Pricing
- Short term (during unwind): Possible price volatility or increases in some premium segments if funding dries up suddenly and weaker players exit. However, the installed base of hardware creates downward pressure as utilization becomes the priority.
- Medium to long term: Significantly lower sustainable token prices. Inference becomes more commoditized. Smaller models, specialized/open-source offerings, and on-prem/edge deployments gain ground. The “true” cost floor — dominated by electricity plus efficient hardware at competitive (not monopoly) prices — becomes visible in the market.
- Winners and losers: Hyperscalers and efficient operators with owned or low-cost infrastructure benefit. Pure-play inference startups may consolidate. NVIDIA’s margins and valuation multiple would likely compress unless it maintains technological leadership. End users and developers see cheaper, more abundant AI capabilities.
This dynamic is not unique to AI. It mirrors past technology cycles (e.g., memory chips, solar panels, or even early cloud computing) where high upstream margins and subsidized downstream adoption eventually gave way to commoditization and lower prices once supply caught up and competition intensified.
The current model of “sell inference below full cost to capture the future” works only as long as capital remains cheap and abundant and hardware pricing power remains concentrated. When either assumption breaks — particularly the hardware premium — the economics of tokens shift decisively toward lower prices driven by real marginal costs rather than inflated averages.
The bubble, if it bursts, would expose and then remove a major artificial cost layer from the stack. That is why many expect token prices to fall further, not rise, once the dust settles.
---
Also read:
- The Anatomy of the AI Bubble: Circular Money, Fragile Foundations, and What Comes Next
- Survey Shows 97.7% of Hiring Leaders Changed Talent Locations Due to AI
- How to Appeal YouTube Demonetization from Inauthentic Content Flags
- HeyGen Open-Sources HyperFrames Framework with Keyframes for AI Video Agents
- Hermes Agent v0.18 Brings Stable MoA and New Control Commands
---
Thank you!
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.