The consensus debate is framed as “AI demand: high or low?” The more useful question is relative: does demand grow faster than effective supply?
Demand can rise tenfold and the answer can still be no. New accelerators add physical capacity while smaller models, sparse routing, quantization, caching, and better serving software multiply the useful work produced by each chip. Those gains stack. Adoption frictions do too, but in the opposite direction.
Supply can outrun demand—even when demand grows 10×.
Effective supply is more than chips.
Installed accelerator hours are only the first layer. The same fleet produces more accepted work when inference engines batch requests better, models activate fewer parameters, prompts reuse cached prefixes, or task-tuned systems need fewer retries. A 2× hardware increase paired with a 3× efficiency gain behaves like 6× effective supply.
That is why “tokens served” can mislead. One cached input token, one fresh prefill token, and one generated reasoning token may all increment a counter while consuming radically different compute. Useful completed tasks are closer to the economic unit.
Demand has organizational latency.
Models can be updated overnight. Enterprises cannot redesign permissions, data access, review policy, procurement, incentives, and accountability at the same speed. Early adopters have already captured the easiest use cases; later adoption increasingly requires workflow change rather than another API key.
Integration
Valuable work sits behind systems, permissions, and process boundaries that a benchmark does not see.
Trust
Human review and liability controls can cap deployment even when model quality clears the technical bar.
Measurement
Teams often cannot price a successful task, making expanding usage difficult to justify internally.
Change
The scarce input may become management attention: deciding which workflows should be rebuilt and who owns the outcome.
Current demand strength does not invalidate the setup.
Google reported its first-party model APIs grew from 16B to approximately 22B tokens per minute in one quarter and said it remained supply constrained. That is strong evidence for near-term demand. It does not tell us whether industry-wide effective supply will grow even faster over the next several years. (Alphabet Q2 2026 remarks)
The Ox Alpha preview is the complementary signal: an anonymous provider was willing to advertise 100T tokens of daily capacity for free, and OpenCode’s observed usage shows 93% of input tokens were cached. Whatever the maker’s identity, the launch demonstrates how cheap, reusable context can be pushed into the market before a durable price is established. (OpenCode model data)
What would falsify the thesis?
| Signal | Why it matters |
|---|---|
| Stable or rising realized price | Efficiency gains are being captured as value rather than competed away. |
| High utilization across vintages | New supply is not stranding older accelerators or forcing idle capacity. |
| Task demand outruns token efficiency | Agents create enough new workflows to absorb smaller, cheaper models. |
| Revenue scales with useful work | Token growth converts into durable gross profit rather than promotional volume. |
This is a market-structure scenario, not a forecast of a specific provider’s utilization or security price. The chart is indexed and intentionally illustrative.