← Research workspace

MARKET STRUCTURE
HORIZON: 12–36 MONTHS

The short-term inference glut.

AI demand can grow explosively and still disappoint the supply stack. Hardware additions, model efficiency, caching, and specialization compound faster than most organizations can redesign workflows around them.

The consensus debate is framed as “AI demand: high or low?” The more useful question is relative: does demand grow faster than effective supply?

Demand can rise tenfold and the answer can still be no. New accelerators add physical capacity while smaller models, sparse routing, quantization, caching, and better serving software multiply the useful work produced by each chip. Those gains stack. Adoption frictions do too, but in the opposite direction.

ILLUSTRATIVE MARKET MODELADW—01

Supply can outrun demand—even when demand grows 10×.

Effective supply DemandIndexed units · illustrative, not a forecast

Effective supply is more than chips.

Installed accelerator hours are only the first layer. The same fleet produces more accepted work when inference engines batch requests better, models activate fewer parameters, prompts reuse cached prefixes, or task-tuned systems need fewer retries. A 2× hardware increase paired with a 3× efficiency gain behaves like 6× effective supply.

EFFECTIVE SUPPLYaccelerator capacity × utilization × model efficiency × task success

That is why “tokens served” can mislead. One cached input token, one fresh prefill token, and one generated reasoning token may all increment a counter while consuming radically different compute. Useful completed tasks are closer to the economic unit.

Demand has organizational latency.

Models can be updated overnight. Enterprises cannot redesign permissions, data access, review policy, procurement, incentives, and accountability at the same speed. Early adopters have already captured the easiest use cases; later adoption increasingly requires workflow change rather than another API key.

01

Integration

Valuable work sits behind systems, permissions, and process boundaries that a benchmark does not see.

02

Trust

Human review and liability controls can cap deployment even when model quality clears the technical bar.

03

Measurement

Teams often cannot price a successful task, making expanding usage difficult to justify internally.

04

Change

The scarce input may become management attention: deciding which workflows should be rebuilt and who owns the outcome.

Current demand strength does not invalidate the setup.

Google reported its first-party model APIs grew from 16B to approximately 22B tokens per minute in one quarter and said it remained supply constrained. That is strong evidence for near-term demand. It does not tell us whether industry-wide effective supply will grow even faster over the next several years. (Alphabet Q2 2026 remarks)

The Ox Alpha preview is the complementary signal: an anonymous provider was willing to advertise 100T tokens of daily capacity for free, and OpenCode’s observed usage shows 93% of input tokens were cached. Whatever the maker’s identity, the launch demonstrates how cheap, reusable context can be pushed into the market before a durable price is established. (OpenCode model data)

What would falsify the thesis?

Signals that demand is absorbing the curve
SignalWhy it matters
Stable or rising realized priceEfficiency gains are being captured as value rather than competed away.
High utilization across vintagesNew supply is not stranding older accelerators or forcing idle capacity.
Task demand outruns token efficiencyAgents create enough new workflows to absorb smaller, cheaper models.
Revenue scales with useful workToken growth converts into durable gross profit rather than promotional volume.

This is a market-structure scenario, not a forecast of a specific provider’s utilization or security price. The chart is indexed and intentionally illustrative.