Until recently, the AI-compute-as-a-service market had room for exactly one obvious challenger to the hyperscalers: CoreWeave. That model is now plural. Nebius — restructured out of the former Yandex group — is building data-center capacity at a pace that puts it firmly in the same conversation, and the public-market response suggests investors believe a second standalone AI cloud is viable.
Key takeaways
- The AI-cloud sub-sector is consolidating around two pure-play challengers plus the hyperscalers' GPU divisions.
- Nvidia's allocation strategy across these players is now the single most important variable for who scales fastest.
- Token pricing for inference is the metric to watch — it determines whether sustained margin emerges or whether the segment commoditizes.
What pure-play AI clouds actually sell
The label "AI cloud" obscures two very different businesses sold under one banner. The first is bare-metal or near-bare-metal GPU rental: the customer brings their own software stack and rents accelerator hours. The second is hosted inference and fine-tuning: the customer brings a model and pays per token. CoreWeave grew up on the first. The hyperscalers focus more on the second, where they bundle their own model-serving stacks. The race is to be credible at both.
The build-out cost reality
Standing up an AI-grade data center is more capital-intensive than a generic cloud region. The relevant components include:
- Liquid cooling — direct-to-chip or rear-door heat exchangers are now standard for top-end GPU racks.
- Power density — racks above 100 kW are common, which forces colocation in markets with available substation capacity.
- Network fabric — non-blocking high-speed interconnect across many GPUs at low latency is non-trivial.
- Allocations — even with the capital, the operator must secure GPU supply, which is the rate limiter for the whole industry.
Why allocation matters more than capital
Capital can be raised; allocations cannot. Nvidia's commercial team chooses where to send the marginal accelerator, and those decisions cascade into who can grow their book. A challenger that secures allocations early can lock in long-duration customer contracts, which then become the collateral for the next financing round. That is exactly the loop CoreWeave used, and it is the same loop Nebius is now running.
In any business where supply is rate-limited by a single vendor, the most valuable asset on the balance sheet is the customer commitment, not the depreciable equipment.
The token-pricing question
Hosted inference has been getting steadily cheaper per token at the model-vendor layer. If that compression continues at the same pace into the infrastructure layer, the AI cloud becomes a low-margin utility business. The counter-case is that quality differentiation — latency, uptime, model selection — supports a premium-tier price that doesn't fully commoditize. Nebius's pitch is built around the second hypothesis: that it can earn premium pricing on quality.
| Buyer type | What they want | Implication for AI clouds |
|---|---|---|
| AI labs / model builders | Raw GPU capacity, peak performance | Bare-metal sales, large multi-year contracts |
| Enterprise customers | Hosted inference, compliance posture | Managed services revenue, higher margin |
| Independent developers | Pay-as-you-go, cheap tokens | Commodity tier; thin margin |
What to watch over the next year
- Allocation share — how much of Nvidia's incremental supply each pure-play receives.
- Long-duration contract value — multi-year, take-or-pay-style commitments from named customers.
- Power capacity — disclosed new-market entries and signed power contracts.
- Margin trajectory at scale — gross margin per occupied rack-equivalent as utilization climbs.
FAQ
Are AI clouds direct competitors of the hyperscalers?
Yes and no. They compete with the GPU divisions of AWS, Azure and Google Cloud for raw GPU rentals; the hyperscalers' broader stack — storage, database, security — is not the comparable competitive surface.
Is this segment a winner-takes-most market?
The infrastructure layer is more likely two-or-three than one. The economics resemble specialized colocation more than they resemble a software platform.
What's the biggest single risk?
Sudden softening of foundation-model capex. If the largest customers slow build-outs, occupancy falls and the financing case for new capacity weakens fast.
The bottom line
The AI cloud is no longer one challenger versus the hyperscalers. It's a real sub-sector with at least two credible pure-plays, and the next year of allocations and contracts will determine which model — bare-metal-heavy or inference-heavy — earns the durable margin.





