Cloud

GPU capacity, not cloud pricing, is the new lock-in

Azure sits at full A100 availability. Google Cloud has been flat at 44% for a year. The vendor you picked for pricing may now be the one that cannot give you the compute your AI roadmap needs.

Cloud lock-in used to be a story about egress fees and proprietary APIs. In 2026 it is increasingly a story about which vendor can physically get you a GPU, and the providers are no longer converging.

The availability gap is real and widening

AWS moved from roughly 50% A100 availability in mid-2024 to full coverage by late 2025, reflecting aggressive capacity investment. Azure has held steady at full availability throughout. Google Cloud, by contrast, has remained flat at around 44% availability with no meaningful year-over-year growth — a gap that has nothing to do with pricing and everything to do with which provider actually has the hardware in the region you need it. AWS has separately committed to deploying more than a million Nvidia GPUs across its regions starting this year, including next-generation Blackwell and Rubin architectures.

Why this changes the multi-cloud calculus

The traditional argument for multi-cloud was risk diversification and negotiating leverage. The argument now is more concrete: GPU procurement is a multidimensional problem involving regional distribution, redundancy, latency, and compliance, and a team locked into a single provider risks hitting a capacity wall that no amount of budget solves on short notice. A vendor relationship chosen two years ago for its pricing or its integration ease may simply not have the compute inventory your current AI workload requires — and unlike a software feature gap, you cannot code around a GPU shortage.

Flexibility now has a price tag attached to it

Teams that can move workloads dynamically across regions and providers are positioned to capture real savings and avoid bottlenecks; teams locked into static contracts or single regions are the ones who will hit capacity constraints precisely when demand is highest. That is a rare case where operational flexibility and cost savings point the same direction rather than trading off against each other — but only for organizations that built the portability in advance, since retrofitting multi-cloud GPU access mid-crisis is a much harder and slower project than negotiating it up front.

The practical move: before signing the next major AI infrastructure commitment, ask not "what does this cost" but "what happens to my roadmap if this provider's regional GPU capacity does not grow as fast as my demand." If you cannot answer that today, you have a pricing strategy, not a capacity strategy — and only one of those two survives contact with a GPU shortage.

What this reacts to

The daily brief

CIOReview, in your inbox before standup

The headlines technology leaders are reading, synthesized and source-linked. One email each morning. No filler.