When compute runs short, Microsoft serves its own products first, so access to AI capacity is now a rationed privilege your contract does not guarantee.
Microsoft's finance chief said the quiet part into a microphone this week. Asked about the GPU shortage around the company's July 29 earnings, CFO Amy Hood confirmed that scarce compute goes to Copilot, GitHub Copilot, and internal research before paying Azure customers get their turn ("AI compute crunch: Microsoft feeds its own products before Azure customers"). The company is even renting capacity from Amazon and Google to cover the gap, while warning that supply will not catch demand before the end of 2026.
Read that from the seat of a startup that chose Azure to sit close to OpenAI's models. Your landlord is now bidding against you for the same racks, and it has said, on the record, who wins.
That one admission unsettles a belief most buyers still carry without inspecting it: that paying for cloud AI puts you in an orderly queue where more money buys more compute. The queue has a priority lane, and the customer is not in it. For anyone deploying or operating AI, that changes what to test before the next planning meeting. Buyers now carry a dependency they need to prove they can survive. Sellers have a concrete audit they can build against it.
The shortage underneath. For three years the open question was demand: would anyone actually buy all this compute. This week answered it. Amazon reported AWS at $42.2 billion for the quarter, its fastest growth in eighteen quarters, and a day later Azure crossed $100 billion in annual revenue ("Cloud AI revenue: AWS tops $42B as Amazon and Microsoft report"). The money is real. What moved is the bottleneck. TSMC raised its 2026 capital budget to as much as $64 billion and still frames advanced packaging and memory, not order books, as the ceiling ("TSMC lifts 2026 capex to $64B as AI's bottleneck shifts from demand to supply"). Nvidia's China-legal H200 makes the point without mercy: buyers ordered more than two million while roughly 700,000 exist ("Nvidia H200: China ordered 2 million chips, only 700,000 exist"). When the limit is set inside a fab, more orders just lengthen the line.
So the change this week is not that AI got more expensive. It is that access to capacity became a decision your provider makes, not a guarantee your contract implies.
The players closest to the supply curve are already acting like it. Nvidia is reportedly weighing a $250 billion guarantee so OpenAI can lease a 10-gigawatt Ohio campus, because OpenAI's credit sits below investment grade and lenders want Nvidia's balance sheet behind the rent ("AI data centers: Nvidia weighs a $250B backstop for OpenAI's Ohio buildout"). The chip seller is cosigning its customer's lease to keep that customer buying chips. Meta lifted the top of its 2026 capex range to $145 billion against $60.8 billion in quarterly revenue ("Meta AI spending: capex guide climbs to $145B against $61B revenue"). AMD locked up more than 530 megawatts through a 15-year Core Scientific lease that does not even fill until 2027 ("AMD data center deal: Core Scientific leases 530MW to seat Instinct GPUs by 2027"). Everyone who can see the curve is buying position years ahead. The customer without that leverage takes what is left after the landlord serves itself.
The relief is real, but only at one layer. It would be easy to read this week's price cuts as the counterweight. DeepSeek says its V4-Flash retrain beats its own larger model on every benchmark it published, at roughly a third of the price, and now speaks OpenAI's Responses API so Codex can call it directly ("DeepSeek V4-Flash: a retrain beats the flagship at a third of the price"). OpenAI says GPT-5.6 tuned its own serving stack to cut costs 20% and pushed the savings into a cheaper API tier at $0.20 per million input tokens ("OpenAI says GPT-5.6 cut its own serving cost 20% and dropped prices"). Both numbers are vendor-reported, and both live at the software layer. A cheaper token still has to run on a chip that someone upstream rationed. Model prices falling and physical capacity tightening are happening at the same time, and they are not the same number. Treat the first as a discount you can bank and the second as a risk you have to manage.
What became risky this week is single-vendor dependency on capacity you do not control. What became valuable is the ability to prove you have somewhere else to go. That is the whole diagnosis, and it splits cleanly by who is reading.
For buyers and operators, treat provider access as a live risk this quarter rather than a fixed line item. Stand up one real fallback before you need it. DeepSeek V4-Flash plugs into Codex through the Responses API, and GPT-5.6's cheapest tier runs at $0.20 per million input tokens, so an alternate is a configuration change, not a rebuild. Route a slice of production traffic to it, measure cost per successful task and error rate against your primary, and get quota commitments and exit terms in writing instead of trusting a status dashboard. If a rationing call reaches your inference tomorrow, the drill you ran today is the difference between a reroute and an outage.
For sellers, consultants, agencies, and software teams, this week wrote your next offer for you. Package a 30-day vendor-dependency scorecard. Map every model call in the client's stack, add fallback routing to a second provider, define the quota and exit terms worth negotiating, and hand back a failure-case report showing exactly what breaks when the primary throttles. The demand is not hypothetical. A named CFO just confirmed the risk out loud, which is the kind of proof that turns a vague worry into a signed statement of work.
The assumption to retire is that cloud compute is a commodity you can summon on demand at a posted price. This week it behaved like a rationed resource whose priority the landlord sets. The more useful principle is to hold access the way you hold any critical supply, with a second source, a written commitment, and a real number for what an outage costs you when the first source says not today. Spend your next planning meeting finding where that single point of failure sits, because your provider has already decided where you sit.
The week in one line: AI capacity is rationed, not bought on demand, so hold a second source and a written quota the way you hold any critical supply.
Sources this week: Microsoft's compute rationing, hyperscaler AI revenue, TSMC's capex jump, the H200 shortage, Nvidia's OpenAI backstop, Meta's capex guide, AMD's Core Scientific lease, DeepSeek V4-Flash, GPT-5.6's self-optimized serving