In March 2026, Google told Meta it could no longer meet the full Gemini capacity Meta had been consuming. Meta was Google Cloud’s largest Gemini customer by volume, and the notification, first reported by the Financial Times and independently confirmed by Reuters, landed while Google Cloud’s backlog sat at $240 billion, nearly double the previous quarter’s figure.
Neither company has commented on the record. The silence is informative: this is not something either party wants to discuss publicly, and that reticence alone makes it worth understanding — and worth examining as part of a broader pattern of AI infrastructure denial reshaping the industry.
What actually happened?
The restriction was a capacity cap, not a binary cut-off. Meta still had Gemini access, just not enough to meet its demand. Reuters described it as Google “putting limits on” Meta’s use; CNBC framed it as “throttled” access. Several other Google Cloud clients were affected, though to a lesser degree, and no other companies were named.
The cap arrived because Google simply did not have enough Gemini capacity to go around. Google Cloud’s backlog then surged to $460 billion by Q1 2026, and CEO Sundar Pichai told investors that compute capacity constraints prevented higher revenue growth. When your largest customer is also your biggest competitor, and you cannot serve everyone, someone gets shorted. Meta drew the short straw.
The cap mattered because of where Gemini was embedded in Meta’s operations.
What Meta was using Gemini for
Meta’s Gemini dependency was not experimental. It was embedded in production across three categories.
First, customer-facing products: customer-service chatbots and advertiser-facing chatbot tools. These are revenue-adjacent workloads (they support revenue but don’t directly transact), not sandboxes. Second, internal developer tooling: Meta’s own coding assistants ran on Gemini, meaning engineering velocity at one of the world’s largest software organisations was partially dependent on a competitor’s model. Third, trust and safety: content moderation pipelines for scam detection and harmful content takedowns across Facebook, Instagram, and WhatsApp relied on Gemini’s classification capabilities because Gemini outperformed Meta’s own Llama models on those specific tasks.
Meta chose Gemini over its own Llama family because Gemini 3.1 Pro ranked first globally on the Artificial Analysis Intelligence Index. The performance gap was not marginal. When you are moderating content for billions of users, the difference between the best model and the second-best model is measured in missed scams and undetected harm.
Capacity, competition, or both?
The capacity shortage was real. Google Cloud’s backlog ballooned. Pichai acknowledged compute constraints as what kept him up at night. Google committed internally to doubling AI serving capacity every six months. And most tellingly, Google signed a deal paying xAI and SpaceX approximately $920 million per month to lease external data centre capacity. Google, with a $175 billion capital expenditure budget, is paying a competitor nearly a billion dollars a month because it cannot build fast enough itself. That is a concession to physics, not a negotiating tactic.
But the competition case is equally strong. Meta and Google compete directly in digital advertising, which is both companies’ core revenue engine. Meta’s consumer AI products increasingly compete with Google’s Search, Workspace, and Cloud AI, all of which consume the same Gemini infrastructure. Prioritising a direct competitor when capacity runs short would be commercially irrational, and Google did not.
Wedbush analyst Matt Bryson distilled the synthesis: the restrictions “underscore the risks companies face when depending on competitors for critical computing resources.” Google likely could not serve everyone at full capacity, but it chose whom to short. The scarcity was real; the allocation was competitive. The two explanations describe the same decision from different angles, and the distinction between “couldn’t” and “chose not to” is what you need to watch for in your own supplier relationships.
The physical constraints behind the shortage
If the capacity shortage were just about money, Google would have written a cheque and fixed it. The problem is physical infrastructure, and physical infrastructure moves at the speed of construction, not the speed of software.
High Bandwidth Memory (HBM), the performance-critical memory technology for AI accelerators, is sold out through 2026 across all major suppliers. SK Hynix controls 62% of the market. Samsung and Micron produce the rest. The advanced packaging that stacks HBM layers is itself capacity-constrained, and new production lines will not reach full-scale mass production until 2027 or later.
Then there is power and cooling. Modern GPU racks draw roughly 370 kW. Most existing data centres were designed for 10 to 15 kW per rack. Retrofitting for liquid cooling requires physical rebuilds, including coolant distribution systems, heat rejection equipment, and in many cases upgraded electrical service from the grid. Roughly half of planned US data centres for 2026 have been delayed or cancelled. Only about a third of planned capacity is under active construction.
Google also ordered 3 million AI chips from Intel to expand its compute footprint. But Intel’s Gaudi chips are not direct substitutes for the NVIDIA GPUs and Google TPUs that run Gemini, so the order broadened capacity rather than solving the Gemini bottleneck.
Each constraint independently limits capacity. Together they create a bottleneck that no amount of capital spending can resolve in under two to three years. That is what Pichai meant when he described compute constraints as the thing that keeps him up at night.
Here is the thing: if these physical constraints mean scarcity is measured in years, not quarters, then depending on a competitor’s infrastructure becomes a structural risk rather than a temporary inconvenience. That is the problem Meta walked into.
What this means for AI vendor dependency
If Meta, with a $115 to $145 billion capex budget, its own frontier model programme, and multi-provider contracts with Google Cloud, Anthropic, CoreWeave, and Oracle, can be materially constrained, then your business isn’t immune, no matter its size.
Meta had the contracts. What it did not have was workload portability: the ability to move specific AI workloads between providers without material degradation or re-engineering. Gemini was embedded in production workflows where alternatives were not drop-in replacements. Here is the practical lesson for your business: multi-provider strategies mitigate dependency risk, but they don’t eliminate it — not when AI infrastructure denial is becoming a competitive battleground.
Apple’s Siri-Gemini partnership presents a parallel case worth watching. Apple chose Google as its preferred cloud provider for Siri’s AI overhaul in a deal reportedly worth $1 billion per year, running on a custom 1.2 trillion parameter Gemini model. Apple competes with Google across operating systems, app stores, and consumer AI. The structural risk is the same as Meta’s, just wearing a different logo. And Apple’s dependency may run deeper: Siri is a flagship consumer product, not an internal tool.
The distinction between renting from a vendor and renting from a competitor becomes the central axis of infrastructure strategy when capacity is scarce. And capacity is now structurally scarce, as those physical constraints make clear.
Meta’s three-pronged response
Meta responded with escalating moves that collectively signal a strategic pivot from AI infrastructure consumer to provider.
The immediate response was token rationing: employees were urged to use shorter prompts, reduce context windows, and consolidate queries. It is the AI equivalent of a data cap. Large enterprises across the board were doing the same thing as AI bills climbed.
The next response was Muse Spark, Meta’s first closed-source AI model, launched April 8, 2026, and updated to version 1.1 in July with a focus on agentic coding. A $14.3 billion investment in Scale AI for a 49% stake provided the data-labelling infrastructure to train frontier proprietary models, and Scale’s co-founder Alexandr Wang, 29, was installed as Meta’s Chief AI Officer. Meta had already been pivoting away from open-weight Llama after the Llama 4 bench-maxxing scandal. The Gemini cutoff accelerated the timeline; it did not create it.
The structural response was Meta Compute, a planned cloud infrastructure business to sell excess AI compute and model access externally. Mark Zuckerberg told shareholders that companies were approaching Meta “every week” asking to buy compute at a premium. Meta Compute transforms the company from a net consumer of external AI infrastructure into a provider, directly challenging AWS, Azure, and Google Cloud. It is designed to reduce the risk of depending on a competitor for something as basic as compute.
The Google-Meta Gemini incident exposed the first visible crack in a system where AI compute is structurally scarce, allocated along competitive lines, and driving the industry’s largest consumers to become self-sufficient — part of a structural shift in how AI compute is allocated. The motivation is not cost savings. It is that building your own infrastructure is the only way to ensure continuity when your supplier is also your rival. Meta’s three-pronged response is not a recovery plan. It is a preview of what your business will eventually face if it builds on someone else’s infrastructure.
Frequently Asked Questions
Why didn’t Meta just use its own Llama models instead of Gemini?
Meta did use Llama for many workloads, but Gemini 3.1 Pro ranked #1 globally on the Artificial Analysis Intelligence Index and materially outperformed Llama on specific production tasks including content moderation and chatbot responsiveness. The performance gap was not marginal. Compounding this, the Llama 4 bench-maxxing scandal had already eroded internal confidence in Meta’s open-weight strategy, accelerating the shift toward proprietary models well before the Gemini cutoff added urgency.
What was the Llama 4 bench-maxxing scandal?
The Llama 4 bench-maxxing scandal involved allegations that Meta optimised its Llama 4 models to score well on public benchmarks rather than perform well on real-world tasks, inflating published performance relative to actual capability. Combined with growing concerns about adversarial distillation (where competitors extract model capabilities through systematic querying), the scandal pushed Meta toward proprietary models like Muse Spark, before the Gemini cutoff added further urgency to that transition.
What exactly is token rationing and how does it work in practice?
Token rationing means imposing internal limits on how many tokens an organisation’s employees can consume. In practice, Meta urged staff to use shorter prompts, reduce context windows, avoid unnecessary regenerations, and consolidate multiple queries into single interactions. It is the AI equivalent of a data cap: when supply is constrained, every token must justify its cost rather than being treated as effectively free. Accenture pursued a similar approach as enterprise AI costs surged across providers.
Could this same thing happen with OpenAI and Microsoft, or Anthropic and Amazon?
Yes, the structural conditions are nearly identical. Microsoft is both OpenAI’s largest infrastructure provider and a direct competitor through Copilot and Azure AI. Amazon has invested heavily in Anthropic while competing through its own Titan and Nova models. When AI compute capacity runs short, any provider that also competes with its own customers faces the same allocation tension Google and Meta demonstrated. The question is when, not if.
What does Google’s $920 million monthly lease to xAI and SpaceX tell us about the compute shortage?
It is the most vivid proof that the shortage is physical, not financial. Google, with effectively unlimited capital and a $175 billion capex budget, is paying a competitor nearly a billion dollars each month just to lease data centre space it cannot build fast enough itself. When a hyperscaler of Google’s scale capitulates to renting from a rival, the constraint is not about money. It is about physics, construction timelines, and power infrastructure.
Why is High Bandwidth Memory such a bottleneck?
HBM sits physically close to AI chips and delivers the extreme memory bandwidth frontier models require, but its supply chain is extraordinarily concentrated: SK Hynix, Samsung, and Micron control virtually all production. The advanced packaging that stacks HBM layers is itself capacity-constrained, and the fabrication facilities take years to build. You cannot spin up more HBM production the way you can increase conventional DRAM output. The bottleneck is structural, not cyclical.
How does the Apple-Siri-Gemini deal put Apple in the same risky position?
Apple chose Google as its preferred cloud provider for the Siri AI overhaul, embedding Gemini into a product used by hundreds of millions of people. The Google-Meta incident demonstrated that Google will ration Gemini capacity when constrained, and Apple competes with Google across operating systems, app stores, and consumer AI. Apple’s dependency is arguably deeper than Meta’s because Siri is a flagship consumer product, not an internal tool.
What is workload portability and how is it different from having multiple AI providers?
Multi-provider means you hold contracts with several AI vendors. Workload portability means you can actually move a specific AI task from one provider to another without material performance degradation or significant re-engineering. Meta had the contracts but not the portability: Gemini was embedded in production workflows where alternatives were not drop-in replacements. Portability is the harder standard, and it is the one that genuinely protects against dependency risk.
Is Google singling out Meta, or are other companies being restricted too?
Reuters reported that other Google Cloud clients were affected “though to a lesser extent,” but no specific companies or workloads were named. The silence is telling: if the restrictions were genuinely neutral and widespread, Google would likely frame them as an even-handed capacity constraint. The fact that Meta, as the largest Gemini customer and a direct competitor, was the primary case suggests competitive allocation rather than uniform rationing across all customers.
What does Meta Compute mean for the broader cloud market?
Meta Compute signals that the largest AI consumers are becoming suppliers, directly challenging the hyperscaler cloud triopoly of AWS, Azure, and Google Cloud. By selling excess compute and model access (including Muse Spark) externally, Meta is building infrastructure that ensures it never again depends on a competitor for critical AI capacity. It represents a new category of entrant: the AI-first cloud provider born from dependency trauma rather than market opportunity.
What is the difference between an AI provider and an AI competitor, and why does it matter?
An AI provider sells you compute or model access with no meaningful competitive overlap. An AI competitor does the same while also competing with you in advertising, consumer products, or enterprise software. The distinction matters because when capacity runs short, a provider will serve you if capacity allows; a competitor will serve you only after its own needs and non-competing customers are met. Vendor selection is now a competitive strategy question, not a procurement exercise.
How long will the AI compute shortage actually last?
The physical constraints (HBM supply concentration, data centre construction timelines, power infrastructure, cooling retrofits) mean the shortage is measured in years, not quarters. Even with Google’s $175 billion capex guidance and industry-wide investment at unprecedented scale, the pipeline for new chip fabrication facilities and purpose-built AI data centres extends two to three years minimum. No amount of spending can compress physics and construction timelines below that horizon.