Google has limited Meta’s access to its Gemini AI models after Meta tried to buy more Gemini capacity than Google could supply, according to a Financial Times report that Reuters relayed but said it could not independently verify. The limit was communicated around March, remains in place, and has disrupted or delayed some internal Meta projects. Google and Meta declined to comment.
The notable part is who the customer is. Meta is one of the world’s largest AI companies, spending hundreds of billions of dollars on infrastructure and building its own Llama and Muse Spark models. Even so, it had reportedly been running a rival’s models internally, for safety automation, content moderation, customer service, ad-support chatbots, internal workflows, and coding, alongside Anthropic’s Claude, because Gemini outperformed Meta’s own models on some tasks.
After the cap, Meta reportedly urged employees to use AI tokens, the units that measure model usage, more efficiently. The detail suggests the issue was less a competitive cutoff than a supply problem: Meta’s internal usage had grown large enough to require rationing. Other Google customers were affected as well, with Meta hit harder because of unusually high demand. Meta has since been shifting more work to Muse Spark, which it launched in April as the first model from its Superintelligence Labs, though that model’s developer API has been delayed.
The cap reflects Google’s own capacity strain. In the first quarter, Google Cloud revenue grew 63 percent and passed $20 billion for the first time, backlog nearly doubled to more than $460 billion, and first-party model API use rose to more than 16 billion tokens per minute from 10 billion a quarter earlier. Alphabet has said constrained capacity held back even faster cloud growth and raised its 2026 capital spending plan to between $180 billion and $190 billion. It has also been buying outside compute, including a deal under which it will pay SpaceX about $920 million a month for capacity that does not fully arrive until later in the year.
The episode points to where the AI bottleneck now sits. Training frontier models remains costly, but running them at scale, for moderation, support, ads, coding, and agents, consumes enormous compute, making inference capacity the binding constraint. It is the same pressure pushing enterprises toward cheaper models and multi-model routing as token bills climb. For Meta, the awkward detail is that even with $600 billion pledged for US AI infrastructure and a multiyear Nvidia buildout, it still leaned on a competitor’s serving capacity for important internal work.
The report leaves open how much capacity Meta sought, which projects slipped, and whether the limits were contractual or technical. The underlying dynamic is clearer: as Gemini adoption outruns supply, access to a frontier model is becoming rationed infrastructure rather than a guaranteed subscription.
Sources: Financial Times, Reuters, Alphabet
–
By the Control Plane Editorial Team