The full-stack cloud bet, and how it differs from Microsoft and AWS#
At Cloud Next, its big annual keynote, Google rolled out the Gemini Enterprise Agent Platform, replacing Vertex AI, which had been its AI platform since 2023.(1) The rename is the tell: Google wants the whole stack, from the chips up to the apps that put agents in front of every employee. As Thomas Kurian, who runs Google Cloud, put it on stage: competitors hand you the pieces, Google hands you the platform.
Microsoft and Amazon are making the same bet. Where they differ is which layer each one owns, so let's go layer by layer.
The four-layer AI Stack#
Every cloud AI business stacks into four layers. At the bottom, silicon: the chips that run AI workloads, and the data centers and power that keep them running. Above silicon, the models: the LLMs that turn a user request into an answer or an action. Above the models, the platform where you build, run, and govern AI application. At the top, distribution: the channels that put AI in front of users.
A few years back, Google, Microsoft, and Amazon each owned a layer or two and rented the rest. Now they are all going for the whole stack.
Layer 1: The Chips#
Every token your agent generates runs on a GPU, and most of that spend goes to Nvidia. A cloud that owns its chips can skip Nvidia's hardware markup and keeps the difference. Nvidia's margins on those chips are the kind of number you don't say out loud in polite company. That is why Google, Microsoft, and Amazon are all building their own chips.
Google has been at this the longest. Its TPU, a chip designed from scratch for matrix math, has shipped a new generation every two years for a decade. At Cloud Next, Google announced two eighth-gen TPUs: the 8t for training, claiming 3x the compute of the last generation, and the 8i for inference, claiming about 80% better performance per dollar.(2) (One letter off from an iPhone.)
Amazon's Trainium chips already carry most of the inference traffic on Bedrock. Microsoft's Maia won't be ready at scale until late 2026, so Azure workloads still run on Nvidia hardware at Nvidia prices.(3)
Even Google, with the most mature custom silicon, still buys Nvidia GPUs in volume. Custom chips only pay off when a single model serves billions of users, like Gemini powering Google Search. A sentence maybe four companies on earth can say with a straight face. Every other workload runs on Nvidia.
Chip supply is the real bottleneck. Google's own DeepMind researchers have reportedly queued for TPU capacity behind paying customers.(4) The people who invented the chip now wait in line behind the people renting it. The cloud with the biggest head start still can't make enough TPUs to meet demand.