CPUs are having a moment: Inside Arm’s hyperscale surge
Two earnings calls, one story. In the span of 24 hours, Microsoft and Amazon each delivered proof that Arm-based custom silicon has moved from experiment to backbone of hyperscale cloud infrastructure.
Microsoft CEO Satya Nadella put the architectural shift in nine words: “When it comes to running agents, CPUs are just as important as GPUs.” Then he backed it up: By the end of July, Microsoft expects Cobalt 200 racks running in more than 25 data centers worldwide.
A day later, Amazon reported its chips business (Graviton, Trainium and Nitro combined) had blown past a $25 billion annual revenue run rate, growing triple digits year over year. Graviton revenue commitments nearly tripled quarter over quarter, and Graviton5 is being adopted almost twice as fast as Graviton4 did.
Different products, different clouds, same direction: the biggest hyperscalers are building Arm-based CPUs as core AI infrastructure, tuned to their own workloads and their own economics.
Agents need more than model compute
Generative AI’s first wave was measured in model size, accelerator supply, tokens out. Agentic AI is more complex and messier. Agents retrieve data, call tools, run code, apply policies, check their own work. That’s a wider job.
GPUs still perform the bulk of model training and inference token generation, but CPUs run everything around them: routing requests, preparing data, orchestrating tools, managing memory and storage, keeping thousands of concurrent interactions moving without falling over.
For Microsoft, its Cobalt VM portfolio already carries first-party Microsoft services plus customer workloads from Adobe, Arm, Elastic, OpenAI, Sprinklr and TomTom, Nadella said during their earnings call. That’s a real production footprint.
AWS is telling the same story from a different angle. Graviton5is built for agentic AI workloads such as real-time reasoning, code generation, and multi-step orchestration. Graviton5 powered M9g instances deliver up to 25% better compute performance than Graviton4-based M8g instances and up to 35% faster machine learning inference.
Two hyperscalers, two paths to the same scale
Cobalt and Graviton are proof that Arm lets hyperscalers optimize silicon for their own fleets while sharing one architecture and one software base underneath. That accelerates time-to-revenue. Microsoft announced Cobalt 200 in November 2025. VM preview opened in June 2026. Less than two months later: racks live across 25-plus data centers, the company said. That’s fast.
Built on Arm Neoverse Compute Subsystems V3, Cobalt 200 is Microsoft’s second-generation Arm cloud CPU — up to 50% better performance than its predecessor, scaling to 128 vCPUs, built for cloud-native, data-intensive, agentic workloads.
AWS has five Graviton generations behind it. Graviton5-powered M9g and M9gd instances went GA in June, delivering up to 25% better compute performance than Graviton4. Across the whole Graviton lineup, AWS claims 30–40% better price-performance than comparable instances.
The adoption numbers are the real headline: 98% of AWS’ top 1,000 EC2 customers use Graviton. Over 120,000 customers build on it. For a third straight year, more than half of new AWS CPU capacity has been Graviton.
Efficiency becomes capacity, capacity becomes revenue
Microsoft drew a straight line from infrastructure efficiency to the bottom line. Azure revenue grew 43% for the quarter, with demand still outrunning available capacity. CFO Amy Hood credited efficiency gains across both the CPU and GPU fleets — plus faster infrastructure turn-up — for capacity that got monetized almost as fast as it came online.
At hyperscale, CPU efficiency isn’t just a lower power bill. When demand outstrips supply, a more efficient fleet squeezes more usable capacity out of the same racks and the same power envelope. That capacity becomes revenue.
Microsoft didn’t break out Cobalt 200’s specific contribution versus other fleet efficiencies. But the logic is now on the record: performance, utilization and power efficiency are a capacity lever and a growth lever at once.
AWS adds a second data point. The $25 billion run rate covers the whole chips business, not Graviton alone, but faster commitments and quicker Graviton5 uptake show a custom CPU line pulling real weight inside a large, fast-growing silicon platform, while giving customers a hard price-performance reason to shift more workloads over.
The stakes keep rising. The IEA says electricity use from AI-focused data centers jumped 50% in 2025. Dell’Oro pegs Q1 2026 capex from Amazon, Google, Meta and Microsoft up 78% year over year. Every watt and every rack matters more than it did a year ago. CPU efficiency isn’t a footnote anymore; it’s how hyperscalers unlock constrained capacity and get better returns on what’s already built.
Additionally, the market is growing past GPUs, with IDC expecting global AI infrastructure spending to hit $497 billion in 2026 — up roughly 56% year over year, with real demand now showing up beyond GPU systems: orchestration, data pipelines, CPU-only inference clusters.
A broader signal for Arm at hyperscale
Microsoft and AWS are today’s headlines, but they’re part of a much broader architectural shift.
Google Cloud built Axion, their custom Arm-based CPU family of processors, serving general purpose CPU underpinning its cloud infrastructure and as well as the AI head node CPU alongside its latest TPU generations, providing the orchestration, networking and infrastructure services that keep AI systems running at hyperscale.
NVIDIA has made the same architectural choice for AI factories. Vera, NVIDIA’s latest Arm based CPU, serves as the compute foundation for next-generation AI systems including Rubin, where CPUs coordinate memory, networking and accelerators across rack-scale deployments.
Taken together, Microsoft Cobalt, AWS Graviton, Google Axion and NVIDIA Vera point to the same conclusion: as AI infrastructure scales from model inference to autonomous agent execution, Arm-based CPUs are becoming the common compute foundation that orchestrates modern AI data centers.
Any re-use permitted for informational and non-commercial or personal use only.






