Microsoft will deploy AMD Helios rack-scale systems on Azure, the two companies said on July 20, 2026, in an expanded partnership that spans GPUs, CPUs, networking, and software. Helios is AMD's integrated rack: seventy-two Instinct MI455X accelerators, sixth generation EPYC Venice processors, Pensando networking, and the open ROCm software stack, wired together to train and serve frontier AI models. For anyone who runs infrastructure, the interesting part is not the raw specification. It is that a hyperscaler is committing to a second, credible, largely open alternative to the dominant GPU stack, with the first Helios shipments to Microsoft and other customers due in the second half of 2026.
The short answer
Microsoft will deploy AMD Helios rack-scale systems on Azure, the two companies said on July 20, 2026. Helios pairs seventy-two Instinct MI455X GPUs with sixth generation EPYC Venice CPUs, Pensando networking, and the open ROCm software stack, all designed to run frontier AI training and inference as one integrated rack. Azure is also adding two new EPYC Venice virtual machine series and broadening its use of AMD Pensando DPUs. Shipments to customers, Microsoft among them, begin in the second half of 2026.
Most AI hardware news is a specification sheet with a bigger number on it. This one is worth a closer look for a different reason. On July 20, 2026, AMD said Microsoft will run its Helios rack-scale systems on Azure, and the reason that matters is not the teraflops. It is that one of the largest cloud operators in the world is putting real capacity behind a second, largely open alternative to the GPU stack that has dominated AI infrastructure. We read this as a supply and portability story first, and a benchmark story second.
What Microsoft and AMD announced
AMD framed the news as an expanded, long-term strategic partnership that reaches across GPUs, CPUs, networking, and software on Azure. The headline commitment is that Microsoft will deploy the AMD Helios rack-scale system to power frontier model AI inference, both for Microsoft's own services and for its cloud customers.
There are two supporting pieces. First, Azure will add two new virtual machine series built on sixth generation EPYC Venice processors, aimed at agentic AI and data pipeline work on one hand and demanding technical workloads on the other. Second, Microsoft will broaden its deployment of AMD Pensando DPUs across the backend AI networking that ties these clusters together, and across select Azure services. AMD said the first Helios systems ship to customers, Microsoft included, in the second half of 2026. Investors noticed: AMD shares rose close to five percent on the day.
What is inside a Helios rack
Helios is best understood as a rack, not a product you slot into an existing server. AMD describes it as an integrated, open rack-scale platform that brings together four things designed to operate as one system.
- Seventy-two Instinct MI455X GPUs. The accelerators that do the training and inference work, pooled at rack scale so a single job can span the whole unit.
- Sixth generation EPYC Venice CPUs. The host processors that feed the GPUs, handle orchestration, and run the parts of a workload that are not GPU bound.
- Pensando networking. The fabric that connects accelerators inside and across racks, offloading data movement so the GPUs spend their time computing rather than waiting.
- The ROCm software stack. AMD's open software layer for programming the GPUs, the piece that decides how easily existing code actually runs on this hardware.
The point of assembling these as one platform is operational. Instead of a team sourcing GPUs, host CPUs, DPUs, and a fabric separately and validating that they cooperate, a Helios rack arrives already engineered to work together. At the scale a hyperscaler operates, that integration is most of the value.
Why a second GPU stack matters to you
If you plan or run AI infrastructure, the useful question is not whether these GPUs beat the incumbent on a given benchmark. It is what a committed second option does to your choices over the next two years.
- Supply gets less fragile. When almost all large-scale AI compute depends on one vendor, availability and lead times bend to that vendor's roadmap. A hyperscaler standing up serious AMD capacity widens the pool, which tends to ease scarcity and, over time, pricing.
- Open software is the real lever. The hardware is only usable if your code runs on it, and Helios is built around the open ROCm stack rather than a closed one. That lowers the switching cost in principle, but it only pays off if you keep your workloads portable. If your stack assumes one vendor's libraries end to end, start noting where those assumptions are baked in.
- Networking is doing more of the work. The Pensando DPUs are not a footnote. As clusters grow, moving data between accelerators becomes the bottleneck, and offloading that traffic to dedicated networking hardware is how these racks keep the expensive GPUs busy. For network engineers, this is the part of AI infrastructure that increasingly looks like your job.
None of this is something to act on this week. Helios ships in the second half of 2026, and cloud availability follows from there. But it is a concrete signal about where the market is heading: toward more than one credible accelerator, more open software, and networking that carries a growing share of the load. Those are all reasons to keep your options open rather than to bet everything on a single stack.
Sources and further reading
- AMD newsroom: Microsoft to deploy next-gen AMD Instinct and EPYC processors
- AMD investor relations: expanded Microsoft Azure partnership
- SiliconANGLE: Microsoft will use AMD's AI-optimized Helios racks in Azure
Frequently asked questions
What exactly did Microsoft and AMD announce on July 20, 2026?
AMD said Microsoft will deploy the AMD Helios rack-scale system on Azure to run frontier model AI inference for Microsoft, its customers, and Azure AI services. The expanded partnership covers AMD GPUs, CPUs, networking, and software. Azure will also add two new virtual machine series powered by sixth generation EPYC Venice processors and will broaden its use of AMD Pensando DPUs across backend AI networking and select services. AMD said Helios shipments to customers, including Microsoft, begin in the second half of 2026.
What is inside an AMD Helios rack?
Helios is an integrated rack-scale platform rather than a single card. AMD describes it as seventy-two Instinct MI455X GPUs, sixth generation EPYC Venice CPUs, Pensando networking, and the ROCm software stack, assembled as one open system for large-scale AI training and inference. The idea is that the compute, the host processors, and the networking fabric arrive designed to work together, so operators deploy a rack as a unit instead of stitching components together themselves.
Why does a second GPU option matter for teams running AI workloads?
Most large-scale AI compute today runs on one vendor's GPUs and one proprietary software stack. A hyperscaler committing to AMD Helios at scale gives the market a second credible path, built around the open ROCm software stack rather than a closed one. For teams that plan capacity, that can mean more supply, more pricing pressure over time, and a real reason to keep workloads portable across accelerators rather than locked to a single stack.
Can I use Helios today?
Not yet. The announcement is a deployment commitment, not a general availability date for customers. AMD said Helios systems begin shipping in the second half of 2026, and Microsoft will bring the capacity into Azure over that period. The two new EPYC Venice virtual machine series and the wider Pensando DPU rollout are part of the same multi-stage plan. Treat this as a roadmap signal worth tracking, not a resource you can provision this week.