NetworkNews

Microsoft Will Run AMD Helios Rack Systems on Azure

On this page
  1. What Microsoft and AMD announced
  2. What is inside a Helios rack
  3. Why a second GPU stack matters to you
  4. Sources and further reading

Microsoft will deploy AMD Helios rack-scale systems on Azure, the two companies said on July 20, 2026, in an expanded partnership that spans GPUs, CPUs, networking, and software. Helios is AMD's integrated rack: seventy-two Instinct MI455X accelerators, sixth generation EPYC Venice processors, Pensando networking, and the open ROCm software stack, wired together to train and serve frontier AI models. For anyone who runs infrastructure, the interesting part is not the raw specification. It is that a hyperscaler is committing to a second, credible, largely open alternative to the dominant GPU stack, with the first Helios shipments to Microsoft and other customers due in the second half of 2026.

The short answer

Microsoft will deploy AMD Helios rack-scale systems on Azure, the two companies said on July 20, 2026. Helios pairs seventy-two Instinct MI455X GPUs with sixth generation EPYC Venice CPUs, Pensando networking, and the open ROCm software stack, all designed to run frontier AI training and inference as one integrated rack. Azure is also adding two new EPYC Venice virtual machine series and broadening its use of AMD Pensando DPUs. Shipments to customers, Microsoft among them, begin in the second half of 2026.

72Instinct MI455X GPUs per Helios rack
H2 2026first Helios shipments, Microsoft included
2new EPYC Venice virtual machine series on Azure
Answer card: Microsoft will deploy AMD Helios rack-scale systems on Azure, combining 72 Instinct MI455X GPUs, EPYC Venice CPUs, Pensando networking, and ROCm software.
A rack, not a card: Helios ships compute, host CPUs, and networking designed to work as one unit. PNG

Most AI hardware news is a specification sheet with a bigger number on it. This one is worth a closer look for a different reason. On July 20, 2026, AMD said Microsoft will run its Helios rack-scale systems on Azure, and the reason that matters is not the teraflops. It is that one of the largest cloud operators in the world is putting real capacity behind a second, largely open alternative to the GPU stack that has dominated AI infrastructure. We read this as a supply and portability story first, and a benchmark story second.

What Microsoft and AMD announced

AMD framed the news as an expanded, long-term strategic partnership that reaches across GPUs, CPUs, networking, and software on Azure. The headline commitment is that Microsoft will deploy the AMD Helios rack-scale system to power frontier model AI inference, both for Microsoft's own services and for its cloud customers.

There are two supporting pieces. First, Azure will add two new virtual machine series built on sixth generation EPYC Venice processors, aimed at agentic AI and data pipeline work on one hand and demanding technical workloads on the other. Second, Microsoft will broaden its deployment of AMD Pensando DPUs across the backend AI networking that ties these clusters together, and across select Azure services. AMD said the first Helios systems ship to customers, Microsoft included, in the second half of 2026. Investors noticed: AMD shares rose close to five percent on the day.

What is inside a Helios rack

Helios is best understood as a rack, not a product you slot into an existing server. AMD describes it as an integrated, open rack-scale platform that brings together four things designed to operate as one system.

  • Seventy-two Instinct MI455X GPUs. The accelerators that do the training and inference work, pooled at rack scale so a single job can span the whole unit.
  • Sixth generation EPYC Venice CPUs. The host processors that feed the GPUs, handle orchestration, and run the parts of a workload that are not GPU bound.
  • Pensando networking. The fabric that connects accelerators inside and across racks, offloading data movement so the GPUs spend their time computing rather than waiting.
  • The ROCm software stack. AMD's open software layer for programming the GPUs, the piece that decides how easily existing code actually runs on this hardware.

The point of assembling these as one platform is operational. Instead of a team sourcing GPUs, host CPUs, DPUs, and a fabric separately and validating that they cooperate, a Helios rack arrives already engineered to work together. At the scale a hyperscaler operates, that integration is most of the value.

Answer card listing the four Helios building blocks: 72 Instinct MI455X GPUs, EPYC Venice CPUs, Pensando networking, and the ROCm software stack.
Four building blocks, one rack: accelerators, host CPUs, the networking fabric, and the ROCm software that ties them to your code. PNG

Why a second GPU stack matters to you

If you plan or run AI infrastructure, the useful question is not whether these GPUs beat the incumbent on a given benchmark. It is what a committed second option does to your choices over the next two years.

  • Supply gets less fragile. When almost all large-scale AI compute depends on one vendor, availability and lead times bend to that vendor's roadmap. A hyperscaler standing up serious AMD capacity widens the pool, which tends to ease scarcity and, over time, pricing.
  • Open software is the real lever. The hardware is only usable if your code runs on it, and Helios is built around the open ROCm stack rather than a closed one. That lowers the switching cost in principle, but it only pays off if you keep your workloads portable. If your stack assumes one vendor's libraries end to end, start noting where those assumptions are baked in.
  • Networking is doing more of the work. The Pensando DPUs are not a footnote. As clusters grow, moving data between accelerators becomes the bottleneck, and offloading that traffic to dedicated networking hardware is how these racks keep the expensive GPUs busy. For network engineers, this is the part of AI infrastructure that increasingly looks like your job.

None of this is something to act on this week. Helios ships in the second half of 2026, and cloud availability follows from there. But it is a concrete signal about where the market is heading: toward more than one credible accelerator, more open software, and networking that carries a growing share of the load. Those are all reasons to keep your options open rather than to bet everything on a single stack.

Sources and further reading

Frequently asked questions

What exactly did Microsoft and AMD announce on July 20, 2026?

AMD said Microsoft will deploy the AMD Helios rack-scale system on Azure to run frontier model AI inference for Microsoft, its customers, and Azure AI services. The expanded partnership covers AMD GPUs, CPUs, networking, and software. Azure will also add two new virtual machine series powered by sixth generation EPYC Venice processors and will broaden its use of AMD Pensando DPUs across backend AI networking and select services. AMD said Helios shipments to customers, including Microsoft, begin in the second half of 2026.

What is inside an AMD Helios rack?

Helios is an integrated rack-scale platform rather than a single card. AMD describes it as seventy-two Instinct MI455X GPUs, sixth generation EPYC Venice CPUs, Pensando networking, and the ROCm software stack, assembled as one open system for large-scale AI training and inference. The idea is that the compute, the host processors, and the networking fabric arrive designed to work together, so operators deploy a rack as a unit instead of stitching components together themselves.

Why does a second GPU option matter for teams running AI workloads?

Most large-scale AI compute today runs on one vendor's GPUs and one proprietary software stack. A hyperscaler committing to AMD Helios at scale gives the market a second credible path, built around the open ROCm software stack rather than a closed one. For teams that plan capacity, that can mean more supply, more pricing pressure over time, and a real reason to keep workloads portable across accelerators rather than locked to a single stack.

Can I use Helios today?

Not yet. The announcement is a deployment commitment, not a general availability date for customers. AMD said Helios systems begin shipping in the second half of 2026, and Microsoft will bring the capacity into Azure over that period. The two new EPYC Venice virtual machine series and the wider Pensando DPU rollout are part of the same multi-stage plan. Treat this as a roadmap signal worth tracking, not a resource you can provision this week.

Advertisement