Microsoft to deploy AMD Helios AI systems on Azure

Share


Microsoft will deploy AMD Helios on Azure.Azure will add AMD EPYC VMs and Pensando networking.

 

Microsoft plans to deploy AMD’s Helios rack-scale computing platform on Azure under an expanded infrastructure partnership covering processors, graphics chips, networking hardware, and software. The systems will support AI inference for Microsoft, Azure customers, and Azure AI services.

Microsoft also plans to offer the Helios infrastructure through its upcoming ND MI455X v7 virtual machines. The instances are designed for production-scale reasoning, search, and agentic inference workloads.

Helios targets large-scale AI inference

Helios combines AMD Instinct MI455X GPUs, sixth-generation EPYC “Venice” processors, Pensando networking technology, and the ROCm software platform. The reference design includes 72 MI455X GPUs, EPYC host processors, and Pensando “Vulcano” networking hardware in a double-wide rack.

AMD developed the system around Meta’s Open Rack Wide specification, which was submitted through the Open Compute Project. The design provides a common rack architecture for high-density computing, power delivery, and liquid cooling.

“AMD and Microsoft have spent years building high-performance infrastructure together, and today we’re extending that partnership across the full stack of AMD AI solutions on Azure,” AMD Chair and CEO Lisa Su said. “Microsoft’s new AMD deployments mark an important milestone as we deliver leadership compute solutions to Azure customers and scale the next generation of AI infrastructure together,” Su said.

Each MI455X GPU includes 432GB of HBM4 memory and provides memory bandwidth of up to 19.6TB per second, according to AMD. A full Helios rack provides up to 31TB of combined HBM4 memory across its 72 accelerators.

AMD rates the system at up to 1.4 exaflops of FP8 compute and 2.9 exaflops of FP4 compute. The figures represent peak manufacturer specifications rather than measured Azure application performance.

The MI455X GPUs handle the primary AI calculations, while the EPYC processors support host computing, workload coordination, and data movement. Microsoft will initially use Helios for frontier-model inference across its own services, Azure AI services, and customer applications.

Although Helios supports both model training and inference, Microsoft’s announced ND MI455X v7 deployment focuses on running trained models at scale. The company has not provided a timetable for customer access to the instances.

“Customers are looking for AI infrastructure that is optimised for a wide range of workloads, from training and inference to data preparation, search, and reinforcement learning,” Microsoft Chairman and CEO Satya Nadella said. “Through our collaboration with AMD, we are expanding the Azure infrastructure portfolio with AMD Helios to give customers the performance, scale, and choice they need to build and run the next generation of AI applications,” Nadella said.

Enterprise customers will also be able to access AMD-based infrastructure through Microsoft Foundry Managed Compute. The service hosts open-source models on dedicated GPU capacity, with Microsoft managing the GPU topology, runtime, container image, and security patching.

Customers select the model, accelerator family, deployment template, and scaling settings. Managed Compute remains in public preview, has no service-level agreement, and is not currently recommended by Microsoft for production workloads.

Azure expands CPU and networking infrastructure

The partnership also covers two Azure virtual machine series powered by AMD’s sixth-generation EPYC “Venice” processors. Azure HDv2 will include nearly 500 physical EPYC CPU cores, 4TB of RAM, 32TB of local NVMe storage, and 400Gb Azure Boost networking.

Microsoft is positioning HDv2 for CPU-intensive AI workloads, including data preparation, search, reinforcement learning, agent coordination, and data pipelines. The company said the VMs will handle processes that prepare data, coordinate workloads, and support GPU-based training and inference.

Azure HXv2 will include 176 sixth-generation EPYC cores running at more than 5GHz, with 50% more addressable cache per core than the previous HX generation. Microsoft plans to offer configurations with nearly 2TB or 4TB of RAM and 800Gb InfiniBand connectivity.

HXv2 will retain AMD’s 3D V-Cache technology, which is also used in the current Azure HX series. Microsoft and AMD introduced the first HX virtual machines in 2023 for electronic design automation workloads.

The HXv2 series will target electronic design automation, scientific simulation, engineering analysis, and distributed-memory computing. Microsoft also identified register-transfer level simulation as a target workload for the new series.

The 800Gb InfiniBand connection is intended to support large-scale Message Passing Interface simulations across distributed computing environments. AMD also uses Azure HX infrastructure for electronic design automation as it develops future EPYC processors and Instinct accelerators, according to AMD Executive Vice-President and CTO Mark Papermaster.

Microsoft has not announced pricing, launch dates, or the Azure regions where HDv2 and HXv2 will initially be available.

AMD and Microsoft are also expanding their work on cloud networking. Azure already uses AMD Pensando data processing units, or DPUs, to handle infrastructure functions separately from a server’s main processors.

Pensando DPUs offload networking, storage, security, and encryption services from host CPUs. Azure Boost also moves selected networking, storage, and virtualisation processing onto dedicated hardware and software.

The expanded agreement will extend AMD’s role in Azure connection processing and backend networking, although the companies have not disclosed how each component will be deployed. Within Helios, UALink-over-Ethernet provides scale-up connectivity among the rack’s GPUs, while Pensando Ethernet components support scale-out networking between racks and clusters.

ROCm provides the software environment used to develop and run workloads on AMD accelerators. Its inclusion gives Helios a common software layer across the rack’s GPU infrastructure.

AMD expects Helios-based systems to enter volume deployment in the second half of 2026. Microsoft has not announced when ND MI455X v7 instances will become available to Azure customers or which regions will receive them first.

 

 

 

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information.

Tech Wire Asia is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.


Source

Visited 1 times, 1 visit(s) today
Share

Recommended For You

Avatar photo

About the Author: News Hound