Microsoft will deploy a new family of Azure virtual machines built on AMD’s Helios rack-scale AI platform in the second half of 2026, the two companies said on July 20. The ND MI455X v7 series, designed specifically for large-scale inference workloads, pairs 72 Instinct MI455X GPUs with AMD’s EPYC “Venice” CPUs and Pensando networking in a single rack. The move marks the first time Microsoft is adopting a complete, integrated AMD hardware stack for AI in Azure, rather than simply offering a GPU-accelerated instance.

What’s Actually Changing

The centerpiece of the expansion is the ND MI455X v7 virtual machine, which Microsoft says is purpose-built for inference—the constant, production-grade serving of AI models behind chatbots, copilots, search, and agentic systems. Each Helios rack contains 72 MI455X accelerators with a combined 31 terabytes of HBM4 memory, delivering up to 2.9 exaFLOPS of FP4 performance and 1.4 exaFLOPS at FP8, according to AMD’s specifications.

But the ND MI455X v7 isn’t just a GPU. It’s a full-stack platform that includes:
- Compute: Sixth-generation EPYC “Venice” host processors
- Networking: AMD Pensando DPUs and AI NICs running Ultra Ethernet Consortium and UALink fabric standards
- Software: AMD’s open-source ROCm stack, with support for PyTorch, TensorFlow, JAX, vLLM, and Triton
- Form factor: An Open Compute Project Open Rack Wide design that Microsoft will integrate into its Azure data centers

Alongside the GPU instances, Microsoft introduced two CPU-only VM families: Azure HDv2 and Azure HXv2. HDv2 targets the data engineering and orchestration work that feeds AI models—data prep, indexing, search, reinforcement learning, and agent coordination. Each VM will offer nearly 500 physical EPYC cores, 4 TB of RAM, 32 TB of local NVMe storage, and 400 Gbps networking. HXv2 is a high-performance computing option with 176 cores clocked above 5 GHz, 2 or 4 TB of memory, 50 percent more addressable cache per core, and 800 Gbps InfiniBand, aimed at electronic design automation, scientific simulation, and large HPC jobs.

What It Means for You

The immediate effect of this announcement is choice, not disruption. If you’re an Azure customer, you won’t see any changes to your existing Nvidia or Maia instances tomorrow. But by late 2026, you’ll have a new class of inference-optimized VMs that could change how you plan capacity and budget for production AI.

For AI Application Developers and Startups

If you build or deploy chat applications, RAG pipelines, or agentic workflows on Azure, the ND MI455X v7 offers a potential alternative to Nvidia H100 or Maia instances. The key advantage is memory capacity: 31 TB of HBM4 per rack means you can run larger reasoning models without sharding across many accelerators, potentially reducing latency and complexity.

The catch is software maturity. AMD’s ROCm stack is open-source and works with popular frameworks, but most production AI code is optimized for CUDA. Your team will need to validate model performance, custom kernels, container images, and observability tooling on the new instances before committing workload. Microsoft says it will support Linux-based containers through Azure Kubernetes Service and Azure Machine Learning; Windows-based development environments (think WSL, VS Code, GitHub tooling) won’t change, but production inference will live in Linux containers.

For IT Infrastructure and Platform Teams

If you manage cloud spending or architect Azure landing zones, this announcement signals where Microsoft is putting its AI infrastructure weight: commodity inference at scale. The ND MI455X v7 isn’t pitched as a training powerhouse—it’s for serving traffic. That suggests Azure wants to be the cheapest, most capacity-rich place to run production AI once models are trained, whether you’re using OpenAI services, open-source models from Hugging Face, or your own fine-tuned weights.

HDv2 and HXv2 fill in the non-GPU parts of the AI stack you actually need: fast CPUs with lots of memory and NVMe storage for the data pipelines, indexing, and HPC simulation work that surrounds GPU clusters.

For Managed Service Providers and Resellers

The AMD partnership doesn’t appear in the Azure price list yet, but it’s worth watching for margins and incentives. Microsoft is increasingly tying Azure growth margins and partner specializations to specific workloads and strategic hardware. The launch of a flagship AMD inference fleet could eventually bring partner incentives around migration, modernization, or new AI workloads—similar to what we’ve seen with Azure Migrate and Modernize programs.

How We Got Here

Three forces pushed Microsoft toward a full-stack AMD commitment:

  1. Inference demand is exploding. Training a model gets headlines, but running it for millions of users every day is what generates Azure revenue. By 2026, agentic AI, retrieval-augmented generation, and real-time reasoning will make inference the dominant AI infrastructure cost. Nvidia has the lion’s share, but Microsoft wants multiple supply chains and cost structures.

  2. Nvidia’s ecosystem is expensive and supply-constrained. Even for a hyperscaler like Microsoft, Nvidia GPUs are costly and hard to get in volume. AMD’s Helios design, with open standards and off-the-shelf networking, could offer a more affordable alternative—especially if AMD prices aggressively to gain cloud footprint.

  3. AMD is finally credible at the rack scale. Helios isn’t a chip; it’s a complete reference platform with Pensando networking, UALink fabric, and a maturing ROCm software stack. Microsoft’s willingness to adopt it as an integrated platform (not just a PCIe card) signals that AMD has reached a threshold of reliability and performance.

Microsoft has been methodically expanding its AMD options for years. The HBv3 and HBv4 HPC VMs used EPYC CPUs with InfiniBand; the Daeafasv7 series, announced earlier, brought general-purpose AMD silicon to Azure. But this is the first time AMD will underpin both the accelerator and the host infrastructure in an AI-targeted VM family.

What to Do Now

For most readers, the ND MI455X v7 is still more than a year away, with volume deployments expected in the second half of 2026 and no public preview announced. That doesn’t mean you should ignore it. Here’s how to prepare:

If You Run Inference Workloads on Azure

  • Add AMD to your capacity planning models. Even if you don’t switch, the arrival of a credible alternative can improve Nvidia pricing and availability. Talk to your Microsoft account team about planned regions and performance targets.
  • Start a low-touch ROCm evaluation. Grab an AMD-powered VM (the existing HBv4 series runs on EPYC and can run ROCm in software mode for lightweight testing). Port a few representative inference scripts—vLLM or PyTorch with Hugging Face models—and identify any gaps in operators or performance.
  • Watch for Azure Machine Learning and AKS support. Microsoft will likely release early-access container images and documentation before the VMs go public. Join the Azure AI preview programs and enable notifications for new GPU instance types.

If You Build Data Pipelines or HPC Simulations

  • HDv2 and HXv2 may arrive first. CPU-only VMs don’t require the same complex bring-up as a GPU platform. If your team works on EDA, computational fluid dynamics, or large-scale search indexing, ask your Microsoft representative about timelines for these families. Their specs—500 cores, 32 TB of local NVMe—could compress pipelines that today run on memory-constrained instances.

If You Manage Azure Costs

  • Don’t budget for AMD instances yet. Without pricing, you can’t model total cost of ownership. But you can benchmark your current inference cost per million tokens and compare it against public Nvidia-based numbers. When AMD pricing arrives, you’ll have a baseline to judge if the switch saves money.

Outlook

The next milestone isn’t a rack diagram—it’s an Azure region list, a public preview date, and a price sheet. Until those arrive, treat the ND MI455X v7 as a credible future option rather than an available tool. Microsoft’s track record with AMD HPC VMs suggests the launch will be technically solid but initially limited to a handful of regions before scaling out. If AMD can deliver competitive performance and if ROCm proves truly interchangeable with CUDA for mainstream inference models, this partnership could reshape Azure’s AI infrastructure economics for the rest of the decade. If not, it’ll remain a niche alternative for customers willing to invest in software porting. Either way, the message from Microsoft is clear: inference is too important to leave to one supplier.