Microsoft will deploy AMD’s new Helios rack-scale AI systems in Azure data centers beginning in the second half of 2026, the companies announced Monday, marking the most significant challenge yet to Nvidia’s dominance in cloud AI infrastructure. Helios combines 72 Instinct MI455X GPUs, 18 next-generation EPYC “Venice” CPUs, Pensando networking, and the ROCm software stack into a single liquid-cooled rack—AMD’s most ambitious data center product to date.

The Deal: Microsoft Commits to Rack-Scale AMD

Forrest Norrod, head of AMD’s data center division, called Helios “our baby.” The system is named after the Greek sun god, and the ambition is fitting: to bring daylight to a market where Nvidia holds a 95% share of data center GPUs. Microsoft’s endorsement instantly makes Helios credible, as Azure engineers must integrate these 7,000-pound racks into real data centers, expose their capabilities through cloud services, and support customer workloads under strict availability requirements.

Microsoft joins Meta, OpenAI, and Oracle as early adopters of Helios. But the Azure deal is especially notable because Microsoft was also the first large-scale user of AMD’s previous-generation MI300X accelerator. That earlier deployment gave the software teams hands-on experience with ROCm, AMD’s open-source GPU programming platform, before committing to a deeply integrated rack-scale design.

Inside Helios: Hardware and Software at a Glance

Helios isn’t a conventional server with several GPUs slapped in. It’s a pre-validated, fully liquid-cooled rack designed to operate as one computational unit. The core components:

  • 72 Instinct MI455X GPUs – Built for training and frontier inference, with high-bandwidth memory to handle massive parameter counts and long context windows. The memory capacity is expected to be particularly attractive for serving large language models without excessive model sharding.
  • 18 EPYC Venice CPUs – Sixth-generation EPYC, based on Zen 6, providing host compute for data preparation, security, and non-accelerated portions of AI pipelines. These chips will also power two new Azure VM families (more on that below).
  • Pensando Vulcano networking – 800Gbps-class scale-out fabric, plus UALink for direct GPU-to-GPU communication inside the rack. UALink is an open standard, contrasting with Nvidia’s proprietary NVLink.
  • ROCm software – AMD’s alternative to CUDA. While support for PyTorch and inference engines has improved, enterprises still report more setup hurdles than with Nvidia’s mature ecosystem.

The rack’s physical bulk—it weighs up to 7,000 pounds and requires liquid cooling—means it won’t fit into traditional enterprise server rooms. Microsoft, however, already builds liquid-cooled data centers for Azure, and has the muscle to deploy them at scale.

What Helios Means for Azure Customers

For most organizations, Helios won’t arrive as a physical rack in their own facility. It will show up as new Azure virtual machine families and managed AI services. Microsoft hasn’t disclosed the exact VM series names, but it confirmed two new Venice-based compute instances:
- Agentic AI and data pipeline VMs – Up to 500 CPU cores, 4TB of memory, 32TB local NVMe storage, and 400Gbps Azure Boost networking. Ideal for the complex, chained workflows that agentic AI demands.
- HPC and semiconductor design VMs – Optimized for simulation, electronic design automation, and other high-throughput computing.

These will complement GPU-heavy instances that leverage the MI455X accelerators. This gives Azure customers a broader menu of options:
- Inference-heavy workloads may see significant cost-per-token improvements on AMD-backed VMs, especially for models that require large memory pools.
- Training clusters will likely remain on Nvidia instances initially, but IT teams can experiment with AMD for select jobs to reduce cost or avoid capacity shortages.
- Developers should start testing their PyTorch or TensorFlow code on ROCm now (Azure already offers MI300X VMs) to gauge portability. Microsoft is likely to abstract away much of the complexity through managed AI services, but custom CUDA kernels will need rewriting.

For home users and small businesses, the effect will be indirect. Services like Microsoft 365 Copilot, Bing AI, and GitHub Copilot run on Azure infrastructure; more accelerator choice could improve service reliability and potentially slow the pace of price increases as demand grows. But don’t expect an overnight transformation—this is a gradual infrastructure play.

How AMD Got Here: From Zen to Helios

AMD’s data center renaissance is one of tech’s most dramatic turnarounds. A decade ago, the company was fighting for relevance in both PCs and servers. Three strategic moves changed the game:
1. Zen architecture (2017) – EPYC CPUs won over cloud providers with high core counts and energy efficiency, building the trust and deployment pipelines that Helios now exploits.
2. Instinct accelerators – The MI300X, adopted early by Microsoft, proved AMD could deliver competitive AI silicon. It gave Microsoft practical experience with ROCm’s strengths and weaknesses at scale.
3. Acquisitions – Buying Xilinx (adaptive computing) and Pensando (networking) gave AMD the pieces to build a full-stack system. The later purchase of ZT Systems’ manufacturing arm added integration expertise.

Helios is the culmination. Instead of selling discrete GPUs, AMD now offers a complete, pre-validated rack that addresses connectivity, cooling, and software—the same “system” approach Nvidia used to corner the market.

What IT Leaders Should Do Now

If your organization uses Azure for AI, take these steps to prepare for Helios-powered VMs:

For IT admins and architects

  1. Audit your AI workloads – Identify services that rely on Nvidia-specific libraries (cuDNN, custom CUDA kernels). These will need the most work to move.
  2. Test AMD’s existing MI300X VMs – Run benchmarks on realistic batch sizes and model architectures. Compare performance, stability, and ease of management with your current Nvidia instances.
  3. Monitor Azure’s regional rollout plans – Helios is likely to land first in liquid-cooled regions (US East, US West, North Europe). Plan capacity accordingly.
  4. Cost-model your inference spending – If AMD instances offer even 20–30% lower per-token costs, a partial migration could free budget for other projects.

For developers and data scientists

  1. Containerize workloads – Use Docker images that support both CUDA and ROCm runtimes. Tag images clearly by accelerator type.
  2. Validate model accuracy – Numerical differences between GPU architectures can creep in. Run regression tests to ensure quality doesn’t degrade.
  3. Learn ROCm debugging tools – Familiarize yourself with AMD’s ROCProfiler and ROCgdb. The community is smaller than CUDA’s, so build internal expertise early.

For business decision-makers

  • Don’t wait on hardware parity – Helios may excel at inference before it challenges Nvidia on training. Target your highest-volume inference services first.
  • Negotiate commitments with flexibility – Azure is likely to offer introductory pricing. Structure contracts so you can shift workloads between AMD and Nvidia as the market evolves.

If you’re not on Azure, the news still matters: cloud competition drives down costs industry-wide. Watch for similar moves from Google Cloud and AWS—both are EPYC customers and may follow suit with rack-scale AMD deployments.

Can AMD Really Rival Nvidia?

Nvidia still commands over 95% of the data center GPU market, and its integrated rack systems like Blackwell NVL72 are battle-tested. CUDA’s vast library of tools and trained developers remains a formidable moat. AMD must prove that ROCm can handle production AI without adding months of integration work.

Timing is another risk. Helios ships in the second half of 2026; by then, Nvidia’s Vera Rubin platform will be ramping up. Unofficial pricing estimates place Helios at $5–5.5 million per rack, compared to $3.5–4 million for Vera Rubin—though such figures are negotiated and depend heavily on configuration. The real differentiator, AMD argues, will be total cost of ownership: lower per-token costs driven by memory efficiency and open networking standards that avoid vendor lock-in.

Analysts see Helios as a potential catalyst for AMD’s AI market share to grow from its current 4.5% to as much as 20–25%. Microsoft’s deployment is the first major proof point. If early production racks perform well, other hyperscalers may accelerate their own AMD adoption.

When Will You Actually See Helios?

Don’t expect a flip of a switch. “Customer shipments” start in the second half of 2026, but initial racks will likely go to internal Microsoft teams and a small set of premier customers for validation. Broad availability across Azure regions may not come until 2027. Microsoft also needs time to integrate the hardware with its managed AI services—Azure OpenAI Service, for example, may not tap Helios until well after installation.

The real test will be performance-per-dollar under sustained, mixed workloads—not benchmark peaks. Independent evaluations will be crucial. For now, the message is clear: Azure’s AI infrastructure is getting a second engine, and that’s good news for anyone who relies on it.