Microsoft's 20 July Azure announcement is easy to misread as another accelerator launch. It is more useful as a rack-design signal. Azure is separating data preparation, CPU-heavy simulation and production AI inference across three forthcoming AMD-backed VM families rather than pretending that one expensive node suits every stage.
The most eye-catching part is ND MI455X v7, an Azure platform based on AMD's Helios rack reference design. Helios puts 72 Instinct MI455X GPUs in an open rack architecture, alongside next-generation EPYC CPUs, Pensando networking and direct liquid cooling. That density changes the procurement question. Buyers are no longer choosing only a GPU; they are choosing a rack power envelope, cooling method, scale-up fabric, scale-out fabric and operating model.
There is an important timing caveat. Microsoft describes the new Azure instances as upcoming, while AMD says volume Helios deployments are expected in the second half of 2026. Neither source gives a general-availability date or Azure price. The specifications quoted below are vendor-published or projected figures, not GPUMachines benchmark results.
Executive Summary
- What Microsoft announced: forthcoming HDv2 instances for data systems, HXv2 for electronic design automation and high-performance computing, and ND MI455X v7 for large AI inference.
- What Helios is: an AMD reference design for a 72-GPU rack. AMD explicitly says Helios is not itself a product for sale, so actual Azure and OEM implementations can differ.
- Why buyers should care: the design couples high-capacity HBM4, a rack-scale GPU fabric, high-speed external networking and liquid cooling. It shows the physical requirements behind very large-model serving.
- Who should wait: teams that have not measured model memory, request concurrency, latency, power or software compatibility should not jump from a small proof of concept to a 72-GPU rack.
- Where GPUMachines fits: we can compare a private AMD cluster, a smaller PCIe GPU server, an HGX system, hosted ownership and Azure capacity using the same workload assumptions.
What Azure Is Actually Adding
Microsoft's announcement covers three different kinds of infrastructure. Keeping them separate is useful because AI projects often spend heavily on the inference tier while ignoring the CPU and data stages that feed it.
| Azure family | Published platform direction | Intended work | Design question for buyers | |---|---|---|---| | HDv2 | Nearly 500 physical 6th Gen AMD EPYC cores, 4 TB RAM, 32 TB local NVMe and 400 Gb Azure Boost networking | Data systems and large CPU-memory workloads | Can data preparation keep the GPU service supplied without creating a separate bottleneck? | | HXv2 | 176 6th Gen AMD EPYC cores, more than 5 GHz frequency, nearly 2 TB or 4 TB RAM and 800 Gb InfiniBand | EDA and tightly coupled HPC | Does the application gain from high per-core speed, memory capacity and low-latency node communication? | | ND MI455X v7 | AMD Helios-based rack platform with MI455X GPUs | Production-scale AI inference | Can the model, serving runtime and facility make sensible use of a rack-scale accelerator domain? |
These are Microsoft descriptions of forthcoming services. Configuration details may change before availability. The table should be read as an architecture map, not a buying specification.
The split also points to a useful operating pattern. Data engineering can run on a memory-rich CPU tier, engineering simulation can sit on a high-frequency HPC tier, and inference can consume a separate GPU pool. A private design may use different products, but the same division of work often produces a better result than forcing each team onto identical nodes.
Helios Is a Reference Design, Not a Server SKU
AMD describes Helios as an open rack-scale reference design. It is not a chassis that a buyer orders directly from AMD. Cloud providers and system builders can implement the design with their own storage, management, service and rack choices.
That distinction matters. A reference design can define the broad compute and fabric arrangement without settling every deployment detail. A quote still needs to state:
- the exact server and rack implementation;
- supported GPU, CPU, NIC, DPU and firmware versions;
- power feeds and redundancy policy;
- coolant supply temperatures, flow and facility interface;
- service access and component replacement process;
- external network ports and optics;
- local boot, cache and shared storage layout;
- cluster management, telemetry and support ownership.
AMD's published Helios specification centres on 72 Instinct MI455X GPUs, next-generation EPYC processors, Pensando Vulcano AI NICs and a Salina DPU. It uses UALink and UALoE for GPU communication and follows the Open Compute Project's Open Rack Wide direction. This is a rack architecture rather than a conventional collection of independent 4U or 8U GPU servers.
For a buyer, the useful question is not whether a reference diagram looks fast. It is whether an available implementation has a support boundary that covers the complete rack, its cooling loop, firmware matrix and network attachments.
HBM4 Capacity Changes the Model-Serving Conversation
AMD lists 432 GB of HBM4 per MI455X and 31 TB across a 72-GPU Helios rack. It also publishes 19.6 TB/s of memory bandwidth per GPU. These are AMD specifications and projections, not independently measured application results.
The capacity is still significant as a planning signal. Large inference services spend GPU memory on more than model weights. They also need KV cache, runtime workspaces, communication buffers and, depending on the serving engine, graph capture or speculative-decoding components. Long context windows and many concurrent requests can consume memory that a simple parameter-count calculation misses.
A rack with a large aggregate memory pool does not turn 31 TB into one transparent address space. The model and serving runtime must shard work across GPUs, and communication cost still matters. Tensor parallelism, pipeline parallelism, expert parallelism and replica placement each produce a different traffic pattern. A team should test the intended runtime and model topology before treating aggregate HBM as usable capacity.
This is also where a smaller system can win. If a model and its target concurrency fit comfortably on one, two or four GPUs, an independent PCIe server may be easier to schedule, upgrade and maintain. A rack-scale platform earns its place when the workload repeatedly needs the memory domain and communication bandwidth, not when a procurement team wants the largest available number.
Scale-Up and Scale-Out Are Different Problems
Helios uses a GPU scale-up fabric within the rack and high-speed NICs for communication beyond it. AMD publishes up to 260 TB/s of aggregate scale-up bandwidth and 43 TB/s of aggregate scale-out bandwidth for the rack. These are platform figures, not guaranteed model throughput.
Scale-up traffic keeps one distributed model working across accelerators in the same rack. Scale-out traffic connects replicas, pipeline stages, storage or multiple racks. Mixing the two in a planning spreadsheet can hide an expensive mistake. A model that spends much of each step exchanging activations or collective data needs a different fabric plan from a service that runs many independent replicas.
Private-cluster buyers should map the workload before selecting switches:
1. Record the model, precision, parallelism strategy and number of GPUs per serving group. 2. Measure collective communication and application traffic separately. 3. Identify whether the storage network shares links with GPU traffic. 4. Add failure cases, including a NIC, switch or rack taken out of service. 5. Price the complete fabric, including optics, cables, switch ports and management.
A high port-speed label does not answer these questions. Topology, oversubscription, congestion control, rail mapping and software support determine whether the fabric behaves as intended.
CPU Infrastructure Still Matters
The HDv2 and HXv2 announcements deserve attention beside the GPU rack. Training and inference pipelines can lose time in tokenisation, retrieval, decompression, feature preparation, simulation, compilation and storage work. Putting every task on the GPU node may waste accelerator time and make fault isolation harder.
Memory-rich CPU systems can hold large indexes, data frames or caches. High-frequency HPC nodes fit applications whose serial stages do not scale across more cores. Local NVMe can stage data close to compute, while a shared filesystem protects the durable dataset and checkpoint history.
The correct balance depends on evidence. A serving trace may show the GPU waiting on retrieval. An HPC trace may show one core holding back a parallel job. A data pipeline may be network-bound rather than CPU-bound. The Azure split is a reminder to size these stages independently and connect them with a storage and network plan that matches actual transfers.
Power and Liquid Cooling Are Part of the Product
AMD's Helios page describes a direct-liquid-cooled rack with a vertical busbar. That is not a cosmetic packaging choice. A 72-GPU rack belongs in a facility design discussion before it belongs in a purchase order.
The deployment team needs confirmed rack power, feed redundancy, PDU or busbar interfaces, cooling distribution units, water quality requirements, supply and return temperatures, flow, leak detection and heat-rejection capacity. Floor loading, rack access, lifting equipment and service clearances also matter. The final numbers must come from the selected implementation and site survey, not from a reference-design summary.
An organisation without liquid-ready space still has options. It can rent Azure capacity when the service becomes available, place owned hardware with a qualified hosting provider, or choose lower-density air-cooled nodes. Hosted ownership can preserve dedicated hardware and a private operating boundary while moving facility work to a data centre built for it.
GPUMachines can review those routes, but the review starts with site facts. No server specification can make an unsuitable room ready for a dense liquid-cooled rack.
Software Readiness Decides Whether MI455X Is Useful
Hardware capacity has value only when the production stack supports it. AMD positions Helios around its ROCm software platform. Buyers should confirm the exact model architecture, framework, inference engine, kernels, quantisation method and orchestration layer they plan to run.
Do not settle for a generic statement that an application supports AMD GPUs. Ask which release was tested, on which accelerator generation, at what precision and with which communication library. Confirm whether the model's custom operators are present, whether containers have a maintained build path and whether monitoring exposes the metrics the operations team needs.
A sensible evaluation uses the real model and request distribution. Record time to first token, output-token rate, end-to-end latency, useful requests per second, memory use, power and error behaviour. Include long contexts, burst traffic and failure recovery. Vendor peak arithmetic can help narrow a shortlist, but it cannot replace that workload test.
Azure, Private Cluster or Hosted Ownership?
The Microsoft announcement does not make one deployment model universally better. It gives buyers another route to the same class of AMD technology.
Azure capacity suits evaluation, variable demand and teams that want cloud integration without owning a rack. Before committing, model reservation terms, storage, data movement, egress, regional availability and quota. As of the source date, Microsoft has not published general availability or pricing for these new families.
A private cluster can suit steady demand, data-location requirements and organisations with suitable facilities and platform staff. Ownership offers control over scheduling, firmware windows and network policy. It also assigns the buyer responsibility for capacity planning, spares, security and lifecycle work.
Hosted ownership sits between those choices. The customer can own dedicated equipment while a specialist data centre supplies power, cooling and physical operations. Contract terms should define remote access, service response, network ownership, data handling and exit arrangements.
Read our existing Azure ND-series versus private AI cluster comparison before reducing the decision to an hourly rate. Useful comparisons include productive utilisation, software support, egress, facility cost, staff time, service risk and the value of capacity on demand.
Who Should Pay Attention
The Helios direction is relevant to teams serving models that genuinely need large aggregate GPU memory or many tightly connected accelerators. That includes large mixture-of-experts inference, very high-concurrency services, multi-tenant model platforms and research groups testing rack-scale parallelism.
Cloud architects should also watch the CPU tiers. A programme that needs large inference may still spend much of its time preparing data, running simulations or compiling models. Separating those jobs can improve queue behaviour and keep costly accelerators focused on work they perform well.
Infrastructure vendors and data-centre operators should treat the announcement as a planning signal for liquid cooling and open-rack power. The physical design work has a long lead time even when the compute product is not yet generally available.
Who Should Not Build Around Helios Yet
A team running a handful of quantised models for internal users does not need a 72-GPU rack. Neither does a research group whose jobs are independent and fit on separate cards. Those users may get better scheduling freedom from a workstation or PCIe GPU server.
Organisations without a confirmed software path should avoid a large commitment. Port the workload, run a representative test and document missing operators before planning at rack scale.
Buyers with no liquid-cooling plan should not assume one can be added after delivery. Facility work can govern schedule and cost. Cloud or hosted deployment may be the more credible first step.
Finally, anyone who needs capacity immediately should treat both the Azure service and Helios volume timing cautiously. The cited vendor pages describe future availability, not a product that GPUMachines can promise today.
A Practical Evaluation Sequence
Start with the workload, not the rack rendering.
1. Fix the model build. Record architecture, precision, framework, serving engine and maximum supported context. 2. Capture demand. Use real request lengths, output lengths, concurrency and service-level targets. 3. Measure memory. Separate weights, KV cache, runtime overhead and spare capacity for failures or traffic spikes. 4. Choose the parallelism. Decide what runs within one GPU, one server, one rack and across racks. 5. Test the software. Confirm ROCm, communication libraries, containers, observability and model operators on relevant AMD hardware. 6. Design data movement. Size model loading, retrieval, logging and checkpoint paths beside GPU traffic. 7. Survey the facility. Confirm power, cooling, rack, network and service access using the exact implementation. 8. Compare deployment routes. Price Azure, private and hosted capacity with the same utilisation and support assumptions.
This sequence may lead to Helios. It may lead to a smaller AMD server, an NVIDIA platform or a cloud pilot. That is a useful result because the choice follows measured constraints rather than a launch headline.
Our Technical View
The important part of Microsoft's announcement is architectural choice. HDv2, HXv2 and ND MI455X v7 target different bottlenecks, while Helios shows how AMD plans to package dense inference at rack scale.
The strongest Helios argument is its combination of HBM4 capacity and a purpose-built rack communication design. For models that need both, the platform could become a serious alternative in large inference deployments. The weak point for a buyer today is uncertainty. Availability, Azure price, final OEM implementations and production software behaviour still need confirmation.
GPUMachines would not treat AMD's projected peak figures as a procurement case on their own. We would ask for the model topology, latency target, memory trace, software bill of materials and facility details. We can then compare AMD rack-scale infrastructure with flexible PCIe nodes, NVIDIA HGX, Azure rental and dedicated hosted equipment.
FAQ
Is AMD Helios a server that GPUMachines can sell?
Helios is an AMD reference design, not a finished product for direct sale. GPUMachines can review and source available AMD-based server and rack implementations as manufacturers bring them to market, subject to confirmed specification, region and availability.
When will Azure ND MI455X v7 be available?
Microsoft calls it an upcoming offering but does not give a general-availability date in the 20 July announcement. AMD says Helios volume deployments are expected in the second half of 2026. Treat both statements as vendor guidance until a service region, quota and order path are published.
Does 31 TB of HBM4 mean one model can use all of it directly?
No. It is aggregate memory across 72 GPUs. The serving software must partition the model and workload, and communication overhead can affect useful performance. Validate the exact parallelism strategy.
Is Helios only for inference?
AMD presents the architecture as an AI rack-scale platform, while Microsoft specifically positions ND MI455X v7 for production-scale AI inference. Suitability for a given training or research workload depends on software support, topology and measured behaviour.
Do we need InfiniBand for a private MI455X cluster?
Not automatically. The required fabric follows the communication pattern, rack count, storage design and supported software. UALink addresses scale-up communication in Helios; external scale-out design still needs separate sizing and validation.
Can we start smaller and migrate later?
Yes. A smaller AMD PCIe server or rented capacity can expose software and model issues before a rack commitment. Keep containers, model artefacts and observability portable, while accepting that scale-up behaviour must still be retested on the final platform.
Verdict
Azure's AMD expansion is meaningful because it connects three parts of the compute path: data work, high-frequency HPC and rack-scale inference. Helios gives MI455X a design built for very large memory and communication demands, but it also brings liquid cooling, fabric engineering and software qualification into the purchase.
The ideal buyer has a measured large-model service, a clear ROCm test plan and a credible facility or hosting route. Smaller or uncertain workloads should begin with fewer GPUs and earn the move to rack scale through evidence.
Ask GPUMachines to compare AMD Helios-class, PCIe, HGX, cloud and hosted AI infrastructure using your model, traffic and site requirements.
Sources and Further Reading
- Microsoft: Azure AI and HPC infrastructure with AMD (20 July 2026). Primary source for the Azure HDv2, HXv2 and ND MI455X v7 announcement.
- AMD: Helios rack-scale reference design. Primary source for AMD's published Helios architecture, projected specifications and availability guidance.
Vendor figures and roadmap statements can change. GPUMachines has not independently benchmarked Azure ND MI455X v7 or a Helios implementation for this article.
