The first trap in comparing AMD Helios with NVIDIA Vera Rubin NVL72 is treating both names as purchasable boxes. Vera Rubin NVL72 is NVIDIA's tightly defined rack-scale system platform. Helios is an AMD reference design that OEM and ODM partners use to build products around the Open Rack Wide form factor.
That difference changes procurement before anyone compares accelerator arithmetic.
The buyer's question is: does the organisation want a tightly co-designed NVIDIA rack and software stack, or an open AMD rack blueprint whose final product, support model and integration depend more heavily on the delivery partner? Neither answer is inherently superior. The better choice is the one the team can qualify, power, cool, operate and keep productive with its actual models.
AMD's MI400 material describes Helios with 72 Instinct MI455X GPUs and 31 TB of HBM4. NVIDIA describes Vera Rubin NVL72 with 72 Rubin GPUs, 36 Vera CPUs, ConnectX-9 SuperNICs, BlueField-4 DPUs and an NVLink 6 scale-up fabric. These are vendor specifications for forthcoming 2026 platforms. GPUMachines has not independently benchmarked production systems, and this article does not convert vendor claims into a cross-platform performance ranking.
The short buying answer
Choose the Helios path when open rack standards, UALink, Ultra Ethernet, ROCm and multi-vendor system integration are strategic requirements, and the organisation is prepared to qualify the exact partner-built implementation.
Choose Vera Rubin NVL72 when CUDA software coverage, NVIDIA's integrated rack architecture, NVLink scale-up behaviour and a more prescribed platform are worth accepting a tighter vendor ecosystem.
Choose neither yet when the facility cannot support dense liquid-cooled racks, the workload has not been measured at rack scale, software qualification is incomplete, or the commercial case can be met by B300, MI350-class or smaller PCIe systems.
The result should come from model completion time, accuracy, availability, power, operational labour and supplier obligations. Peak low-precision figures alone cannot answer the purchase.
What AMD Helios actually is
AMD calls Helios its first rack-scale AI reference design. It is based on Meta's Open Rack Wide design submitted to the Open Compute Project. The double-wide format creates room for high-density trays, power delivery and liquid cooling.
AMD's published design combines:
- 72 Instinct MI455X GPUs;
- EPYC server CPUs;
- Pensando DPUs and networking;
- UALink for scale-up connectivity;
- Ultra Ethernet for scale-out networking;
- the ROCm software stack;
- 31 TB of HBM4 across the rack, according to AMD.
AMD explicitly says Helios is a reference design, not a product it sells directly. OEM and ODM partners determine the final system, support package, serviceability details, firmware matrix and commercial availability. AMD says partner volume deployments are expected in the second half of 2026.
That openness is valuable, but it moves work into procurement. A buyer must name the partner product, not merely “one Helios”. Two Helios-derived systems can follow the same architecture and still differ in power shelves, coolant distribution, management, cabling, service access, warranty and software image.
What Vera Rubin NVL72 actually is
Vera Rubin NVL72 is NVIDIA's next rack-scale compute engine after Blackwell NVL72. NVIDIA describes one rack with:
- 72 Rubin GPUs;
- 36 Vera CPUs;
- 18 compute trays;
- nine NVLink switch trays;
- ConnectX-9 SuperNICs for scale-out traffic;
- BlueField-4 DPUs for infrastructure services;
- an NVLink 6 copper spine that joins the GPUs into a rack-scale domain.
NVIDIA states 3.6 TB/s of NVLink bandwidth per GPU and 260 TB/s across the rack. These are vendor platform specifications. Delivered application performance depends on software, model partitioning, communication pattern and system condition.
The NVL72 proposition is integration. NVIDIA designs CPU, GPU, scale-up fabric, scale-out endpoint, DPU and software as one platform. The buyer still chooses an OEM system and support route, but the architecture and software expectations are more prescribed than in an open reference design.
For a general explanation of the system, read NVIDIA Vera Rubin NVL72 explained. The purpose here is to decide between platform strategies.
The architectural comparison that matters
| Decision area | AMD Helios | NVIDIA Vera Rubin NVL72 | |---|---|---| | Commercial form | Reference design implemented by partners | NVIDIA-defined rack-scale system platform delivered through partners | | GPU count | 72 MI455X in AMD's design | 72 Rubin GPUs | | Published HBM4 capacity | 31 TB per Helios rack | Confirm against final Rubin product configuration | | Scale-up direction | UALink and open rack standards | NVLink 6 integrated scale-up domain | | Scale-out direction | Ultra Ethernet with Pensando networking | ConnectX-9 with Spectrum-X Ethernet or Quantum InfiniBand options | | Infrastructure processing | Pensando DPUs | BlueField-4 DPUs | | Main software ecosystem | ROCm | CUDA and CUDA-X | | Buyer burden | Qualify the partner implementation and open-stack integration | Qualify a tightly integrated NVIDIA platform and its ecosystem dependencies |
The table avoids a synthetic “winner” column. The two platforms have not been tested by GPUMachines on the same production models, software and power envelope. A responsible comparison keeps architectural facts separate from performance claims.
Memory capacity: useful, but not a verdict
AMD's 31 TB HBM4 figure is a material part of the Helios proposition. More accelerator memory can reduce model partitioning, hold larger expert sets or KV caches, and support memory-heavy scientific workloads.
Capacity at rack level is not automatically a flat pool available to one process. The scale-up topology, software communication model and memory placement determine what the application can use efficiently. A model that fits across 72 devices can still spend too much time moving data between them.
Ask both suppliers for the memory available per accelerator, the practical collective bandwidth for the intended parallelism strategy, and the behaviour when a device or link is degraded. Measure the exact model at the required context length and precision.
For inference, include KV-cache growth and concurrency. For training, include optimizer state, activation checkpointing and restart behaviour. For HPC, include FP64 needs, communication pattern and host-memory movement.
Do not let a rack-total number replace a placement plan.
ROCm versus CUDA is a workload question
CUDA has broad support across AI frameworks, libraries and commercial software. Many teams already have CUDA kernels, containers, monitoring and staff experience. Vera Rubin extends that environment into a new platform generation.
ROCm has matured substantially and supports common frameworks including PyTorch, JAX, TensorFlow, vLLM and Triton according to AMD's current material. Helios also appeals to organisations that want open APIs and less dependence on one scale-up and networking stack.
Software compatibility is not binary. A framework can launch successfully while a required operation falls back, a model uses an inefficient kernel, or an observability tool lacks support. The acceptance test should include:
- the exact model commit and tokenizer;
- training, fine-tuning or serving framework version;
- custom kernels and extensions;
- quantisation and low-precision path;
- collective library and topology settings;
- checkpoint, restart and fault-recovery workflow;
- profiler, scheduler and monitoring integrations;
- container build and security scanning.
Porting cost belongs in the commercial comparison. So does the long-term value of portability. A team with a mature CUDA estate may rationally pay for continuity. A new national or research programme may value open standards enough to invest in ROCm qualification.
Scale-up openness versus integrated behaviour
UALink is intended to provide an open scale-up interconnect for accelerators. That can broaden the supplier ecosystem and reduce dependence on a proprietary fabric. Helios uses this direction as part of an open rack design.
NVLink is a mature NVIDIA scale-up technology tied closely to NVIDIA GPUs and software. Vera Rubin NVL72 uses a defined NVLink 6 spine so the rack behaves as one large accelerator domain for suitable workloads.
Openness does not remove integration work. Connector, switch, firmware, topology, collective software and failure recovery still have to work as one system. Conversely, integration does not guarantee a workload will scale efficiently. Parallelism strategy and software still matter.
Buyers should request topology diagrams and measured collective performance across the full rack. The test must include degraded operation, because a scale-up fabric is only useful if the system can recover predictably from component failure.
Scale-out networking after the first rack
Neither platform ends at one rack. Multi-rack training and large inference services need a scale-out fabric that matches endpoint bandwidth and traffic patterns.
AMD aligns Helios with Ultra Ethernet and Pensando networking. This supports the open-stack story and can appeal to buyers who want Ethernet operations and multi-vendor standards. NVIDIA offers Spectrum-X Ethernet and Quantum InfiniBand, with ConnectX-9 endpoints in Vera Rubin.
The practical work is port arithmetic, rail mapping, switch tiers, optics, cable length, oversubscription and congestion control. A nominal port speed says little about collective goodput when many jobs share the fabric.
The Spectrum-X and InfiniBand buying comparison explains the wider trade-off. For current Spectrum architecture, see NVIDIA Spectrum-6 AI Ethernet.
Ask vendors to run distributed jobs across the proposed rack count, not one tray or one rack. Record p99 iteration time, collective bandwidth, retransmission, congestion signals and job completion under a failed link.
Facility fit can decide before software does
Both platforms belong in purpose-built data-centre environments. The final power figure, rack width, coolant requirements and service clearances must come from the shipping OEM design.
Helios uses a double-wide Open Rack Wide format. That may fit facilities designed around open hyperscale racks, but it is not a drop-in replacement for every standard 19-inch rack row. Vera Rubin NVL72 uses NVIDIA's single-wide MGX rack direction, yet it is still a dense liquid-cooled system with substantial weight and power demands.
Before approving either platform, verify:
- delivered rack dimensions and loaded weight;
- floor loading and transport route;
- facility power at the required voltage;
- branch circuit and busway arrangement;
- coolant supply temperature, flow, pressure and water quality;
- coolant distribution unit redundancy;
- heat-rejection capacity during degraded operation;
- front, rear and side service clearances;
- leak detection and emergency procedure;
- network and management cable routes;
- spare tray handling and remote-hands capability.
A platform that wins a model benchmark but cannot be maintained in the selected site is not the better system.
The AI cluster power requirements guide provides a starting point for facility questions. Final engineering must use the OEM's current site-planning documents.
Who should favour Helios
Helios deserves serious evaluation when an organisation has a strategic commitment to open rack and interconnect standards, wants an AMD accelerator roadmap, or needs a route away from an exclusively CUDA estate.
It may suit hyperscalers, sovereign AI programmes, advanced research centres and operators with enough engineering staff to qualify a partner implementation. The large published HBM4 capacity can be relevant to memory-heavy training, inference and scientific computing.
The buyer should be comfortable treating the OEM or ODM as a major part of the solution. Support boundaries between AMD, the system builder, networking vendors and software providers must be written into the contract.
Who should favour Vera Rubin NVL72
Vera Rubin is attractive to organisations with a mature NVIDIA software estate, a need for the tightly integrated NVLink domain, or a preference for one prescribed platform architecture.
It may reduce some integration choices because NVIDIA defines CPU, GPU, DPU, SuperNIC, scale-up and software relationships together. It also offers continuity for teams already operating DGX, HGX or Blackwell NVL systems.
That continuity has a cost. The buyer becomes more dependent on NVIDIA's roadmap, software and networking integration. Procurement should account for licences, support, upgrade paths and the ability to move workloads elsewhere.
Who should buy neither in the first phase
A team with an unproven application should not begin at a 72-GPU rack. Use smaller systems or cloud capacity to establish model compatibility, scaling behaviour and useful utilisation.
An organisation without liquid-cooling operations should not treat commissioning as an accessory. Facility work can dominate schedule and risk.
A workload that fits on eight GPUs and has modest growth may be cheaper and easier to operate on HGX or PCIe servers. Spare capacity in a rack-scale system is not free simply because it may be used later.
A mixed research workload can also prefer flexible PCIe GPU servers. Rack-scale systems are strongest when the work can exploit their interconnect and runs often enough to justify the estate.
GPUMachines Buy & Host can help when ownership is desirable but the local site is not ready for high-density liquid-cooled infrastructure.
An RFQ-ready comparison test
Require both proposals to answer the same workload and operational questions.
1. Freeze the model and software bill. Name commits, frameworks, kernels, precision, container images and driver versions. 2. Run the full job. Measure time to train, fine-tune or serve the target workload. Do not accept extrapolation from a small layer unless it is clearly labelled. 3. Report accuracy and convergence. Low-precision throughput is useful only when the required output quality is preserved. 4. Measure useful energy. Record completed valid work per kilowatt-hour at the rack input, including CPUs, networking and cooling allocation where measurable. 5. Test scale-up. Run the intended tensor, pipeline, expert or data parallelism across the rack. Record collective behaviour and idle time. 6. Test scale-out. Repeat across the planned rack count with representative background traffic. 7. Break components. Remove a link, GPU tray and service. Observe job survival, checkpoint recovery and repair procedure. 8. Validate the software estate. Run profilers, scheduler, monitoring, security scanning, backup and incident tooling. 9. Inspect serviceability. Time a documented tray replacement with the proposed remote-hands and spare strategy. 10. Price the operating term. Include facility work, support, licences, port optics, spares, engineering and expected upgrade labour.
Keep raw results and configuration files. A vendor chart without the exact software and power conditions cannot settle the comparison.
Contract questions for Helios proposals
Because Helios is a reference design, the contract should identify the actual manufacturer and configuration. Ask who owns firmware qualification, UALink and UEC integration, coolant components, rack controls and first-line support.
Require one escalation path for faults that cross GPU, DPU, switch and system boundaries. Define spare availability and the period for which the partner will maintain the validated software matrix.
Confirm which Helios elements are fixed by AMD's design and which are partner choices. Record any deviation from the reference architecture.
Contract questions for Vera Rubin proposals
Ask which parts are NVIDIA reference components and which are OEM-specific. Confirm software subscriptions, support levels, allowed firmware combinations and the process for receiving critical fixes.
Document the scale-out choice, switch software and optical bill. Confirm whether a proposed rack can join future systems without replacing the first fabric.
Ask how the supplier handles failed trays, NVLink components, coolant hardware and BlueField services. Integrated architecture needs integrated support rather than several disconnected warranties.
Questions buyers ask
Is AMD Helios a server I can order from AMD?
No. AMD describes Helios as a reference design for OEM and ODM partners. The order will be for a partner-built implementation with its own configuration and support terms.
Does 31 TB of HBM4 make Helios faster than Vera Rubin?
No general conclusion follows. Capacity can help memory-heavy workloads, but delivered performance depends on model fit, interconnect, kernels, precision, parallelism and software maturity.
Is UALink automatically more flexible than NVLink?
UALink's open standard can support a broader ecosystem. The delivered Helios system still needs qualified hardware, firmware and software. NVLink offers a tightly integrated NVIDIA path. Flexibility and integration should be tested as operational outcomes.
Can existing CUDA workloads move directly to Helios?
Not necessarily. Common frameworks may support ROCm, but custom CUDA code, extensions and operations need qualification or porting. Measure that work before setting a migration date.
Which platform is better for scientific computing?
It depends on FP64 needs, memory footprint, communication pattern and software libraries. MI455X is aimed mainly at frontier AI; AMD also positions MI430X for converged HPC and AI. The exact scientific application should be tested.
Can GPUMachines source both platform directions?
GPUMachines can work with OEM, distributor and hosting channels to specify suitable AMD or NVIDIA rack-scale infrastructure, subject to product availability and partner terms. Final quotations must name the shipping system, not only the reference architecture.
Sources and Further Reading
- AMD Instinct MI400 Series and Helios
- AMD CDNA architecture
- AMD Advancing AI 2026 platform announcement
- AMD: Azure Expands AI Infrastructure Choice with Helios
- AMD: AI Networking Built for Scale
- NVIDIA: Vera Rubin Pod Architecture
- NVIDIA: Inside the Vera Rubin Platform
- Hot Chips 2026 conference programme
Verdict
Helios and Vera Rubin NVL72 express two different procurement philosophies. Helios uses an open reference design and asks the delivery ecosystem to turn it into a supported product. Vera Rubin uses NVIDIA's tighter co-design and asks the buyer to commit more deeply to one integrated platform.
Choose Helios when open standards, ROCm, memory capacity and partner-level customisation are strategic and the team can qualify the full implementation. Choose Vera Rubin when CUDA continuity, NVLink integration and a prescribed rack architecture carry more value. Delay both when the workload or facility evidence is not ready.
GPUMachines can compare the final OEM proposals, map rack and fabric requirements, and define a workload acceptance test before the order is fixed. Use the GPU Cluster Configurator to expose power, networking and rack assumptions, then require each supplier to prove the same useful workload.
