Three 800Gb links per GPU sound tidy until somebody counts the switch ports. AMD's Pensando Vulcano 800 AI NIC can support that arrangement, giving the accelerator up to 2.4Tbps of published scale-out bandwidth. The number describes an endpoint design, not application throughput, and it carries a large physical bill: adapters, planes, optics, cables, switch radix, power and software all have to agree.
AMD presented Vulcano 800 in detail on 23 July 2026 as the next AI NIC for its Helios rack-scale platform and other large GPU clusters. The company describes an 800Gbps adapter with PCIe 6.0 and UALink attachment options, programmable transport and congestion functions, support for multi-plane back-end configurations, and Multipath Reliable Connection (MRC).
This is a timely product because 400GbE is no longer an exotic cluster link, while 800GbE and 1.6TbE are moving into new rack designs. It is also easy to misread. Installing one faster NIC doesn't create a balanced AI fabric, and installing three per GPU can multiply cost faster than useful throughput. GPUMachines hasn't tested Vulcano 800. AMD's job-completion and switching-cost comparisons rely on its own simulations and stated assumptions.
Executive Summary
- What it is: an 800Gbps programmable Ethernet AI NIC for scale-out and what AMD calls scale-across networking.
- Headline configuration: up to three NICs per GPU, giving 3 x 800Gbps = 2.4Tbps of nameplate scale-out bandwidth per accelerator.
- Why it matters: more endpoint bandwidth can reduce communication stalls in large collective-heavy training and distributed inference, but only when switches and topology carry the traffic.
- What buyers must verify: physical NIC count, attachment mode, breakout, rail mapping, transport profile, switch support, optics and failure behaviour.
- When it is too much: small clusters, independent inference replicas and environments where storage or software already limits GPU use.
GPUMachines can plan an Ethernet GPU cluster or build the endpoint, switch and rack arithmetic through the GPU Cluster Configurator.
What AMD Announced
Vulcano follows the 400Gbps Pensando Pollara AI NIC. AMD says the new adapter supports 800Gbps throughput and can attach through PCIe 6.0 or UALink, depending on the system. The public Helios material places it in four-GPU compute trays and uses Vulcano for rack scale-out.
AMD also says the adapter can present back-end configurations such as 4 x 200Gbps or 8 x 100Gbps. Those modes affect the cable and switch design. One 800Gbps adapter may therefore produce several logical network legs; the connector, cable assembly and switch-side breakout must come from the qualified option list.
The transport story includes MRC, a multipath reliable transport developed with contributions from AMD, OpenAI and other industry participants. AMD describes support for SRv6 path control as well as ECMP or dynamic load-balancing approaches. Its programmable P4 engines allow transport and congestion behaviour to change through software.
Vulcano also sits next to Ultra Ethernet rather than replacing it. The Ultra Ethernet Consortium's UET specification defines a wider communications stack for AI and HPC, including multipath operation, congestion control, delivery services, telemetry and security. AMD markets its AI networking as UEC-ready, but a buyer should ask which UEC profile, software release and compliance status apply to the quoted system. MRC and UET shouldn't be treated as interchangeable names.
The 2.4Tbps Arithmetic
The calculation is simple:
3 NICs x 800Gbps = 2.4Tbps per GPU
The procurement consequences aren't.
Suppose a system design genuinely assigns three 800Gbps NICs to each of four GPUs. That creates twelve 800Gbps endpoint links and 9.6Tbps of summed nameplate bandwidth for the four-GPU group. If every adapter breaks into four 200Gbps legs, the logical link count becomes 48. This example is arithmetic, not a claim about a particular server.
Each endpoint needs a supported termination. A non-blocking design must provide enough leaf downlinks and uplinks for the chosen planes. Dual- or triple-plane layouts divide the links across separate switches so one failure doesn't remove every path, but they also create more devices to configure and monitor.
Nameplate bandwidth is not payload rate. Protocol overhead, message size, congestion, collective behaviour, PCIe or UALink attachment and application synchronisation all reduce what the job observes. The slowest rank can delay a collective even when average link use looks healthy.
Any proposal using "2.4Tbps per GPU" should therefore include an endpoint table:
| Field | What must be recorded | |---|---| | GPU association | Which accelerator or compute tray the NIC serves | | Physical adapter | Exact Vulcano SKU and form factor | | Host attachment | PCIe 6.0, UALink or another qualified interface | | Network presentation | 800G, 4 x 200G, 8 x 100G or supported alternative | | Plane or rail | The independent fabric path carrying the link | | Switch port | Exact switch, cage, breakout and speed | | Media | DAC, active copper, optical module or cable assembly | | Software | Driver, firmware, transport and congestion profile |
If one of those cells is unknown, the design isn't ready for a bill of materials.
Scale-Out and Scale-Across Are Different Problems
Scale-out connects GPU nodes or racks inside a cluster. The fabric carries collectives, remote memory operations, model shards, control traffic and sometimes storage. Distance is usually limited to a data hall or campus, so latency can remain low enough for tightly coupled work.
AMD uses "scale-across" for extending connectivity between data centres. A programmable NIC can provide path control, telemetry and recovery features across that boundary, but physics still applies. Fibre distance and intermediate networks add latency. Synchronous training collectives that perform well within one site may stall across a long-haul link.
Cross-site designs often suit asynchronous replication, checkpoint movement, model distribution, service failover or loosely coupled inference. They can support some distributed jobs when sites are close and the network is engineered for them. Buyers should not read "scale-across" as proof that two distant GPU clusters become one low-latency machine.
Ask AMD or the system supplier for measured distance, latency, failure and workload limits for the proposed deployment. A transport feature isn't a WAN service.
Multi-Plane Fabrics Need Physical Independence
AMD describes Vulcano with a multi-plane architecture. Several independent paths can improve bandwidth and allow traffic to continue when a link, adapter or switch fails.
Logical planes only help when their dependencies are separate. Two switch planes fed from one power circuit, one cooling loop or one fibre tray can fail together. Rail diagrams need to show rack location, PSU feed, cooling dependency, patch panel and upstream path.
For every GPU or tray, confirm that endpoint links distribute evenly across the intended planes. Then remove one plane during testing and measure:
- Remaining application throughput.
- Collective completion time and tail behaviour.
- Route convergence and retransmission.
- Jobs that fail rather than slow down.
- Operator steps needed to identify the fault.
- Whether storage and management access remain available.
A fabric can be redundant on paper yet unusable during maintenance because the remaining plane is oversubscribed. Failure-state oversubscription should be stated alongside normal-state ratios.
What AMD's 13% Claim Actually Measures
AMD says Vulcano can complete AI jobs up to 13% faster and reduce switching costs by up to 33%. The footnotes matter more than the headline.
The 13% figure comes from AMD engineering-silicon modelling and synthetic simulation of an 8,000-GPU MI455X system. It compares three Vulcano NICs with two for an FP8 mixture-of-experts training case. AMD says the model covers decode layers, includes collective costs, and excludes evaluation and checkpointing. It also assumes ideal behaviour that may differ from released products.
That is evidence that a third network path can help a communication-sensitive simulated workload. It isn't evidence that every cluster will gain 13%, nor that a smaller estate should buy 50% more adapters.
The switching-cost claim also depends on AMD's chosen multi-plane comparison and assumptions about cables and transceivers. A buyer needs the full port and media list for both designs. Compare the same oversubscription, resilience and reach; otherwise a cheaper topology may simply provide less.
GPUMachines would keep both figures in the test plan, not the financial model. The acceptance target should use the buyer's training or inference workload.
MRC, UET and Existing RoCE
AI Ethernet is no longer one transport choice. Many current clusters run RoCEv2 with ECN, PFC and vendor-specific congestion controls. Ultra Ethernet defines UET as a newer transport aimed at multipath operation, scale and varied AI/HPC delivery semantics. AMD's MRC work addresses reliable multipath transport and can use SRv6 or familiar load-balancing routes.
The names are less important than the deployed software path:
1. Which collective library and framework release supports the transport? 2. Does the adapter require a particular kernel, driver or firmware? 3. Which switch telemetry and congestion features are mandatory? 4. Can the transport share a fabric with storage or tenant traffic? 5. How are keys, isolation and policy handled? 6. What happens during mixed-version upgrades? 7. Who supports a fault that crosses NIC, switch and framework boundaries?
Existing RoCE knowledge remains useful, but operators shouldn't assume every UET or MRC control maps directly to a RoCE setting. Runbooks and observability need to follow the installed stack.
The GPUMachines RoCE versus InfiniBand comparison covers the established fabric choice. Vulcano adds another Ethernet endpoint option; it doesn't make that comparison obsolete.
PCIe 6.0 and Direct GPU Attachment
An 800Gbps full-duplex adapter demands a substantial host interface. PCIe 6.0 x16 provides enough theoretical bandwidth for one 800Gbps network link, but actual throughput depends on encoding, protocol overhead and implementation. If the NIC falls back to a narrower or older link, the host attachment can cap it.
Direct GPU attachment through UALink can reduce reliance on host paths in supported systems. It also ties the NIC to a specific platform design, firmware and service model. A Vulcano card listed as UALink-capable isn't proof that it can be installed beside any GPU.
Request the server block diagram. Confirm lane width, NUMA ownership, retimers, PCIe switches, direct-attachment mode and peer-to-peer support. Then measure GPU memory traffic over the selected path rather than relying on a bus calculation.
Adapter cooling matters too. Several 800Gb NICs per accelerator add heat near GPUs, memory, cables and liquid connections. The chassis needs validated airflow or cold-plate coverage for the full option set.
Switch Radix, Optics and Cable Count
Endpoint bandwidth consumes switch radix rapidly. A 64-port 800GbE leaf can attach 64 single-link endpoints if no ports are reserved for uplinks. A non-blocking leaf-spine arrangement commonly divides capacity between downlinks and uplinks, which lowers endpoint count per leaf. Exact ratios depend on topology.
Breakout changes the logical count but not the switch silicon budget. Four 200Gb lanes still consume 800Gbps of the parent port. They may use one fan-out cable or several optical legs, depending on the connector and reach.
Build the media bill separately:
- Adapter-side cages or cable interfaces.
- Switch-side cages.
- Optical modules at each end where required.
- Fibre or copper cable assemblies.
- Breakout legs and patch-panel cassettes.
- Spares by failure rate and service target.
Do not use logical-link count as cable count. A single assembly can carry multiple lanes; some optical designs separate modules and fibre. The qualified hardware list decides.
At 800GbE, reach and power choices have facility consequences. Short copper links can reduce optical cost inside a rack, while longer connections need active copper or optics. Cable diameter, bend radius and front-panel density can make a mathematically valid rack impossible to service.
Workloads That May Benefit
Large mixture-of-experts training can generate heavy all-to-all communication as tokens route between expert partitions. A wide scale-out fabric may reduce the period GPUs spend waiting for those transfers.
Distributed inference can also become network-heavy when model stages, expert layers, KV cache or request components sit on different nodes. Traffic may be less regular than training, making congestion and tail latency important.
Classic data-parallel training uses large reductions. Extra links help when the current fabric limits collective completion, but gains depend on overlap with compute and the model's communication-to-compute ratio.
Independent inference replicas are different. If each request stays on one GPU or one server, user and storage traffic may need far less than 800Gbps per accelerator. More NICs won't improve a workload that rarely crosses nodes.
Measure bytes transferred per GPU, collective types, message sizes and link use at peak load. That trace should decide whether the third NIC earns its place.
Storage Traffic Should Not Be an Afterthought
Checkpoint writes and dataset reads can share the back-end fabric or use a separate storage network. Either choice needs a capacity model.
If storage shares Vulcano-connected switches, assign traffic classes and prove that checkpoint bursts don't delay collectives. If it uses separate NICs, those adapters consume more I/O lanes, switch ports and rack power. Management access should remain available even when the compute fabric is impaired.
The endpoint table should label network roles explicitly: compute, storage, in-band services and out-of-band management. Combining roles is allowed; hiding the combination isn't.
For large training estates, GPUMachines' guide to AI training network bandwidth provides a starting point for workload measurements.
Ethernet Versus InfiniBand After Vulcano
Vulcano strengthens AMD's case for programmable Ethernet. It provides an open-standard path alongside a wide switch and routing ecosystem, and it gives Helios an AMD-controlled scale-out endpoint.
InfiniBand still has a mature AI/HPC operating model, well-understood collective behaviour and integrated fabric management. Some research and training teams will prefer that consistency, particularly when their staff and software already support it.
The decision isn't "open versus closed" in one line. Buyers should compare:
- Supported GPU and collective stack.
- Congestion and failure behaviour at target scale.
- Switch, optic and cable availability.
- Automation and telemetry.
- Team experience.
- Multi-vendor support boundaries.
- Cost in normal and resilient topologies.
- Migration from the installed network.
The Ethernet versus InfiniBand guide for AI training explains those trade-offs without assuming one fabric wins every workload.
Who Should Consider Vulcano 800
Vulcano 800 is relevant to organisations planning AMD Helios racks, very large MI400-series clusters or dense distributed inference where communication traces already show a scale-out bottleneck. Cloud providers may value its programmable transport and multi-plane operation for varied tenant workloads.
Large research centres can consider it when they want an Ethernet-based back end and have staff who can validate a new transport and adapter generation. It also deserves attention for tightly located multi-site designs, provided latency tests support the workload.
The best candidate has a measured problem, a switch plan and an acceptance workload. It isn't buying 2.4Tbps because the number looks future-proof.
Who Should Wait or Choose Less Bandwidth
A cluster with a few servers can often use 200 or 400GbE effectively. Even 800GbE may be unnecessary when training stays within one scale-up domain or inference requests remain local.
Teams without network engineering capacity should be cautious about early multi-plane deployments with new transports. A familiar 400GbE RoCE or InfiniBand fabric can produce more useful work than an advanced network nobody can diagnose.
Buyers should also wait when the exact NIC SKU, supported switch combinations, optics or general-availability dates remain unclear. AMD's product modelling isn't a substitute for released hardware and qualification data.
If workload demand is uncertain, hosted GPU capacity or a smaller cluster can establish traffic patterns before a large fabric purchase.
A Vulcano 800 Procurement Checklist
Require these items before approval:
1. Exact NIC SKU, form factor, firmware and driver. 2. PCIe 6.0 or UALink attachment mode with a server block diagram. 3. NICs per GPU and rail assignment. 4. Connector, breakout and supported link modes. 5. Leaf and spine port arithmetic per plane. 6. Oversubscription in normal and failed states. 7. Switch NOS, transport, congestion and telemetry configuration. 8. Optic, cable and patching part numbers with reach. 9. Power, cooling, rack U and service clearances. 10. Framework, collective library and operating-system support. 11. Link, NIC, switch and plane failure tests. 12. A representative job-completion benchmark including checkpoints where relevant.
The hardware order should follow this design pack. Starting with an adapter quantity and improvising the topology later is expensive.
Our Technical View
Vulcano 800 is interesting because AMD treats the network adapter as a programmable part of the AI platform rather than a generic Ethernet endpoint. The 800Gb rate, MRC support and multi-plane model address real problems in large clusters.
The three-NIC option needs restraint. It can help a communication-heavy 8,000-GPU model in AMD's simulation, yet smaller systems may receive little benefit while paying for half again as many adapters, switch ports and media. Buyers should calculate from traffic, not from the maximum supported arrangement.
We also want clearer product-level qualification data before treating Vulcano as a general server option. Helios integration is one path. PCIe deployment in other systems needs exact chassis, lane, cooling, driver and switch confirmation.
GPUMachines would design the compute endpoints, planes, switches, optics, storage links and rack placement together. If a two-link or 400Gb design meets the job, we would say so.
FAQ
Does Vulcano 800 provide 2.4Tbps through one adapter?
No. AMD's headline uses up to three 800Gbps NICs per GPU. The summed nameplate bandwidth is 2.4Tbps.
Is 2.4Tbps the application throughput?
No. It is aggregate link rate. Protocol overhead, host attachment, congestion, message sizes and synchronisation reduce useful job throughput.
Does every GPU server support three Vulcano NICs per GPU?
No. That requires a qualified platform with enough attachment bandwidth, physical space, cooling and switch connectivity. Confirm the exact server option list.
Is MRC the same as Ultra Ethernet Transport?
They are related parts of the AI Ethernet discussion but shouldn't be treated as identical. Ask which transport, UEC profile and software release the deployed system uses.
Can Vulcano connect GPU clusters across data centres?
AMD positions it for scale-across networking, but distance adds latency. Test the actual sites and workload; many synchronous jobs won't behave like a single-site cluster.
Do we need three NICs per GPU for inference?
Only if distributed inference traffic can use them and the current network is limiting throughput or latency. Independent replicas may need much less back-end bandwidth.
How should Vulcano be compared with InfiniBand?
Compare the complete fabrics under the same workload, resilience and support assumptions. Include switches, optics, software, operations and failed-state performance.
Can GPUMachines design a Vulcano-based network?
GPUMachines can review AMD Ethernet cluster requirements and build endpoint, switch, optic, cable, rack and storage-network plans. Final sourcing depends on product availability and supported vendor configurations.
Verdict
AMD Vulcano 800 raises the endpoint ceiling for Ethernet AI clusters, but its real contribution is the combination of high bandwidth, programmable transport and multi-plane operation.
The strongest use case is a large, communication-bound AMD GPU estate with traffic traces and network staff. The weakest is a modest cluster buying three adapters per GPU without proving that one or two are insufficient.
Ask GPUMachines to design or review an Ethernet scale-out fabric for an AMD GPU cluster.
Sources and Further Reading
- AMD: Vulcano 800 AI NIC, scale-out and scale-across (23 July 2026). Primary source for the product announcement and AMD's published claims.
- AMD: AI Networking Built for Scale (23 July 2026). Primary source for the Salina, UALoE and Vulcano platform roles.
- AMD: Next-generation transport for large-scale AI training. Primary source for MRC, breakout modes, SRv6 and qualification statements.
- AMD Pensando networking solutions. Primary source for AMD's product positioning and simulation footnotes.
- AMD Helios rack-scale platform. Primary source for Vulcano's role inside Helios.
- Ultra Ethernet Consortium specification history. Standards source for the current UEC specification revision.
- Ultra Ethernet 1.0 specification. Standards source for UET objectives, profiles and deployment model.
AMD's job-completion, cost, performance and efficiency figures are vendor-reported or vendor-modelled. The 13% job-completion comparison uses AMD's stated 8,000-GPU simulation and doesn't include every production workload stage. GPUMachines hasn't independently tested Vulcano 800.
