GPUmachines

AI Cluster Network Architecture: 32 to 1,024 GPUs

A practical guide to AI cluster fabrics from 32 to 1,024 GPUs, covering rails, leaf-spine sizing, oversubscription, storage, management and acceptance tests.

AI Cluster Network Architecture: 32 to 1,024 GPUs

An AI cluster network should be sized from the server outward. Start with the number and speed of network interfaces on each node, map each interface to its GPU or PCIe root, then build leaf and spine capacity that can carry the intended collective traffic. Starting with “400GbE or 800GbE?” skips the topology question that decides whether those links stay busy.

The design also needs more than one network. Distributed training traffic, storage reads, checkpoints, customer access and out-of-band management have different failure and congestion behaviour. They can share equipment at small scale if bandwidth and isolation are explicit; at larger scale, combining them creates a fabric that looks economical in a diagram and becomes impossible to diagnose under load.

This guide uses NVIDIA's current enterprise reference patterns as worked examples. It doesn't assume that every deployment must use one vendor or protocol. The arithmetic applies to Ethernet and InfiniBand: count endpoints, choose a topology, decide acceptable oversubscription, preserve failure domains and test the result with the workload's communication pattern.

Five traffic classes belong in the diagram

Draw these paths separately before selecting switches:

| Traffic class | Typical flows | What failure looks like | | --- | --- | --- | | GPU compute, east-west | AllReduce, All-to-All, tensor/pipeline traffic, KV transfer | GPUs wait at collectives; tail latency rises | | Storage | Dataset reads, checkpoint writes, model distribution | Jobs starve or checkpoint windows expand | | Customer, north-south | User access, APIs, inference requests, data ingress/egress | Service is unreachable or throttled | | In-band management | Provisioning, images, telemetry, registries, control-plane APIs | Nodes cannot join, update or report health | | Out-of-band management | BMC, switch management, console, power control | Operators lose recovery access during an incident |

NVIDIA's reference-architecture guidance treats east-west, north-south, storage, customer uplink and management as distinct functions. That doesn't force five physical fabrics. It forces the designer to name the capacity, isolation and recovery path for each one.

Read the node code before the switch brochure

NVIDIA describes reference nodes with a CPU-GPU-network-bandwidth pattern. Two current examples show how different the network can be for servers that both contain eight GPUs:

| Reference node | CPUs | GPUs | Network adapters | Average east-west bandwidth per GPU | | --- | ---: | ---: | ---: | ---: | | RTX PRO AI Factory 2-8-5-200 | 2 | 8 | 5 total: 4 east-west, 1 north-south | 200 GbE | | HGX B300 AI Factory 2-8-9-800 | 2 | 8 | 9 total: 8 east-west, 1 north-south | 800 GbE |

The first pattern commonly uses four 400 Gb/s east-west interfaces for eight GPUs. The second uses eight ConnectX-8 SuperNICs, one per GPU, with as much as 800 Gb/s per adapter. Substituting the same switch count for both nodes would be a fourfold bandwidth error before storage is considered.

Check physical topology too. A NIC can have an impressive line rate while sitting behind the wrong CPU socket or PCIe path for its assigned GPU. The server vendor's topology drawing, nvidia-smi topo -m, PCIe inventory and NUMA map should agree with the rail plan.

Scale in complete units

Current NVIDIA RTX PRO and HGX enterprise designs use a four-node scalable unit. At eight GPUs per node, one unit contains 32 GPUs. The following table turns common project labels into node and unit counts:

| Installed GPUs | Eight-GPU nodes | Four-node units | Comment | | ---: | ---: | ---: | --- | | 32 | 4 | 1 | Smallest complete unit | | 128 | 16 | 4 | A useful first multi-unit fabric | | 256 | 32 | 8 | Dedicated compute fabric becomes easier to justify | | 512 | 64 | 16 | Spine capacity and failure domains need close review | | 1,024 | 128 | 32 | Top scale of NVIDIA's HGX B300 enterprise reference design |

For HGX B300, node-side east-west adapter counts equal eight times the node count. A 128-GPU cluster therefore presents 128 adapters; a 1,024-GPU cluster presents 1,024. Count connectors and links, not only adapter cards, because dual-plane breakout can change the physical switch-port requirement.

Rail-optimised networking

A rail groups corresponding interfaces from multiple nodes onto the same portion of the fabric. In an eight-GPU HGX node, rail 0 may connect the interface nearest GPU 0 across every node, rail 1 serves GPU 1, and so on. Collective libraries can then keep communication aligned with GPU and NIC locality.

Rail optimisation reduces unnecessary cross-fabric movement, but it creates placement rules. A job spanning several nodes needs compatible rail connectivity on each node. A failed leaf, cable or NIC may affect one rail across many servers, so the scheduler and monitoring system need topology awareness.

NVIDIA's Spectrum-X guidance now supports single-, dual- and quad-plane ConnectX-8 designs. More planes can increase scale and spread failure domains; they also multiply port, optic, cable and operational counts. Choose the plane count from the reference version and workload, not from a desire to make the diagram symmetrical.

Leaf-spine arithmetic without hand-waving

A two-tier leaf-spine fabric has server-facing leaf ports and leaf-to-spine uplinks. Oversubscription is:

Oversubscription ratio = total leaf downlink bandwidth / total leaf uplink bandwidth

For a nominally non-blocking 1:1 leaf, aggregate uplink bandwidth equals aggregate server-facing bandwidth. Suppose a leaf receives 32 x 400 Gb/s server links, or 12.8 Tb/s. It also needs 12.8 Tb/s of usable spine capacity. With 400 Gb/s uplinks, that means 32 uplink ports. The leaf therefore consumes 64 ports before management or spare capacity.

At 2:1 oversubscription, the same leaf uses 16 x 400 Gb/s uplinks. That may work for independent inference jobs whose traffic rarely synchronises; it can damage distributed training when many nodes enter a collective at once. Record the ratio and the workload assumption instead of calling both designs “high performance”.

High-radix 800 Gb/s switches change port counts but not the method. NVIDIA's Quantum-X800 Q3400 family exposes 144 ports of 800 Gb/s and supports a two-level fat tree connecting as many as 10,368 NICs. Real designs still reserve ports for topology, serviceability and growth, and cable breakout determines how advertised switch ports map to server links.

100-, 200-, 400- or 800-Gb/s per link?

Link speed isn't the same as useful bandwidth per GPU. Start with the node pattern and workload communication ratio.

100GbE can serve management, customer access, storage or small independent inference estates, but it becomes difficult to justify as the primary compute fabric for current multi-node training. 200GbE remains useful for north-south and storage paths, and it appears in NVIDIA's 2-8-5-200 design as the per-GPU average after four 400 Gb/s adapters serve eight GPUs.

400GbE is a common building block for Spectrum-X and Quantum-2-era clusters. It supports direct 400 Gb/s server links and 800 Gb/s adapters presented as two 400 Gb/s ports. 800 Gb/s becomes relevant with ConnectX-8 and current B300/GB300 systems, where the server and reference design can expose that rate end to end.

Don't down-rate a current node without checking the result. Connecting an eight-port B300 platform through a much narrower fabric may be a valid inference choice, but it no longer represents the 2-8-9-800 design and should be priced, benchmarked and sold as a different performance class.

Ethernet or InfiniBand

Both can support large GPU clusters. The choice should follow the operating environment and reference architecture.

Spectrum-X Ethernet combines NVIDIA Spectrum switches, ConnectX or BlueField interfaces, lossless RoCE behaviour, telemetry and congestion controls. It fits organisations that want Ethernet operations and integration while deploying an AI-specific fabric rather than ordinary datacentre Ethernet with a few priority settings.

Quantum InfiniBand provides adaptive routing, congestion control and in-network computing such as SHARP. Quantum-X800 offers 800 Gb/s links and high-radix switches for very large fabrics. Teams with established InfiniBand skills and tightly coupled training workloads may prefer it.

The dangerous choice is an unsupported mixture assembled from compatible connector speeds. Firmware, transceivers, switch software, NIC mode, Network Operator and reference-architecture versions need to form one tested bill of materials.

Storage should not borrow whatever ports remain

Training data and checkpoints can compete with collectives if they share an undersized path. Size storage from measured or modelled throughput:

Required storage bandwidth = active jobs x per-job sustained read/write target x concurrency factor

Then check client interface count, leaf capacity, storage-controller ports and the filesystem or object layer. Large sequential reads, metadata-heavy input pipelines and checkpoint bursts stress different parts of the system.

NVIDIA's RTX PRO reference gives a useful ratio example: at 256 GPUs it allocates 32 storage connections and states a minimum of 12.5 Gb/s per GPU under its assumptions. That isn't a universal storage target, but it demonstrates that storage ports belong in the architecture table alongside compute links.

Management needs a physical escape route

Out-of-band management should survive a compute-fabric failure or bad network configuration. Connect BMCs, switch management, rack PDUs and console services to dedicated management switches, then provide redundant uplinks to a protected operations network.

In-band services such as provisioning, registries, DNS, NTP, telemetry and schedulers need redundancy too. They can use the north-south fabric if bandwidth, VLAN/VRF boundaries and incident access remain clear. Don't place the only monitoring path behind the component it monitors.

Failure domains and spare ports

List what disappears when each component fails:

  • One server-facing link.
  • One leaf switch or one rail.
  • One spine.
  • One management switch.
  • One storage controller or storage leaf.
  • One rack power feed.

The scheduler may need to drain a node, a rail-aligned block or an entire scalable unit. Keep that operational unit aligned with monitoring and spares. A fabric with enough aggregate bandwidth can still produce poor availability if a single leaf removes an awkward selection of GPUs that no remaining job can use efficiently.

Reserve ports for failures and expansion before cable lengths are ordered. A spare port in the wrong rail, rack or switch role may not help.

Cabling and optics are part of topology

Generate a cable schedule from source port to destination port, with device, rack, rail, speed, medium, length, optic, breakout and label. Count transceivers at both ends unless the cable type integrates them. Include spares by exact part, not a generic percentage that mixes incompatible lengths and modules.

Copper can suit short in-rack links; active copper and optical options extend reach with different power, bend-radius and service constraints. Cable paths, patch panels and rack doors must handle the planned density. An elegant logical design can fail installation because OSFP modules, fibre trunks or bend radii weren't modelled physically.

Acceptance tests that mean something

Link-up is the beginning of acceptance. Test in layers:

1. Verify switch, NIC and optic inventory against the approved firmware matrix. 2. Check PCIe, NUMA and GPU-to-NIC topology on every node. 3. Measure single-link bandwidth and latency, then test all links concurrently. 4. Run collectives across one node, one scalable unit and multiple units. 5. Exercise AllReduce and All-to-All patterns that resemble the intended workload. 6. Load storage and compute fabrics together to expose shared bottlenecks. 7. Remove links and switches in controlled tests; confirm detection, routing and scheduler behaviour. 8. Save counters, versions and baselines so later incidents have a comparison point.

NVIDIA's NCCL guidance recommends separating fabric faults from library tuning and points to tools such as ib_write_bw, ib_write_lat, nvbandwidth and topology inspection. Don't tune obscure NCCL variables until the physical and routing layers pass their own tests.

Common design errors

Buying switch capacity before choosing the server topology creates adapter and port mismatches. Sharing compute, storage and management without a traffic model hides contention. Treating 1:1 port counts as proof of a non-blocking fabric ignores spine capacity and breakout. And benchmarking one pair of nodes says almost nothing about simultaneous collective traffic across the cluster.

The most expensive error is leaving no growth state between today's cluster and the final target. Expansion needs reserved ports, rack locations, fibre paths, power, addressing, software licences and a maintenance method. Otherwise the “next scalable unit” triggers a fabric replacement.

Planning outputs GPUMachines should produce

A quote-ready network design should include a node connectivity diagram, logical traffic-plane drawing, physical topology, switch and adapter BOM, port map, cable schedule, rack placement, power estimate, firmware matrix, failure-domain table and acceptance plan.

Use the GPU cluster configurator to frame compute quantities, then connect the result to NVIDIA networking options and scale-out storage. GPUMachines can review whether the proposed bandwidth, rail count and rack layout fit the selected servers and workloads.

FAQ

Does every AI cluster need a non-blocking fabric?

No. Independent inference or batch jobs can tolerate oversubscription when measurements support it. Large synchronous training and communication-heavy distributed inference are less forgiving. Publish the ratio and benchmark the intended workload.

How many networks should the cluster have?

There is no fixed count. Separate compute, storage, customer, in-band management and out-of-band functions logically; decide which may share physical switches only after bandwidth, security and failure behaviour are defined.

Is 800Gb/s required for B300?

HGX B300's current 2-8-9-800 reference pattern supports 800 Gb/s average east-west bandwidth per GPU through ConnectX-8. A narrower design can be built for a specific workload, but it should not be represented as the same performance class.

Can standard Ethernet switches run RoCE?

Basic protocol support doesn't make a validated AI fabric. Buffering, congestion control, telemetry, routing, NIC firmware, switch software and operational tooling all matter. Use a supported architecture and version matrix.

What should be tested before hand-over?

Test topology, every link, concurrent fabric load, multi-node collectives, storage interaction, failure handling and monitoring. Save the results as the cluster baseline.

Sources

Ask GPUMachines to review the node, fabric and rack arithmetic.

← Back to blog