GPUmachines

How to Design a Non-Blocking Network for GPU Clusters

Calculate GPU-node injection, leaf and spine ports, rail topology, failure-state bandwidth and acceptance tests before calling an AI network non-blocking.

How to Design a Non-Blocking Network for GPU Clusters

Non-blocking is one of the easiest claims to put on a GPU cluster diagram and one of the hardest to prove. A switch may be non-blocking internally while the fabric is oversubscribed at its uplinks. A healthy topology may provide full bandwidth but lose half of it when one spine fails. Eight 400 Gb/s NICs in a server do not guarantee 3.2 Tb/s of useful GPU communication.

A defensible design starts with node injection bandwidth, traffic pattern and failure policy. It then works through leaf ports, spine ports, rails, optics, routing and software before testing the complete system with collective operations.

The short answer

A GPU cluster network can be called non-blocking only after its scope is stated:

  • Which endpoints and traffic class are covered?
  • Is the claim for one rail, one scalable unit or the whole cluster?
  • Is it valid only when every link is healthy?
  • Does it apply to arbitrary all-to-all traffic or a specific job placement?
  • Are storage and management traffic separate or included?

For a two-tier leaf-spine fabric, a basic healthy-state check is:

oversubscription ratio = aggregate leaf downlink bandwidth / aggregate leaf uplink bandwidth

A 1:1 result means the leaf has as much uplink capacity as endpoint-facing capacity. That is necessary for a full-bandwidth design at that boundary, but it is not sufficient. The spine layer, endpoint PCIe paths, routing and failure state must also pass.

Use the GPU cluster configurator to define compute nodes, then compare InfiniBand clusters, Ethernet clusters, network switches and rack planning as one system.

Separate scale-up from scale-out

Modern GPU systems can contain two different high-speed networks.

Scale-up

NVLink and NVSwitch connect GPUs inside a server or rack-scale domain. They provide GPU-to-GPU communication without using the external Ethernet or InfiniBand fabric for every transfer. The domain, bandwidth and topology depend on the GPU platform.

Scale-out

Ethernet or InfiniBand connects separate server nodes and sometimes separate scale-up domains. Distributed training, fine-tuning and inference use this fabric for NCCL collectives, point-to-point transfers and control traffic.

Do not add NVLink bandwidth to NIC bandwidth and call the result network capacity. They serve different paths. The scale-out design begins at each node's supported NIC interfaces.

Identify every traffic plane

NVIDIA's DGX BasePOD reference architecture separates compute, management/storage and out-of-band management networks. Its deployment guide also identifies an external network. A custom cluster may differ, but the traffic classes still need names and owners.

  • Compute fabric: east-west GPU traffic between nodes.
  • Storage fabric: datasets, checkpoints and model artefacts.
  • In-band management: provisioning, scheduler, telemetry and user access.
  • Out-of-band management: BMCs, switch management and recovery access.
  • External connectivity: corporate, campus, WAN or service ingress.

Convergence can reduce switch and cable count, but it changes the capacity and failure model. If storage and compute share links, checkpoint bursts must be tested alongside collectives. BMC access should not depend on the failed production fabric it is meant to recover.

Start at the node

Record the exact server model and its network map:

  • Number and speed of compute NIC ports.
  • Which GPU or NVSwitch domain is nearest to each NIC.
  • PCIe generation, width, switch and CPU root complex for each adapter.
  • NUMA placement and GPU Direct RDMA support.
  • Storage and management interfaces.
  • Supported cables, optics and transceivers.
  • Firmware, driver and network software release.

NVIDIA's enterprise reference architectures use a C-G-N-B notation for CPU sockets, GPUs, network adapters and average east-west bandwidth per GPU. The useful lesson is to describe the node as a communication pattern, not a server name alone.

An eight-GPU node with one 400 Gb/s NIC has a very different scale-out ratio from a node with eight 400 Gb/s NICs. More adapters are useful only if the server exposes enough PCIe bandwidth and the workload maps communication across them.

Calculate node injection bandwidth

For equal-speed ports:

node injection bandwidth = active compute ports x port speed

An eight-port node at 400 Gb/s has 3.2 Tb/s of raw one-way line rate. Dividing by eight gives 400 GB/s as a raw byte-rate equivalent. This is not application payload and not a prediction of NCCL bandwidth.

NVIDIA's DGX BasePOD NDR400 example uses eight compute connections per DGX H100, H200 or B200 node. Treat that as a platform-specific reference, not a rule for every eight-GPU server.

For a 32-node cluster with eight 400 Gb/s compute ports per node:

32 nodes x 8 ports x 400 Gb/s = 102.4 Tb/s

That total describes endpoint injection across all rails. Every leaf and spine path cannot necessarily carry all 102.4 Tb/s to any arbitrary destination simultaneously unless the complete topology and routing are built for it.

Calculate leaf oversubscription

Consider one simplified 400 Gb/s rail. A leaf has 32 endpoint-facing ports and 32 spine-facing ports, all at 400 Gb/s.

  • Downlink capacity: 32 x 400 Gb/s = 12.8 Tb/s
  • Uplink capacity: 32 x 400 Gb/s = 12.8 Tb/s
  • Oversubscription: 12.8 / 12.8 = 1:1

If the same leaf has only 16 uplinks:

  • Uplink capacity: 16 x 400 Gb/s = 6.4 Tb/s
  • Oversubscription: 12.8 / 6.4 = 2:1

The 2:1 design may be acceptable for independent inference replicas or jobs kept within a rack. It is not full-bisection bandwidth for arbitrary traffic leaving that leaf.

Repeat the arithmetic for every port group and breakout mode. A switch faceplate count is not enough: ASIC boundaries, supported breakout combinations and licensed speeds can constrain the usable layout.

Size the spine layer

Every leaf uplink must terminate on a spine port. The spine tier needs enough total ports and switching capacity for all leaves, with a routing pattern that exposes the paths to endpoints.

For L leaves, U uplinks per leaf and P spine ports available for leaf connections:

minimum spine count by port quantity = ceiling((L x U) / P)

That is only the quantity floor. Uplinks should be distributed across enough spines to meet the resilience policy. Port speed, cable reach, switch ASIC capacity and routing limits must also be checked.

Avoid a design where each leaf has nominally enough uplink bandwidth but most links terminate on the same failure domain. Non-blocking arithmetic and fault isolation have to be solved together.

Design the rails deliberately

A rail-optimised topology maps corresponding NICs from each node into the same network plane. The intent is to keep GPU and NIC locality predictable and provide multiple parallel paths for collective traffic.

Label rails consistently across servers, switches and cables. A useful record includes:

  • Node name and NIC PCI address.
  • Physical port and rail number.
  • Nearest GPU or scale-up domain.
  • Leaf, spine and switch port.
  • Cable or optic identifier.
  • Firmware and link configuration.

Mis-cabling one rail can preserve link-up status while creating an asymmetric topology. NCCL may route around it, but collective performance can become uneven and difficult to diagnose.

Current Spectrum-X architectures include single-plane and multiplane modes for different GPU platforms. Follow the reference architecture version that matches the NIC, switch, GPU system and Network Operator release. Mixing tuning profiles from different releases is not a safe shortcut.

Ethernet or InfiniBand

Non-blocking does not select a protocol. Both Ethernet and InfiniBand can be built as full-bandwidth leaf-spine or fat-tree fabrics. The operational stack differs.

Spectrum-X Ethernet and RoCE

Spectrum-X combines NVIDIA Spectrum switches with supported BlueField or ConnectX SuperNICs for east-west RoCE traffic. RoCE needs coordinated endpoint, switch, routing and congestion-control configuration.

NVIDIA Cumulus Linux exposes lossless RoCE modes using Priority Flow Control and Explicit Congestion Notification. PFC operates per traffic priority; ECN marks congestion for endpoint response. Buffer and cable-length settings matter, so configuration should follow the matching reference architecture rather than a copied switch template.

Ethernet can share skills and tooling with the wider data-centre network, but an AI RoCE fabric is not made reliable by turning on jumbo frames and pause globally.

InfiniBand

InfiniBand provides an RDMA fabric with a subnet manager and routing features designed for high-performance computing. NVIDIA documents adaptive routing that can move traffic away from temporarily congested paths when multiple paths exist.

InfiniBand still requires topology, routing, partitioning, telemetry and firmware discipline. A credit-based fabric does not correct missing uplinks or an incorrect rail map.

Choose from operations and workload

Compare the complete supported stack, staff experience, scale, software, failure response and commercial lifecycle. Protocol preference cannot compensate for wrong port arithmetic.

Congestion and incast

Distributed jobs do not generate smooth traffic. AllReduce, AllGather and ReduceScatter create coordinated transfers. Checkpoints and data shuffles add bursts. Many senders may target one receiver or leaf at once.

NCCL defines collectives across ranks, and every rank must participate with matching operation parameters. One slow path can hold the complete collective. Average link utilisation may look low while short queue spikes dominate iteration time.

Monitor at least:

  • Port utilisation and queue occupancy.
  • ECN marking, PFC pause and discard counters on Ethernet.
  • Congestion and routing counters on InfiniBand.
  • Link retries, symbol or bit errors and lane health.
  • NCCL algorithm and bus bandwidth.
  • Per-rank completion time and stragglers.

Adaptive routing and congestion control improve path use; they do not create bandwidth that was omitted from the topology.

Healthy-state and failure-state bandwidth

State both. Suppose a leaf has 32 endpoint ports and 32 uplinks divided evenly across four spines. In the healthy state it is 1:1. If one spine and eight uplinks fail, 24 uplinks remain:

12.8 Tb/s down / 9.6 Tb/s up = 1.33:1

The cluster remains connected but is no longer non-blocking at that leaf boundary. If the requirement is full bandwidth after one spine failure, extra healthy-state capacity is needed.

Apply the same reasoning to failed links, failed leaves, switch maintenance and dual-rail loss. Decide whether the policy is:

  • Connectivity after failure.
  • Reduced but bounded performance.
  • Full injection bandwidth after any single failure.

Those are different designs and budgets.

Job placement can reduce or expose contention

A scheduler may keep a job inside one leaf, one scalable unit or one rail group. This can reduce spine traffic and improve repeatability. It can also strand GPUs if topology-aware placement is too rigid.

Document the intended allocation sizes. A fabric designed for four-node local jobs may behave differently when one 64-node job crosses every leaf. Test the largest expected gang-scheduled job and two or more concurrent jobs with competing paths.

Topology-aware placement is an optimisation, not permission to mislabel a 2:1 fabric as physically non-blocking.

Cabling and optics are part of the network

At 400 and 800 Gb/s, the link is a combination of switch port, adapter, firmware, cable or transceiver, fibre type, breakout and reach. Build a port-level bill of materials.

Include:

  • DAC, AOC or optical module by validated part number.
  • Fibre and connector type.
  • Breakout mapping and polarity.
  • Maximum reach and patch-panel loss budget.
  • Spare percentage and replacement procedure.
  • Cable pathway, bend radius and service access.
  • Power and cooling for switches and optics.

High-density GPU racks can have hundreds of compute-fabric connections. Rack elevation and cable routing must be reviewed before equipment ships.

Do not forget PCIe and NUMA

A perfect leaf-spine cannot fix a congested path inside the node. Validate NIC and GPU placement with manufacturer topology diagrams and operating-system tools.

Look for:

  • NICs sharing an undersized PCIe switch or CPU uplink.
  • Traffic crossing CPU sockets unnecessarily.
  • IOMMU or ACS settings that block expected peer-to-peer paths.
  • GPUs mapped to remote NICs instead of local rails.
  • Firmware or BIOS settings that reduce link width.

NCCL uses system topology information to select transports. Incorrect or incomplete topology exposure can produce a valid but slower path.

Acceptance testing

Port arithmetic is the design review. Workload testing is the proof.

1. Verify every physical link

Check negotiated speed, width, errors, optic telemetry, rail mapping and firmware. Save a known-good topology inventory.

2. Test point-to-point paths

Measure latency and bandwidth within a leaf, across leaves, across rails and across the farthest intended path. Look for asymmetric pairs.

3. Run NCCL tests at several message sizes

NVIDIA's nccl-tests includes all_reduce_perf and other collective tests. Small messages expose latency; large messages expose bandwidth. Test one node, one scalable unit and the full intended job size.

Read both algorithm bandwidth and busbw. NVIDIA documents busbw as a normalised way to compare collective use of the hardware across rank counts.

4. Test competing jobs

Run collectives on separate and overlapping paths. A network that performs well for one job may become unfair or unstable under concurrency.

5. Add storage traffic

If fabrics are converged, run training reads and checkpoint writes while collectives continue. Measure iteration time, not only network counters.

6. Remove a link or spine

Repeat the tests during the supported failure case. Record convergence time, job survival and steady-state bandwidth after rerouting.

7. Run an application workload

Use the production framework, model, precision and node count. Report scaling efficiency, step time distribution and completed work. A network microbenchmark is not the final acceptance criterion.

Common design mistakes

  • Calling a switch non-blocking and applying the claim to the complete fabric.
  • Counting full-duplex transmit and receive bandwidth twice in a one-way injection calculation.
  • Matching downlinks and uplinks at each leaf but under-sizing the spines.
  • Summing NVLink and external NIC bandwidth.
  • Assuming eight NICs map cleanly to eight GPUs without checking PCIe and NUMA.
  • Using healthy-state 1:1 arithmetic when the requirement applies after a failure.
  • Converging storage and compute without a mixed-traffic test.
  • Applying an RoCE configuration from a different switch, NIC or software release.
  • Buying switches before producing a port-level optic and cable schedule.
  • Testing only large AllReduce messages and missing latency or incast problems.

Frequently asked questions

What does 1:1 oversubscription mean?

At a stated boundary, aggregate uplink bandwidth equals aggregate endpoint-facing bandwidth. It does not prove the rest of the fabric, routing or endpoints can sustain arbitrary full-rate traffic.

Is a 2:1 network always unsuitable for AI?

No. It may suit independent inference, localised jobs or cost-sensitive clusters with topology-aware placement. It should be described as 2:1 and tested with the intended job mix.

How many spine switches are required?

Divide total leaf uplink ports by usable leaf-facing ports per spine and round up, then add the distribution needed for the failure policy. Port speed, ASIC capacity and cable design must also pass.

Does every eight-GPU server need eight NICs?

No. NIC count depends on the server platform, workload and target bandwidth per GPU. NVIDIA reference systems with eight NDR400 links are specific designs, not a universal minimum.

Can Ethernet be non-blocking for distributed training?

Yes. A correctly sized leaf-spine Ethernet fabric can be 1:1. RoCE performance also depends on supported NICs, switches, congestion control, routing and software configuration.

Is InfiniBand automatically non-blocking?

No. InfiniBand provides an HPC-oriented RDMA stack, but port counts and topology still determine oversubscription. Routing cannot replace absent capacity.

What should the acceptance target be?

Set targets for collective bandwidth and latency, job scaling efficiency, tail step time, link errors and failure-state behaviour. Tie them to the intended node count and software release.

Turn the diagram into a bill of materials

A purchase-ready network plan should include node NIC mapping, every leaf and spine port, optics and cables, rail and plane IDs, firmware, management switches, spare ports, rack locations and power. It should also state healthy-state and failure-state oversubscription.

GPUMachines can translate the compute design into InfiniBand or Ethernet fabrics and place the complete system in the rack planner. The final approval should follow port arithmetic and a written acceptance test, not a topology label.

Sources

← Back to blog