The best network for a 1,000-GPU cluster is not simply the fastest switch available. It is the fabric that can accept the supported injection bandwidth of every compute node, preserve predictable paths through failures, carry storage traffic without delaying collectives and still be cabled, powered and operated in the chosen data centre.
At this scale, the phrase 1,000 GPUs also needs interpretation. Eight-GPU servers produce a neat 125-node total, but proven architectures are built in repeatable blocks. NVIDIA's H100 DGX SuperPOD planning table describes four 32-node scalable units, or 128 systems and 1,024 GPUs. Its detailed deployment pattern reserves one system position for Unified Fabric Manager connectivity, leaving 127 DGX systems and 1,016 GPUs. That distinction affects rack space, switch ports, cables, power and the scheduler.
This guide uses the documented H100 SuperPOD as a worked example. It is not a parts list for every Hopper, Blackwell or OEM cluster. A B200, B300, GB200 NVL72 or mixed PCIe estate must use the reference architecture and qualified components for that exact platform and software release.
The short answer
For tightly coupled training across roughly 1,000 GPUs, use a rail-aligned, full-bandwidth compute fabric with separate management and recovery networks. NVIDIA's documented H100 design uses NDR400 InfiniBand, eight compute connections per DGX system, 32 leaf switches and 16 spine switches for the 127-node pattern.
Spectrum-X Ethernet can also be the right answer when the supported server, SuperNIC, switch, congestion-control and Network Operator stack are designed together. The protocol name is not the proof. Whichever route is selected, require:
- a port-level topology and cable schedule;
- stated healthy-state and failure-state oversubscription;
- GPU-to-NIC and NUMA mapping for every node type;
- separate ownership of compute, storage, in-band and out-of-band traffic;
- tests using the intended NCCL collectives, job sizes and concurrency;
- a growth unit that can be added without recabling the existing cluster.
Start with the GPU cluster configurator, then develop the fabric alongside the rack planner, network switches and storage design. The server count cannot be finalised in isolation.
What NVIDIA's 1,000-GPU reference point contains
The H100 DGX SuperPOD reference architecture is useful because NVIDIA publishes the fabric quantities instead of describing the system only as non-blocking.
| Item | Four-SU planning table | Detailed deployed pattern | | --- | ---: | ---: | | DGX systems | 128 | 127 | | GPUs | 1,024 | 1,016 | | GPUs per system | 8 | 8 | | Compute links per system | 8 x NDR400 | 8 x NDR400 | | Compute leaf switches | 32 | 32 | | Compute spine switches | 16 | 16 | | Node and UFM to leaf connections | 1,024 nominal node links | 1,020 compute and UFM cables | | Leaf-to-spine connections | 1,024 | 1,024 |
The detailed fabric guide explains that one DGX position is removed for UFM connectivity. NVIDIA's component schedule then lists 48 Quantum QM9700 switches for the compute fabric, 2,040 NDR400 compute cables, 1,536 switch OSFP transceivers and 508 system OSFP transceivers. These are configuration-specific counts, but they show why the network bill of materials cannot be a percentage allowance on the server quote.
NVIDIA also recommends building by complete scalable units. If budget or site constraints require a different node count, the fabric should still support the full unit, with the unused positions left available. That avoids an irregular first phase that must be rewired when the final nodes arrive.
Calculate endpoint injection before choosing switches
For equal-speed compute interfaces:
node injection = active compute links x link speed
The reference DGX H100 node exposes eight NDR400 compute connections:
8 x 400 Gb/s = 3.2 Tb/s per node
For 127 nodes:
127 x 8 x 400 Gb/s = 406.4 Tb/s
That is raw, one-way endpoint line rate across all rails. It is not NCCL application bandwidth, and it does not mean any one destination can receive 406.4 Tb/s. Protocol overhead, collective algorithm, message size, PCIe and NUMA paths, routing and contention all affect useful throughput.
The same calculation for one 32-node scalable unit is:
32 x 8 x 400 Gb/s = 102.4 Tb/s
Write these figures down before selecting a switch. Then prove that every leaf and spine stage provides the intended capacity for the traffic pattern. A switch with enough faceplate bandwidth can still produce an oversubscribed fabric when too few ports are reserved for uplinks.
Why rail alignment matters
An eight-link GPU node naturally maps into eight network rails when the platform topology supports one compute path per GPU. Corresponding interfaces from each node join the same rail, keeping GPU-to-NIC locality predictable and giving collective libraries parallel paths.
In the H100 SuperPOD design, each group of 32 nodes is rail aligned. Traffic on a rail stays one switch hop away from the other 31 nodes in its scalable unit. Traffic between scalable units traverses the spine layer.
This has three practical consequences:
1. Cabling must preserve the rail number from server to leaf and from leaf to spine. 2. The scheduler should understand topology, but job placement must not be used to disguise physical oversubscription. 3. Acceptance tests must include jobs inside one scalable unit and jobs that cross units.
A single crossed cable can leave every link showing up while creating an asymmetric path. Maintain a source-of-truth record containing node, PCI address, GPU locality, rail, leaf port, spine port, cable ID and firmware release.
Leaf and spine arithmetic
At each leaf, calculate:
oversubscription = aggregate endpoint-facing bandwidth / aggregate uplink bandwidth
A 1:1 result means the leaf has equal downlink and uplink capacity in the healthy state. It does not prove the entire cluster is non-blocking. The spine layer must have enough ports and switching capacity for all leaf uplinks, and the routing design must expose those paths.
For L leaf switches, U uplinks per leaf and P usable leaf-facing ports per spine:
minimum spine count = ceiling((L x U) / P)
That formula gives only the quantity floor. Spread each leaf across multiple spines so that one failed switch does not remove most of its capacity. Document the result after a failed link, failed spine and planned maintenance event.
If a leaf is 1:1 across four equal spines, losing one spine usually leaves three quarters of its uplink capacity. The cluster may remain connected, but that boundary becomes 1.33:1 oversubscribed. Full bandwidth after one spine failure requires spare healthy-state capacity, not just multipath routing.
The companion guide on designing a non-blocking GPU network explains the arithmetic and acceptance tests in more detail.
InfiniBand or Spectrum-X Ethernet
The H100 SuperPOD reference uses Quantum-2 NDR400 InfiniBand in a rail-optimised full fat-tree. That is a strong fit for large synchronous training because the topology, adaptive routing, congestion controls, UFM telemetry and certified component set are published together.
Spectrum-X Ethernet is a credible alternative when the deployment follows a matching NVIDIA architecture. It combines Spectrum switches with supported BlueField or ConnectX SuperNICs and a coordinated RoCE stack. Priority Flow Control, Explicit Congestion Notification, routing, buffer settings, cable reach, endpoint firmware and Network Operator versions must be treated as one supported system.
Choose between InfiniBand clusters and Ethernet clusters from the workload, staff experience, support model and exact platform qualification. Do not choose from link speed alone. An 800 Gb/s port may be one endpoint, a breakout into two 400 Gb/s rails or part of a dual-plane design; the server and switch documentation decide which.
Keep four traffic planes visible
NVIDIA's SuperPOD design separates four networks:
- Compute fabric: NCCL collectives and east-west GPU traffic.
- Storage fabric: datasets, checkpoints and model artefacts.
- In-band management: provisioning, scheduler, software repositories and service access.
- Out-of-band management: BMCs, switch management, PDUs and recovery access.
This separation is operational, not cosmetic. A failed compute fabric should not remove BMC access needed to recover it. A checkpoint burst should not create unpredictable pauses in an AllReduce. Software deployment and monitoring should not compete with the highest-value traffic in the cluster.
Convergence can be engineered, but its capacity and failure model must be explicit. Test compute and storage traffic together before accepting a converged design.
Storage is a network workload too
The H100 SuperPOD reference states that storage I/O per node must exceed 40 GB/s. Across 127 nodes, the simple aggregate is:
127 x 40 GB/s = 5.08 TB/s
That is a reference-architecture target, not a claim that every workload continuously reads at 5.08 TB/s. Training data access, metadata, model loading and checkpoint writes have different shapes. The storage system must sustain the required mix, not just a sequential benchmark.
NVIDIA's documented storage fabric connects storage at a 1:1 port-to-uplink ratio while DGX system connections are near 4:3 oversubscription. This is a useful reminder that the compute and storage fabrics do not need identical designs. Each should be sized from its own endpoints, burst profile and recovery objectives.
Use the AI data pipeline guide and storage for AI training to turn dataset and checkpoint behaviour into throughput, metadata and capacity requirements.
Cable reach can determine the room layout
A thousand-GPU fabric contains thousands of high-speed connections. NVIDIA's H100 component schedule lists 2,040 NDR400 compute cables before storage, management and out-of-band cabling are counted. The data-centre design guide also constrains InfiniBand cable path distance, so rack placement cannot be postponed until after the switch order.
The port schedule should include:
- switch and server port;
- rail and plane ID;
- validated cable or optic part number;
- fibre and connector type;
- length and routed path;
- breakout and polarity;
- spare quantity;
- power and cooling load for optics and switches.
Model these paths in the rack plan. Dense GPU rows, overhead trays, patch panels and service loops can consume more reach than a straight-line floor drawing suggests.
Do not copy H100 counts into Blackwell
B200, B300 and rack-scale GB200 or GB300 systems change GPU power, memory, scale-up domains, NIC generation and fabric patterns. Some current NVIDIA architectures use dual-plane Spectrum-X designs and 800 Gb/s ports broken into 400 Gb/s GPU-facing connections. NVL72 also changes the boundary between scale-up and scale-out: 72 GPUs can communicate inside one NVLink domain before external networking is used.
The design method remains valid, but the quantities do not transfer automatically:
1. Select the exact compute node or rack-scale unit. 2. Record its qualified NIC count, speed and GPU locality. 3. Use the matching NVIDIA and OEM reference architecture release. 4. Recalculate endpoint injection, leaves, spines, rails and failure state. 5. Rebuild the cable and optic schedule.
Treating every eight-GPU server as eight interchangeable 400 Gb/s endpoints can create invalid PCIe, cooling and support assumptions.
Acceptance tests for the completed fabric
A port count is a design review. The cluster still needs proof under load.
Physical and topology checks
Verify negotiated speed, lane width, optic telemetry, error counters, firmware and every rail mapping. Compare the live topology with the cable source of truth.
NCCL collectives
Run NVIDIA nccl-tests at small and large message sizes. Test one node, one scalable unit and a job that crosses all intended units. Record algorithm bandwidth, bus bandwidth, per-rank completion time and tail behaviour.
Concurrent jobs
Run at least two collectives on overlapping paths. A fabric that looks excellent for one benchmark may become unfair under real scheduler concurrency.
Storage contention
Read training data and write checkpoints while collectives are active. Measure model step time and completed work, not only switch counters.
Failure state
Remove a supported link or spine and repeat the tests. Record convergence time, job survival and steady-state bandwidth. Traffic rerouted is not the same result as performance remained within target.
Application run
Finish with the actual framework, model, precision and node count. The purchase acceptance criteria should describe scaling efficiency and time to useful work, not an isolated fabric microbenchmark.
Procurement checklist
Before approving a 1,000-GPU network, require these deliverables:
- exact node count and expansion unit;
- per-node GPU, NIC, PCIe and NUMA map;
- healthy and failed-state topology diagrams;
- leaf, spine and management switch port schedules;
- compute, storage, in-band and out-of-band separation;
- optic and cable bill of materials with lengths;
- firmware and network software compatibility matrix;
- rack elevation, power and cooling design;
- monitoring, spare and replacement plan;
- written acceptance tests and pass thresholds.
GPUMachines can turn the selected GPU systems into a complete fabric and rack plan. The useful output is not a switch recommendation on its own. It is a configuration in which every endpoint, path, cable, failure case and test has an owner.
Frequently asked questions
Is 1,000 GPUs exactly 125 eight-GPU servers?
It can be, but a proven architecture may use a different installed count. NVIDIA's four-SU H100 planning table shows 128 systems and 1,024 GPUs, while the detailed SuperPOD deployment pattern uses 127 systems and 1,016 GPUs after reserving a position for UFM connectivity.
Does every GPU need a 400 Gb/s NIC?
No universal rule applies. The H100 reference node uses eight NDR400 compute connections for eight GPUs. Other PCIe, HGX and rack-scale systems use different NIC counts and scale-up boundaries. Follow the exact server topology and reference architecture.
Is InfiniBand always better than Ethernet?
No. InfiniBand and Spectrum-X Ethernet can both support large AI fabrics. The right choice depends on the qualified platform, topology, operational skills, software stack and support model. Port arithmetic and workload tests still apply to both.
Should storage share the compute network?
Only when convergence is designed and tested. The H100 SuperPOD reference uses a separate storage fabric. If traffic is combined, checkpoint and dataset bursts must be included in compute acceptance tests.
What does non-blocking mean at this scale?
It must name a boundary and state. A healthy 1:1 leaf is only one part of the proof. The spine capacity, routing, node paths and required failure state must also preserve the promised bandwidth.
Can an initial phase contain fewer nodes?
Yes, but design the leaf and spine topology for the complete repeatable unit. NVIDIA specifically advises leaving unused positions in a full scalable-unit fabric when a smaller initial node count is required.
Sources
- NVIDIA DGX SuperPOD H100 architecture and scale table
- NVIDIA DGX SuperPOD H100 network fabrics
- NVIDIA DGX SuperPOD H100 major components
- NVIDIA DGX SuperPOD H100 key components
- NVIDIA DGX BasePOD network overview
- NVIDIA DGX SuperPOD data-centre planning guide
- NVIDIA NCCL documentation
- NVIDIA nccl-tests
.jpg)