GPUmachines

How to Build a 1,000-GPU Cluster: 2026 Design Guide

Turn a 1,000-GPU target into node blocks, fabrics, storage, racks, power, management and acceptance gates that can actually be built.

How to Build a 1,000-GPU Cluster: 2026 Design Guide

A 1,000-GPU cluster is not a thousand copies of a server added to a spreadsheet. With eight-GPU nodes, the exact arithmetic is 125 compute nodes. Most validated architectures use repeatable units, so the build may land at 124 nodes and 992 GPUs or 128 nodes and 1,024 GPUs instead.

That first decision affects switch counts, racks, power, cable routes, storage, management services and commissioning. Treat 1,000 GPUs as a capacity target, then choose the nearest supportable architecture rather than forcing an awkward node count.

This guide covers the complete platform. The separate 1,000-GPU network guide goes deeper into leaf, spine, rail and cable arithmetic.

Resolve the GPU count before designing anything else

For eight-GPU HGX or DGX nodes:

| Design point | Compute nodes | GPUs | What it means | | --- | ---: | ---: | --- | | Exact target | 125 | 1,000 | clean arithmetic, but may split a validated unit | | 31 four-node units | 124 | 992 | eight GPUs below the headline target | | 32 four-node units | 128 | 1,024 | one complete unit above the target | | NVIDIA H100 SuperPOD deployment pattern | 127 | 1,016 | one nominal node position is used for UFM connectivity |

The term scalable unit is not universal. NVIDIA's current HGX H100, H200 and B200 Spectrum-X enterprise reference uses four-node units. The DGX SuperPOD H100 reference uses 32-node units and combines up to four of them in the documented 128-node design. Do not mix component counts from one architecture with the unit size of another.

The practical choice is usually 1,024 GPUs when the workload and facility justify the final four-node block. It preserves the repeatable design and leaves a small amount of operational spare capacity. A 992-GPU design may be reasonable when the power or commercial ceiling is firm.

Define what the thousand GPUs will run

The cluster cannot be sized from accelerator count alone. Record at least four workload cases:

1. The largest synchronous training or fine-tuning job. 2. The busiest period of concurrent research or inference jobs. 3. The largest checkpoint and the required write window. 4. The service state after one node, one link or one switch path is unavailable.

For each case, capture model precision, GPU memory, GPUs per job, nodes per job, dataset location, checkpoint size, target completion time and queue policy. This tells the design team whether the estate is one tightly coupled machine, a shared batch platform, an inference fleet or a mixture of all three.

If demand is still speculative, stop at a smaller repeatable block and measure it. Scaling an unmeasured workload to 1,000 GPUs creates expensive uncertainty in every subsystem.

Standardise the compute node

Large training clusters usually use an eight-GPU HGX or DGX node because NVLink and NVSwitch create a strong scale-up domain inside the server. NVIDIA's HGX H100, H200 and B200 enterprise reference sets the following baseline for its eight-GPU design:

  • two CPU sockets;
  • at least 48 physical cores per socket, with 56 recommended;
  • at least 1.5 TB of system memory;
  • at least 500 GB/s of host-memory bandwidth;
  • symmetrical DIMM population across sockets and memory channels;
  • eight east-west SuperNICs in the recommended one-NIC-per-GPU pattern;
  • one north-south BlueField-3 DPU;
  • balanced PCIe connectivity across CPU sockets and root ports.

These are reference values, not a substitute for the OEM qualification list. The selected server still determines supported CPU TDP, DIMM types, local NVMe, NIC form factors, power supplies and cooling.

Create one approved node profile per workload class. A thousand-GPU estate should not contain dozens of small CPU, memory and NIC variations unless the scheduler and operations team have a clear reason to support them.

Browse HGX server platforms for the compute block. Use PCIe GPU servers for separate inference or development pools rather than mixing them invisibly into the training partition.

Translate nodes into a fabric

An eight-GPU node with one 400 Gb/s east-west interface per GPU contributes 3.2 Tb/s of raw network injection:

8 x 400 Gb/s = 3,200 Gb/s = 3.2 Tb/s

Across 125 nodes, the endpoint total is 400 Tb/s, or 50 TB/s, before protocol overhead:

125 x 3.2 Tb/s = 400 Tb/s

That figure describes endpoint injection, not application throughput. The realised requirement depends on collective algorithms, parallelism, message size and how many jobs run together.

NVIDIA's 127-node DGX H100 SuperPOD bill of materials shows the physical scale of a reference build:

  • 48 Quantum QM9700 compute-fabric switches;
  • 16 QM9700 storage-fabric switches;
  • eight 100 Gb/s in-band management switches;
  • eight 1 Gb/s out-of-band management switches;
  • 2,040 NDR 400 Gb/s compute-fabric cables;
  • 254 100 Gb/s in-band cables for the DGX nodes, plus management and inter-switch links.

Those counts belong to the documented H100 SuperPOD topology. They should not be copied into a Spectrum-X Ethernet design or a newer rack-scale architecture. They do show why the fabric must be designed with the rack layout: cable count and reach are procurement and installation constraints, not an afterthought.

Separate the traffic roles:

  • east-west compute traffic;
  • high-performance storage traffic;
  • user and in-band management traffic;
  • isolated out-of-band management;
  • north-south application or customer traffic.

Use InfiniBand cluster design or Spectrum-X Ethernet according to the reference architecture, workload and operating skills. Then calculate oversubscription and failure-state capacity for every leaf and spine layer.

Size storage for single-node and aggregate demand

NVIDIA's H100 SuperPOD reference separates high-performance storage from user storage. The high-performance tier is built for parallel reads and writes over InfiniBand, while user storage handles metadata-heavy work and provides a secondary Ethernet path.

The reference gives two measurements that must not be confused:

  • a GDS-enabled application should be able to read at more than 40 GB/s from one DGX H100 node;
  • the best four-unit aggregate guideline is 500 GB/s read and 250 GB/s write.

Multiplying 40 GB/s by every node would produce a 5 TB/s headline for 125 nodes, but that is not the published aggregate sizing point. The storage design needs both a strong single-client path and an aggregate target derived from real concurrency.

Checkpoint arithmetic is more useful than capacity alone. If a distributed job writes 20 TB in five minutes, its application-level average is:

20,000 GB / 300 seconds = 66.7 GB/s

Add filesystem, protection and contention overhead. Then model what happens when another job reads training data while the checkpoint runs.

Plan separate roles for local boot, local cache, shared training data, checkpoints, model repositories, user homes and archive. The AI data-pipeline guide explains how preprocessing, caching and checkpointing change the I/O profile. Review scale-out storage before the compute bill of materials is frozen.

Calculate the facility envelope

The H100 SuperPOD data-centre guide lists 10.2 kW maximum power for one DGX H100. Using that specific system as an example:

125 nodes x 10.2 kW = 1.275 MW

That is compute-node power only. Fabric switches, storage, management servers and UFM appliances sit above it. The 127-node reference lists roughly 1.295 MW for compute nodes before those supporting systems are added.

At the documented density of four DGX H100 systems per rack, 125 nodes require 32 compute racks because the final rack is not full. NVIDIA's 127-node representative bill of materials lists 38 racks for the complete deployment, together with 108 rack PDUs across two models. Newer B200, B300, GB200, GB300 and Rubin systems have different power and cooling profiles; do not reuse the H100 arithmetic for them.

The site review must cover:

  • available utility and UPS capacity;
  • A/B feed design and breaker limits;
  • rack PDU count, connector type and service access;
  • air, rear-door or direct-liquid-cooling design;
  • water temperature, flow, pressure and water quality where liquid cooling is used;
  • rack weight, floor loading, lift and delivery route;
  • hot-aisle and cold-aisle containment;
  • cable-tray capacity and maximum fibre reach;
  • expansion space for the next unit;
  • maintenance operation with one feed or cooling component unavailable.

Use the rack planner with complete-system power values. A sum of GPU TDPs is not a facility design.

Build the management plane as a supported service

The 127-node H100 SuperPOD reference includes five management servers and four UFM appliances in addition to GPU nodes. The exact implementation will vary, but a thousand-GPU platform needs dedicated capacity for:

  • provisioning and image management;
  • Slurm, Kubernetes or another approved scheduler;
  • identity, projects, quotas and accounting;
  • container and package registries;
  • telemetry, alerts and log retention;
  • fabric management;
  • firmware, driver and software lifecycle;
  • backup and recovery of management state;
  • secure user entry points and data transfer.

NVIDIA notes that the referenced SuperPOD supports multiple teams through Base Command Manager but is not a general multitenant design. A service provider or strongly isolated internal platform therefore needs an explicit security architecture rather than assuming the reference creates tenant boundaries automatically.

Name the operating owner before purchase. Procurement can buy hardware; it cannot supply queue policy, incident response, change control or capacity management.

Use deployment gates, not one delivery date

Gate 1: workload and software proof

Run the target framework, precision and model on the chosen GPU generation. Measure memory, collective traffic, storage traffic and node power. Confirm that software support is mature enough for the intended production date.

Gate 2: node qualification

Approve one complete OEM configuration. Record CPU, DIMM, GPU, NIC, NVMe, firmware, BIOS settings, PCIe topology and maximum measured power. Freeze the profile before volume ordering.

Gate 3: four-node integration block

Validate multi-node collectives, storage, scheduling, monitoring and failure recovery on a small block. For a four-node HGX block, this is 32 GPUs. It is large enough to expose topology and data-path problems without putting the complete estate at risk.

Gate 4: first production rack or unit

Install the exact rack, PDU, cooling and cable method intended for production. Check service clearances, fibre routes, labelling, remote management and thermal behaviour at sustained load.

Gate 5: repeatable expansion

Add identical blocks only after acceptance thresholds are met. Track firmware and configuration drift so later racks do not quietly become a different platform.

Write acceptance criteria into the order

Power-on is not acceptance. Define measurable gates for:

  • hardware inventory and sensor health;
  • PCIe, NVLink, NIC and storage topology;
  • single-node GPU and memory tests;
  • multi-node collective bandwidth and latency;
  • rail and switch-path outliers;
  • single-client and aggregate storage throughput;
  • checkpoint completion time under concurrent reads;
  • scheduler placement, quotas and accounting;
  • node drain, reprovisioning and return to service;
  • failed-link and failed-node behaviour;
  • rack power and cooling at sustained load;
  • monitoring and alert delivery;
  • restoration of management services from backup.

Keep the results as the reference baseline for firmware, driver and network changes. For large estates, continuous acceptance testing is more useful than a one-off benchmark during handover.

Procurement checklist

Before issuing a final order, confirm:

  • the target is 992, 1,000, 1,016 or 1,024 GPUs;
  • node and scalable-unit definitions come from one architecture;
  • the OEM configuration is qualified as a complete system;
  • leaf, spine, optic and cable counts match the rack layout;
  • storage has single-node and aggregate acceptance targets;
  • complete IT load fits the power and cooling design;
  • management, UFM, storage and service racks are included;
  • software subscriptions and support ownership are priced;
  • delivery, staging and secure storage space are available;
  • spares and failed-part procedures are agreed;
  • the cluster can be commissioned in repeatable blocks.

Use the GPU cluster configurator to establish the first compute block. GPUMachines can then review the server, fabric, storage, rack, power, cooling and hosting plan together.

Frequently asked questions

Is a 1,000-GPU cluster exactly 1,000 GPUs?

It can be 125 eight-GPU nodes, but validated units often make 992 or 1,024 GPUs cleaner. NVIDIA's H100 SuperPOD deployment pattern also documents 127 active compute nodes and 1,016 GPUs because one nominal node position is used for fabric-management connectivity.

How many racks does a 1,000-GPU cluster need?

It depends on node size, rack density, power and cooling. At four DGX H100 systems per rack, 125 nodes need 32 compute racks. The complete 127-node reference lists 38 racks after supporting infrastructure is included.

How much power does it use?

Using DGX H100's documented 10.2 kW maximum as an example, 125 compute nodes total 1.275 MW before fabric, storage and management. Other GPU generations and OEM platforms must be calculated from their own complete-system values.

Does it need a non-blocking network?

Large synchronous training jobs benefit from a balanced fabric with predictable failure-state behaviour. An inference estate may tolerate more oversubscription. Design from the communication domain and concurrent job mix rather than applying the same ratio to every thousand-GPU cluster.

How much storage throughput is required?

NVIDIA's H100 reference includes more than 40 GB/s for a strong single-node GDS path and up to 500 GB/s aggregate read in its best four-unit guideline. The final target must reflect datasets, caching, checkpoints and concurrent users.

Should the cluster be on-premise or hosted?

Choose on-premise only when the site and operations team can support the power, cooling, network and service model. A hosted private cluster or Buy & Host deployment can keep dedicated hardware while avoiding a large facility project.

Sources

← Back to blog