A “1,000-GPU cluster” is usually a planning label, not the number that should appear on the purchase order. With eight NVIDIA B300 GPUs per HGX node and four nodes per NVIDIA scalable unit, the clean reference-architecture endpoint is 128 nodes and 1,024 GPUs. Buying exactly 1,000 would mean 125 nodes, which breaks the four-node expansion pattern and makes network, spare-capacity and rack planning needlessly awkward.
That distinction changes the budget before anyone prices a server. A credible estimate has to cover 128 complete compute nodes, their east-west fabric, storage and north-south networks, control-plane systems, optics, cabling, racks, power delivery, cooling, software, support and site work. Multiplying a quoted GPU price by 1,000 misses most of the engineering that keeps those GPUs productive.
This guide supplies the arithmetic and the questions needed for a budgetary quote. It doesn't publish a universal dollar total because accelerator supply, OEM platform, network choice, storage target, support term, deployment country and facility scope can move the result by millions. The method remains useful: every quantity is visible, assumptions can be replaced, and procurement can see which figures still need supplier confirmation.
Start with 1,024 GPUs, not 1,000
NVIDIA's current HGX B300 Enterprise Reference Architecture uses a four-node scalable unit and supports as many as 32 scalable units. Each node contains eight B300 GPUs, so the top reference point is:
| Item | Calculation | Planning quantity | | --- | ---: | ---: | | GPUs per HGX B300 node | Fixed by platform | 8 | | Nodes per scalable unit | NVIDIA reference design | 4 | | GPUs per scalable unit | 4 nodes x 8 GPUs | 32 | | Scalable units | 1,024 GPUs / 32 | 32 | | Compute nodes | 32 units x 4 nodes | 128 | | Installed GPUs | 128 nodes x 8 GPUs | 1,024 |
An exact 1,000-GPU design requires 125 nodes. It saves three nodes on paper, but it also leaves one incomplete scalable unit and no natural allowance for maintenance capacity. Unless a workload or contract demands exactly 1,000 active GPUs, 1,024 installed GPUs makes the cleaner engineering and commercial baseline.
There is another interpretation worth separating: 1,000 available GPUs. If the service-level target requires 1,000 GPUs to remain schedulable during planned maintenance or a node failure, 1,024 installed GPUs provide only a 2.4% margin. The spare policy may call for extra nodes, cold spares, or reserved capacity outside the main pool. State that policy before issuing the RFQ.
Define the node before asking for a total
The compute-node quote carries much more than eight accelerators. NVIDIA's HGX B300 specification calls for eight B300 GPUs with 2,304 GB of HBM3e per node, eight ConnectX-8 SuperNICs for east-west traffic, one BlueField-3 DPU for north-south traffic, dual host CPUs, at least 2 TB of system memory and local NVMe sized to the workload.
Across 128 nodes, the reference minimums produce the following estate:
| Resource | Per node | 128-node total | | --- | ---: | ---: | | B300 GPUs | 8 | 1,024 | | GPU memory | 2.304 TB HBM3e | 294.912 TB HBM3e | | Host memory | 2 TB minimum | 256 TB minimum | | East-west adapters | 8 ConnectX-8 | 1,024 adapters | | North-south adapters | 1 BlueField-3 minimum | 128 DPUs | | Training NVMe | 2 TB per CPU socket plus 1 TB boot | At least 640 TB with dual-socket hosts |
The NVMe figure is a floor, not a storage design. It covers node-local boot, cache and working space under NVIDIA's published recommendation. Dataset repositories, checkpoints, model artefacts, tenant volumes, backups and archive usually sit on shared storage that needs a separate capacity and throughput model.
OEM differences matter. The ASRock Rack 8U16X-GNR2 B300, for example, uses an 8U chassis, dual Intel Xeon 6 sockets, 32 DIMM slots, twelve front NVMe bays, four PCIe 5.0 x16 slots and eight 800 Gb/s OSFP connections. A DGX B300 uses a different 10U mechanical and storage design. Compare complete, supportable nodes; don't compare an HGX baseboard price with a finished system quote.
Build the budget from eight cost pools
Use a worksheet in which every line has a quantity, unit cost, currency, quote date, support term and named source. The total can then be expressed as:
Project budget = compute + fabric + storage + control plane + facility work + software + support + contingency
1. Compute nodes
Multiply the delivered node price by 128, then add the agreed spare policy. Confirm what “node price” includes. The quote should say whether CPUs, host RAM, boot drives, data NVMe, ConnectX adapters, BlueField, rails, power cords, firmware entitlement, operating system and support are present.
For high-value platforms, record the GPU serialisation and warranty route as part of acceptance. A low headline figure can become expensive if it excludes memory, network adapters or the service arrangement needed to replace a failed baseboard.
2. East-west compute fabric
HGX B300's 2-8-9-800 reference configuration means two CPU sockets, eight GPUs, nine network adapters in total and as much as 800 GbE of east-west bandwidth per GPU. Eight ConnectX-8 adapters per node create 1,024 node-side east-west connections across 128 nodes.
The switch and optic count depends on the chosen Spectrum-X or InfiniBand topology, port breakout, rail design and failure domains. Count both ends of every link. A cable schedule should distinguish node-to-leaf links, leaf-to-spine links, storage links, customer uplinks and management connections; “network included” isn't a bill of materials.
At this scale, optics can form a major line item. Include spare transceivers, fibre lengths, patch panels, cleaning equipment and installation testing. Cheap optics that produce intermittent errors cost far more once the cluster enters acceptance testing.
3. Storage and data movement
Size storage from the workload rather than a percentage of GPU spend. Training needs sustained read throughput, checkpoint write bandwidth and enough capacity for active datasets plus replicas. Production inference may place more weight on model loading, artefact distribution, logs and retrieval data. Multi-tenant cloud service adds quotas, snapshots and tenant isolation.
Request four numbers from the storage design: usable capacity after protection, sustained read throughput, sustained write throughput and metadata performance. Then add the storage network, client licences, support and growth shelves. GPUMachines' scale-out storage overview gives the starting categories, but the final bill should follow measured dataset and checkpoint behaviour.
4. Control plane and service nodes
The GPUs don't run the estate by themselves. NVIDIA's HGX reference material describes as many as seven control-plane nodes when Base Command Manager, Slurm and Kubernetes run together: two management nodes, two Slurm heads and three Kubernetes control-plane nodes. A production design may also need login nodes, registry services, telemetry, logging, identity integration, DNS, NTP, licence services and backup.
Keep these systems redundant and outside the accelerator count. Losing a scheduler or container registry can idle the whole cluster even when every GPU remains healthy.
5. Racks, power and cooling
Power turns a server quote into a facility project. NVIDIA lists DGX B300 consumption at 14.5 kW, with an estimated peak of 19.7 kW for the AC configuration in its data-centre planning guide. Those figures don't automatically describe every OEM HGX B300 server, but they provide a defensible facility benchmark until the selected OEM supplies measured and peak values.
For 128 nodes, that benchmark gives:
| Facility measure | Calculation | Compute-node estimate | | --- | ---: | ---: | | Typical IT load | 128 x 14.5 kW | 1.856 MW | | Peak design check | 128 x 19.7 kW | 2.522 MW | | Heat at 14.5 kW per node | 128 x 49,476 BTU/h | About 6.33 million BTU/h |
Those totals exclude switches, storage, control-plane servers and losses. Utility and cooling demand must also account for the site's measured or target PUE. For illustration, a 1.856 MW compute load at a PUE of 1.25 implies 2.32 MW at the facility boundary before the non-compute IT estate is added. This is a derived planning example, not a promise of operating consumption.
Rack density needs its own decision. Four 8U OEM nodes consume 32U and, using the 14.5 kW benchmark, about 58 kW. Five consume 40U and about 72.5 kW. The denser option reduces compute-rack count from 32 to 26, but it raises busbar, PDU, cooling, weight and cable-density demands. Many sites will choose fewer nodes per rack because power or cooling, not free rack units, sets the limit.
6. Software and platform operations
Include operating-system support, NVIDIA AI Enterprise where required, scheduler or Kubernetes tooling, observability, security, workload accounting and any commercial storage or virtualisation licences. For a cloud service, add provisioning, metering, billing, tenant IAM, image management, quota enforcement and an operator portal.
Licence metrics differ. Some charge per GPU, some per node, socket, capacity or support tier. Model the complete term rather than the first year alone.
7. Deployment, validation and support
Budget for rack integration, firmware baselining, cable installation, link validation, burn-in, collective-communication testing, storage tests, scheduler acceptance and documentation. Define the acceptance criteria before shipping starts. Useful measures include the number of healthy GPUs, link error thresholds, NCCL or equivalent collective results, storage throughput, job completion, failover behaviour and telemetry coverage.
Support scope should name response targets, parts location, remote diagnostics, onsite labour and the boundary between OEM, network, storage, software and GPUMachines responsibilities. A cluster can have several valid warranties and still leave the operator to arbitrate between suppliers during an outage.
8. Contingency and currency exposure
Keep contingency visible rather than hiding it in unit prices. It may cover exchange-rate movement, optic-length changes, extra PDUs, fibre remediation, customs, delayed site work and acceptance defects. Procurement should know which items are firm, budgetary, indexed or subject to allocation.
A quote-ready worksheet
The first budget review should contain quantities even when prices remain blank.
| Cost pool | Quantity basis | Supplier input still required | | --- | --- | --- | | HGX B300 compute | 128 production nodes plus agreed spares | Delivered node price, configuration and warranty | | Compute fabric | 1,024 node-side east-west ports plus topology links | Switch, optic, cable and support BOM | | North-south network | At least 128 node attachments plus storage and customer uplinks | Redundancy, port speed and security design | | Shared storage | Workload-derived usable TB and throughput | Appliance/software price and growth plan | | Control plane | Redundant management, scheduler and Kubernetes services | Server count, licences and support | | Compute racks | About 26-32 for 8U nodes, subject to facility density | Rack, busbar/PDU, cooling and installation | | Power | 1.856 MW compute benchmark; 2.522 MW peak check | OEM measured values and facility headroom | | Deployment | One integrated acceptance plan | Labour, travel, commissioning and handover | | Software and support | Full planned term | Per-GPU, per-node and platform subscriptions |
Three budget mistakes that cause late redesigns
Quoting 1,000 accelerators without a node topology. Procurement receives a GPU figure while engineering later discovers that server, scalable-unit and spare boundaries require a different count.
Treating PSU capacity as measured consumption. Nameplate power helps size protection, but the rack plan needs OEM typical, maximum and transient figures for the chosen configuration. Until those arrive, keep the benchmark and its uncertainty visible.
Leaving optics and facility work outside the AI budget. They don't disappear; they return as separate projects with different owners and longer lead times.
What GPUMachines needs for a budgetary design
A useful request includes the accelerator generation, target active GPU count, workload mix, training scale, inference concurrency, storage volume, checkpoint pattern, preferred fabric, deployment country, rack power limit, cooling method, software stack, support term and desired go-live date. Add any requirement for tenant isolation, data sovereignty or hosted operation.
GPUMachines can then compare suitable HGX server platforms, build the rack and network quantities in the GPU cluster configurator, or review a Buy & Host deployment when the customer's site can't accept a multi-megawatt estate.
FAQ
Is a 1,000-GPU cluster really 1,024 GPUs?
For an eight-GPU HGX B300 design following NVIDIA's four-node scalable units, 1,024 is the natural supported endpoint. Exactly 1,000 GPUs equals 125 nodes and leaves an incomplete scalable unit.
Can the cluster start smaller?
Yes. NVIDIA defines four HGX nodes, or 32 GPUs, as one scalable unit. The network and facility plan should still reserve a credible route to the next units; expansion without spare switch ports, fibre paths or power becomes a rebuild.
How much power should the site reserve?
Obtain the selected OEM's configuration-specific figures. As an early benchmark, 128 DGX B300-class nodes at NVIDIA's published 14.5 kW consumption equal 1.856 MW for compute nodes alone, while its 19.7 kW estimated peak yields a 2.522 MW check. Add network, storage, control plane, losses and operating margin.
Why isn't there one public dollar figure?
No single figure stays honest across OEM nodes, accelerator allocation, fabric type, optics, storage, country, support and facility scope. A supplier-backed worksheet is more useful than a headline number whose exclusions surface after approval.
What should be priced first?
Price the complete compute node and settle the topology at the same time. Network port count, rack density, power delivery and storage all depend on that choice.
Sources and calculation notes
- NVIDIA HGX B300 Enterprise Reference Architecture overview
- NVIDIA HGX B300 components and node requirements
- NVIDIA reference-architecture design tenets and scalable units
- NVIDIA DGX B300 physical and power specifications
- ASRock Rack 8U16X-GNR2 B300 product specification
All totals in this guide are derived from the stated quantities. They are planning calculations, not a GPUMachines quotation. Send the workload and facility assumptions for a reviewed cluster design.
.jpg)