A 42U cabinet is not a promise that 42U of GPU equipment will fit. Modern AI servers can exhaust the rack's electrical, cooling, weight or cable capacity long before the last mounting position is used. Rack density therefore starts with the facility envelope, not a target such as “eight servers per rack”.
The practical rule is simple: design around measured or vendor-published system draw at the intended configuration, add the required operational margin, and have qualified facilities engineers approve the feeds and heat-rejection path. Do not add power-supply nameplate ratings and call the result consumption. Redundant PSUs describe available capacity and resilience; they do not report what the server draws from the wall.
Current NVIDIA reference material illustrates the scale of the change. NVIDIA says many existing data centres still operate below 20 kW per rack and lack liquid cooling. By contrast, its GB300 NVL72 component guide states that a full rack can require up to 142 kW and uses integrated liquid cooling with leak detection. That is a different facility class, not a denser version of an ordinary air-cooled rack.
GPUMachines supplies and hosts GPU infrastructure. This guide helps buyers prepare a rack and facility brief; it does not replace electrical design, cooling design, structural review, local regulations or manufacturer installation instructions. Start with the GPU cluster configurator, then use the AI Factory profiles to compare server-level and rack-scale deployment patterns.
The decision in one page
Use four limits for every proposed rack:
| Limit | Question to answer | Evidence required | | --- | --- | --- | | Electrical | Can the approved A/B feeds carry the expected load with the specified redundancy and margin? | Measured or published equipment draw, feed design and qualified electrical approval | | Thermal | Can the room and rack remove the heat in normal and failure conditions? | Air and liquid cooling capacities, temperatures, flow, pressure and failure response | | Physical | Can the floor, cabinet, rails and delivery route carry and admit the equipment? | Weights, dimensions, load limits, lift route and rack compatibility | | Operational | Can staff cable, monitor and replace equipment without taking unsafe shortcuts? | Service clearances, cable schedule, leak plan, spares and maintenance procedure |
The lowest limit sets the rack density. If the feeds allow 80 kW but the cooling path removes 45 kW, the usable envelope is not 80 kW. If both support 100 kW but the raised floor, rack or delivery route cannot accept the loaded cabinet, the design still fails.
Power draw is not PSU capacity
GPU server specifications often list several power supplies, for example six 3.2 kW units in a redundant arrangement. Multiplying them produces installed PSU capacity. It does not tell you the steady load, peak input or how the load redistributes after a PSU or feed failure.
For each configured system, collect:
- the manufacturer's maximum or rated system input for that exact configuration;
- measured wall power from a representative workload where available;
- the expected sustained draw for the buyer's software and power policy;
- start-up, transient and failure behaviour supplied by the OEM;
- the supported redundancy mode and feed mapping.
Use these as separate numbers. A planning estimate may apply diversity when workloads cannot peak together, but the assumption must be written and approved. Do not hide diversity inside a spreadsheet cell with no owner.
GPU thermal design power is also incomplete. Host CPUs, memory, NVMe devices, NICs, DPUs, fans, pumps, motherboard regulators and conversion losses all consume power. Dense air-cooled systems can devote a substantial share to fans as static pressure rises. Liquid cooling reduces the air-side burden for captured components, yet pumps, coolant distribution equipment and remaining air-cooled parts still need energy and capacity.
Convert electrical load into heat
Nearly all electrical energy consumed by IT equipment becomes heat in the room or liquid loop. For a planning conversion, one kilowatt equals about 3,412 BTU per hour. A 100 kW IT load therefore represents roughly 341,200 BTU/h of heat rejection before adding non-IT plant losses.
Using NVIDIA's published “up to 142 kW” figure for a GB300 NVL72 rack gives about 484,500 BTU/h:
142 kW x 3,412 BTU/h per kW = 484,504 BTU/h
That arithmetic does not size a chiller, CDU or pipe. It only makes the scale visible. Cooling engineers still need supply and return temperatures, coolant chemistry, flow, pressure drop, allowable approach temperatures, redundancy target and environmental conditions.
For air-cooled racks, ask whether the published room capacity is usable at the proposed inlet temperature and pressure. Nominal cooling tonnage can coexist with local recirculation, weak underfloor delivery or a hot-air return path that cannot carry the rack's exhaust.
Air cooling has a practical ceiling
Air cooling remains a sensible choice for many workstations, PCIe GPU servers and smaller HGX deployments. It is familiar, avoids water connections at the server and can fit existing colocation halls. The difficulty is that heat removal depends on moving large volumes of air through restrictive chassis and rack paths.
At higher density, verify:
- front intake temperature at the top, middle and bottom of the rack;
- hot-aisle containment and return-air capacity;
- blanking panels and sealing around cable openings;
- fan power, acoustic limits and OEM airflow direction;
- behaviour after a cooling unit, fan wall or containment component fails.
Do not mix front-to-back and back-to-front equipment in one rack without an engineered airflow plan. Network switches frequently create this problem. Cable bundles can also obstruct exhaust, especially where many 400 or 800 Gb/s links leave the rear of a GPU rack.
High density may force fewer servers per cabinet even when U space remains. Spreading systems across two racks can improve cooling and service access, but it changes fabric port counts, cable lengths and floor footprint. Treat that as an architecture trade, not wasted space.
Direct liquid cooling changes the facility contract
Direct liquid cooling carries heat from cold plates into a facility or technology cooling loop. It can support densities that air cooling cannot handle, but the boundary between server vendor, rack integrator, CDU supplier, colocation provider and facilities team must be explicit.
NVIDIA's GB300 NVL72 documentation describes a liquid-cooled rack with integrated leak detection. That does not mean every site can accept it. The project still needs compatible coolant conditions, connections, isolation, monitoring, drainage or containment policy, commissioning and an agreed response to alarms.
Record these items before a purchase order:
- required supply temperature, flow and pressure range;
- permitted coolant chemistry and water quality;
- CDU location, capacity, control ownership and redundancy;
- manifold, hose and quick-disconnect compatibility;
- leak detection, alert routing and shutdown policy;
- maintenance access and the method for isolating one rack or tray;
- who carries responsibility at every connection boundary.
A liquid-cooled rack often retains an air-cooled fraction for networking, power conversion and control electronics. The room therefore needs both liquid heat removal and residual air cooling. Ask for the split rather than assuming “liquid cooled” means zero room heat.
Rack-scale systems are bought as rack-scale systems
NVIDIA's reference guidance treats RTX PRO and HGX clusters as repeatable groups of servers, while NVL72 scales by full racks. That distinction affects procurement. A buyer can often distribute air-cooled PCIe or HGX nodes across available cabinets; an NVL72-class system arrives with an integrated rack architecture, power shelves, compute trays, switch trays and cooling dependencies.
The published GB300 NVL72 component guide lists eight 33 kW power shelves, each with six 5.5 kW PSUs, while the complete rack can require up to 142 kW. The large installed PSU capacity supports architecture and redundancy; it should not be mistaken for expected wall draw. Use the vendor's rack input requirements and approved installation design.
At larger scale, row and block planning matters more than the single cabinet. NVIDIA's DGX GB200 SuperPOD reference architecture describes an eight-rack scalable unit with a 1.2 MW TDP and hybrid liquid/air cooling. That figure shows why switch racks, management racks, storage, CDUs and facility plant belong in the same power model as compute.
Feeds, redundancy and failure state
AI racks usually use A and B feeds, but “dual fed” can describe several different failure outcomes. Ask what happens when one feed, PDU, PSU group or upstream protective device is unavailable. Can the remaining path carry the load without overload? Does equipment power-cap, shut down cleanly or continue at full output?
Do not specify breakers, conductor sizes or protective settings from a blog article. A qualified electrical engineer must design and approve them for the site, voltage, code and equipment. The buyer's job is to supply accurate load data, redundancy requirements and future-growth assumptions.
Meter at useful boundaries. Rack PDUs should expose per-feed and preferably per-outlet data; CDUs should report thermal and hydraulic conditions. Keep time-synchronised telemetry so a job slowdown can be compared with power capping, inlet temperature, flow or a feed event.
Peak demand also affects commercial power contracts and generator or UPS capacity. A rack that fits the hall may not fit the site's reserved utility capacity. Confirm this before fixing a delivery date.
Weight, depth and the delivery route
Dense GPU racks can challenge floor loading, rack static capacity and handling equipment. Obtain operating weight for the populated cabinet, not an empty-rack figure. Include power shelves, switches, coolant, manifolds, cable management and any rear-door equipment.
Check concentrated and distributed floor loads with a structural or facilities specialist. Raised floors, ramps, lift thresholds and loading-bay routes may have different limits. The narrowest doorway or weakest lift can decide whether a factory-integrated rack reaches its final position.
Server depth, rail travel and rear cable clearance matter after installation. A cabinet may accept the chassis but leave no room for fibre bends, power whips or service access. For heavy trays, document the lifting device and safe removal procedure. Staff should not improvise around a live liquid manifold or densely cabled rear door.
Network cabling consumes space
An eight-GPU node can carry many high-speed fabric links. Multiply that by a rack and the rear becomes a cable-management project in its own right. Breakout cables add branches; optical modules add heat and handling rules; rail-optimised fabrics punish incorrect port mapping.
Create the cable schedule alongside the rack elevation. It should include source and destination ports, media type, length, route, label, spare policy and allowed bend radius. Confirm that the chosen top-of-rack or end-of-row layout keeps every run within supported limits.
Network switches add electrical and thermal load, sometimes with a different airflow direction from the servers. Do not hide them in “miscellaneous”. Include management switches, console access and monitoring appliances too.
Serviceability sets a useful density
Maximum density and operable density are not the same number. If technicians cannot reach a failed PSU, drain and isolate a cooling branch, identify a cable or remove a tray, recovery takes longer and human error becomes more likely.
Run a paper service exercise before installation:
1. Identify a failed GPU tray or server. 2. Trace its power, network and cooling connections. 3. State which workloads and neighbouring equipment must stop. 4. Describe safe isolation and removal. 5. Show where the replacement and lifting equipment will stand. 6. Restore service and prove that monitoring has cleared.
If that sequence cannot be completed within the planned clearances, reduce density or change the rack arrangement. A spare U can be cheaper than repeated maintenance risk.
Three practical deployment classes
Conventional air-cooled racks
Use these for workstations, edge systems and GPU servers whose combined measured load fits the existing hall with margin. They suit sites that value familiar maintenance and do not need extreme density. Do not force a high-power eight-GPU system into a low-density room merely because one electrical outlet appears available.
High-density air or hybrid rows
These use stronger containment, higher air volume, rear-door heat exchangers or a mix of air and liquid capture. They can support substantial HGX and PCIe estates, but the exact ceiling depends on the equipment and hall. Model the whole row, because one dense rack can disturb its neighbours.
Direct-liquid-cooled rack-scale systems
NVL72-class equipment belongs here. Plan facility water or an approved secondary loop, CDUs, residual air cooling, rack delivery, commissioning and failure procedures as one project. If the site is not ready, a hosted deployment may be faster and lower risk than modifying an occupied room.
Acceptance tests before production
Commissioning should prove the promised envelope rather than merely power on the servers. Agree tests for:
- feed balance and metering under representative load;
- inlet temperatures and residual air-cooling behaviour;
- coolant temperatures, flow, pressure and leak alarms where applicable;
- GPU, CPU and memory stress while network and storage are active;
- one approved feed or cooling-component failure scenario;
- monitoring, alert routing and escalation;
- safe service access and component replacement.
Record ambient conditions, firmware, power policy, workload and duration. A short synthetic test may miss thermal soak, so the run time should reflect the equipment and acceptance objective. The OEM and facilities specialists should approve the procedure.
Questions for a rack-density review
Before GPUMachines fixes the server count, provide the rack voltage and feed limits, redundancy policy, cooling type, available liquid conditions, cabinet model, floor limits, delivery constraints and preferred maintenance window. Add the intended GPU, CPU, memory, storage and network configuration for each node.
If those inputs are uncertain, start with a smaller repeatable block or consider Buy & Host. Hosting does not remove the need for power engineering; it moves that responsibility to a site built to carry it.
FAQ
How many GPU servers fit in one rack?
There is no safe answer from U height alone. Divide the approved rack envelope by the measured or published system load, then check cooling, weight, network and service limits. The lowest result wins.
Can I use PSU ratings to calculate rack draw?
No. PSU ratings show available output capacity, often with redundancy. Use the OEM's system input data and representative wall measurements for the intended configuration.
Does liquid cooling remove all air-cooling demand?
Usually not. Fans, networking, power conversion and other components can leave a residual air-cooled fraction. Ask the system vendor for the liquid and air split.
Is a 142 kW rack expected to draw 142 kW all the time?
No. NVIDIA describes the GB300 NVL72 requirement as up to 142 kW. Actual draw depends on workload and policy, while the facility must follow the approved design envelope and failure assumptions.
Should a first deployment use maximum rack density?
Often no. Lower initial density can simplify an existing site, leave service room and provide measurements for the next block. Extreme density pays off when facility readiness and floor-space constraints justify it.
Verdict
AI rack density is the minimum of electrical, thermal, physical and operational limits. Count real configured-system loads, distinguish consumption from PSU capacity, convert the heat honestly and test failure behaviour before equipment arrives.
For a buildable plan, configure the proposed GPU block and ask GPUMachines to review it with the site's qualified electrical, cooling and facilities teams.
.jpg)