GPUmachines

ASRock Rack 4U16X-GNR2/ZC B300 Technical Review

A source-checked review of ASRock Rack's 4U HGX B300 server, including its waterless two-phase cooling, 12-drive topology and eight ConnectX-8 fabric links.

ASRock Rack 4U16X-GNR2/ZC B300 Technical Review

The two letters after the slash are the first thing to check on this server. ASRock Rack's 4U16X-GNR2/ZC B300 uses ZutaCore HyperCool, a two-phase dielectric cooling system with no water in the server-side loop. The similar 4U16X-GNR2/DLC B300 is a separate single-phase design. Treating them as one product would hide the most important facility and service decision.

Behind that cooling choice sits a fixed NVIDIA HGX B300 NVL8 baseboard with eight Blackwell Ultra GPUs, fifth-generation NVLink and NVSwitch, and eight ConnectX-8 fabric adapters. This is not a 4U PCIe chassis into which buyers add a preferred set of cards. The accelerator domain arrives as the platform around which the Xeon, memory, storage, network and rack design must be built.

This technical review uses current ASRock Rack, NVIDIA and ZutaCore sources. GPUMachines has not claimed hands-on benchmark testing of the unit.

Configure the ASRock Rack 4U16X-GNR2/ZC B300 or compare the wider HGX server range.

Verified specification

| Area | Verified detail | | --- | --- | | Chassis | 4U rackmount; current product page lists 919 x 435 x 175 mm | | Cooling | ZutaCore HyperCool two-phase waterless direct-to-chip cooling | | GPU platform | NVIDIA HGX B300 NVL8 with 8 fixed Blackwell Ultra GPUs | | Scale-up fabric | Fifth-generation NVLink and NVSwitch; 14.4 TB/s aggregate NVLink bandwidth in NVIDIA's HGX reference | | Host processors | 2 x LGA4710; qualified Xeon 6700P, 6500P and 6700E families | | Host memory | 32 DDR5 slots, 16 per socket | | RDIMM | Up to 128GB per module; up to 6400 MT/s at 1DPC and 5200 MT/s at 2DPC | | RDIMM-3DS | Up to 256GB per module; up to 6400 MT/s at 1DPC and 5200 MT/s at 2DPC | | MRDIMM | Up to 64GB per module and 8000 MT/s at 1DPC | | Front storage | 10 x 2.5-inch Gen5 NVMe through the PCIe switch plus 2 x Gen5 NVMe from CPU1 | | Internal storage | 1 x M.2 22110/2280, PCIe Gen5 x2 from CPU0 | | Expansion | 2 x FHHL PCIe Gen5 x16 from the PCIe switch | | GPU fabric | 8 x OSFP through NVIDIA ConnectX-8 SuperNICs | | Host network | 2 x Intel i350-AM2 1GbE RJ45 plus dedicated IPMI | | Power | 5+5 3000 W 80 PLUS Titanium CRPS modules | | Cooling fans | 2 x 80 mm PWM plus 8 x 60 mm PWM fans | | Management | ASPEED AST2600 and IPMI |

ASRock Rack's current product page and its Q2 catalogue disagree on chassis length and width. The page gives 919 x 435 x 175 mm; the catalogue gives 900 x 448 x 175 mm. That is not a harmless editorial difference when rail fit, rear clearance and manifold position are being designed. Use the current page for initial planning and request the ordered system's mechanical drawing before approving the rack.

/ZC means a different thermal architecture

ZutaCore HyperCool uses a dielectric heat-transfer fluid at the cold plate. Heat causes the fluid to boil, the vapour moves away from the chip, and the cooling system condenses and recirculates it. The phase change carries concentrated GPU and CPU heat without water inside the server cooling circuit.

That distinction can be useful in facilities concerned about bringing a water loop close to electronics. It does not make the node independent of the data centre. Heat still has to reach a cooling distribution unit and then leave the building. Pumps, condensers, controls, alarms, quick disconnects, manifolds and the final heat-rejection path all need a design and an owner.

The exact scope of the ordered ZutaCore equipment should be written into the quote. Confirm whether the cooling distribution unit, rack manifold, monitoring software, hoses, commissioning and service contract are included. A server shown with cold plates and connectors is not yet an operational cooling plant.

The /DLC system uses a single-phase direct liquid design. It may suit an established water-based facility, while /ZC may suit a site that wants a dielectric server loop. That is a facility choice, not a performance ranking. The correct answer depends on the site's existing cooling, service team, water policy, heat-reuse plan and deployment schedule.

Eight B300 GPUs are fixed, not selectable

The HGX B300 baseboard contains eight Blackwell Ultra GPUs. NVIDIA's current enterprise reference architecture lists up to 2,304 GB of HBM3e across the node, derived from eight 288 GB devices, and up to 8 TB/s of memory bandwidth per GPU. NVIDIA's DGX marketing pages sometimes use a rounded 2.1 TB total. Buyers should use the exact baseboard and software documentation supplied with the ordered system instead of merging numbers from adjacent DGX and partner pages.

Fifth-generation NVLink and two NVSwitch devices form the local scale-up domain. NVIDIA lists 14.4 TB/s of aggregate NVLink bandwidth and 1,800 GB/s of GPU-to-GPU bandwidth for HGX B300. This is the fast path for collectives and tensor movement within one node. It is separate from the OSFP cluster links used between nodes.

The architecture suits training, fine-tuning, high-throughput inference and HPC applications that can use an eight-GPU memory and communication domain. It is not automatically faster for every job. Small models, CPU-heavy pipelines, single-user development and software that scales poorly across GPUs can leave much of the platform idle.

The configurator should therefore present one fixed NVIDIA HGX B300 NVL8 accelerator, not a menu containing B100, B200 or GB-series PCIe choices. Those are different products and topologies. Replacing the fixed board would change the server rather than configure it.

CPU selection stops at the published Xeon families

ASRock Rack specifies dual LGA4710 sockets for Xeon 6700P, 6500P and 6700E processors. The platform page does not list Xeon 6900P. A shared socket label is insufficient evidence to add that family.

Host CPUs handle much more than booting the GPUs. Data loading, decompression, tokenisation, augmentation, storage queues, network interrupts, checkpoint handling, orchestration and post-processing can all create host-side limits. A training build with eight 800 Gb/s fabric ports should not be paired with an arbitrary minimum CPU simply because the model runs on GPUs.

Core count is only one variable. P-core models can suit host work that values stronger per-core throughput, while E-core processors can supply high thread density for services and data preparation. The exact CPU pair must appear on ASRock Rack's support list and match the memory and thermal build.

NVIDIA's HGX B300 enterprise guidance specifies at least 48 physical cores per socket and recommends 56. That is a reference-architecture requirement for a balanced certified node, not proof that every listed ASRock Rack CPU is equally suitable for every workload. Use it as a design floor when building to that reference architecture.

Thirty-two DIMM slots need a channel plan

The motherboard exposes 16 DIMM slots per CPU. RDIMMs and RDIMM-3DS modules can run at up to 6400 MT/s with one DIMM per channel and 5200 MT/s at two DIMMs per channel. MRDIMMs are limited to 64GB per module and up to 8000 MT/s at one DIMM per channel.

Those limits matter in the configurator. A 128GB MCR DIMM option exceeds the model's published MRDIMM capacity and should be removed. Modules sold with an 8800 MT/s rating may be able to operate at a lower platform speed, but the label does not prove QVL support. Confirm the exact Micron part, BIOS and negotiated speed rather than promising 8800 MT/s operation.

NVIDIA starts its HGX B300 reference node at 2 TB of system memory and asks for symmetric population across sockets and memory channels. Training input pipelines, large CPU-side caches, retrieval services and many concurrent inference workers may justify more. A narrow HPC application could need less. Capacity and bandwidth should follow measured working sets rather than a round number.

Twelve front drives use two PCIe paths

The server has 12 front 2.5-inch Gen5 NVMe carriers. Ten connect through a PEX89104 PCIe switch. Two connect directly to CPU1. They are all NVMe bays, but they are not topologically identical.

That distinction can affect device affinity, latency, aggregate throughput and failure analysis. A workload that stages data through CPU0 may see a different path to the two CPU1-attached devices. Benchmark the final filesystem and placement strategy rather than treating the front row as twelve equivalent labels.

One M.2 22110/2280 connector hangs from CPU0 through PCIe Gen5 x2. It is separate from the front carriers and can hold a boot or service device. One drive does not provide boot redundancy. If the deployment needs mirrored boot media, define a supported arrangement instead of assuming the single M.2 slot supplies it.

Local NVMe should be assigned a job: boot, active dataset cache, checkpoints, shuffle, spill, logs or temporary scratch. A 12-drive node does not replace shared storage, data protection or an ingest path. The design should show how data enters the node, how it is reused across jobs and what happens after a drive or host failure.

ConnectX-8 is an integrated GPU fabric

Eight OSFP ports connect through NVIDIA ConnectX-8 SuperNICs. NVIDIA's HGX AI Factory reference calls this a 2-8-9-800 node: two CPUs, eight GPUs, nine network adapters and up to 800 Gb/s of east-west bandwidth per GPU. The ninth adapter in the reference design is a separate north-south DPU. ASRock Rack's public specification does not state that such a DPU is included.

Do not read 8 x 800 Gb/s as eight ordinary user-selected NICs. These links are part of the GPU fabric design. The exact adapter OPN, Ethernet or InfiniBand mode, breakout, optics, cables, switch platform, rail assignment and topology must be confirmed together.

The public ASRock Rack page identifies ConnectX-8 and OSFP but does not publish the precise OPN or GPU-to-port affinity. The product can be verified at model-family level while those network subfields remain limited. Inventing a rail map would be worse than leaving a clear engineering check.

Two Intel i350-AM2 1GbE RJ45 ports provide host connectivity. ASRock Rack shows the same pair at the front and rear as shared I/O, so they are two physical ports, not four. A dedicated IPMI path reaches the AST2600 BMC. Production, storage, cluster and out-of-band traffic should remain separate where the network design permits it.

Two FHHL PCIe Gen5 x16 slots remain for qualified north-south networking, a BlueField DPU, storage connectivity or another supported card. Check card power, auxiliary connectors, airflow, firmware and slot sharing before assigning them.

Power planning starts near 15 kW, not at the PSU sum

ASRock Rack fits ten 3000 W Titanium supplies in a 5+5 arrangement. Ten multiplied by 3000 W is installed nameplate capacity, not a 30 kW normal load. The feed, redundancy and failover design needs confirmation from the ordered electrical documentation.

NVIDIA specifies 14.5 kW consumption and a 15 kW system maximum for its own DGX B300. That is a useful planning analogue, not an ASRock Rack measurement. The two systems use the same class of eight-GPU baseboard but have different chassis, storage and cooling.

GPUMachines currently assigns 13,200 W to the fixed ASRock Rack platform before selectable CPUs, memory and drives. Adding two high-power Xeons, 32 DIMMs, NVMe media and conversion losses places a complete build in the expected 14 to 15 kW class. Keep the value labelled as an engineering estimate until the delivered server supplies measured idle, representative sustained-load and peak readings.

At 14.5 kW, one node rejects about 49,500 BTU/h before any external cooling overhead, using the NVIDIA DGX figure as an order-of-magnitude cross-check. Rack design must account for both IT power and the cooling system that moves that heat out of the white space.

Redundancy must extend beyond the chassis. Separate A and B feeds, suitable PDUs, breaker coordination, cable ratings and enough capacity on the surviving path matter more than the count of PSU modules. An upstream single point of failure makes 5+5 less useful.

Rack and service checks

The current product page lists 919 x 435 x 175 mm. The Q2 catalogue lists 900 x 448 x 175 mm. Request the mechanical drawing, rail part, loaded weight and centre-of-gravity information for the exact order. Also reserve space for cooling connectors, power cords, OSFP cables and their bend radius.

Service paths deserve equal attention. Identify which components can be replaced without draining or opening the cooling loop, who is authorised to handle the dielectric fluid, and how a failed pump, cold plate or quick disconnect is diagnosed. Record firmware ownership for the BMC, CPUs, NVSwitch, GPUs, ConnectX-8 adapters and storage devices.

A single 4U node can be a substantial commissioning project. A multi-node row adds switch ports, optics, cable trays, cooling capacity, power balancing, orchestration and acceptance testing. Validate one representative node and fabric path before multiplying the design.

Workloads that fit

Large-model training and fine-tuning are direct uses when model and batch strategy can exploit eight GPUs and the cluster network can feed multiple nodes. High-throughput inference can also fit, particularly where many requests, long contexts or large models need a broad memory and communication domain.

HPC applications with strong GPU scaling and substantial peer-to-peer exchange can benefit from NVLink. Research groups can consolidate several demanding projects on one system, but only if scheduling, isolation, storage and user access are designed alongside the hardware.

The platform is too large for most proof-of-concept work. A PCIe GPU server, tower workstation or hosted instance is often a cleaner first step when utilisation, software scaling or facility readiness is uncertain. Buying fewer GPUs is sound engineering when the workload cannot use eight.

Procurement checklist

Before approving a 4U16X-GNR2/ZC B300 build, confirm:

  • The order is the /ZC two-phase waterless system, not the /DLC single-phase model.
  • The exact HGX B300 board memory figure, firmware baseline and supported software stack are documented.
  • Both Xeon processors appear on the current ASRock Rack support list.
  • The DIMM plan follows the 1DPC or 2DPC speed limits and the 64GB MRDIMM maximum.
  • The ten switch-attached, two CPU1-attached and one M.2 storage paths are represented correctly.
  • ConnectX-8 OPNs, protocol mode, rail mapping, optics, cables and switch ports are part of the network bill of materials.
  • Any BlueField or other north-south adapter is explicitly included and fits one of the two FHHL slots.
  • Electrical feeds, PDU outlets, breaker capacity and failover can support the measured node load.
  • ZutaCore cooling distribution, manifolds, controls, heat rejection, commissioning and service ownership are specified.
  • The mechanical drawing resolves the published dimension conflict before rails and rack space are ordered.

Frequently asked questions

Is this an air-cooled B300 server?

No. The /ZC model uses ZutaCore HyperCool two-phase waterless direct-to-chip cooling. It still contains fans for components and airflow, but the high-heat devices use the liquid cooling system.

Can the eight B300 GPUs be changed to other cards?

No. They are fixed on the HGX B300 NVL8 baseboard. A server with selectable PCIe accelerators is a different product class.

Does waterless mean that no cooling infrastructure is needed?

No. Water is absent from the server-side dielectric loop, but a cooling distribution and heat-rejection path is still required. Confirm exactly which ZutaCore equipment and services are included.

Why are there two groups of front NVMe bays?

Ten bays connect through the PCIe switch and two connect to CPU1. Keeping the groups separate makes host affinity and performance testing clearer even though all 12 accept 2.5-inch Gen5 NVMe media.

Is each ConnectX-8 link guaranteed to run at 800 Gb/s?

NVIDIA's reference architecture supports up to 800 Gb/s per adapter. The exact protocol, breakout, optics, cable and switch configuration determine the ordered operating mode.

Is 13.2 kW the final power draw?

No. It is a GPUMachines fixed-platform planning allowance. The completed node also includes CPUs, memory, storage and conversion losses. Use measured data from the delivered system for rack capacity management.

Buying assessment

The ASRock Rack 4U16X-GNR2/ZC B300 is a specialised way to deploy an eight-GPU Blackwell Ultra domain where waterless two-phase cooling fits the facility. Its value is not simply 4U density. The product combines the HGX scale-up fabric, integrated 800 Gb/s-class east-west networking and a thermal design that differs materially from ordinary DLC systems.

That density makes omissions expensive. The quote must resolve cooling-distribution scope, ConnectX-8 mode and rail mapping, electrical feeds, storage affinity and the conflicting chassis dimensions. It should also remove generic PCIe GPU selections and unsupported 128GB MCR DIMMs from the configurator.

Choose this model when the workload can use eight tightly coupled B300 GPUs and the site can operate the complete power, fabric and cooling stack. Choose a smaller server or hosted deployment when those conditions are not yet proven.

Configure the ASRock Rack 4U16X-GNR2/ZC B300 with an exact CPU, memory, NVMe, network and facility plan. Compare other fixed-baseboard systems in the HGX server category.

Official sources

Sources were checked on 22 September 2026. Confirm the exact mechanical drawing, cooling-distribution scope, CPU and DIMM QVL, storage topology, ConnectX-8 OPNs, protocol mode, optics, cables, power feeds and firmware on the final quote.

← Back to blog