Eight Blackwell Ultra GPUs are the easy part of this buying decision. The harder question is whether the rack, fabric, storage path and operating team can keep an ASRock Rack 8U16X-GNR2 B300 productive after the installation engineers leave.
ASRock Rack builds this 8U system around NVIDIA HGX B300 NVL8. It combines eight B300 SXM GPUs, dual Intel Xeon 6 sockets, 32 DDR5 DIMM slots, twelve front NVMe bays and eight ConnectX-8 network interfaces. That makes it a serious scale-up node for model training, fine-tuning, distributed inference and GPU-heavy scientific work. It does not make it a complete AI cluster.
GPUMachines sells and configures GPU systems, so this is a commercial technical review rather than an independent laboratory test. We have checked the platform details against ASRock Rack's product material and NVIDIA's HGX B300 reference architecture. No benchmark result in this article comes from GPUMachines testing.
Open the ASRock Rack 8U16X-GNR2 B300 configurator if the fixed eight-GPU platform already matches your requirement. Keep reading if you still need to decide whether HGX B300 is the right class of machine.
Executive summary
Buy the 8U16X-GNR2 B300 when a workload needs eight GPUs to act as one tightly connected scale-up domain and the deployment team can support a dense HGX node. Large language-model training, post-training, model-parallel inference, AI-for-science and selected HPC codes fit that description.
Choose something smaller when jobs fit comfortably on one or two PCIe GPUs, when demand remains intermittent, or when the site cannot provide an agreed power and cooling envelope. A 4-GPU or 8-GPU PCIe server can cost less and offer more freedom to mix accelerator models. A hosted system may also be the more sensible route if the building cannot accept an 8U HGX platform.
The specification has four points that deserve close attention:
- NVIDIA's HGX B300 reference gives the node eight B300 GPUs and up to 2,304 GB of HBM3e, with fifth-generation NVLink and NVSwitch inside the baseboard.
- ASRock Rack uses dual Socket E2 (LGA4710) for Intel Xeon 6700P, 6500P and 6700E series processors. Earlier copy on this page incorrectly referred to AMD EPYC Turin; that was wrong.
- Eight ConnectX-8 SuperNIC interfaces supply the east-west data path. They are part of a fabric design, not eight ordinary LAN ports waiting for any available switch.
- Twelve hot-swap 2.5-inch PCIe 5.0 NVMe bays provide useful local capacity, but shared datasets and checkpoint protection still need a separate storage plan.
Verified specification
| Area | Published configuration | | --- | --- | | Form factor | 8U rackmount, 930 x 448 x 353.6 mm | | GPU platform | NVIDIA HGX B300 NVL8 with eight B300 SXM GPUs | | GPU memory | 288 GB HBM3e per GPU; 2,304 GB per eight-GPU node according to NVIDIA | | Scale-up fabric | Fifth-generation NVLink and NVSwitch; NVIDIA publishes 1,800 GB/s GPU-to-GPU bandwidth and 14.4 TB/s aggregate bandwidth for HGX B300 | | CPU platform | Two Socket E2 (LGA4710) sockets for Intel Xeon 6700P, 6500P or 6700E series processors | | System memory | 32 DIMM slots, arranged as 16 + 16; DDR5 RDIMM, RDIMM-3DS and MRDIMM support | | Front storage | 12 hot-swap 2.5-inch PCIe 5.0 x4 NVMe bays | | Internal storage | One M.2 PCIe 5.0 x2 position and two M.2 PCIe 5.0 x4 positions | | PCIe expansion | Four FHHL PCIe 5.0 x16 slots | | East-west networking | Eight OSFP interfaces through NVIDIA ConnectX-8 SuperNICs, up to 800 Gb/s per adapter | | Management and base LAN | Two 1GbE RJ45 ports through Intel i350-AM2 plus IPMI | | Power subsystem | 6 + 6 redundant 3,000 W CRPS modules, 80 PLUS Titanium | | Cooling hardware | ASRock Rack's current product brief lists twenty-six 80 mm fans and six 60 mm fans |
Specifications can change by revision and market. A final quote should name the server revision, CPU SKUs, memory population, drives, DPU, optics, cables, software entitlement and support package rather than relying on the family name alone.
What HGX B300 changes inside the node
A PCIe GPU server gives each add-in card a path back to the host and may support peer-to-peer traffic, but its topology still depends on risers, switches, CPU root complexes and the chosen cards. HGX B300 starts from a different premise: eight SXM GPUs sit on a purpose-built baseboard and communicate through NVLink and NVSwitch.
NVIDIA lists 288 GB of HBM3e per B300 GPU, giving 2.304 TB across the node. Memory remains physically distributed across the GPUs, so software still has to partition or shard models correctly; the headline capacity is not a single flat pool that every application can consume without planning. The advantage appears when a framework can keep the eight-GPU domain busy and exchange tensors quickly enough for model parallelism, collective operations or large inference services.
That is why GPU count alone is a poor selection rule. Suppose eight independent inference replicas each fit on one 96 GB PCIe card and rarely exchange data. An HGX baseboard may add cost and operational density without solving a real constraint. If one job needs hundreds of gigabytes of accelerator memory and constant GPU-to-GPU exchange, the scale-up fabric becomes much more valuable.
CPU and host-memory sizing
ASRock Rack specifies Intel Xeon 6, not AMD EPYC, for this model. The dual-socket layout supports the Xeon 6700P, 6500P and 6700E families on LGA4710. CPU choice should follow host-side work: data decoding, preprocessing, compilation, orchestration, network services and any CPU-resident stage in the pipeline.
NVIDIA's current HGX B300 reference calls for at least 48 physical CPU cores per socket and recommends 56, with a minimum 2.0 GHz base clock. Treat those figures as a platform baseline, then profile the application. A job that reads pre-tokenised data from fast shared storage may need less host compute than a multimodal pipeline decoding video or generating features at runtime.
Memory population matters just as much as nominal capacity. The chassis has 32 DIMM slots, sixteen per CPU. ASRock Rack lists RDIMM and 3DS RDIMM operation up to 6,400 MT/s at one DIMM per channel, with lower published speed at two DIMMs per channel; MRDIMM support reaches higher transfer rates in the supported one-DIMM-per-channel arrangement. The exact speed depends on CPU, DIMM type, rank, population and firmware.
NVIDIA sets 2 TB of host memory and 500 GB/s of aggregate host-memory bandwidth as minimums for its HGX B300 system reference. That is useful guidance because an under-populated memory subsystem can leave an expensive GPU node waiting on host work. Fill channels symmetrically across both processors and price the capacity you actually need, not the cheapest way to make the server boot.
Local NVMe is a working tier, not the whole data platform
Twelve PCIe 5.0 x4 NVMe bays make the ASRock Rack chassis much more useful than a design with only boot media. They can hold active shards, model files, container caches, temporary checkpoints and scratch data close to the GPUs. Three internal M.2 positions give the system room for boot volumes or other local services.
NVIDIA's reference guidance asks for at least 2 TB of local NVMe per CPU socket for training and deep-learning servers, plus a 1 TB boot drive. That is a floor, not a capacity recommendation. A real storage worksheet should record dataset size, read pattern, checkpoint size, checkpoint interval, retention, model-distribution traffic, restore target and expected simultaneous jobs.
Twelve drives also raise failure and endurance questions. Training scratch and read cache can tolerate a different protection policy from code, metadata or checkpoints. Don't quietly treat the front bays as shared durable storage simply because the raw capacity looks large. A multi-node deployment still needs a shared data tier with measured throughput and a recovery design.
Eight ConnectX-8 interfaces need a complete fabric plan
Each HGX B300 GPU has a paired ConnectX-8 path on NVIDIA's reference baseboard. ASRock Rack exposes eight OSFP interfaces, with up to 800 Gb/s per SuperNIC. NVIDIA describes a 1:1 GPU-to-NIC arrangement and publishes minimum and recommended aggregate compute-network bandwidth targets for multi-node systems.
The connectors don't select the topology for you. The project still has to decide between supported Ethernet and InfiniBand designs, allocate rails, count switch and spine ports, specify optics or copper assemblies, set cable lengths and leave enough spare capacity for failure or expansion. A fabric quote that lists switches but omits endpoint mapping and cable media is unfinished.
Keep east-west collective traffic separate from the north-south path used by storage, customer traffic and platform services when the scale warrants it. NVIDIA recommends a BlueField-3 DPU for the north-south role in HGX B300 reference systems. The chassis has four FHHL PCIe 5.0 x16 slots, but card placement, auxiliary power, airflow and CPU-root balance still need checking before the bill of materials is approved.
For one standalone node, all eight fabric ports may not be required on day one. That doesn't mean the network can be ignored. Decide whether the system will remain single-node, join a second node, or become the first member of a four-node scalable unit. Recabling a live HGX cluster later is rarely the cheap option.
Power, cooling and physical installation
ASRock Rack lists twelve 3,000 W CRPS modules in a 6 + 6 arrangement. Those PSU ratings describe the power subsystem; adding their labels together does not produce the server's expected wall draw. The configured CPUs, GPU operating mode, memory, drives, adapters, fan speed, inlet temperature and redundancy policy all affect consumption.
Before purchase, request the OEM power profile for the quoted configuration and translate it into the site's electrical design. That work should cover steady-state and peak demand, A/B feed allocation, connector and PDU requirements, breaker loading, rack position and the loss of one feed. The 930 mm chassis depth, 8U height and service mass also affect delivery, rails, lift equipment and maintenance space.
Cooling deserves the same treatment. ASRock Rack's product material lists a dense fan wall because eight B300 GPUs, two server CPUs, memory and network adapters generate a large thermal load. Confirm airflow direction, allowable inlet conditions and the data centre's pressure and heat-rejection capacity. A room described as “high density” may still lack the per-rack airflow or electrical distribution needed by several systems of this class.
Workloads that fit
The strongest case is a job whose runtime depends on scale-up communication or very large accelerator-memory capacity. Examples include large-model pre-training, substantial post-training, model-parallel fine-tuning, distributed inference with long contexts, scientific simulation and GPU-heavy analytics where the software stack supports the platform.
Service providers may also use HGX B300 for private AI or dedicated tenant environments. In that case, utilisation policy matters. Partitioning, scheduling, access control, telemetry and maintenance windows determine whether a large node behaves like a useful service or an expensive shared workstation.
One system can also serve as a development target for code intended to scale across a larger HGX estate. That can be a sound purchase, provided the team accepts that single-node results won't expose every multi-node fabric and storage failure mode.
Who should not buy it
Skip this model when the main workload consists of independent single-GPU jobs, occasional experiments or modest inference endpoints. A PCIe GPU server can offer a cleaner price and power curve for separable work, while a tower GPU workstation keeps local development close to its user.
Don't place the system in a facility that still has unresolved power or cooling assumptions. Buying first and asking the data-centre team to “make it fit” later creates ugly compromises: reduced operating modes, stranded fabric ports, delayed commissioning or a server that has to move before it has done useful work.
Teams without a scheduler, platform owner or maintenance plan should pause too. Hardware does not create an operating model.
Sensible configuration paths
Single-node model development
Use balanced Xeon 6 CPUs, populate every memory channel and provide enough host RAM for the intended data pipeline. Size local NVMe for active models, repeatable datasets and checkpoints, then retain a protected copy elsewhere. The network can start smaller than a cluster fabric, but document the upgrade path and reserve the required slots, ports and rack space.
Four-node HGX B300 block
NVIDIA defines four HGX nodes as one scalable unit in its enterprise reference architecture. At this point, the east-west fabric, north-south DPU path, control plane and storage endpoints belong in the same design. Specify switch redundancy, rail mapping, optics, cable schedule, rack placement and acceptance tests before ordering the first node.
Inference platform
Start with model topology, precision, context distribution, concurrent requests and latency target. Eight B300 GPUs may suit large model-parallel services, but smaller replicas can favour PCIe GPUs or several less dense nodes. Include model loading, key-value cache strategy, observability and failure recovery in the sizing discussion.
Hosted deployment
Hosting can remove the immediate building constraint, but it does not remove architecture work. Agree on network access, data transfer, remote hands, security boundary, monitoring, support ownership and the process for replacing failed components. GPUMachines can review the on-premises and hosted routes against the same workload data.
Cluster and AI Factory fit
One 8U16X-GNR2 B300 is a compute node. An AI Factory adds the east-west fabric, storage, management services, user access, software, security, power, cooling and acceptance process that turn multiple nodes into a production platform.
NVIDIA's enterprise reference architecture scales HGX B300 from four nodes to 128 nodes in four-node increments, representing 32 to 1,024 GPUs. That published range doesn't mean every organisation should aim for it. It gives designers a repeatable growth unit and a point at which dedicated network planes, spine-leaf switching and operational controls become unavoidable.
For deployments beyond a single node, review the GPUMachines AI Factory profiles alongside the server configuration. The output should be a versioned bill of materials and topology, not a loose bundle of GPU servers and switches.
Procurement questions worth answering
Ask these before approving a quote:
- Which exact Xeon 6 processors and DIMMs are qualified, and what memory speed results from the proposed population?
- What is the configured system power profile, and which feed, PDU, connectors and redundancy assumptions support it?
- Which ConnectX-8 ports will carry each rail, and what switches, optics, fibre lengths and spare ports complete the fabric?
- Where do training data, model artefacts and checkpoints live; how quickly can a failed job restore them?
- Which BlueField configuration handles north-south traffic, and who owns its firmware and operating mode?
- What hardware, firmware, fabric, storage and workload tests define acceptance?
Written answers protect both buyer and supplier. They also expose an imbalanced design before it reaches the loading bay.
Alternatives
Compare the HGX server range when the workload needs scale-up GPUs but the preferred CPU vendor, cooling method, service model or chassis layout differs. Compare PCIe GPU servers for independent inference, rendering, visual computing or mixed accelerator pools.
A deskside AI system or workstation can be a better engineering tool for one team, even though it has much less throughput. Public GPU Cloud or a dedicated hosted node can be preferable for uncertain demand. The right alternative is the one that removes the actual bottleneck without importing a facility problem.
FAQ
Does the 8U16X-GNR2 B300 use AMD EPYC CPUs?
No. This exact model uses two Intel Socket E2 (LGA4710) processors from the Xeon 6700P, 6500P or 6700E series. ASRock Rack has other GPU systems built around AMD EPYC, which is probably how the earlier error entered the article.
Does 2.3 TB of HBM3e behave as one ordinary memory pool?
No. Eight B300 GPUs provide 2,304 GB in total, but software must still account for the distributed GPU topology. NVLink and NVSwitch provide the fast scale-up path; they don't remove framework, sharding or model-parallel design work.
Are all twelve NVMe bays suitable for checkpoint storage?
They can provide a fast local tier, subject to drive qualification and the chosen protection policy. Durable checkpoints should also have an agreed shared or replicated destination so a node failure does not take the only useful copy with it.
Does every deployment need all eight 800 Gb/s ports connected?
No, particularly for a standalone node. Multi-node performance depends heavily on the fabric, though, so the port plan must follow the workload and expansion target. Leaving the decision blank is not a design.
How much host RAM should be installed?
NVIDIA's HGX B300 reference sets a 2 TB minimum and a 500 GB/s aggregate host-memory-bandwidth minimum. Some data pipelines and CPU-heavy workflows need more. Populate both sockets and their memory channels symmetrically.
Can GPUMachines host the server?
GPUMachines can scope a hosted or Buy & Host deployment, subject to current capacity and a technical review. The quote should state power, network, remote access, support and data-transfer assumptions rather than treating hosting as a generic add-on.
Verdict
ASRock Rack's 8U16X-GNR2 B300 is a credible HGX B300 node for buyers who need an eight-GPU Blackwell Ultra scale-up domain and can operate it properly. Its dual Xeon 6 host, 32 DIMM slots, twelve front NVMe bays and eight ConnectX-8 paths give an architect enough room to build a balanced system.
That balance is the purchase. If the workload does not need NVLink-connected GPUs, choose a less dense platform. If it does, treat the network, storage and facility plan as part of the server rather than work to be finished later.
Configure the ASRock Rack 8U16X-GNR2 B300 and send GPUMachines the model topology, precision, dataset path, checkpoint pattern, user count and deployment location for a technical review.
Sources
Specifications and reference guidance checked 21 September 2026:
.jpg)