The ASRock Rack 6U8X-TURIN2 SYN H200 is a 6U NVIDIA HGX H200 server with eight fixed SXM GPUs, NVLink and NVSwitch, dual AMD EPYC processors, 24 DDR5 DIMM slots and twelve front Gen5 NVMe bays. It is an NVIDIA-qualified scale-up platform for workloads that need all eight GPUs to exchange data at NVLink speed.
This is not an eight-card PCIe server and it does not use H200 NVL add-in boards. The eight H200 SXM GPUs are part of one HGX baseboard. Each GPU has 141 GB of HBM3e, giving the node about 1.1 TB of aggregate GPU memory, while NVLink and NVSwitch form the internal GPU fabric.
Configure the ASRock Rack 6U8X-TURIN2 SYN H200 on GPUMachines only after the network, storage and facility design are understood. An HGX server without enough data movement or rack power is an expensive underused node.
Technical review summary
The fixed accelerator subsystem is NVIDIA HGX H200 8-GPU. NVIDIA publishes 141 GB of HBM3e and 4.8 TB/s of memory bandwidth per H200 SXM GPU, or roughly 1.1 TB and 38.4 TB/s across the eight-GPU node. The memory remains physically distributed across eight GPUs; software must shard models and data across the NVLink fabric.
ASRock Rack pairs HGX H200 with two socket SP5 processors. The system page lists AMD EPYC 9005, 9004, selected 3D V-Cache and EPYC 97x4 processors. The current NVIDIA-qualified entry in ASRock Rack's support list names EPYC 9005, so a buyer requiring the qualified reference configuration should confirm the exact CPU generation rather than relying on the broader motherboard capability.
Twenty-four DDR5 sockets provide twelve memory channels per processor at one DIMM per channel. Host memory is important for data preparation, caches, services and checkpoint handling even though the main model state sits in GPU HBM.
The SYN suffix refers to PCIe switch synthetic mode. ASRock Rack positions it to improve the paths among GPUs, host CPUs, GPUDirect RDMA network adapters and GPUDirect Storage devices. It does not replace NVSwitch: NVSwitch carries GPU-to-GPU scale-up traffic, while the PCIe fabric connects host, storage and external network devices.
Local storage comprises eight front Gen5 x4 NVMe bays connected through a PCIe switch and four front Gen5 x4 NVMe bays connected from the CPUs. Two internal M.2 sockets support Gen3 NVMe or SATA at x2 and x4 widths. These paths should be measured separately because their locality and contention are different.
Eight HHHL Gen5 x16 positions and one FHFL Gen5 x16 position can take qualified network, DPU or storage adapters. Onboard networking is only dual 1GbE plus IPMI, so production fabric NICs are a required design choice rather than an included high-speed network.
Power comes from eight Titanium CRPS units in a 4+4 arrangement. ASRock Rack publishes 3,002.4 W per supply at 220 to 240 V and 2,900 W at 200 to 220 V. The final rack design needs the exact feed and redundancy interpretation approved for the ordered build.
Verified system specification
| Area | ASRock Rack 6U8X-TURIN2 SYN H200 specification | | --- | --- | | Form factor | 6U rackmount | | Dimensions | 930 x 448 x 264.7 mm | | GPU subsystem | NVIDIA HGX H200 8-GPU baseboard | | GPU memory | 141 GB HBM3e per GPU; about 1.1 TB across eight GPUs | | GPU memory bandwidth | 4.8 TB/s per GPU; 38.4 TB/s aggregate | | GPU interconnect | Fourth-generation NVLink and third-generation NVSwitch | | Host processors | 2 x AMD EPYC 9005 or 9004, socket SP5 | | Current NVIDIA-qualified support-list entry | AMD EPYC 9005 | | Host memory | 24 x DDR5 DIMM, twelve channels per CPU, 1DPC | | Front NVMe through PCIe switch | 8 x hot-swap 2.5-inch Gen5 x4 | | Front NVMe from CPUs | 4 x hot-swap 2.5-inch Gen5 x4 | | Internal storage | 1 x M.2 Gen3 x2 or SATA; 1 x M.2 Gen3 x4 or SATA | | Expansion | 8 x HHHL Gen5 x16 and 1 x FHFL Gen5 x16 | | Onboard network | 2 x 1GbE RJ45 through Intel i350-AM2 | | Management | Dedicated IPMI and BMC management | | Power supplies | 4+4 x 3,000 W-class 80 PLUS Titanium CRPS | | Published PSU output | 3,002.4 W at 220-240 V; 2,900 W at 200-220 V |
These are platform specifications, not a complete rack bill of materials. Fabric adapters, switches, optics, cables, storage endpoints, PDUs and cooling are separate decisions.
HGX H200 is one eight-GPU scale-up domain
HGX H200 integrates eight SXM GPUs on a baseboard with NVLink and NVSwitch. NVIDIA states up to 900 GB/s bidirectional NVLink bandwidth per H200 GPU, far above its PCIe host link. NVSwitch lets every GPU reach the other GPUs through the internal fabric.
This matters for workloads that repeatedly exchange tensors, gradients, experts or key-value cache across the node. Examples include:
- large language model pre-training;
- full or parameter-efficient fine-tuning across several GPUs;
- inference using tensor or pipeline parallelism;
- mixture-of-experts models with heavy communication;
- scientific simulation with tightly coupled GPU domains; and
- high-throughput data analytics written for multi-GPU execution.
Eight GPUs do not behave as one ordinary 1.1 TB memory allocation. A framework such as PyTorch, JAX or an inference engine must divide model state and work. The aggregate memory figure is useful for capacity planning, but communication, precision, activations and KV cache still determine whether a model fits and runs efficiently.
HGX earns its cost when the software uses the fabric. Independent single-GPU jobs can run on the node, but a dense PCIe server or several smaller nodes may provide a better commercial and scheduling fit if jobs rarely communicate.
H200 SXM versus H200 NVL
The names are easy to confuse. H200 SXM is the GPU used on the HGX baseboard. H200 NVL is a dual-slot PCIe card that can use two- or four-way NVLink bridges in supported systems.
6U8X-TURIN2 SYN H200 uses eight H200 SXM GPUs and NVSwitch. It is not built from eight H200 NVL cards. The difference affects:
- GPU form factor and service procedure;
- scale-up topology;
- power and cooling;
- supported software stack; and
- the way the server is priced and configured.
The HGX accelerator subsystem ships as a fixed platform. Customers should not be offered an individual GPU quantity selector for this product. CPU, memory, storage and network remain configurable around the fixed eight-GPU baseboard.
What synthetic-mode PCIe switching does
The HGX NVLink fabric handles GPU-to-GPU traffic. Data still has to enter and leave the node through CPUs, NVMe drives and network adapters. Those devices connect through PCIe.
ASRock Rack describes PCIe switch synthetic mode as an optimisation for GPU-to-CPU, GPUDirect RDMA and GPUDirect Storage paths. In practical terms, the platform presents and routes devices through its PCIe switch fabric so that qualified NICs and storage can communicate efficiently with the GPU subsystem.
Synthetic mode should not be treated as a universal performance guarantee. Results depend on device placement, firmware, IOMMU policy, operating system, driver, GPUDirect support and application behaviour. The delivered topology should be inspected and benchmarked with the exact NICs and drives.
Acceptance should verify:
- GPU Direct Storage from each front NVMe pool;
- GPUDirect RDMA through every fabric adapter;
- CPU-to-GPU bandwidth and NUMA locality;
- peer-to-peer access and Fabric Manager state; and
- performance with the intended container and framework versions.
This is one of the areas where a source-backed design review cannot replace commissioning tests.
CPU generation and qualification
The system page supports EPYC 9005 and 9004 families, including selected high-cache and dense-core models. ASRock Rack's GPU support list currently records this H200 system as NVIDIA-qualified with EPYC 9005.
That distinction matters for buyers who require the exact NVIDIA-qualified configuration for support or procurement policy. A processor may be electrically supported by the motherboard without belonging to the current qualification entry.
CPU choice should follow host work:
- high-frequency models can help preprocessing and serial orchestration;
- high core counts suit heavy data transformation, compilation and many services;
- large cache can benefit selected simulation and analytics tasks; and
- power-efficient host choices can leave more rack and cooling capacity for the fixed GPUs.
The GPU investment is large enough that the host should be measured, not guessed. Profile tokenisation, data loading, decompression, feature processing and service concurrency. A CPU bottleneck can leave H200 GPUs waiting even when the NVLink fabric is correctly configured.
Twenty-four DIMMs and host memory
Twelve channels per EPYC processor give the platform the full SP5 memory interface. Populate both sockets symmetrically with modules from the current memory QVL.
Host RAM is used for:
- dataset and filesystem cache;
- preprocessing and augmentation;
- model loading and staging;
- CPU-based services and databases;
- checkpoints and fault recovery;
- containers, orchestration and monitoring; and
- virtualisation where supported.
NVIDIA's reference designs often use substantial host memory beside HGX, but capacity should follow the deployed software. A fixed ratio to GPU HBM is only a starting point. Memory bandwidth can matter as much as capacity, so avoid sparse channel population on a 24-slot board.
EPYC 9005 and 9004 have different supported memory speeds, and the selected DIMM rank and population influence the result. Use one qualified module family across the node unless the vendor explicitly approves another pattern.
Twelve front Gen5 NVMe bays
The front storage is divided into two groups:
- eight hot-swap Gen5 x4 NVMe bays reached through the PCIe switch; and
- four hot-swap Gen5 x4 NVMe bays connected from the host CPUs.
Both groups accept NVMe devices, but they are not topologically identical. CPU-direct drives may suit operating data or services tied closely to a host socket. Switch-attached drives can provide a large local tier close to the accelerator I/O fabric. The best use depends on the synthetic-mode topology and application.
Possible roles include:
- local training shards;
- model and container cache;
- checkpoint staging;
- inference model repositories;
- high-speed scratch; and
- spill or preprocessing space.
Twelve drives do not remove the need for shared storage. Important datasets and checkpoints need replication, backup and movement between nodes. Multi-node training also needs a storage system that can feed many servers at once.
Use enterprise NVMe devices with appropriate endurance, power-loss protection and thermal qualification. Size write endurance from checkpoints and data preparation, not only from model size.
Two M.2 boot paths
The internal M.2 devices have different PCIe widths: one Gen3 x2 and one Gen3 x4, with SATA mode also listed. They are suitable for boot, recovery or service images.
Do not combine their capacities and present them as identical high-speed scratch. The interfaces differ, and internal access is less convenient than front hot-swap service. A mirrored operating-system design should be confirmed against the controller and firmware behaviour of the selected devices.
Keep active training data on the front NVMe tier or shared storage. The internal M.2 pair should remain a dependable platform service rather than a hidden performance dependency.
Nine Gen5 x16 expansion positions
Eight half-height, half-length Gen5 x16 positions and one full-height, full-length Gen5 x16 position provide room for fabric adapters, DPUs and storage cards. The HGX GPUs are already fixed on their baseboard, so these slots are not customer GPU positions.
A common scale-out design assigns one high-speed network rail per GPU or per topology group. The exact number of NICs depends on adapter generation, port mode and the surrounding switch fabric. The chassis has physical capacity, but every adapter still consumes power, airflow and cables.
Draw the complete slot map with:
- NIC or DPU model and firmware;
- which PCIe switch or CPU path serves it;
- Ethernet or InfiniBand mode;
- port speed and breakout;
- optic, cable and switch-port count; and
- rail-to-GPU affinity.
The single full-length position may be useful for a large DPU or storage adapter, subject to qualification. Do not reserve all nine slots abstractly and decide their purpose after the server arrives.
Onboard 1GbE is management-grade
Dual Intel i350 1GbE ports are useful for provisioning, service traffic and low-rate application access. They cannot feed twelve NVMe drives or eight H200 GPUs under production load.
High-speed network adapters are therefore part of the system design. Depending on deployment, the node may need:
- east-west GPU fabric for multi-node training;
- a storage fabric for datasets and checkpoints;
- client or inference traffic; and
- a separate management network.
These can share physical adapters in smaller installations, but the choice should be deliberate. For AI factories, separate rails and redundant leaf switches are common. Switch count, oversubscription, optics and cable reach must be sized with the server ports.
GPUDirect RDMA needs a supported NIC, driver, firmware and topology. Installing a fast adapter is not enough. Verify peer-memory access and end-to-end throughput during commissioning.
Power, cooling and rack deployment
Eight 3,000 W-class Titanium supplies are arranged 4+4. The published per-supply output is 3,002.4 W at 220 to 240 V and 2,900 W at 200 to 220 V. The server belongs on high-voltage data-centre power with a documented A/B feed and redundancy plan.
Do not add eight PSU nameplates and call the result server consumption. Some supplies provide redundancy, and actual draw depends on GPU, CPU, memory, NVMe, NIC and fan load. Use the manufacturer's configured maximum and measured acceptance data for rack planning.
At 930 mm deep and 6U high, the chassis needs a suitable rack, rail set, rear cable clearance and service space. Eight fabric adapters can add a dense cable bundle. Plan bend radius and label both ends before installation.
Cooling review should include:
- cold-aisle inlet temperature;
- sustained H200 and CPU load;
- fan and PSU failure behaviour;
- rack-level heat rejection;
- blanking and containment;
- the effect of optics and NICs on rear temperature; and
- neighbouring high-density equipment.
This is not an office server and should not be placed in a lightly cooled general IT cupboard. The node, switches and storage have to be designed as one rack system.
H200 in the Blackwell era
H200 is no longer NVIDIA's newest HGX generation, but that does not make it obsolete. Its 141 GB per GPU, mature Hopper software stack and broad application support can make it a sensible production platform when the workload is validated on H200 and acquisition economics are favourable.
HGX B200 and B300 offer more memory, newer numeric formats and higher platform capability. They also change power, cooling, network and purchase cost. A workload that meets its target on H200 does not automatically benefit enough from Blackwell to justify migration.
Compare measured job time, supported precision, memory fit, software readiness and rack cost. Existing H100/H200 operational knowledge and container images can reduce deployment risk. Conversely, a new long-lived AI factory may favour Blackwell if the models and facility are ready for it.
The decision should not be reduced to “newer is better” or “H200 is cheaper”. Use the complete workload and infrastructure model.
Who should buy it?
6U8X-TURIN2 SYN H200 is a good candidate for:
- teams training or fine-tuning models across eight tightly connected GPUs;
- inference platforms using tensor or pipeline parallelism;
- HPC applications written for NVLink-scale GPU communication;
- private AI clusters standardised on Hopper;
- buyers who need twelve local Gen5 NVMe bays; and
- data centres able to supply high-voltage power, dense cooling and fast network fabrics.
When not to buy it
Do not buy this server for occasional single-GPU experiments, lightly used independent inference workers or a site without the required power and cooling. A PCIe GPU server, workstation or hosted service may use capital more efficiently.
It is also a poor fit when the organisation has no plan for high-speed NICs, switches and storage. Onboard 1GbE cannot support the accelerator subsystem. Buyers needing the newest Blackwell formats or much larger per-node HBM should compare current B200 and B300 systems.
Configuration checklist
Before approving the system, record:
1. the exact ASRock Rack model and NVIDIA-qualified configuration; 2. the EPYC generation and CPU OPNs; 3. a symmetric 24-DIMM population; 4. the fixed HGX H200 baseboard and software entitlement; 5. use of the eight switch-attached and four CPU-attached NVMe bays; 6. the two M.2 boot devices; 7. all NICs or DPUs and their PCIe locality; 8. switch, optic, cable and rail mapping; 9. shared-storage capacity and throughput; 10. rack voltage, PDUs, A/B feeds, depth and cooling; and 11. firmware, Fabric Manager, driver, container and acceptance versions.
Commissioning should include NVLink health, Fabric Manager, GPU errors, GPU-to-GPU collectives, GPUDirect RDMA, GPUDirect Storage, NVMe throughput, host-memory bandwidth, thermals and power under sustained load.
Frequently asked questions
How many GPUs are in the 6U8X-TURIN2 SYN H200?
Eight NVIDIA H200 SXM GPUs on one HGX H200 baseboard. The accelerator count is fixed, not a customer-selectable PCIe quantity.
How much GPU memory does it provide?
Each H200 has 141 GB of HBM3e, giving about 1.1 TB across eight GPUs. The memory is distributed and must be used through multi-GPU software.
Is this the same as eight H200 NVL cards?
No. It uses H200 SXM GPUs with NVLink and NVSwitch on an HGX baseboard. H200 NVL is a PCIe add-in card with a different topology.
Which CPUs are supported?
The system page lists dual EPYC 9005 or 9004 processors, including selected 3D V-Cache and 97x4 parts. The current NVIDIA-qualified support-list entry names EPYC 9005, so confirm the required qualification scope.
How many DIMM slots are available?
Twenty-four, with twelve memory channels per processor at one DIMM per channel.
How many front NVMe drives fit?
Twelve Gen5 x4 drives: eight through the PCIe switch and four from the CPUs. The groups have different paths and should be tested separately.
Does the server include high-speed GPU networking?
The base specification lists dual 1GbE and multiple Gen5 x16 expansion positions. Production Ethernet or InfiniBand adapters, switches and media must be selected as part of the deployment.
What does SYN mean?
It identifies the PCIe switch synthetic-mode design intended to optimise GPU-to-CPU, GPUDirect RDMA and GPUDirect Storage paths. NVSwitch remains the separate internal GPU-to-GPU fabric.
Verdict
ASRock Rack 6U8X-TURIN2 SYN H200 is a complete Hopper-era scale-up node: eight H200 SXM GPUs, about 1.1 TB of HBM3e, NVSwitch, dual EPYC hosts, 24 memory channels and twelve front Gen5 NVMe bays.
Its distinctive feature is the synthetic-mode PCIe fabric around HGX. That gives the integrator many paths for NICs, storage and host traffic, but it also raises the bar for topology validation. The onboard 1GbE ports are nowhere near a production data path, and the 4+4 power system belongs in a properly engineered rack.
Configure the ASRock Rack 6U8X-TURIN2 SYN H200 as one node in a complete storage, network, power and software design.
Technical sources
- ASRock Rack 6U8X-TURIN2 SYN H200 product page
- ASRock Rack Q2 2026 GPU server guide
- ASRock Rack GPU support list
- NVIDIA H200 product specifications
- NVIDIA HGX H100, H200 and B200 reference components
- NVIDIA NVSwitch platform support
- AMD EPYC 9005 processor family
Specifications, qualification lists and software support can change. Confirm the exact system revision, CPUs, DIMMs, NVMe devices, NICs, switches, optics, power feeds and software stack before purchase. This is a source-backed technical review, not a record of hands-on GPUMachines testing.
.jpg)