Two NVIDIA adapters can sit in the same AI server and still solve different problems. ConnectX-9 carries the workload traffic that keeps distributed GPUs busy. BlueField-4 creates a separate infrastructure domain for security, storage, policy, provisioning and telemetry. Buying both because they appear on a reference architecture is not a network design.
The practical question is narrower: does this server need infrastructure services that must remain available and trustworthy when the host operating system is busy, compromised or controlled by a tenant? If the answer is yes, a DPU can earn its place. If the server is part of a private, single-tenant cluster with a simple trust boundary, a SuperNIC and well-sized host CPUs may be the better purchase.
NVIDIA introduced more detail about BlueField-4 and its “Scale-In” role around Hot Chips 2026. The announcement matters because Vera Rubin systems put BlueField-4 and ConnectX-9 in the same compute tray. It does not make the devices interchangeable, nor does it prove that every smaller AI platform needs the same arrangement.
GPUMachines has not independently benchmarked production BlueField-4 or ConnectX-9 hardware. Specifications below are attributed to NVIDIA and should be confirmed against the exact OEM platform, firmware and software release offered at quotation.
The short buying answer
Choose ConnectX-9 without a DPU when the priority is a fast east-west data path for distributed training or inference, the cluster has one trusted operator, and host-based networking and security are acceptable.
Add BlueField-4 when the platform must provide host-independent tenant isolation, inline security, storage protocol offload, remote provisioning or telemetry that cannot depend on the tenant OS. It is strongest in shared GPU clouds, regulated environments, large inference estates and systems where storage or north-south services consume material host CPU time.
Do not treat BlueField-4 as an upgrade that makes ConnectX-9 unnecessary. NVIDIA assigns the adapters different paths in Vera Rubin: ConnectX-9 carries scale-out tenant workload traffic, while BlueField-4 runs infrastructure services around the server. A design may need one, the other or both.
ConnectX-9 and BlueField-4 in plain terms
A ConnectX SuperNIC is an accelerated network endpoint. It handles high-rate packet movement, RDMA and congestion-aware transport close to the GPU workload. In a distributed training job, this is the path used by collectives and model traffic between compute nodes. Latency, effective bandwidth and predictable behaviour under congestion determine whether expensive accelerators spend their time computing or waiting.
A BlueField DPU combines network interfaces, programmable processing cores and fixed-function accelerators into a separate computer on the adapter. It can run services independently of the host CPU. That separation is the point. Policy and infrastructure work can continue even when the host is overloaded, being reimaged or operated by a tenant that should not control the enforcement layer.
NVIDIA describes BlueField-4 with a 64-core Grace CPU, LPDDR5X memory, PCIe Gen6 connectivity and up to 800 Gb/s of network throughput. The company positions it for networking, security, storage and telemetry offload through DOCA software. These are vendor specifications, not a guarantee of application throughput.
For Vera Rubin, NVIDIA describes ConnectX-9 as the endpoint for the scale-out network and quotes up to 1.6 Tb/s per GPU in its platform material. That figure belongs to the full reference design. A real server may expose a different port count or topology, and delivered job bandwidth will depend on switches, rail mapping, optics, cabling, collective libraries and oversubscription.
| Question | ConnectX-9 SuperNIC | BlueField-4 DPU | |---|---|---| | Primary job | Move AI workload traffic across the scale-out fabric | Run and enforce infrastructure services outside the host | | Typical traffic | GPU collectives, distributed training and inference | North-south access, storage, security, provisioning and control | | Processing model | Accelerated data path controlled as part of the host and fabric | Separate programmable infrastructure processor plus inline accelerators | | Main buyer benefit | Low-latency, high-throughput GPU communication | Isolation, offload and host-independent control | | Main cost | Ports, optics, switches and fabric engineering | Extra hardware, DOCA lifecycle, policy design and another operating domain |
The table is intentionally functional. Product names and maximum line rates matter only after the jobs have been separated correctly.
What a DPU changes in the trust model
Host-based security has an awkward weakness in a shared AI service: the operating system being protected also participates in enforcing the rules. A privileged tenant, a kernel fault or a bad update can affect both workload and control.
BlueField places selected controls outside that host. NVIDIA describes using it for virtual private cloud policy, encryption, firewall processing, runtime threat detection, secure bootstrapping and telemetry. In this model, the tenant can use the accelerated server without gaining authority over the infrastructure processor that applies network and storage policy.
That is commercially useful when GPUs are rented to multiple customers, when research groups need strict separation, or when regulated data crosses shared hardware. It can also reduce the number of privileged agents installed in the host OS.
It is not automatically safer. The DPU adds firmware, an operating environment, credentials, DOCA services and update procedures. A team that cannot inventory and patch those components has created another neglected control plane. The security case requires named owners, supported versions, key management, log export and a recovery procedure.
Ask the supplier to show which rules continue to work while the host is down or hostile. “DPU-enabled security” is too vague for an acceptance test.
Storage is often the deciding workload
AI servers do more than exchange gradients. They pull checkpoints, training shards, embeddings, model weights and generated artefacts from shared storage. Network storage can consume host CPU cycles for protocol handling, encryption, virtualisation and data movement.
BlueField-4 is designed to offload NVMe over Fabrics, RDMA and other storage-path work. NVIDIA also describes file and object access services in its Scale-In architecture. The possible benefit is not merely a lower CPU percentage. It is more predictable access to data while the host is busy with tokenisation, preprocessing, scheduling or tool execution.
Whether that benefit appears depends on the storage stack. A slow metadata service, undersized array or congested switch will remain slow behind a DPU. A local-NVMe training job may have little storage work to offload. A server that already has ample CPU headroom may not recover enough useful GPU time to justify the adapter and software.
Measure the path rather than assuming it. Record storage throughput, p99 I/O latency, host CPU utilisation, GPU idle time and completed training steps before and after offload. Repeat with encryption and tenant isolation enabled, because those are often the reasons for buying the DPU.
For the wider storage decision, see the GPUMachines guide to designing AI storage infrastructure. It explains why headline gigabytes per second cannot replace workload-level evidence.
When ConnectX-9 alone is the right answer
A private cluster used by one trusted team may not need host-independent policy. If the server OS, scheduler and network are operated by the same organisation, a separate infrastructure processor can add more work than risk reduction.
ConnectX-9 alone is a credible design when:
- distributed GPU communication is the main traffic class;
- tenants do not receive privileged host access;
- storage traffic is modest or handled by dedicated interfaces;
- host CPUs retain headroom during the target job;
- provisioning through the baseboard management controller and existing automation is sufficient;
- security policy does not need to survive a compromised host;
- the operations team does not already manage DOCA at fleet scale.
This is especially relevant to smaller research clusters. The money and switch ports assigned to a DPU may buy more storage, a less contended fabric or spare optics instead. Those improvements can have a clearer effect on completed work.
The Ethernet versus InfiniBand guide covers the next decision: which fabric should carry the scale-out traffic once the endpoint role is clear.
When BlueField-4 earns its place
The DPU becomes easier to justify when several requirements arrive together.
Shared GPU infrastructure
Cloud and hosting providers need isolation that tenants cannot disable. They also need repeatable provisioning, metering and telemetry across a fleet. Running those services on a separate processor can protect host capacity and reduce dependence on each tenant image.
Regulated or sensitive workloads
Healthcare, financial, government and sovereign AI deployments may require a stronger boundary between workload administrators and infrastructure enforcement. The exact compliance outcome still depends on architecture, policies and evidence. A DPU is a mechanism, not a certification.
Heavy networked storage
If host CPU time and jitter in the storage path are measurable constraints, offload may improve consistency. The procurement case should state the before-and-after workload and not rely on an isolated throughput maximum.
Bare-metal lifecycle automation
BlueField can participate in bootstrapping, network configuration and policy before the host OS starts. That is valuable when hundreds of nodes must be reimaged or reassigned without manual intervention.
Independent observability
Telemetry collected outside the host can remain available during a host crash or tenant incident. This aligns with the need for evidence described in the agent-native telemetry guide: operational automation should be able to prove what it observed and changed.
The operational bill can exceed the hardware bill
Adding a DPU creates another software estate. Someone must own firmware compatibility, DOCA versions, container images, policy rollout, keys, logs, high availability and incident response. The change calendar now includes NIC firmware, DPU software, switch software, host drivers, CUDA and orchestration components.
That complexity can be justified at scale. It can also overwhelm a small platform team.
Include these items in total cost of ownership:
- DPU hardware and any platform-specific risers or cooling requirements;
- switch ports and optics for the Scale-In path;
- DOCA qualification and update testing;
- telemetry storage and integration;
- security policy development and audit evidence;
- spare adapters and replacement procedure;
- training for network, Linux and security teams;
- performance testing after every material software change.
A design that saves host CPU cycles but introduces fragile manual operations may be a net loss.
Do not buy a DPU for these reasons
Do not add BlueField-4 merely because it appears in an NVIDIA rack diagram. Reference designs cover demanding deployment patterns and may include functions that a smaller buyer does not use.
Do not buy it to repair a poor scale-out topology. A DPU cannot remove oversubscription, correct bad rail mapping or compensate for insufficient switch capacity.
Do not buy it on the promise of “zero CPU overhead”. Some control work, drivers and orchestration remain. Measure host utilisation on the planned software path.
Do not buy it before the team has identified the services that will run on it. A line item without an operational owner becomes expensive dormant capacity.
Do not assume it replaces the SuperNIC. In Vera Rubin, NVIDIA uses BlueField-4 and ConnectX-9 for complementary roles.
An RFQ-ready acceptance test
Turn the business reason for the DPU into pass or fail evidence.
1. Name each offloaded service. List networking, security, storage, provisioning and telemetry functions. Record where each service runs and who owns it. 2. Baseline the host-only design. Measure application throughput, p99 latency, GPU idle time, storage performance and host CPU use without DPU offload. 3. Repeat with policy enabled. Turn on the encryption, isolation and inspection required in production. A benchmark with those controls disabled is not representative. 4. Test host independence. Reboot the host, stop its network agents and simulate a tenant with administrative access. Confirm which policy and telemetry functions remain. 5. Load the storage path. Run the intended file, object or NVMe-oF traffic beside GPU communication. Inspect tail latency and congestion, not only an average transfer rate. 6. Fail one path. Remove a link or service and observe workload behaviour, policy continuity and recovery time. 7. Patch and roll back. Upgrade the DPU software and firmware using the proposed production procedure. Prove that a known-good release can be restored. 8. Export evidence. Confirm logs, counters and policy decisions reach the organisation's existing monitoring and security systems.
The test should use the exact server, firmware, switch software and DOCA release intended for delivery. A lab result from a neighbouring platform is useful context, not acceptance.
How this affects the server quotation
The adapter choice changes more than a part number. The server needs the correct PCIe connectivity, airflow or liquid-cooling assumptions, firmware support and management integration. The fabric needs port and optic capacity for each traffic path. Rack power and cabling must include every endpoint.
Ask the supplier to show a port map that separates:
- scale-up links inside the rack-scale system;
- scale-out GPU workload links;
- Scale-In or infrastructure links;
- storage links where physically separate;
- management and service-processor links.
The names may vary by platform, but the traffic classes should not disappear into one line labelled “networking”. The GPU Cluster Configurator can be used to map node count, fabrics and rack assumptions before a final bill of materials is fixed.
Questions buyers ask
Is BlueField-4 faster than ConnectX-9?
That is the wrong comparison. ConnectX-9 is the high-rate scale-out endpoint for workload traffic. BlueField-4 is an infrastructure processor for services around the host. Their quoted bandwidth and processing resources describe different jobs.
Does every Vera Rubin server contain both?
NVIDIA's Vera Rubin compute-tray descriptions include ConnectX-9 SuperNICs and BlueField-4. OEM configurations and smaller systems can differ, so confirm the exact platform bill of materials.
Can BlueField-4 replace a storage server?
No. It can accelerate and virtualise parts of the storage path, but capacity, metadata, durability and data services still come from the storage system.
Will a DPU improve training speed?
It can help when host processing, storage services or security work are causing measurable stalls. It will not improve a job that is limited by GPU compute, collective algorithms or an undersized fabric.
Is a DPU useful in a four-GPU or eight-GPU server?
Sometimes, particularly for shared hosting, regulated data or storage offload. GPU count alone is not the trigger. Trust boundaries and infrastructure workload are more useful criteria.
Can GPUMachines source BlueField and ConnectX configurations?
GPUMachines can help specify and source compatible networking options as part of an AI server or cluster design, subject to platform support and distributor availability. The quotation should name adapter, optics, switch ports, firmware and software assumptions together.
Sources and Further Reading
- NVIDIA: BlueField-4 Powers New Scale-In Infrastructure
- NVIDIA BlueField DPU product platform
- NVIDIA: Scaling Agentic AI Factories with BlueField
- NVIDIA: Inside the Vera Rubin Platform
- Hot Chips 2026 conference programme
Verdict
ConnectX-9 and BlueField-4 are not premium and standard versions of the same adapter. ConnectX-9 is there to move AI workload traffic. BlueField-4 is there to make infrastructure services independent of the host and to accelerate selected network, security and storage work.
For a private cluster with one trusted operator, start from the SuperNIC and prove that host CPU or policy limits require more. For a shared, regulated or heavily automated AI estate, test BlueField-4 against named isolation, storage and lifecycle requirements. Buy the DPU only when those tests show a useful operational boundary, not because a reference diagram has an empty slot to fill.
GPUMachines can translate that decision into a compatible server, adapter, switch, optic and rack design. Start with the GPU Cluster Configurator or discuss the trust boundary and traffic profile before requesting hardware.
