ASUS ExpertCenter Pro ET900N G3 puts a GB300 Grace Blackwell Ultra Desktop Superchip into a 27kg tower. Its reason for existing is not difficult to spot: 748GB of coherent memory gives an AI team far more local model capacity than a conventional workstation with one or two PCIe GPUs. The harder question is whether that capacity is worth accepting an Arm host processor, a fixed 1,600W system power pool and limited internal storage.
This is a source-backed technical review, not a hands-on benchmark. We have checked the current ASUS and NVIDIA specifications and interpreted them for procurement and deployment. GPUMachines sells and configures this system, so our commercial interest is clear; our recommendation below still includes the cases where a cheaper RTX PRO tower or an eight-GPU server makes more sense.
Configure the ASUS ExpertCenter Pro ET900N G3 to review its current planning price and supported options.
Executive summary
ET900N G3 suits AI developers who need several hundred gigabytes of local model memory but don't need an eight-GPU rack server. The fixed GB300 module joins a 72-core Grace CPU to a Blackwell Ultra GPU through NVLink-C2C, then presents 496GB of LPDDR5X CPU memory and 252GB of HBM3e GPU memory as a coherent address space.
That design removes a hard capacity boundary, but it doesn't turn all 748GB into equally fast GPU memory. HBM3e runs at a stated 7.1TB/s, while LPDDR5X reaches 396GB/s. Placement, framework support and memory traffic will decide how close a workload gets to the platform's potential.
Our view is straightforward. Buy ET900N G3 when local model capacity, data control and deskside access matter more than raw multi-GPU throughput. Don't buy it merely because ASUS and NVIDIA quote up to 20 PFLOPS or support for trillion-parameter models. Those figures refer to sparse FP4 operation and a carefully fitted workload, not an across-the-board application result.
ASUS ET900N G3 specifications
| Component | Verified specification | |---|---| | Compute module | NVIDIA GB300 Grace Blackwell Ultra Desktop Superchip | | CPU | 72-core NVIDIA Grace, Arm Neoverse V2 | | GPU | One NVIDIA Blackwell Ultra GPU | | Coherent memory | 748GB total | | GPU memory | 252GB HBM3e at 7.1TB/s | | CPU memory | 496GB LPDDR5X at 396GB/s | | CPU-GPU link | NVIDIA NVLink-C2C, 900GB/s stated by NVIDIA | | AI compute claim | Up to 20 PFLOPS sparse FP4 | | RTX expansion | One optional RTX PRO Blackwell card | | PCIe | One Gen 5 x16 slot and two Gen 5 x8 electrical slots | | Internal storage | Four M.2 positions; two 2TB Gen 5 OS drives in RAID 1, plus two Gen 6 training-data slots | | High-speed network | Two 400Gb/s QSFP112 ports via ConnectX-8 | | Other network | One 10GbE RJ45 host port and one dedicated 1GbE BMC port | | Management | ASPEED AST2600, IPMI 2.0, Redfish and TPM 2.0 | | Power | 1,600W Titanium ATX PSU and fixed system power budget | | Size and weight | 584 x 232 x 565mm; 27kg net | | Shipping software | Ubuntu with NVIDIA AI Developer Tools |
ASUS notes that regional configurations and supplied items can vary. QSFP transceivers and the IEC C19 power cord appear as optional items in the current specification, so they belong on the quotation checklist rather than in an assumption.
The real attraction: model capacity at the desk
A 96GB PCIe workstation GPU runs many useful inference and fine-tuning jobs, but it sets a firm memory ceiling. Tensor parallelism across two cards can extend that ceiling, although software support and PCIe traffic complicate the result. ET900N G3 takes a different route: one coherent address space spans the Blackwell Ultra GPU and the Grace CPU memory.
The headline 748GB figure deserves careful reading. Only 252GB sits in HBM3e beside the GPU. The remaining 496GB is LPDDR5X attached to Grace. NVLink-C2C gives the processors coherent access and NVIDIA states 900GB/s for the link, yet the underlying pools retain different bandwidth and latency characteristics. A workload that constantly pulls weights or KV-cache data from the slower pool won't behave like one contained in HBM3e.
NVIDIA says the platform can handle models with up to one trillion parameters. Some arithmetic puts that claim in context. One trillion parameters stored at four bits each occupy about 500GB before quantisation scales, runtime state, activations and KV cache. Eight-bit weights alone would need about 1TB. The one-trillion figure therefore describes a narrow low-precision capacity case, not a promise that any trillion-parameter model will run with a long context window and useful throughput.
That distinction doesn't diminish the machine. A coherent 748GB address space is still unusual in a tower, and it can remove awkward model sharding for inference, retrieval-augmented generation, agent development and memory-heavy data work. It just needs a workload test with the intended model, precision, context length, concurrency and framework.
What 20 PFLOPS does and doesn't tell you
ASUS and NVIDIA quote up to 20 PFLOPS of AI compute. NVIDIA's specification marks this as FP4 Tensor Core performance with sparsity; without sparsity, the listed FP4 figure drops to 15 PFLOPS. The same platform has different rates at FP8, BF16, TF32 and FP32.
Procurement teams shouldn't compare the 20 PFLOPS number with an unqualified figure from another GPU. Precision, sparsity assumptions, software kernels, memory movement and batch shape all matter. We haven't seen independent ET900N G3 application benchmarks that justify turning the vendor peak into a tokens-per-second claim, so we won't invent one.
A sensible acceptance test uses the buyer's actual stack. Load the planned model at the intended precision, set a realistic context, measure latency and throughput under expected concurrency, then watch how memory placement and power allocation behave. Synthetic peaks can support architecture comparisons; they can't approve a purchase on their own.
One optional RTX PRO card, with an important catch
ASUS provides three physical PCIe slots but supports one optional graphics card in the first slot. Its current technical page lists RTX PRO 6000 Blackwell Max-Q Workstation Edition, RTX PRO 4000 Blackwell SFF and RTX PRO 2000 Blackwell. NVIDIA's wider DGX Station specification also lists RTX PRO 6000 Workstation Edition, though it warns that OEM support varies. Confirm the exact card, firmware and regional ASUS configuration on the final quote.
The extra GPU has a specific job. It can handle display, ray-traced visualisation, simulation or a separate CUDA task while the GB300 module runs a large-memory AI workload. It doesn't create an HGX-style scale-up fabric between multiple accelerators, and three physical slots don't mean three supported GPUs.
Power complicates the choice. NVIDIA's vsloshd service shares one fixed 1,600W budget between the GB300 module and the optional RTX card. In dynamic mode, the service moves power allowance between them according to live draw and prioritises the RTX card when it needs more. A 600W RTX card doesn't turn the workstation into a 2,200W system; it takes part of the existing pool, which can reduce the power available to GB300.
For pure large-model work, start with the base system. Add RTX PRO only when a named visualisation, display or physical-AI workflow needs it. Buying the largest card by habit can spend money while constraining the module that justified the workstation in the first place.
Storage is useful, but it is not a server backplane
Two pre-installed 2TB M.2 PCIe 5.0 drives form the operating-system RAID 1 pair. ASUS then provides two PCIe 6.0 M.2 positions for training data. Four internal M.2 slots can hold active models, container layers, indexes, checkpoints and a working dataset, but they don't replace shared storage for a research group.
There are no front hot-swap U.2 bays. Capacity, service access and sustained write planning therefore look more like a specialised workstation than a rack server. Teams training on large corpora will normally stage hot data locally and keep the authoritative dataset, model registry and backup elsewhere.
Before ordering, define the data path. Checkpoint writes can stall expensive compute if the local tier is undersized or if the networked tier can't absorb them. The two 400Gb/s ports give the system ample network potential, but a fast port alone doesn't guarantee storage throughput; the remote array, switch, optics, filesystem and client tuning all have to keep up.
ConnectX-8 is serious networking, not a free cluster
ET900N G3 carries two QSFP112 ports at 400Gb/s each through NVIDIA ConnectX-8. A separate 10GbE RJ45 port suits ordinary host traffic, while the AST2600 BMC gets its own 1GbE management connection. That separation is welcome for a machine expected to run unattended.
ASUS says two DGX Station systems can be linked to increase model capacity and performance. Treat that as a distributed deployment, not as one automatically coherent 1.5TB computer. The framework must support the chosen parallel method; the design also needs suitable cables or transceivers, addressing, security and monitoring. ASUS lists the 400G transceivers as optional.
The BMC matters more than its modest 1GbE port suggests. IPMI 2.0 and Redfish let IT staff monitor and manage the workstation without relying on its main operating system. Put that interface on a restricted management network. A deskside chassis doesn't make out-of-band administration optional.
Power, heat and the office question
The 1,600W rating deserves facility planning. At full draw, almost all of that electrical power becomes heat: roughly 5,460BTU/h. One machine can add the heat load of a small electric heater to a room, before monitors, storage and networking enter the calculation.
ASUS specifies a 115-240V input range and an IEC C19 power connection, with current depending on voltage and regional cord assembly. At 230V, 1,600W represents about 7A before allowing for conversion and circuit design; at 120V it is roughly 13.3A. Have the site owner or a qualified electrician confirm the circuit, connector, sustained load and local requirements. Don't treat a spare socket under a desk as a deployment plan.
Noise also needs a real conversation. ASUS describes a closed-loop liquid-cooling system and data-centre-grade thermal design, but it doesn't publish an acoustic figure on the pages reviewed. We haven't measured the machine. A 27kg, 1,600W AI tower may be physically desk-side while still belonging in a lab, equipment room or acoustically separated office.
ASUS recommends an ambient temperature below 30 degrees C during heavy workloads with the RTX PRO 6000 Max-Q option. Room cooling, intake clearance and exhaust recirculation need checking before sustained operation.
Arm software compatibility can decide the purchase
Grace uses 72 Arm Neoverse V2 cores. Ubuntu with NVIDIA AI Developer Tools supplies the intended environment, and much of the current AI container ecosystem supports Arm. Existing x86 software should not be assumed compatible, though.
NVIDIA's DGX Station development guide calls out instruction-set differences, SIMD porting work and Arm memory-ordering behaviour. Pure Python code inside a supported container may move easily. Proprietary binaries, old x86-only containers, hand-written assembly, uncommon plugins or native extensions can hold up a project.
Run a software inventory before purchase. Check container manifests for linux/arm64, rebuild native dependencies where needed, and test licence services or security agents that sit outside the AI stack. If an essential application remains x86-only, a conventional workstation or rack server can be the better machine even when it offers less coherent memory.
Best-fit workloads
ET900N G3 earns its place in four situations:
- Local inference for models that don't fit comfortably in 96GB or 192GB of aggregate PCIe GPU memory.
- Fine-tuning, quantisation and evaluation where a large coherent address space simplifies model handling.
- Agent and retrieval systems that keep sensitive data, indexes and services on premises.
- AI research that needs deskside iteration before moving a workload to a larger data-centre platform.
It can also suit physical-AI and visualisation teams when the optional RTX PRO card has a defined role. The buyer should validate simultaneous GB300 and RTX performance because both devices draw from the same power pool.
Who should not buy ET900N G3
Skip it if your models and datasets fit on one RTX PRO 6000. A conventional tower GPU workstation will cost less, consume less power and keep an x86 software environment.
It is also the wrong shape for scale-up training that depends on several tightly connected GPUs. An HGX server gives that work an eight-GPU baseboard, rack serviceability and a clearer path into shared cluster infrastructure, although cost, power and cooling rise sharply.
Storage-heavy teams may dislike four internal M.2 positions and the lack of front hot-swap drives. Organisations without a suitable power circuit, room cooling or Arm-compatible software should solve those constraints before considering an order.
Recommended configuration paths
Base GB300 for large-memory AI development
Keep the first deployment simple: GB300, the mirrored OS pair, two data M.2 drives sized for the active project and the existing ConnectX-8 interfaces. This preserves the whole dynamic power pool for the main module and avoids paying for an RTX card that sits idle.
GB300 plus a lower-power RTX PRO card
An RTX PRO 4000 Blackwell SFF or RTX PRO 2000 can provide display and visualisation capability with less pressure on the shared power pool. Confirm that the required application supports Arm on the host and the chosen GPU on the device side.
GB300 plus RTX PRO 6000 Max-Q
Choose the larger card for a measured physical-AI, rendering or simulation requirement. Expect the power manager to trade GB300 headroom for RTX performance under simultaneous load. This path deserves workload testing rather than a paper-only approval.
Two linked ET900N G3 systems
Two stations can make sense for a small research group that wants distributed jobs without installing an eight-GPU rack platform. Specify the 400GbE media, software topology, shared storage and management network as part of the purchase. Don't assume every model or framework scales cleanly across two hosts.
Our technical view
ASUS has built a convincing machine around a strange but useful proposition: far more addressable AI memory than a normal tower, without the electrical and mechanical commitment of an HGX server. The memory design is the reason to buy it. ConnectX-8, BMC management and the mirrored OS drives make the system easier to place in a professional environment than a home-built multi-GPU workstation.
The compromises are equally specific. Only one third of the coherent memory is HBM3e. The Grace host is Arm, internal storage is modest, and an optional RTX card competes with GB300 inside the same 1,600W limit. Those aren't footnotes; they decide whether ET900N G3 feels unusually capable or badly matched.
Our verdict: shortlist it for local AI projects that already hit GPU-memory limits and can prove Arm compatibility. For ordinary workstation inference, buy less. For sustained multi-GPU training, buy or host a rack platform built for that job.
Frequently asked questions
Does ASUS ET900N G3 have 748GB of GPU memory?
No. It has 748GB of coherent memory across two pools: 252GB of HBM3e attached to the Blackwell Ultra GPU and 496GB of LPDDR5X attached to the Grace CPU. Software can address both coherently, but their bandwidth differs substantially.
Can it run a one-trillion-parameter model?
NVIDIA states support for models up to one trillion parameters. That claim depends on low-precision weights, model structure, runtime overhead, context length and framework support. Four-bit raw weights for one trillion parameters need about 500GB before overhead, so buyers should test the exact model rather than rely on parameter count alone.
Which RTX PRO cards does ASUS support?
The current ASUS technical page lists RTX PRO 6000 Blackwell Max-Q Workstation Edition, RTX PRO 4000 Blackwell SFF and RTX PRO 2000 Blackwell. NVIDIA's wider platform list includes another RTX PRO 6000 edition, but OEM support varies. Confirm the requested card on the ASUS quotation.
Does an RTX PRO card raise total system power above 1,600W?
No. NVIDIA's power service shares a fixed 1,600W budget between the GB300 module and the optional RTX card. Under dynamic operation it reallocates power according to demand, with the RTX card receiving priority when it needs more.
Is this a replacement for an eight-GPU HGX server?
No. ET900N G3 offers one Blackwell Ultra GPU, large coherent memory and one optional RTX PRO card. HGX systems use multi-GPU baseboards and high-bandwidth GPU fabrics for scale-up workloads. The workstation answers a different capacity and deployment problem.
Can it sit in a normal office?
Physically, yes, but plan around a 27kg chassis, a 1,600W electrical rating and roughly 5,460BTU/h of heat at full draw. ASUS doesn't publish an acoustic figure on the pages reviewed. Confirm the circuit, cooling and acceptable noise before installation.
Buying through GPUMachines
GPUMachines can quote ET900N G3 with the required M.2 storage, supported RTX PRO option, 400GbE media and deployment services. The online figure is a planning price until a formal supplier quote confirms the regional system, warranty, power cord, transceivers and delivery terms.
Open the ASUS ExpertCenter Pro ET900N G3 configurator to build the initial specification, then use the shared configuration link for internal review.
Sources
- ASUS ExpertCenter Pro ET900N G3 product page
- ASUS ET900N G3 technical specifications
- NVIDIA DGX Station platform specifications
- NVIDIA DGX Station dynamic power guide
- NVIDIA DGX Station Arm CPU guidance
Sources checked 28 August 2026. Specifications and regional options can change; confirm the final configuration before ordering.
