GPUmachines

How to Design an AI Factory: Decisions Before the BOM

A GPU count is not an AI Factory design. Workload shape, scalable units, network planes, storage tests, power, cooling and operating ownership must be fixed before the bill of materials.

How to Design an AI Factory: Decisions Before the BOM

A target of 64 GPUs is not an AI Factory specification. It doesn't say whether jobs communicate across nodes, how checkpoints reach storage, which rack can accept the heat, who patches the fabric or what the buyer will test before signing acceptance.

Start with the production service, then choose the compute pattern. NVIDIA's current enterprise reference architectures give buyers three useful families: RTX PRO PCIe clusters for flexible enterprise workloads, HGX B300 clusters for communication-heavy jobs, and rack-scale NVL72 systems for the largest tightly connected workloads. Each family has a defined expansion unit, network model and facility consequence.

GPUMachines designs and sells this infrastructure. This guide explains the evidence a buyer should supply and the decisions a project team should record before hardware selection. It does not replace an OEM design, data-centre electrical study, cooling design or qualified network and storage review.

Review the GPUMachines AI Factory profiles for current solution patterns, or use the GPU cluster configurator to explore node and fabric counts. Neither tool can rescue vague workload data, so settle the questions below first.

Executive answer

An AI Factory design needs a workload envelope, a repeatable compute block, separate traffic roles, measured storage targets, a facility envelope and an operating model. The finished proposal should join those parts in one versioned bill of materials with topology, cable schedule, software baseline, dependencies, exclusions and acceptance tests.

NVIDIA's September 2026 enterprise guidance covers deployments from 32 to 1,024 GPUs, depending on architecture. It defines four-node scalable units for RTX PRO and HGX systems, while NVL72 grows by full 72-GPU racks. Those patterns help buyers avoid arbitrary node counts, but they don't decide what the organisation needs.

A sensible project works through twelve decisions before the BOM:

1. Define the workload and service objective. 2. Calculate the useful capacity rather than the purchased capacity. 3. Select the right AI Factory family. 4. Fix the scalable unit and growth boundary. 5. Draw every network plane. 6. Size the fabric from traffic, not port marketing. 7. Set storage and checkpoint acceptance targets. 8. Design control-plane, scheduler and user access. 9. Prove electrical and cooling capacity. 10. Assign security, support and lifecycle ownership. 11. Name the failure cases and spare strategy. 12. Write acceptance tests before ordering.

Skipping one usually reappears later as idle GPUs, stranded hardware or an expensive change order.

1. Define the workload as a service

“We need an AI cluster” is too broad. Training, post-training, batch inference, online inference, retrieval, synthetic data and scientific computing stress different parts of the system. Even two LLM projects can disagree sharply: one may run a week-long synchronous training job, while another serves thousands of small requests with strict tail latency.

Write down the model families, parameter counts, precision, parallelism plan, context distribution, batch behaviour and expected concurrent users. For training, add dataset size, read pattern, checkpoint size, checkpoint interval and target time to recover a failed run. For inference, record prompt and output token distributions, latency objective, availability target, model-loading method and growth curve.

Then name the operating constraint. Sensitive data may keep the platform on premises. A fixed launch date may favour hosted capacity. A research team may value flexible scheduling over maximum utilisation, while a service provider needs tenancy, metering and repeatable isolation.

The design should answer a service question such as: “Support two concurrent 70B fine-tuning jobs and one inference environment, with four-hour checkpoints restored within thirty minutes.” That sentence can drive storage, network and scheduler choices. “Buy 64 GPUs” cannot.

2. Convert GPU quantity into useful capacity

Purchased capacity counts accelerators. Useful capacity counts completed work under the required service level.

Estimate how many GPUs each job occupies, how long it runs and how much time it loses to data loading, collective communication, compilation, checkpointing and queue gaps. Add maintenance and failure assumptions. A platform that reports 95 per cent allocation while jobs wait on storage isn't well utilised; the scheduler has simply reserved expensive devices.

For early designs, keep the model honest. Use a range for throughput and occupancy until application measurements exist. Don't turn vendor peak arithmetic into a job-completion forecast. If possible, collect traces from a smaller system or cloud run: GPU utilisation, host CPU load, network counters, storage reads, checkpoint writes and memory use reveal more than a generic benchmark.

Capacity planning should also include the low end. What happens during holidays, between research projects or while a new model pipeline is still unstable? A smaller first block, a hosted system or mixed owned-and-cloud plan can beat an oversized facility build when demand remains uncertain.

3. Choose the compute family by communication pattern

NVIDIA's current enterprise reference material describes RTX PRO, HGX B300 and NVL72 AI Factory families. They overlap in workload labels, but their physical and network assumptions differ.

| Family | Published starting unit | Published scale | Where it fits | | --- | --- | --- | --- | | RTX PRO AI Factory | Four PCIe GPU nodes | Four to 32 nodes; up to 256 GPUs in the 8-GPU reference shape | Distributed inference, fine-tuning, perception, visual computing, analytics and jobs that value PCIe flexibility | | HGX B300 AI Factory | Four 8-GPU HGX nodes | Four to 128 nodes; 32 to 1,024 GPUs | Training, substantial fine-tuning, distributed inference, analytics and scientific work that benefits from NVLink inside each node | | NVL72 AI Factory | One rack with 18 compute trays | One to eight racks; 72 to 576 GPUs in the current GB300 RA | Rack-scale training, fine-tuning, inference and scientific work requiring one dense NVLink domain per rack |

RTX PRO systems can be the strongest answer for independent or loosely coupled workloads. NVIDIA publishes an 8-GPU foundational pattern with 200 GbE of east-west bandwidth per GPU, plus lower-cost 2-GPU and 4-GPU designs. It also publishes north-south-only variants for work that doesn't communicate across nodes.

HGX B300 raises the scale-up and scale-out density. Its reference node has two CPU sockets, eight B300 GPUs, eight east-west adapters plus one north-south adapter, and 800 GbE of average east-west bandwidth per GPU. Choose it because the workload needs that communication structure, not because B300 sits higher on a product list.

NVL72 makes the rack the unit of purchase. It brings much higher facility demands and a fixed expectation of dedicated network planes. Most enterprise AI work does not need it. Buyers considering NVL72 should already have application evidence, liquid-cooling plans and a rack-scale operating model.

4. Fix the scalable unit and expansion boundary

A scalable unit gives the project a repeatable block for compute, network, power, cooling and commissioning. NVIDIA defines one RTX PRO or HGX scalable unit as four nodes. The NVL72 scalable unit is one complete rack.

Choose an initial unit and a planned final boundary. If the first phase contains four HGX nodes but the programme may reach sixteen, reserve the spine ports, rack positions, fibre routes, storage headroom and electrical capacity required by that boundary. Buying a switch that fits phase one perfectly can make phase two need a second topology.

Expansion assumptions belong in the commercial document. State which components arrive in phase one, which capacity is reserved, what must be replaced at each growth point and which acceptance tests repeat. Avoid vague “scales to” claims that omit the cost or disruption of scaling.

The same logic applies to spares. A four-node cluster loses 25 per cent of its nodes when one fails. A 128-node cluster has a different spare and scheduling problem. Match the strategy to service impact rather than copying a fixed percentage across every project.

5. Draw separate network planes

One diagram should show every endpoint and traffic role. NVIDIA's reference guidance separates east-west compute traffic, north-south services, storage connectivity, customer uplinks and management/support services. Some smaller RTX PRO and HGX designs can consolidate roles, but the design team should still name them.

East-west traffic carries collectives and distributed workload data between GPUs. North-south paths connect storage, external services and infrastructure. The management network handles BMC access, provisioning, monitoring and lifecycle operations. Customer or application traffic may need a separate security boundary from storage and cluster control.

For each plane, record endpoint ports, switch ports, link speed, media, cable length, VLAN or subnet, routing boundary, redundancy and expected traffic. A switch count without those assignments is not a network design.

NVIDIA's current HGX B300 reference pairs eight ConnectX-8 SuperNICs with the eight GPUs for east-west communication and adds a BlueField-3 DPU for north-south traffic. RTX PRO reference nodes use fewer adapters and lower per-GPU bandwidth in the published patterns. NVL72 uses dedicated north-south and dual-plane east-west fabrics throughout its published scale range.

6. Size fabric bandwidth from workload behaviour

Port speed matters, but topology and contention decide what the application sees. Training collectives can generate synchronized bursts that expose oversubscription. Checkpoint traffic can collide with data reads. Inference services may need predictable north-south latency while batch jobs consume the east-west fabric.

Decide whether the project uses Spectrum-X Ethernet, InfiniBand or another supported pattern with a clear reason. Operations skill, existing standards, collective behaviour, telemetry, congestion control, failure recovery and ecosystem support all belong in that choice. “Ethernet is familiar” and “InfiniBand is fast” are not enough.

The fabric workbook should include endpoint count, rails, leaf and spine links, blocking ratio, usable ports, spare ports, optics, cables and failure domains. Test the arithmetic twice: once for the initial phase and once at the expansion boundary. If a dual-plane design is required, keep the planes genuinely independent through switches, paths and power where the availability target demands it.

Plan observability at the same time. Port errors, congestion, retransmissions, buffer pressure and path imbalance need named dashboards and alert owners. A fast fabric that nobody can diagnose will spend its worst day as a meeting rather than a tool.

7. Turn storage into measurable tests

Capacity in terabytes tells only a fraction of the story. An AI Factory storage plan needs read throughput, write throughput, metadata behaviour, checkpoint burst, restore performance, model distribution, retention and protection targets.

Map each data class:

  • source datasets and governed records
  • prepared training shards
  • model repositories and container images
  • active checkpoints and restore points
  • scratch, temporary outputs and caches
  • logs, metrics, audit records and long-term archive

Each class can use a different tier. Local NVMe may cache active data and absorb scratch writes, while shared storage supplies the authoritative dataset and protected checkpoints. Object storage can suit model artefacts or archive but may need a caching or parallel-file layer for demanding training reads.

Write acceptance tests in workload terms. “Sustain X GB/s for N minutes while 32 nodes read the target shard size” is testable. “High-performance AI storage” is not. Include a checkpoint write and restore test under realistic fabric load, then state what happens when a storage path or controller fails.

NVIDIA's enterprise reference architecture reserves network endpoints for certified storage, but certification does not replace sizing. The workload still decides capacity, throughput, metadata and recovery.

8. Design the control plane and user path

Users need a way onto the platform, workloads need a scheduler, and operators need a supported method to rebuild nodes. Those services consume compute, memory, storage and network capacity of their own.

Choose how Slurm, Kubernetes or another orchestrator divides responsibility. Define identity, role mapping, quotas, container and image policy, secrets, notebook access, batch submission, model endpoints and audit logging. If both Kubernetes and Slurm appear in the same proposal, document the boundary instead of assuming they will cooperate by goodwill.

Control-plane high availability should match the service target. NVIDIA's HGX reference describes separate control-plane nodes and gives an example combining Base Command Manager, Slurm and Kubernetes roles. The exact count can differ, but the principle holds: don't run the platform's only scheduler, registry or monitoring service on an arbitrary GPU node because it was available during installation.

Provisioning and firmware baselines matter too. Record BIOS, BMC, NIC, DPU, switch, driver, CUDA and container versions. Decide how the team stages updates, proves rollback and detects drift. Day-two work starts before day one.

9. Prove the facility envelope

Power and cooling can veto an otherwise sound architecture. NVIDIA notes that many enterprise data centres still operate below 20 kW per rack and lack a liquid-cooling path. That environment may suit selected air-cooled RTX PRO designs, but dense HGX or NVL72 infrastructure needs a different site.

Build a rack schedule from the quoted equipment rather than a generic GPU TDP sum. Include compute, switches, storage, management, losses, redundancy mode and realistic operating power. Map every device to A and B feeds, PDU outlets, connector types, breaker limits and upstream capacity. State which equipment loses redundancy after a feed failure.

Cooling design needs inlet or coolant conditions, flow, pressure, water quality, heat rejection, CDU and manifold layout, leak response and maintenance isolation. Air-cooled deployments need airflow direction, rack pressure, blanking, return-air path and room capacity at the planned density.

Physical details can stop a project just as effectively: floor loading, rack depth, door and lift dimensions, delivery route, cable bend radius, overhead containment and safe service space. Ask the facility team to sign the interface schedule before hardware ships.

10. Assign security and lifecycle ownership

Draw the trust boundaries around users, datasets, models, BMCs, DPUs, switches, storage and external services. Decide where encryption starts and ends, how administrators reach out-of-band management, how tenants remain separated and which logs support an investigation.

Then put names against operational duties. Someone owns firmware, driver qualification, scheduler policy, failed jobs, capacity review, security updates, backups and vendor cases. If the project relies on an integrator or hosted provider, the support contract should show the hand-off points and response commitments.

Software entitlement needs equal care. Identify the NVIDIA and third-party licences included with each system, the subscription term, support level and renewal assumption. A hardware price without the required software and support does not describe the operating cost.

Lifecycle planning also covers component availability. Record approved substitutions, spare optics, cables, drives, fans, PSUs and at-risk items. Large clusters should avoid one undocumented “temporary” replacement creating a permanent compatibility branch.

11. Name failure cases before they happen

Pick failures the design must survive: one compute node, one leaf switch, one fabric plane, one storage controller, one power feed, one control-plane service and one bad software release. For each, state the expected service impact and recovery owner.

Some jobs will fail when a node disappears; the platform's job is to detect the failure, preserve useful state and make recovery predictable. Other services may need continuous availability. Don't buy duplicate hardware everywhere without deciding which failures actually justify it, but don't call a design resilient without a written failure model.

Spares and support follow from this exercise. A long replacement lead time can justify an on-site spare even when the statistical failure rate is low. Conversely, a hosted system with firm remote-hands and replacement terms may need fewer customer-owned parts.

12. Write acceptance tests before ordering

Acceptance protects the buyer from receiving a collection of powered-on components instead of a working platform. It also protects the supplier by fixing what “done” means.

Cover inventory, firmware baseline, physical installation, cable map, management access, fabric health, storage throughput, scheduler operation, user access, monitoring, security controls and failure handling. Then run at least one production-intent workload with agreed measurements. For training, that may include collective performance and checkpoint recovery. For inference, use representative model loading, latency and concurrency.

Record versions, test inputs, pass criteria and results. If a figure comes from an OEM diagnostic rather than the customer's workload, label it. A failed test should have an owner, remedy and retest date.

Do not leave acceptance until commissioning week. Tests influence the design, software licences, instrumentation and schedule; writing them early catches missing pieces while the BOM can still change.

What a complete proposal should contain

A useful AI Factory proposal is versioned and specific. It should include:

  • named compute systems, revisions, CPUs, memory, GPUs, local drives and adapters
  • fabric topology with endpoint and switch-port assignments, optics, cables and spares
  • storage design with capacity, performance and recovery targets
  • rack elevations, power map and cooling interface schedule
  • management, control-plane and software architecture
  • installation, commissioning and acceptance plan
  • support ownership, licences, exclusions, dependencies and expansion boundary

Pricing can follow that structure later. Publishing an attractive total before the dependencies are fixed creates false precision.

When to start smaller

Start with a workstation, one server or hosted GPU capacity if model choice, demand or operations remain unsettled. The aim is to measure memory use, job duration, data movement and user behaviour without committing the organisation to a large facility programme.

A PCIe GPU server suits many independent inference, visual-computing and development jobs. An HGX server lets a team validate scale-up behaviour inside one eight-GPU node. Cloud or Buy & Host capacity can expose utilisation and software requirements while the on-premises site is prepared.

Starting smaller is not a failure of ambition. It is the right answer when the missing evidence could change the architecture family.

Information to send GPUMachines

Prepare one short pack before requesting a design:

  • model and application list, precision, context, parallelism and target job time
  • training versus inference share, concurrent users and service levels
  • dataset size, read pattern, checkpoint size, frequency and recovery target
  • current measurements from cloud, workstation or pilot runs
  • delivery country, data location, security and tenancy requirements
  • site power, cooling, rack, floor-loading, delivery and fibre information
  • preferred scheduler, identity source, monitoring tools and support hours
  • initial capacity, expansion boundary, budget stage and required delivery window

Unknowns are acceptable when they are labelled. The design can include tests or assumptions to close them. Pretending they are settled is more dangerous.

FAQ

How small can an NVIDIA enterprise AI Factory start?

NVIDIA's current RTX PRO and HGX reference architectures start at one four-node scalable unit, representing 32 GPUs in their foundational 8-GPU node shapes. NVL72 starts at one full 72-GPU rack. A buyer can still run a smaller pilot, but it sits below those published multi-node RA starting points.

Does an AI Factory have to use HGX servers?

No. NVIDIA publishes RTX PRO PCIe, HGX B300 and NVL72 families. Independent inference, perception, analytics and visual-computing work may fit RTX PRO systems better than HGX.

Should storage share the compute fabric?

Sometimes, but make the contention and failure implications explicit. Larger designs often benefit from separate or carefully partitioned paths. The answer depends on checkpoint bursts, dataset reads, network scale and the chosen reference topology.

Is Ethernet suitable for distributed AI?

Yes, when engineered for the workload. NVIDIA's current enterprise RAs publish Spectrum-X Ethernet patterns. InfiniBand remains relevant for supported workloads and environments. Pick the fabric with topology, congestion control, operations and application behaviour in view.

What causes expensive AI Factory redesigns?

The usual causes are weak workload data, a phase-one network with no growth path, underestimated power or cooling, storage sized by capacity alone and no named operating owner. Those problems look separate during procurement and collide during commissioning.

Can GPUMachines design a hosted AI Factory?

GPUMachines can scope on-premises, hosted and Buy & Host routes. A hosted design still needs workload, storage, network, security, support and acceptance requirements; the facility interface moves to the provider rather than disappearing.

Verdict

Good AI Factory design begins before the product shortlist. Define the service, select a repeatable compute unit, prove every data path and make the facility and operating teams part of the architecture.

NVIDIA's reference patterns give the project sensible boundaries: four-node RTX PRO or HGX units, and full-rack NVL72 units. Use them to avoid arbitrary growth, then adapt the final system to measured workloads and the real site.

Select an AI Factory profile when the workload and facility envelope are known, or build an initial cluster model to expose the node, switch and rack arithmetic that still needs review.

Sources

Reference material checked 21 September 2026:

← Back to blog