GPUmachines

Giga Computing GAIFA: What Buyers Should Demand From a Rack-Scale Proving Ground

Giga Computing has announced a six-rack, 700 kW AI infrastructure proving ground. Here is the acceptance evidence a buyer should demand before a rack-scale deployment ships.

Giga Computing GAIFA: What Buyers Should Demand From a Rack-Scale Proving Ground

A GPU rack can pass component checks and still fail as a system. Power quality, coolant behaviour, switch configuration, firmware drift and storage pressure only meet one another when the equipment is assembled and exercised together.

Giga Computing's newly announced GAIFA facility is interesting for that reason. The company describes a six-rack, direct-liquid-cooled proving ground in New Taipei City with a planned initial power capacity of 700 kW. Its stated purpose is to validate compute, storage, networking, cooling and deployment as one operating environment before customer infrastructure reaches production.

The buyer's answer is less promotional than the announcement: a proving ground is useful when it reproduces the ordered configuration, test load and failure conditions, then hands over evidence that the customer's site can use. It is not a substitute for site acceptance. A green factory report cannot prove that the destination power feeds, water loop, network, storage or operating team are ready.

This article explains what buyers should ask Giga Computing, GPUMachines or any rack-scale supplier to prove before accepting a high-density AI deployment. It is based on public vendor documentation; GPUMachines has not independently tested GAIFA.

What Giga Computing has announced

Giga Computing announced GAIFA, short for GIGABYTE AI Factory Accelerator, on 22 September 2026. The first phase is planned around six racks for compute, storage and networking, with 700 kW of power capacity and direct liquid cooling.

The announced compute equipment includes two NVIDIA GB300 NVL72 racks, one G4L4-SD3-LAX7 server rack built around NVIDIA HGX B300, RTX PRO 6000 systems and a liquid-cooled AMD EPYC blade platform. That mix gives the facility several kinds of load to exercise rather than one homogeneous GPU rack.

Giga says it intends to test compute performance, rack power density, liquid-cooling behaviour, thermal management, network performance, component compatibility and deployment procedures. Those are claims about the planned role of the facility, not published results. The announcement also says detailed specifications and operating information will follow when GAIFA is complete.

That last sentence matters. At launch, buyers have a description of the proving ground, not a public test methodology, acceptance threshold or dataset of measured results. Procurement teams should treat GAIFA as a useful capability announcement and ask for project-specific evidence when placing an order.

Why rack-scale validation has become a buying issue

An eight-GPU server can be tested as a chassis. A GB300 NVL72 rack is a different object: NVIDIA documents 18 compute trays, nine NVSwitch trays, power shelves, liquid-cooling manifolds and out-of-band management switches inside one rack-scale system. It depends on facility services and several internal networks before an application can use all 72 GPUs as intended.

The same systems also sit at power levels that change the installation conversation. Giga's current GIGAPOD material lists a design figure of 140 kW per GB300 NVL72 rack. That number is not proof of measured consumption at GAIFA, nor should it be copied into a customer's electrical design without the ordered configuration and engineering documents. It does show why a supplier needs somewhere to test the rack, power path and cooling loop together.

Validation at this scale should answer four different questions:

| Question | Evidence a buyer should receive | | --- | --- | | Was the ordered hardware assembled correctly? | Serialised inventory, cable map, firmware bill, photographs and inspection record | | Can the rack operate under an agreed load? | Workload definition, duration, telemetry, alarms, throttling record and pass criteria | | Can the rack recover from expected faults? | Failure-injection record, failover behaviour, alert trail and recovery time | | Can the destination site reproduce the conditions? | Power, coolant, network, floor-loading and environmental interface requirements |

If a supplier cannot answer all four, the word "validated" is too broad to help with acceptance.

Factory acceptance and site acceptance solve different problems

A factory acceptance test, or FAT, checks the system before shipment. It can find wrong firmware, damaged cabling, mis-seated components, coolant leaks, failed sensors and network errors while the supplier still has engineers, spare parts and test equipment nearby.

A site acceptance test, or SAT, begins after installation at the customer's location. It proves the actual utility feeds, power distribution, cooling-water conditions, uplinks, storage, identity systems, scheduler and monitoring stack. It also checks whether the customer's staff can operate and recover the environment.

The two tests should share an evidence model, but they should not be merged into one vague sign-off. A rack that ran for 24 hours in Taiwan may behave differently behind another site's PDU, CDU, fibre plant or storage system. Shipping can also loosen connections and introduce damage after factory testing.

For a high-density installation, write the FAT and SAT before the purchase order is signed. Name the test owner, witness, data-retention period and dispute process. If acceptance language appears only after delivery, every failed threshold becomes a commercial negotiation.

What the factory test should reproduce

The test configuration must match the order closely enough that its results remain useful. "Same GPU platform" is not enough.

Record the exact compute trays, CPUs, GPUs, DIMM population, local drives, NICs, DPUs, switches, optics, cables, firmware and software image. Capture switch operating systems and configuration hashes. For liquid-cooled equipment, include CDU model, control firmware, coolant chemistry, supply temperature, flow, pressure and facility-water conditions.

The workload matters just as much. A short synthetic GPU run may expose dead devices but miss network congestion, checkpoint stalls, power excursions or thermal equilibrium. The supplier should explain why each workload exists and what failure it is intended to reveal.

A sensible sequence may include:

  • hardware inventory and sensor sanity checks before load;
  • node-level diagnostics and memory or PCIe error review;
  • collective communication tests across the intended fabric;
  • storage reads, writes and checkpoint patterns that resemble production;
  • an application or model workload chosen for the customer's intended use;
  • an extended run long enough for coolant and room temperatures to settle;
  • controlled loss of a redundant component, followed by recovery and evidence capture.

The list is not universal. A training cluster needs different storage and network stress from an inference service. The point is traceability: every test should connect to a buyer requirement or a credible failure mode.

Power evidence needs more than a peak number

Procurement documents often ask for maximum rack power and stop there. Operations teams need a time series.

Request input power per feed, rack and major subsystem at a useful sampling interval. The report should show steady state, start-up, workload transitions and the highest short-duration excursion. It should also record voltage, current, power factor, phase balance and redundancy mode where the instrumentation supports them.

Nameplate totals are not measurements. The G4L4-SD3-LAX7 product page, for example, lists a 5+5 arrangement of 3,000 W power supplies. Adding those labels would describe installed supply capacity, not the server's actual draw, expected rack demand or usable output at every input voltage. Test reports must keep those concepts separate.

Ask what happened when a feed, power shelf or monitored outlet became unavailable. A redundant design should not merely stay on; it should remain within the rating of the surviving path, raise the correct alert and preserve enough telemetry to explain the event.

The site team can then compare the factory trace with breaker settings, PDU ratings, upstream redundancy and any facility power cap. If the numbers do not reconcile before shipment, the rack is travelling toward a known problem.

Cooling evidence must describe the whole loop

Direct liquid cooling moves the heat problem rather than deleting it. Cold plates, manifolds, hoses and rack CDUs transfer heat to the facility-water system; fans may still cool NICs, drives, memory, power electronics and switches.

The factory report should state coolant supply and return temperature, flow, pressure, heat removed, pump state and any approach-temperature assumptions. It should record room conditions because residual air cooling still matters. Alarm thresholds, leak-detection tests and behaviour during pump or facility-water changes belong in the evidence pack.

Buyers should also ask how the supplier controlled condensation risk, water quality, filtration, materials compatibility and maintenance access. These details depend on the rack and CDU design, so project documents outrank a generic article.

NVIDIA's published GB200/GB300 deployment checklist explicitly calls for adequate power, cooling, network infrastructure and verified cabling. That is a useful minimum structure, but the buyer still needs values tied to the exact delivered rack and destination facility.

Network and storage tests should expose topology mistakes

A fabric can report every link as up while delivering poor application behaviour. Wrong breakout settings, mixed firmware, bad optics, oversubscribed paths or congestion control errors may remain hidden until many nodes communicate at once.

Request a physical and logical topology, port map, cable identifiers and firmware versions. The test should show link rate, error counters and collective communication results across the same paths the workload will use. When the design includes separate management, storage and GPU fabrics, prove that traffic follows the intended network rather than an accidental route.

Storage deserves a workload-shaped test. Reading a cached file at impressive speed says little about first-epoch dataset access or checkpoint writes. State the dataset size, block or object pattern, client count, cache state, mount options and duration. Preserve latency distributions and error logs rather than one headline throughput figure.

For buyers still planning the fabric, the GPUMachines GPU cluster configurator can turn node count and topology choices into a preliminary bill of materials. It is a planning tool, not an acceptance certificate; the final design still needs cable, optic, switch and rack review.

Firmware and software must be part of the configuration record

Rack validation loses value if the delivered system runs another firmware set or software image.

Ask for a machine-readable inventory covering BMC, BIOS, CPLD, GPU, NVSwitch, NIC, DPU, storage and switch firmware. Add driver, CUDA or ROCm, communication libraries, container runtime, workload manager and test-tool versions. Record configuration files or hashes where licensing and security policy permit it.

Then define change control between FAT and SAT. Security fixes may require an update before delivery, but the supplier should identify the change, retest the affected functions and update the evidence pack. Silent version drift is not a harmless administrative detail.

NVIDIA's Mission Control installation documentation contains separate verification stages for control-plane setup, rack import, provisioning, power-on, firmware and deployment summary. Buyers do not need to copy one vendor's procedure verbatim, but the staged approach is sound: a final green dashboard cannot replace the checks that led to it.

Failure testing is where a proving ground earns its cost

Normal-operation benchmarks are easy to demonstrate. Recovery evidence is more useful.

Agree which faults can be introduced without risking equipment or voiding support. Candidates may include loss of one redundant power path, a stopped pump in a redundant cooling arrangement, a disabled network link, a failed storage path or a compute node removed from the scheduler. The exact method must follow the manufacturer's procedure.

For each test, record four things: the initiating event, what the operator saw, whether the workload stayed within its agreed service level, and how the system returned to a known state. A fault that triggers an alert but leaves stale topology or degraded performance after recovery is not fully resolved.

Do not demand destructive demonstrations simply to make the acceptance plan look rigorous. Some failures are better proven through component certificates, design review or a non-production mock-up. The evidence method should match the risk.

Questions to put into the request for quotation

Buyers considering a rack-scale system should ask the supplier to answer these before order acceptance:

1. Which exact customer configuration will be assembled in the proving ground? 2. Which facility power and coolant conditions will the test use, and how do they compare with the destination site? 3. Which workloads, data sizes, network paths and run durations define the test? 4. What thresholds decide pass, conditional pass or failure? 5. Which fault and recovery scenarios will be witnessed? 6. What telemetry, logs, configuration files and inventory will the customer receive? 7. Which changes are allowed between factory test, shipment and site acceptance? 8. Who owns unresolved defects, retesting, travel and schedule impact?

Those questions turn "validated at GAIFA" or any other lab into a defined deliverable.

When this level of validation is excessive

Not every GPU purchase needs a six-rack proving ground.

A single air-cooled server, or a pair of independent PCIe GPU systems, may be covered by documented burn-in, diagnostics, network checks and a shorter workload run. If the application does not depend on a scale-up fabric or shared high-throughput storage, a full simulated AI-factory environment may add cost without changing the decision.

The dividing line is integration risk. Rack-scale validation becomes more valuable when several high-power liquid-cooled systems share facility services; when the training job spans many nodes; when storage and network performance determine useful GPU time; or when a fixed launch date makes on-site debugging unusually expensive.

Organisations without a suitable high-density site also have another option. Buying the equipment and placing it in an established facility can move part of the power, cooling and remote-hands problem to a hosting provider. GPUMachines describes that route on its Buy & Host page. The hosting contract still needs acceptance evidence and clear responsibility boundaries.

For conventional eight-GPU systems, compare the available HGX server platforms before assuming that an NVL72 rack is the right unit of purchase. A smaller system is often easier to power, cool and operate, and it may fit independent workloads better.

What buyers should take from GAIFA

GAIFA signals that server manufacturing is moving closer to facility and cluster integration. A planned 700 kW environment gives Giga Computing space to test interactions that cannot be seen on an isolated production line.

The announcement does not yet publish enough method or result detail for a buyer to accept a rack on the strength of the GAIFA name alone. That is normal for a newly announced facility. It also sets the correct next question: what exact test record will accompany my system?

A useful proving ground produces portable evidence. The rack inventory should match the order; the power and cooling traces should reconcile with the site design; network and storage tests should resemble the workload; failures should trigger known responses; and software versions should survive the handover without mystery.

If those records exist, factory validation can remove expensive uncertainty before shipment. If they do not, the buyer has received a demonstration rather than acceptance evidence.

Sources and further reading

← Back to blog