GPUmachines

EuroHPC AI Gigafactories: What Infrastructure Buyers Must Specify

A procurement guide to EuroHPC AI Gigafactory compute, networking, storage, power, cooling, access and acceptance testing.

EuroHPC AI Gigafactories: What Infrastructure Buyers Must Specify

A facility built around more than 100,000 advanced AI processors is not merely a larger GPU cluster. It is a power project, a network project, a storage project and a shared research service rolled into one procurement.

That distinction matters as the EuroHPC Joint Undertaking develops the framework for European AI Gigafactories. The European Commission describes these as very large facilities intended to train next-generation models, backed by an InvestAI facility designed to mobilise €20 billion for up to five sites. EuroHPC has also warned that information shared before the formal call is preliminary and that only the published tender documents will be legally binding.

For bidders, suppliers and research consortia, the sensible response is not to guess the final specification. It is to prepare an evidence-backed architecture that can survive changes in accelerator generation, power availability, software policy and user demand.

This guide sets out what such an architecture should specify and where proposals commonly become vague. It is aimed at organisations planning infrastructure, joining a consortium or preparing workloads that may later use a Gigafactory. It is not legal advice and does not replace EuroHPC tender documents.

The procurement question behind the headline

Public discussion tends to focus on the processor count. A buyer cannot procure useful AI capacity by counting accelerators alone. The facility must deliver completed training and inference work to authorised users while meeting energy, reliability, security and access obligations.

The tender therefore needs to answer six linked questions:

  • Which workloads and user groups will the service admit?
  • How will accelerators communicate within a node and across the facility?
  • How will data, checkpoints and model artefacts reach the processors without long stalls?
  • What electrical and cooling envelope can the site deliver through each deployment phase?
  • How will capacity be allocated, measured and recovered after failures?
  • Which parts of the platform can change without forcing a redesign of the whole service?

If any one of these remains a slogan, the processor specification cannot rescue the project. A strong response converts policy goals into measurable acceptance tests.

GPUMachines can help prospective users and delivery partners model smaller reference environments through the GPU cluster configurator, including compute, network and storage assumptions that can be tested before a full-scale proposal is fixed.

Start with the service, not the bill of materials

An AI Gigafactory will serve more than one project. Training a frontier model, fine-tuning a multilingual model, running safety evaluations and hosting high-volume inference place different demands on memory, network collectives, storage metadata and job duration.

The procurement should define service classes. A tightly coupled training partition may need non-blocking scale-out communication and strict placement. An inference partition may favour independent replicas, predictable tail latency and rapid model loading. Research access may require flexible containers and short queue times, while a production tenant may need reserved capacity and change control.

These classes affect hardware selection. They also affect scheduling, accounting and support. A single average utilisation target hides whether the platform is delivering the work users need. Tender metrics should include queue time, job completion, failed-run recovery, accelerator occupancy and service availability by class.

Data governance belongs in the same service definition. The facility may handle public research data, commercially sensitive training sets and regulated information. Identity, encryption, key custody, tenant isolation, audit retention and data deletion cannot be added after the storage system has been selected.

Accelerator procurement without a frozen architecture

A multi-year facility cannot assume that one accelerator generation will define its entire life. Procurement should separate immediate deployment from the interfaces that later phases must preserve.

The initial platform may use NVIDIA, AMD or another accelerator architecture. The evaluation needs evidence from representative workloads and supported software, not peak arithmetic alone. Memory capacity and bandwidth, intra-node interconnect, host-to-device paths, precision support, collective performance and fault behaviour all matter.

Software is part of the acceptance boundary. A bid should state which compilers, communication libraries, container runtimes, schedulers, model frameworks and observability tools are supported at launch. It should also set an update policy. A fast-moving research service cannot wait months for a security fix or framework release, but uncontrolled upgrades can break long training campaigns.

Heterogeneous phases are possible. That does not mean every job must run on every device. It means the service should avoid unnecessary coupling between workload intake, data format, scheduler policy and one accelerator API. Where a platform-specific optimisation is chosen, the cost and exit path should be recorded.

Procurement should ask bidders to demonstrate a real model workflow from data ingest to checkpoint recovery. Synthetic accelerator tests have a place, but they do not show whether the system works as a service.

Scale-up and scale-out are different contracts

Inside an accelerator server, high-bandwidth links and switching allow devices to exchange data with lower overhead than standard host paths. Across servers, the facility relies on a scale-out fabric. Those two domains must be specified separately.

An InfiniBand cluster design may suit tightly coupled training where mature RDMA operations, collective libraries and predictable fabric behaviour are priorities. A carefully engineered Ethernet AI fabric can provide a broader operational ecosystem and may align with existing network skills. Neither technology name guarantees results.

The tender needs topology, endpoint and traffic assumptions. State the number of accelerator-facing network interfaces per server, link speed, rail mapping, switch radix, oversubscription ratio and failure domains. Include management, storage and service traffic rather than pretending that one fabric carries nothing but collectives.

Acceptance should use multi-node workloads at the intended scale. Measure all-reduce behaviour, congestion response, tail latency and performance after a link or switch failure. A network that performs well in an empty lab may degrade sharply when storage reads, checkpoints and tenant traffic share the fabric.

Optics and cables deserve their own lifecycle plan. At Gigafactory scale, installation quality, cleaning, telemetry, spare ratios and replacement procedures can determine availability. A proposal that prices switches but treats the physical layer as an accessory is unfinished.

Storage must be tested against training behaviour

The processor count implies a storage problem of equal seriousness. Training data may consist of large shards, billions of smaller objects or a mixture. Checkpoints can produce large coordinated writes. Model loading can create bursts when many workers start together. Metadata traffic can dominate even when aggregate capacity appears sufficient.

The scale-out storage design should define performance by workload phase:

  • Dataset ingestion and transformation.
  • Repeated training reads across many workers.
  • Checkpoint write and restart time.
  • Model registry and artefact access.
  • Inference model loading and cache behaviour.
  • User home, container and small-file metadata operations.

Do not accept one sequential throughput figure as proof. Specify concurrency, file or object size distribution, read/write mix, cache state, metadata operations and failure conditions. Measure useful data delivered to jobs as well as storage-side counters.

Tiering can reduce cost, but the policy must be explicit. A fast tier that is too small may produce unpredictable recalls from capacity storage. An object store can be effective for durable datasets and model artefacts, while a parallel file system may better suit active training. The facility may need both, joined by controlled data movement and a catalogue users can understand.

Data durability and job availability are related but not identical. A system can protect every byte while taking hours to restore a usable namespace. Recovery objectives should cover access to work, not only media protection.

Power and cooling define the deployment pace

The EU framework calls attention to power capacity, efficiency and reliable supply. Those requirements cannot be reduced to a facility-wide power usage effectiveness target.

A bidder needs rack-level figures: normal and peak IT load, allowable step changes, redundancy mode, busway capacity and the number of racks that can be commissioned in each phase. Accelerator systems can create concentrated loads that exceed the assumptions of general-purpose halls.

Cooling design should state inlet conditions, liquid temperatures, flow, pressure, water quality, heat rejection and responsibility at each interface. Direct liquid cooling may be required for dense systems, but the cold-plate loop is only part of the plant. Coolant distribution units, secondary loops, controls, leak detection, maintenance isolation and spare capacity all need acceptance tests.

Energy reuse can be valuable where a suitable heat customer exists. It should not be used to disguise uncertain heat-rejection capacity. The tender should distinguish guaranteed operation from an environmental benefit that depends on a separate network or seasonal demand.

Power availability may force staged deployment. A staged plan is credible when each stage provides a useful service and the network, storage and cooling designs have clear expansion points. Installing accelerator rows faster than the supporting plant can be accepted simply converts capital into idle equipment.

Supply chain and serviceability

At this scale, no single shipment is the system. Accelerators, hosts, switches, optics, storage media, power equipment and cooling components will arrive on different schedules. Firmware and manufacturing revisions may change during the build.

The procurement should require a configuration-control process. Every accepted variation needs compatibility evidence, updated documentation and a rollback route. Serial-number and firmware inventory should feed the operations platform from day one.

Spares policy must be based on failure rate, replacement time and common-mode risk. Keeping a few complete servers is not enough if a scarce optical module or coolant component can stop a row. Suppliers should identify field-replaceable units, repair locations, response targets and parts ownership.

Serviceability also affects architecture. Can a compute tray be isolated without draining a large cooling zone? Can a switch be replaced without recabling an entire row? Are storage rebuilds contained? These questions have direct effects on delivered capacity.

Sovereignty needs measurable definitions

“Sovereign AI” can refer to data location, ownership, legal control, software visibility, supply diversity or operational autonomy. A tender should state which of these is required.

European data residency does not by itself create technical independence. A facility may still depend on foreign accelerators, firmware, cloud services or support personnel. Conversely, using an international supplier does not automatically prevent European control of data and operations.

Useful requirements include control of encryption keys, documented administrator access, auditable software supply chains, incident response within the relevant jurisdiction and the ability to operate during a supplier outage. Open interfaces and exportable data formats can reduce switching cost, but open-source software still needs maintainers and support.

The strongest procurement position is transparent about unavoidable dependencies and funds mitigation. Pretending a complex accelerator stack has no external dependencies makes the risk harder to manage.

How bidders should prove performance

A phased acceptance programme should begin below production scale and grow. Each phase should use the same instrumentation and pass criteria.

Start with component and node validation. Confirm thermals, memory errors, storage paths, network interfaces and management controls. Move to a rack test with realistic power transitions and concurrent traffic. Then validate multi-rack collectives, checkpointing, scheduler recovery and tenant isolation.

At facility scale, test degraded conditions. Remove a link, a switch, a storage target and a compute node. Measure how long the service takes to identify the fault, reschedule work and restore capacity. Planned maintenance should be tested as carefully as unexpected failure.

Benchmark disclosure matters. Record software versions, model, dataset shape, precision, batch size, node count, topology, storage state and excluded time. A result without these details cannot be reproduced or compared.

Performance guarantees should be tied to workload classes, not one hero benchmark. The facility is successful when users complete accepted work within target time and cost.

What prospective users should prepare now

Organisations hoping to use EuroHPC capacity do not need to wait for every procurement detail. They can prepare their workloads for a shared service.

Containerise the application, document external dependencies and separate data preparation from accelerator execution. Test checkpoint restart after a forced interruption. Measure storage behaviour and network communication rather than recording accelerator utilisation alone. Define the data classification and access controls the project needs.

Estimate compute requests in completed work: model size, token count, run duration, repetitions and evaluation. EuroHPC’s Large Scale Access programme, for example, is designed for proposals requiring more than 50,000 GPU hours and expects applicants to justify both the allocation and their capacity to use it. A credible request explains why smaller resources are insufficient.

Teams can use owned or hosted infrastructure to de-risk software before applying for shared capacity. This does not reproduce a Gigafactory, but it can expose unsupported dependencies, storage bottlenecks and weak restart procedures.

What should not be bought early

Do not commit to a specific accelerator quantity based on an informal presentation. EuroHPC explicitly states that preliminary material is non-binding and that the formal call controls.

Do not buy network components before endpoint counts, topology and optics are agreed. Do not select storage from capacity and peak GB/s alone. Do not promise a rack density the site cannot cool. Do not treat software support as a future integration task.

Pilot equipment can still be useful when it answers a defined question and can be redeployed. A small cluster for application readiness, network testing or storage qualification has a clear purpose. A speculative miniature of an unknown final design does not.

Frequently asked questions

Is an AI Gigafactory the same as an AI Factory?

No. The Commission describes Gigafactories as a larger class of facility intended for next-generation models and more than 100,000 advanced AI processors. AI Factories already provide AI-optimised EuroHPC resources and support services to European users.

Has the final EuroHPC Gigafactory specification been published?

Prospective bidders must check the current EuroHPC call documents. Material from the June 2026 information session was described by EuroHPC as preliminary and non-binding. Only the formal tender establishes the final legal requirements.

Must a Gigafactory use one accelerator vendor?

The architecture should follow the tender and selected design. From an engineering perspective, heterogeneous phases are possible, but each platform adds software, scheduling and support work. Diversity should solve a stated risk rather than exist for appearances.

Is InfiniBand required for a facility this large?

Not automatically. InfiniBand and engineered Ethernet fabrics can both support AI clusters. Endpoint design, topology, congestion control, collective behaviour, operations and acceptance results matter more than the label.

How should storage be sized?

Capacity is only the start. Buyers need workload-specific throughput, small-file and metadata behaviour, checkpoint time, concurrency, recovery and tiering tests. The storage system must be assessed with the planned network and compute clients.

Can GPUMachines supply an entire AI Gigafactory?

GPUMachines should not be represented as the prime contractor for a EuroHPC Gigafactory without a specific agreement. It can support component sourcing, reference clusters, workload pilots and design work for compute, storage and networking within an appropriately governed consortium.

Verdict

EuroHPC’s AI Gigafactory initiative should be read as a service procurement, not an accelerator shopping exercise. The strongest bids will connect public goals to measurable workload, reliability, energy and access outcomes. They will also admit where technology and call requirements may change.

For suppliers and prospective users, the useful work now is evidence gathering: validate applications, quantify data movement, test recovery and document dependencies. That preparation remains valuable even if the final tender changes a component choice.

To build a pilot or reference environment, ask GPUMachines to review the compute, fabric, storage and hosting assumptions before they enter a consortium proposal or capacity request.

Sources and Further Reading

← Back to blog