GPUmachines

AMD and Core Scientific Plan 2.5GW of AI Capacity: What GPU Buyers Should Learn

A purchase order cannot create megawatts. AMD's Core Scientific agreement shows why power, cooling and deployment-ready space must be settled before a large GPU order.

AMD and Core Scientific Plan 2.5GW of AI Capacity: What GPU Buyers Should Learn

A purchase order cannot create megawatts.

That sounds obvious until a GPU project reaches the point where the selected accelerators, servers and network design are ready, but the building cannot supply the required electrical capacity or remove the heat. The order then becomes an expensive queue: hardware waits for switchgear, cooling plant, utility work or a suitable site.

AMD and Core Scientific have put that constraint in unusually large numbers. Their 28 July 2026 agreement gives AMD access to more than 500MW of US data-centre capacity from 2027, with an option to expand the relationship to 2.5GW. Core Scientific's second-quarter results describe 15-year agreements covering roughly 530MW across five sites and more than $14 billion of potential base contracted revenue.

For most buyers, the headline figure is not a template to copy. It is a warning about sequence. Secure the facility, deployment route and operational model before treating the GPU bill of materials as final.

The Answer for Buyers

If your organisation is planning a dense AMD Instinct deployment, answer these questions before ordering the full cluster:

  • Where will the first production rack run, and what power can that room deliver continuously rather than on paper?
  • Does the cooling design match the server or rack configuration that will actually be purchased?
  • Which network and storage paths must be present on day one?
  • Can the site accept staged expansion without rebuilding the electrical and mechanical design?
  • Who owns commissioning, remote access, monitoring, spares and incident response after the systems arrive?

No single GPU specification settles those points. A pilot can run in a workstation, PCIe server or hosted rack while the wider facility plan matures. That is often the sounder route than holding expensive accelerators in storage.

The GPUMachines GPU cluster configurator can help frame compute, networking and storage requirements, but the output should be tested against a real site survey or hosted-capacity plan before procurement.

What AMD and Core Scientific Announced

The joint infrastructure announcement says AMD will secure more than 500MW of US capacity for customer deployments beginning in 2027. The parties plan to work together on physical infrastructure design and deployments involving AMD Instinct GPUs, EPYC CPUs and ROCm software. AMD also holds an expansion option that could take the arrangement to 2.5GW.

Core Scientific's second-quarter release gives more commercial detail. It refers to approximately 530MW across five sites, 15-year agreements and more than $14 billion of potential base contracted revenue. Those figures describe planned capacity and contractual potential; they are not a record of 530MW already operating as AMD AI compute.

That distinction matters. Data-centre announcements often combine existing buildings, capacity under conversion, future construction and expansion rights. Buyers should ask when each block becomes ready for IT load, what density it supports, which cooling method it uses and what acceptance tests define delivery.

GPUMachines has not inspected these five sites or independently checked the commercial forecasts. The useful fact is narrower: AMD is treating deployment-ready capacity as part of its customer supply chain, not as a problem left entirely to whoever buys the accelerators.

Why AMD Wants Capacity Rather Than Another Hardware Agreement

Accelerator supply is only one gate. Large installations also depend on utility interconnection, transformers, switchgear, backup power, cooling equipment, fibre routes, rack systems and people who can operate the facility.

A silicon vendor can have a competitive GPU and still lose a deployment if the customer's chosen location will not be ready for eighteen months. Capacity reservations reduce that mismatch. They give customers somewhere to place the systems while AMD and its partners coordinate a platform that spans Instinct accelerators, EPYC hosts, Pensando networking and ROCm.

AMD's Helios rack-scale reference design illustrates the direction. AMD describes Helios as an open rack-scale blueprint combining MI455X GPUs, sixth-generation EPYC CPUs, Pensando networking and ROCm. It is a reference design rather than a server sold directly by AMD, and volume deployments are expected through system partners.

The Core Scientific agreement should therefore be read as part of a wider deployment route. It does not mean every reserved megawatt will contain the same rack, nor does it identify the precise accelerator mix for every customer.

Do Not Convert 2.5GW Straight into a GPU Count

News coverage often turns a power figure into an implied accelerator count. That calculation looks tidy and can be badly wrong.

The facility number may refer to gross site power, critical IT load, contracted customer power or a future expansion ceiling. Those are not interchangeable. A GPU's board power also tells only part of the story. Host CPUs, memory, network adapters, storage, fans, pumps and power-conversion losses consume energy. Cooling and electrical overhead sit above the IT load.

Even the rack unit changes. A PCIe server pool, an eight-GPU scale-up system and a liquid-cooled rack-scale platform can place very different demands on distribution and cooling. Redundancy policy changes usable capacity again because the site must survive a failed feed, UPS module or cooling component.

For a buyer-sized project, use a bottom-up model:

1. Start with the finished node or rack configuration and its vendor power envelope. 2. Add switches, storage, management equipment and realistic growth space. 3. Apply the site's redundancy and derating rules. 4. Confirm whether the quoted facility figure means usable IT capacity. 5. Check the cooling design at expected load, not only nameplate maximum. 6. Repeat the calculation for a failed component or maintenance state.

That model will not produce a glamorous social-media number. It will tell you whether the installation can run.

Power Planning Starts Upstream of the Rack

Rack PDUs are the visible end of a long electrical chain. Utility service, substations, transformers, generators, UPS systems, busways, branch circuits and rack distribution all need enough capacity and the right redundancy.

The first practical question is voltage. Dense accelerator systems may require input conditions that a normal server room does not provide. A room with spare floor space can still be unsuitable because its branch circuits, connectors or upstream distribution were designed for much lower rack loads.

Then check failure behaviour. If a rack uses two feeds, can either side carry the required load during maintenance or a fault? Does a dual-corded server spread load as expected? What happens if a UPS module or generator is unavailable? A design that only works with every component healthy has no operational margin.

Power quality and commissioning matter as well. New equipment should not become the test instrument for a hurried electrical build. Meter the feeds, verify protection settings, test transfer behaviour and record the accepted configuration before production jobs begin.

The GPUMachines article on power requirements for AI clusters covers the calculation in more detail. The short version is simple: specify the rack and the room together.

Cooling Is a Platform Choice

High-density AI systems turn cooling from a facilities footnote into part of product selection.

Air-cooled PCIe servers remain practical for many deployments, particularly where jobs can be distributed across independent GPUs and rack density stays within the room's airflow limits. They still need correct front-to-back airflow, containment and enough fan capacity at the chosen accelerator power.

Direct liquid cooling supports higher density, but it adds another system to own. Buyers must consider coolant distribution units, facility-water conditions, hoses, drip detection, isolation valves, service procedures and what happens when a loop needs maintenance. A liquid-ready building is not necessarily ready for every vendor's rack.

Some teams should lower density rather than rebuild a room. Spreading the same compute across more racks can reduce local cooling pressure, though it consumes more floor space and may lengthen network cables. Others will find that colocation or hosted ownership costs less than retrofitting an office or laboratory.

AMD's Helios material identifies direct liquid cooling for its rack-scale design. That is a planning signal, not permission to assume that any liquid-cooled site can accept the platform without engineering review.

Networking Must Be Reserved with the Space

A large AI hall without the intended network is unfinished capacity.

Within a rack, the design may include scale-up connections, PCIe paths and local switching. Across racks, distributed training and large inference services need a scale-out fabric sized for their communication pattern. Storage traffic, management access and user traffic may need separate paths so a checkpoint storm or model load does not interfere with cluster collectives.

AMD's new platform direction includes Pensando networking. The GPUMachines analysis of the AMD Vulcano 800 AI NIC explains why NIC count, rail mapping, switch ports and congestion behaviour must be designed before ordering cables. A late network decision can leave accelerators waiting for ports or force awkward oversubscription.

Capacity planning should include:

  • switch positions and their power draw;
  • optic and cable quantities, reach and replacement stock;
  • NIC placement in every server;
  • management ports, service processors and out-of-band access;
  • storage-client bandwidth under checkpoint and model-loading pressure;
  • fabric expansion without recabling the first phase.

Fibre availability outside the building matters too. A hosted cluster that cannot ingest datasets or serve customers at the required rate has only moved the bottleneck.

Storage Can Delay a Cluster That Is Otherwise Ready

GPU projects often place the storage order after compute. That is risky when the workload needs shared datasets, frequent checkpoints or a production model repository.

Training storage must feed many clients, handle metadata pressure and recover cleanly from failures. Inference storage has a different pattern: model loading, retrieval corpora, logs, cache tiers and tenant data may dominate. A mixed platform needs policy for scratch data, durable datasets, backups and archive rather than one large capacity number.

Local NVMe can absorb temporary work and reduce pressure on shared systems. It does not replace a data lifecycle. If a node fails, the team must know which data can disappear and which state needs another copy.

For larger designs, GPUMachines scale-out storage guidance helps separate bandwidth, metadata, durability and operational requirements. Storage should be commissioned with representative clients before the GPU acceptance run, not after utilisation problems appear.

ROCm Readiness Belongs in the Capacity Plan

Physical space does not make software ready.

The AMD and Core Scientific release names ROCm alongside Instinct GPUs and EPYC CPUs. That is sensible because an accelerator deployment succeeds only when the intended frameworks, models, kernels, containers, monitoring and orchestration work on the selected platform.

Before reserving a large production block, run the actual stack on the accelerator generation you plan to buy. Check model support, precision, distributed communication, container images, driver management and the failure cases your operations team will see. Measure end-to-end task completion, not a single kernel.

The pilot also reveals host requirements. Tokenisation, data preparation, retrieval, storage clients and control-plane services use CPU and memory. A rack designed around accelerator arithmetic can underperform because the surrounding work was treated as negligible.

Do not assume that code which runs on one ROCm system will scale unchanged across many racks. Collective communication, scheduling, fault recovery and job restart deserve their own tests.

Build the First Useful Phase, Not the Final Dream

The 2.5GW option is a reminder that capacity can be staged. Most buyers should do the same at a far smaller scale.

A sensible first phase is large enough to run representative production work but small enough to change. It should prove the software image, network rails, storage path, monitoring, support process and power measurements. Expansion follows measured queue pressure and utilisation rather than a forecast made before users arrived.

Four common routes emerge:

Existing room, modest density

Use a workstation or a small number of PCIe GPU servers where power, acoustics and cooling are known. This suits research, evaluation and independent inference jobs. Do not force a dense rack-scale platform into a room that was never designed for it.

Colocation with owned hardware

The GPUMachines Buy & Host service fits organisations that want to own the systems but do not want to build the facility layer. The commercial comparison should include remote hands, cross-connects, power charging, support ownership and hardware access.

Dedicated hosted pilot

Hosted capacity is useful when model choice, concurrency or software support remains unsettled. It buys measurement time before the organisation commits to a fixed facility and full cluster.

Purpose-built on-premise cluster

This route makes sense when data locality, operating control or sustained utilisation justifies the building work. It also places responsibility for power, cooling, security, monitoring and spares on the owner. Budget for that responsibility.

Who Should Not Follow the Gigawatt Story

A gigawatt agreement can make a normal enterprise project feel too small or too late. Ignore that pressure.

Do not reserve a large hardware block because a vendor announced a large campus. If the workload is a pilot, rented capacity or one configurable server may answer the important questions. Do not rebuild a machine room before confirming that the selected model and software stack meet the business requirement. Do not buy rack-scale hardware for jobs that run independently on PCIe GPUs.

Some organisations need no private cluster at all. Short campaigns, uncertain demand and teams without infrastructure staff may be better served by cloud or hosted capacity. Ownership becomes attractive when use is sustained and the buyer can operate the system.

The right scale is the smallest one that satisfies the workload, service and growth plan.

What GPUMachines Would Ask Before Quoting

Bring evidence rather than a round GPU count:

  • target models, precision and expected context or batch behaviour;
  • training, fine-tuning, inference and non-AI workloads;
  • expected users, peak concurrency and job duration;
  • dataset size, checkpoint pattern and model-loading requirements;
  • current facility power, cooling method and rack limits;
  • preferred deployment route and data-location restrictions;
  • network topology, site connectivity and management separation;
  • an expansion horizon tied to measured demand.

GPUMachines can then review the server class, accelerator choice, CPU and RAM, local NVMe, storage, switches, optics, rack plan and hosted alternatives. The review is configuration-dependent; it does not replace a qualified electrical or mechanical site assessment.

Questions Buyers Are Likely to Ask

Does AMD now own 2.5GW of operating data-centre capacity?

The announcement describes more than 500MW of US capacity beginning in 2027 and an option to expand the relationship to 2.5GW. Core Scientific reports roughly 530MW under 15-year agreements across five sites. Buyers should distinguish contracted or planned capacity from capacity already commissioned with live AI systems.

Does the agreement specify which AMD Instinct GPUs will be installed?

It names AMD Instinct GPUs, EPYC CPUs and ROCm, but the public announcement does not assign one precise accelerator and rack configuration to every megawatt.

Should a smaller buyer choose AMD because this capacity exists?

Capacity access can help AMD's deployment route, but platform choice should still follow workload support, software readiness, model performance, operational skills, commercial terms and facility fit.

Is liquid cooling mandatory for every AMD Instinct server?

No. Cooling depends on the chosen system and density. AMD's Helios rack-scale design uses direct liquid cooling, while other Instinct deployments may use different chassis and thermal designs. Verify the exact server.

Can GPUMachines host an AMD GPU system?

GPUMachines can review hosted ownership and other deployment routes for suitable configurations. Final availability, power density and compatibility need confirmation during quoting.

What should be tested in a pilot?

Test the real model or application, ROCm stack, network path, storage behaviour, monitoring, restart process and end-to-end task time. Record power and thermal behaviour under representative load.

Sources and Further Reading

Verdict

AMD's agreement with Core Scientific is less interesting as a race to the largest number than as evidence of what now holds AI projects back. Accelerators need a prepared building, a working software stack and a commissioned data path. Miss any one of those and the GPU delivery date stops mattering.

For a GPUMachines buyer, the practical response is to settle deployment location early, prove the workload on a small but representative platform, and expand only after the facility and operations model have survived real use.

Plan an AMD AI cluster, its network and storage through the GPUMachines cluster configurator.

← Back to blog