GPUmachines

A100 vs H100: Which NVIDIA GPU Platform Should You Buy?

Compare NVIDIA A100 and H100 across memory, bandwidth, FP8, MIG, NVLink, power and server design before choosing a PCIe or HGX platform.

A100 vs H100: Which NVIDIA GPU Platform Should You Buy?

NVIDIA A100 and H100 are both data-centre accelerators, but they are not drop-in alternatives. A100 is an Ampere-generation platform available in PCIe and SXM forms. H100 moves to Hopper, adds FP8 support and the Transformer Engine, increases memory bandwidth, and changes the surrounding server, power and scale-up requirements.

The first buying question is not simply whether H100 is faster. It is whether the workload, software and facility can use the extra capability. A mature A100 estate may still be economical to expand. A new transformer training cluster usually has a stronger case for H100 or a later platform. A small inference service may need neither.

The short answer

  • Choose A100 80 GB PCIe when you need a proven 300 W data-centre GPU, your software is already qualified on Ampere, or you are extending a compatible server estate.
  • Choose A100 80 GB SXM when an existing HGX A100 platform and its NVSwitch topology remain operationally useful. Do not assume an SXM module can be moved into a PCIe server.
  • Choose H100 PCIe or H100 NVL for newer PCIe inference and mixed workloads that can benefit from Hopper features without moving to an eight-GPU HGX baseboard.
  • Choose H100 SXM in HGX H100 for tightly coupled multi-GPU training, HPC and large-model workloads that can use 900 GB/s NVLink per GPU and the HGX scale-up fabric.
  • Choose neither when a smaller PCIe GPU, current-generation workstation, hosted trial or newer Blackwell platform fits the workload and buying horizon better.

For a current system, compare PCIe GPU servers, HGX systems and GPU Cloud. The GPU choice should follow the deployment model, not precede it.

A100 and H100 are families, not single specifications

A quotation that says only A100 or H100 is incomplete. Ask for the memory capacity, form factor, power limit, interconnect and server topology.

A100 80 GB PCIe

NVIDIA specifies 80 GB HBM2e, 1,935 GB/s memory bandwidth, a 300 W maximum TDP and PCIe Gen4. It is a passive, dual-slot card that depends on validated chassis airflow. Two cards can be connected with NVLink bridges in supported systems.

A100 80 GB SXM

The SXM version also has 80 GB HBM2e, but NVIDIA lists 2,039 GB/s memory bandwidth and a standard 400 W TDP. It belongs on an HGX A100 server board. NVIDIA notes that a custom thermal solution version can operate at up to 500 W, so an exact platform specification matters.

H100 80 GB PCIe

The original H100 PCIe card is a passive, dual-slot accelerator with 80 GB HBM3 and a 350 W board design. It brings Hopper compute features to conventional PCIe servers. GPU spacing, power connectors, CPU lane allocation and airflow still need model-specific validation.

H100 NVL

H100 NVL is a 94 GB HBM3 PCIe variant. NVIDIA publishes 3.9 TB/s memory bandwidth, a configurable 350 to 400 W TDP and a 600 GB/s NVLink connection. A supported two-card pair provides 188 GB of aggregate HBM3 for large-model inference, but software still has to partition work across the two GPUs.

H100 SXM

H100 SXM has 80 GB HBM3, 3.35 TB/s memory bandwidth, up to 700 W configurable TDP and 900 GB/s NVLink per GPU. It is used in four- or eight-GPU HGX H100 systems with NVSwitch. That scale-up topology, rather than the module alone, is the reason to buy the platform.

Specification comparison

| Specification | A100 80 GB PCIe | A100 80 GB SXM | H100 NVL | H100 SXM | | --- | ---: | ---: | ---: | ---: | | Architecture | Ampere | Ampere | Hopper | Hopper | | Memory | 80 GB HBM2e | 80 GB HBM2e | 94 GB HBM3 | 80 GB HBM3 | | Memory bandwidth | 1.935 TB/s | 2.039 TB/s | 3.9 TB/s | 3.35 TB/s | | Maximum/configurable power | 300 W | 400 W standard | 350-400 W | Up to 700 W | | Host interface | PCIe Gen4 | HGX/SXM | PCIe Gen5 | HGX/SXM with PCIe Gen5 host path | | Published NVLink bandwidth | 600 GB/s for a supported two-GPU bridge | 600 GB/s per GPU | 600 GB/s per GPU | 900 GB/s per GPU | | Maximum MIG instances | 7 | 7 | 7 | 7 | | Native FP8 Tensor Core support | No | No | Yes | Yes |

The numbers are not a promise of application speed. NVIDIA's peak figures may include sparsity, while training and inference performance depend on model dimensions, precision, kernel support, communication and data delivery. Compare the same software release and workload on both platforms.

What Hopper changes

H100 is not only a higher-clocked A100. Hopper introduces fourth-generation Tensor Cores and native FP8 matrix operations. NVIDIA Transformer Engine can choose between FP8 formats and higher-precision accumulation to reduce memory traffic while managing numerical range.

That is especially relevant to transformer training and inference, but it is not automatic acceleration for every program. The framework, libraries and model path must use compatible kernels. A legacy FP32 or custom CUDA application may see a very different gain from a current Transformer Engine workload.

Hopper also raises the scale-up ceiling. H100 SXM increases published NVLink bandwidth from A100's 600 GB/s per GPU to 900 GB/s. HGX H100 pairs those modules with NVSwitch, giving collective operations a different topology from independent PCIe cards.

For HPC, H100 adds DPX instructions aimed at dynamic-programming workloads such as sequence alignment and graph path calculations. Again, application support decides whether the hardware feature matters.

Memory capacity is not the same as usable model capacity

Both A100 80 GB and H100 SXM expose 80 GB per GPU. H100 NVL increases that to 94 GB. The model still needs more than its weight file: runtime workspace, activations, temporary tensors and KV cache also consume memory.

A rough dense-weight floor is straightforward:

parameter count x bytes per parameter = weight memory

At BF16 or FP16, a 70-billion-parameter model starts near 140 GB before runtime overhead. At eight bits, the weight floor is near 70 GB. At four bits, it is near 35 GB. Quantisation metadata, cache and execution workspace make the real requirement higher.

This arithmetic explains why a two-GPU H100 NVL pair can be useful, but aggregate capacity does not create one physically unified memory pool. Tensor or pipeline parallelism, a supported inference engine and enough inter-GPU bandwidth remain necessary.

Memory bandwidth and data movement

H100's HBM3 bandwidth is a substantial platform advantage. H100 SXM's published 3.35 TB/s is about 64 percent above A100 80 GB SXM's 2.039 TB/s. H100 NVL's 3.9 TB/s is roughly double A100 80 GB PCIe's 1.935 TB/s.

Memory-bound attention, embedding, scientific and inference kernels may benefit, but GPU memory bandwidth is only one link in the path. A server can still stall on host memory, PCIe topology, storage reads, checkpoint writes or a congested network.

Before buying, trace one representative job:

1. Measure dataset and model load time from the intended storage tier. 2. Record host-to-device transfer and preprocessing time. 3. Measure GPU utilisation and memory-controller activity during steady state. 4. Test checkpoint write and restart behaviour. 5. Repeat at the intended multi-GPU and multi-node scale.

The result is more useful than comparing theoretical Tensor TFLOPS from unrelated precision modes.

MIG: similar count, different capacity and operations

Both families support Multi-Instance GPU, which partitions one physical GPU into isolated GPU instances with dedicated compute and memory resources. NVIDIA documents up to seven instances on A100 80 GB and H100 80 GB.

On A100 80 GB, the smallest common profile is 1g.10gb. H100 80 GB also supports seven 1g.10gb instances, while the 94 GB H100 variant uses 12 GB-class profile names. Larger profiles expose different fractions of SMs, cache and copy engines.

MIG is useful for predictable tenancy, development queues and smaller inference services. It does not make every workload elastic. Profile changes, driver requirements, scheduler integration and application-level testing must be planned. Check the current NVIDIA MIG guide against the exact GPU and driver branch before promising a tenant profile.

PCIe versus SXM changes the whole server

PCIe cards offer broader chassis choice and can suit independent jobs, inference replicas and mixed accelerator fleets. They still require a qualified platform. Slot width alone does not prove support: card length, passive airflow, auxiliary power, firmware, retimers, CPU root complexes and NIC placement all matter.

SXM modules are not expansion cards. They are soldered or socketed onto an HGX baseboard designed around high-power GPUs and NVSwitch. An eight-GPU H100 SXM server can place up to 5.6 kW into GPU modules alone before CPUs, memory, drives, NICs, fans and conversion losses are counted. Rack power and cooling must be designed from measured system data and the manufacturer's input specifications.

This is why an H100 SXM quote should include the complete server model, PSU configuration, input-voltage requirements, network adapters, rack allocation and cooling method.

When A100 still makes sense

A100 remains defensible in four situations.

Expanding a validated Ampere estate

If the team already operates A100 servers, adding the same platform can reduce image qualification, scheduler changes and operational training. Confirm spare compatibility and support life before assuming old and new nodes will be identical.

Mature software with no Hopper-specific benefit

Some scientific, analytics and inference workloads are stable on Ampere and do not use FP8 or Transformer Engine. Benchmark the actual code. A lower acquisition cost can be meaningful when throughput, support and energy targets are still met.

PCIe deployments with constrained board power

A100 80 GB PCIe's 300 W TDP may fit servers or facilities that cannot accommodate 700 W SXM modules. That does not make the complete A100 server automatically efficient; compare measured system power per completed job.

Secondary markets with controlled risk

Available A100 hardware can be attractive, but condition, provenance, firmware, cooling parts, warranty and remaining support matter. A low card price does not cover the cost of an unsuitable host.

When H100 earns its premium

H100 is easier to justify when the workload can use Hopper rather than merely run on it.

  • Transformer training or inference uses FP8-capable libraries and has been validated for acceptable accuracy.
  • Jobs are memory-bandwidth bound and profiling shows that HBM throughput constrains useful work.
  • Multi-GPU training can use HGX H100's NVLink and NVSwitch topology.
  • MIG-backed services need Hopper capacity, media engines or current software support.
  • The facility can support the complete platform's electrical, cooling and network design.

Do not use vendor headline speedups as a universal business case. NVIDIA's published comparisons specify models, precisions, software and system configurations. Reproduce a representative workload or obtain a benchmark with the same assumptions.

Cases where neither is the right purchase

A100 and H100 are older platform generations in 2026. A new long-life deployment should also compare H200, Blackwell and current RTX PRO options where software and supply permit.

Neither is a good default for a developer who needs local display output, an office-friendly acoustic profile or occasional experimentation. A tower GPU workstation may be easier to operate. Short projects and uncertain utilisation can start in GPU Cloud before capital is committed.

Inference services that fit comfortably on smaller GPUs may achieve better cost per request by scaling replicas. Conversely, models that exceed one server may need an architecture review rather than a larger card shopping list.

A practical evaluation plan

Use the following sequence before requesting a final bill of materials.

1. Define the workload envelope

Record model size, precision, context length, batch or concurrency target, dataset size, checkpoint frequency and required latency. Include the software versions used in production.

2. Choose the scale-up model

Decide whether GPUs will run independent jobs, communicate in bridged pairs, or participate in an HGX NVSwitch domain. This choice separates many PCIe and SXM designs immediately.

3. Benchmark the constrained path

Measure end-to-end tokens, samples or simulations per second. Record GPU utilisation, HBM use, host memory, storage and network counters. A synthetic GEMM result is not an application acceptance test.

4. Price the complete system

Include CPUs, correctly populated RAM, local NVMe, shared storage, NICs, optics, rack space, power, cooling, support and software. Compare cost per useful unit of work, not accelerator price alone.

5. Test recovery and operations

Validate job restart, checkpoint restore, failed-GPU handling, monitoring, firmware procedures and scheduler behaviour. Production value depends on completed work and recoverability.

Common purchasing mistakes

  • Comparing A100 PCIe with H100 SXM without stating the platform difference.
  • Treating total memory across GPUs as one transparent address space.
  • Quoting sparse peak compute as guaranteed application performance.
  • Assuming any dual-slot server supports a passive data-centre GPU.
  • Ignoring the extra rack power, fan power and network equipment around HGX.
  • Buying H100 for FP8 while leaving the production stack on kernels that never use it.
  • Buying discounted A100 cards before checking host qualification and support.

Frequently asked questions

Is H100 always faster than A100?

H100 has much higher published compute and memory capabilities, but application speed depends on form factor, precision, software and topology. A workload that cannot use Hopper features will not reproduce a transformer benchmark headline.

Can an A100 be replaced by an H100 in the same server?

Not without exact manufacturer confirmation. PCIe generation, power connector, board power, firmware and cooling requirements differ. SXM modules require their own HGX baseboards and are not field substitutes for PCIe cards.

Do A100 and H100 both support seven MIG instances?

Yes, supported 80 GB A100 and H100 configurations can expose up to seven MIG instances. Available profile sizes and operational requirements differ, so check the current MIG guide for the exact product and driver.

Is H100 NVL the same as H100 SXM?

No. H100 NVL is a 94 GB PCIe card intended to support a two-GPU NVLink configuration. H100 SXM is an 80 GB module used on HGX/DGX baseboards with NVSwitch and a higher power envelope.

Which platform is better for LLM inference?

H100 usually has the stronger technical case when FP8, HBM bandwidth and current inference engines are useful. A100 can remain economical for validated models and lower-cost capacity. Model fit, KV cache, concurrency, latency and power decide the outcome.

Which platform is better for multi-node training?

H100 is normally the stronger starting point because of Hopper compute, H100 SXM NVLink and newer cluster fabrics. The network, collective communication efficiency, storage and software configuration still determine scaling.

Verdict

For a greenfield transformer training or high-throughput inference platform, H100 is the more capable architecture. For an existing Ampere environment, an acquisition-sensitive project or software already proven on A100, A100 can still be a rational operational choice.

The quote must name the exact variant and server. Compare A100 PCIe with H100 PCIe for flexible card-based systems, or compare HGX A100 with HGX H100 for tightly coupled scale-up. Mixing those categories produces a misleading result.

GPUMachines can size the complete platform around model memory, GPU topology, CPU and RAM population, NVMe, high-speed networking, rack power and hosting. Start with PCIe GPU servers, HGX systems or the GPU cluster configurator, then validate a representative workload before purchase.

Sources

← Back to blog