GPUmachines

Llama 3.1 405B Hardware Requirements: GPU and Memory Sizing

Size Llama 3.1 405B from its 810 GB BF16 weight floor, then add quantisation, KV cache, concurrency, GPU topology, host RAM, NVMe and failover.

Llama 3.1 405B Hardware Requirements: GPU and Memory Sizing

Llama 3.1 405B is an eight-GPU-class model at full precision, not a large workstation model. Its theoretical weight floor is about 810 GB at BF16, 405 GB at 8-bit and 202.5 GB at 4-bit. Runtime workspace, quantisation metadata, allocator margin and KV cache come on top.

An eight-GPU H200, B200 or B300 platform is the cleanest single-node starting point for production evaluation. A carefully tested FP8 deployment can fit smaller GPU sets by weight, but capacity alone does not make the topology fast or resilient. Multi-node serving becomes appropriate when one node cannot provide the required cache, concurrency, throughput or failure headroom.

The correct quotation must name Llama 3.1 405B, the exact checkpoint and precision. Llama 405B without those fields is not enough to size the system.

What Meta actually released

Meta's official model card lists Llama 3.1 in 8B, 70B and 405B pretrained and instruction-tuned versions. The 405B model is text-in/text-out, supports a 128K context window and uses Grouped-Query Attention. Meta names eight supported languages and distributes the model under the Llama 3.1 Community License and Acceptable Use Policy.

The model card is also clear about scale. Meta reports 30.84 million H100-80GB GPU hours for 405B pretraining. That figure is not a hardware recommendation for inference, but it prevents a common category error: serving the released model and reproducing its original training run are entirely different projects.

Most organisations evaluating 405B are planning inference, synthetic-data generation, distillation or a narrowly scoped adaptation. Full pretraining or full-parameter fine-tuning needs a separate distributed-training design.

Weight-memory floor

Use the simple first-pass calculation:

weight memory = 405 billion parameters x bits per weight / 8

| Weight format | Theoretical weight floor | Practical implication | | --- | ---: | --- | | BF16 or FP16 | About 810 GB | Eight high-memory GPUs are the practical single-node class | | FP8 or INT8 | About 405 GB | Four H200/B200-class GPUs can fit weights nominally, but cache and runtime need validation | | 4-bit | About 202.5 GB | Two or more high-memory GPUs can fit weights, but quality and kernel efficiency decide value |

These numbers exclude scales, zero points, runtime workspace, CUDA graphs, KV cache and fragmentation. A checkpoint that fits with two gigabytes to spare is not a production configuration.

Quantisation is also a model-quality decision. Test the chosen artefact on the organisation's multilingual, coding, reasoning, tool-use and safety workloads. Keep its exact hash with the benchmark. 4-bit Llama 405B is not one reproducible model.

Single-node platform options

Eight H200 SXM GPUs

Eight H200 GPUs provide about 1.1 TB of aggregate HBM3e. BF16 weights fit with a nominal margin for the engine and cache, while FP8 leaves much more room for batching. The HGX H200 scale-up domain uses NVLink and NVSwitch at 900 GB/s per GPU.

NVIDIA publishes a Llama 3.1 405B inference result using eight H200 GPUs, FP8 and TensorRT-LLM. Treat the published throughput as evidence that this software/hardware combination is viable, not as a quote for a different prompt length, batch policy or service objective.

Eight B200 GPUs

Eight B200 GPUs provide 1.44 TB of HBM3e, up to 8 TB/s of memory bandwidth per GPU and 1.8 TB/s of NVLink bandwidth per GPU. The extra memory gives BF16 deployments more cache and workspace room, while Blackwell low-precision support can improve throughput when the chosen TensorRT-LLM or other engine release supports the exact path.

B200 has a higher facility requirement. NVIDIA allows each GPU to be configured up to 1,000 W. Eight accelerators can account for 8 kW before the host, NICs, drives, fans and power-conversion losses.

Eight B300 GPUs

B300 supplies 288 GB per GPU, or 2.3 TB across eight GPUs. That capacity is useful for longer contexts, larger batches, multiple replicas or models beyond 405B. It is not required merely to make Llama 3.1 405B weights fit.

A B300 platform should be justified by accepted throughput, future model plans or consolidation. Do not buy unused HBM as a substitute for a traffic forecast.

Four H200 NVL or SXM GPUs

Four 141 GB H200 GPUs provide 564 GB of aggregate memory. FP8 weights fit by arithmetic, leaving roughly 159 GB before runtime and cache. H200 NVL supports two- or four-way NVLink bridges in certified PCIe/MGX systems; H200 SXM uses the HGX scale-up design.

This can be a sensible inference profile when FP8 quality and the intended context pass testing. The topology must match the engine. Four PCIe cards that cannot use the expected bridge arrangement are not equivalent to a validated four-GPU system.

Four B200 GPUs

Four B200 GPUs provide 720 GB, enough for FP8 weights with more headroom than four H200s. BF16 weights still do not fit by capacity. B200 is normally purchased as an eight-GPU HGX platform, so verify that the proposed OEM configuration and software support the intended four-GPU operating model rather than assuming half a baseboard behaves like a separate product.

Why 128K context changes capacity

Meta lists a 128K context window and Grouped-Query Attention. GQA reduces KV-cache memory compared with storing a separate key/value head for every query head, but cache still grows with retained tokens, active sequences, layers and cache datatype.

A service promising one 128K request at a time has a different capacity profile from one promising 32 concurrent 128K sessions. Most production deployments need an operational limit below the model maximum, plus a queue or admission policy for unusually long prompts.

Measure at least four distributions:

  • input tokens at median, P95 and maximum accepted length;
  • generated output tokens;
  • simultaneous active sequences;
  • prefix-cache hit rate, where the serving engine uses one.

Then record peak HBM and P95 latency. Do not multiply the advertised context window by a user count and call that a specification; real requests vary, batching changes occupancy, and the engine may page or quantise cache.

Tensor and pipeline parallelism

Inside one HGX node, tensor parallelism can split each layer across the eight-GPU NVSwitch domain. That keeps frequent collective traffic on the scale-up fabric. Pipeline parallelism divides layers into stages and is often used across nodes when a model or its production headroom no longer fits one server.

NVIDIA NIM documents Llama 3.1 405B as a multi-node example with tensor parallel size eight per node and pipeline parallelism across two eight-GPU nodes. This is a deployment pattern, not a mandatory topology. A single eight-GPU node can serve 405B when the selected precision, cache and throughput fit.

Crossing a node boundary changes the network requirement. Activations and control traffic now depend on the east-west fabric. A nominal NIC speed is not enough; GPU-direct paths, rail mapping, switch oversubscription and software configuration must all be validated.

Host CPU and RAM

The GPUs perform the model compute, but the host still handles tokenisation, request scheduling, storage, networking, monitoring and model preparation. Choose the CPU platform from PCIe lanes and balanced I/O rather than peak core count alone.

NVIDIA's HGX H100/H200/B200 reference architecture specifies dual-socket x86 hosts, at least 1.5 TB of system RAM and at least 500 GB/s aggregate memory bandwidth. It calls for symmetric DIMM population and balanced PCIe root connectivity. Those figures are reference-design guidance, but they are a sensible warning against pairing eight GPUs with an under-populated host.

Operational teams may need a complete checkpoint in system memory during conversion or startup. A BF16 405B artefact is about 810 GB before packaging overhead, so 1.5 to 2 TB of host RAM can be reasonable for a node that performs local conversion. If the process streams weights from NVMe instead, benchmark the cold-start path and keep swapping disabled.

NVMe and model distribution

A 405B repository is not one file. Production needs the approved checkpoint, tokenizer and configuration, an engine or converted format, a rollback release, container images and temporary conversion space. Quantised variants can add hundreds of gigabytes each.

Use local enterprise NVMe for the active engine and predictable restarts. A practical starting point is 4 to 8 TB per node, then increase it from the release policy. Two complete BF16 generations already consume roughly 1.6 TB before converted engines and scratch files.

Keep a source copy in shared or object storage, but test a simultaneous rollout. If eight nodes each pull an 800 GB checkpoint, the storage tier sees 6.4 TB of reads before the service becomes healthy. Staggering or peer distribution may be necessary.

Record these operational times:

  • download to local NVMe;
  • engine build or conversion;
  • load from NVMe to GPUs;
  • health-check completion;
  • rollback to the previous version.

Model load is part of availability, not a one-off installation detail.

Network design

One eight-GPU node does not need an external AI fabric for its internal tensor-parallel traffic; NVLink and NVSwitch handle the scale-up domain. It still needs client/API, storage and management networks.

A multi-node deployment needs a high-speed east-west fabric. Start at 400 Gb/s-class networking per rail for current HGX reference designs and calculate ports from the exact number of NICs in each server. Newer 800 Gb/s adapters can reduce port count or increase per-node bandwidth, but only if the server, switches, optics and software support them as one design.

Separate or explicitly isolate:

  • model-parallel traffic;
  • model and dataset distribution;
  • user/API traffic;
  • provisioning and telemetry;
  • out-of-band management.

Use the AI cluster network architecture guide for leaf, spine and rail arithmetic. Do not buy switches from total GPU count alone.

Power, cooling and rack planning

Eight H200 SXM GPUs can represent up to 5.6 kW of accelerator power. Eight B200 GPUs can represent up to 8 kW. Complete server input is higher once CPUs, RAM, NICs, NVMe, fans or pumps and PSU losses are included.

Obtain the OEM maximum-input figure and power-feed requirement for the exact server. Then reserve space and power for network switches, storage and any coolant distribution equipment. A four-node rack can cross a facility threshold even when one server appears manageable.

Run a sustained replay and measure power at the rack PDU. Inference demand can be bursty, while long-context or high-batch tests may hold the GPUs near their configured limit. The rack planner needs both maximum draw and realistic steady-state data.

Inference, fine-tuning and training are different designs

For inference, weights, KV cache, latency and availability define capacity. For parameter-efficient tuning, activations, optimiser choice, sequence length and which layers are trainable change memory demand. Full-parameter fine-tuning adds gradients and optimiser states that can multiply the weight footprint.

Do not take an eight-GPU inference node and promise full 405B training. Meta's reported 30.84 million H100 GPU hours illustrates the original pretraining scale. A full training project needs data engineering, checkpoint storage, a large high-speed fabric, reliability engineering and a separate budget.

Many organisations will get more value by using 405B as a teacher for synthetic data or distillation, then serving a 70B or smaller model. Benchmark the smaller model against the target task before building permanent 405B capacity.

Production availability

One eight-GPU node is one failure domain. A driver update, failed component or model reload removes the entire service. Production usually needs a second replica, a degraded-service model or a documented maintenance window.

Two full 405B replicas may be expensive. Alternatives include:

  • an active/passive pair with the standby used for testing;
  • two active replicas sized so one can carry reduced traffic;
  • a smaller fallback model for outages and overload;
  • public or hosted burst capacity while owned hardware recovers.

State the failure behaviour in the quotation. Throughput measured with every GPU occupied is not failover capacity.

Acceptance test before purchase

Use a prompt replay that preserves real token-length and output distributions. Remove confidential data, but keep languages, tool schemas and request mix. Record:

1. exact model revision, precision and quantisation hash; 2. engine, container, CUDA and driver versions; 3. tensor and pipeline-parallel layout; 4. time to first token and output tokens per second at P50 and P95; 5. peak HBM at the agreed context and concurrency; 6. inter-GPU and inter-node bandwidth utilisation; 7. complete-system power and throttling; 8. model load, restart and failover time; 9. quality results on the organisation's own task set.

NVIDIA's published benchmark used eight H200 GPUs, FP8, TensorRT-LLM 0.19.0 and short 128-token input/output lengths. It is useful evidence, but it is not a substitute for a 32K-document or multi-user replay.

FAQ

How much GPU memory does Llama 3.1 405B need?

Weights alone need about 810 GB at BF16, 405 GB at 8-bit or 202.5 GB at 4-bit. Add runtime, quantisation and KV-cache headroom. An eight-GPU HGX system is the practical full-precision single-node class.

Can four H200 GPUs run Llama 405B?

Four H200s provide 564 GB, so FP8 weights fit nominally. The exact engine, topology, context and concurrency must still fit. BF16 weights do not fit four H200s.

Can one B300 run Llama 405B?

One 288 GB B300 cannot hold BF16 or FP8 weights by simple arithmetic. A 4-bit checkpoint can fit by weight, but runtime, cache, kernel support and output quality still need validation.

Does Llama 405B require multiple nodes?

Not necessarily. One eight-GPU H200, B200 or B300 node can hold common production formats. Multiple nodes are used when precision, cache, concurrency, throughput or availability exceeds one node.

Is 128K context available to every user?

No hardware design should promise that without a concurrency test. KV cache grows with retained tokens and active sequences. Publish a tested operational limit and an admission policy.

Which serving engine should be used?

TensorRT-LLM, NVIDIA NIM, vLLM and SGLang can all be relevant depending on the GPU and release. Use a version whose support matrix includes the exact model and hardware, then freeze it for benchmarking.

Recommendation

Start with an eight-GPU high-memory node for a serious Llama 3.1 405B evaluation. H200 is a proven FP8 route; B200 adds HBM and a faster NVLink domain; B300 adds far more memory for context, batching or larger future models. Move to multiple nodes only after a single-node replay identifies the limiting resource or the availability plan requires another failure domain.

GPUMachines can compare HGX server platforms, model multi-node networking through the GPU cluster configurator, or place dedicated hardware through Buy & Host.

Sources

← Back to blog