GPUmachines

Quantized vs Full-Precision LLMs: Memory, Speed and Quality

Compare BF16, FP8, INT8 and 4-bit LLM deployment by GPU memory, runtime support, speed and measured quality, not file size alone.

Quantized vs Full-Precision LLMs: Memory, Speed and Quality

Quantisation can turn an LLM that needs several GPUs into one that fits on a single accelerator. It can also produce a smaller checkpoint that is slower than expected, unsupported by the chosen runtime or measurably worse on the questions that matter to the business.

The useful comparison is not simply quantized versus full precision. A deployment decision must identify what is quantized, which numerical format is used, which kernels execute it, the GPU architecture, the model and the target workload. A weight-only 4-bit model with BF16 activations is a different proposition from FP8 weights and activations, while KV-cache quantisation affects a different part of the memory budget again.

For most LLM inference projects, "full precision" means the released BF16 or FP16 checkpoint. It rarely means FP32. This guide uses BF16/FP16 as the unquantized baseline and explains where FP8, INT8, INT4, NF4, AWQ and GPTQ fit.

The short answer

  • Choose BF16 or FP16 when fidelity, broad runtime support and a clean evaluation baseline matter more than GPU count.
  • Test FP8 when the selected model, runtime and Hopper or newer NVIDIA platform provide a validated path. It can halve weight memory relative to BF16 while retaining a floating-point representation.
  • Consider INT8 when the serving stack has mature kernels and a measured quality result for the exact checkpoint.
  • Use 4-bit weight-only formats when fitting a model into fewer GPUs is the main constraint, but benchmark quality and speed before buying around the result.
  • Keep BF16 plus QLoRA or another adapter method separate from inference-only quantisation decisions. Training memory includes model weights, activations, gradients and optimiser state.
  • Quantize the KV cache only after testing long-context quality and the runtime's implementation. It reduces per-request memory rather than the checkpoint's parameter count.

The winning format is the lowest precision that passes the organisation's quality, latency, throughput and operational tests. It is not automatically the format with the smallest download.

Start with the weight-memory floor

The first sizing calculation is simple:

weight memory = parameter count x bits per weight / 8

For a dense 70-billion-parameter model, the theoretical floors are:

| Weight representation | Theoretical weight memory | Relative to BF16 | | --- | ---: | ---: | | BF16 or FP16 | 140 GB | 100% | | FP8 or INT8 | 70 GB | 50% | | 4-bit | 35 GB | 25% |

These are decimal planning numbers, not promised GPU requirements. A working engine also needs quantisation scales or metadata, runtime workspace, CUDA graphs, allocator headroom and KV cache. Some model components may remain at higher precision. The checkpoint format and the in-memory engine representation may differ.

A nominal 35 GB model can therefore be uncomfortable on a 48 GB GPU once a useful context length and concurrency target are applied. Likewise, an FP8 70B checkpoint may fit by weight on an 80 GB GPU but leave too little cache for the service objective. The Llama 70B production cost guide shows why traffic assumptions matter as much as the model file.

Sparse mixture-of-experts models require another field: total parameters versus active parameters per token. Total parameters drive weight storage, while active parameters help explain compute per token. Do not size an MoE model from the active count alone.

What exactly can be quantized?

Weight-only quantisation

Weight-only methods store weights at lower precision but usually compute with higher-precision activations. Names such as W4A16 and W8A16 describe this split: 4-bit or 8-bit weights with 16-bit activations.

AWQ and GPTQ are established post-training approaches for low-bit weight compression. The AWQ paper uses activation statistics to identify important weight channels. GPTQ uses approximate second-order information to compensate quantisation error as weights are processed. Both names describe a method, not a universal performance level. Group size, calibration data, packing format, kernel and conversion implementation still matter.

Weight-only 4-bit is attractive when memory bandwidth and capacity are the main limits. It does not guarantee a fourfold end-to-end speed increase. The GPU may need to unpack or dequantize weights, and an inefficient kernel can erase the expected gain.

Weights and activations

Formats described as W8A8, FP8 or W4A4 can reduce both weight traffic and activation precision. They depend more heavily on hardware support, calibration and optimised matrix-multiplication kernels.

FP8 is especially relevant on recent data-centre GPUs because it retains a floating-point exponent. That does not make every FP8 conversion lossless. The exact FP8 format, scaling strategy, model architecture and runtime path determine the result.

NVIDIA TensorRT documents INT8, FP8, INT4 and NVFP4 quantisation schemes, while TensorRT-LLM exposes LLM-specific engine building and serving optimisations. Support changes by release and GPU generation, so a quote should pin the runtime, container, driver and engine configuration rather than saying only "supports FP8".

KV-cache quantisation

The KV cache holds attention state for active sequences. Its memory grows with model architecture, token count and concurrent requests. Quantizing it does not shrink the model's stored weights, but it can free substantial capacity in long-context or high-concurrency serving.

This is a separate quality and performance decision. Validate retrieval accuracy, instruction following, multi-turn consistency and long-context behaviour at the actual cache format. A short perplexity test does not represent every production conversation.

Optimiser state, gradients and adapters

Inference weight formats do not describe full training memory. Full-parameter training also carries gradients, optimiser states, saved activations and communication buffers.

Hugging Face documents bitsandbytes support for 8-bit inference and 4-bit QLoRA-style adapter training. Its 4-bit path can store a base model in NF4 or FP4 while calculating in a higher precision such as BF16 and training additional adapter parameters. That can make fine-tuning accessible on smaller systems, but it is not equivalent to updating every base-model parameter.

Precision options in practical terms

BF16 and FP16

BF16 or FP16 is the reference configuration for most open-weight LLMs. It offers broad framework compatibility and avoids the extra variable of post-training quantisation. BF16 has a wider exponent range than FP16, which is useful for numerical stability, although the model and runtime decide what is supported.

Use this baseline when:

  • the model already fits the available GPU topology;
  • quality regressions carry a high commercial or safety cost;
  • the team needs an unambiguous reference for evaluating lower precision;
  • the workload includes fine-tuning or research that changes frequently;
  • the serving engine lacks a mature low-precision path for the model.

The cost is memory. A 70B model generally moves into a multi-GPU class at BF16, while a 405B model becomes an eight-high-memory-GPU project. See the Llama 3.1 405B hardware guide for that larger example.

FP8

FP8 halves the theoretical weight footprint from BF16. On supported accelerators and engines, it can also improve throughput. It is a strong production candidate when a model provider or serving framework supplies a tested FP8 artefact or conversion path.

Check whether FP8 applies to weights only, weights and activations, or the KV cache as well. Also check which layers remain at BF16. Two models labelled FP8 may have different memory use and quality.

FP8 is not a reason to skip evaluation. Retain the BF16 engine and compare both under the same prompt set, sequence lengths, concurrency and decoding parameters.

INT8

INT8 can offer a similar headline weight footprint to FP8. Integer quantisation often relies on scales and careful handling of outliers. Hugging Face's bitsandbytes documentation describes an 8-bit path that keeps sensitive computations at higher precision rather than applying naive integer conversion everywhere.

INT8 remains useful where the selected runtime has mature kernels or where the deployment must cover older GPU generations. Verify the compatibility table for the actual accelerator and software release.

INT4, AWQ and GPTQ

Four-bit weights provide the largest common reduction without entering more experimental sub-4-bit territory. AWQ and GPTQ can preserve useful model quality, but the outcome is checkpoint- and workload-specific.

Use 4-bit when it changes the feasible hardware class, such as moving a 70B model from several GPUs to one high-memory GPU, or when a workstation needs to hold a larger local model. Do not assume the smallest GPU that can load the weights will deliver the required tokens per second.

The original GPTQ and AWQ results demonstrate that carefully designed post-training methods can outperform simple rounding. They do not prove that every community conversion of a newer model will match its BF16 parent on legal drafting, code generation, multilingual support or tool use.

NF4 and QLoRA

NF4 is a 4-bit data type designed for normally distributed weights and is commonly associated with QLoRA. It is useful when the objective is memory-efficient adapter training. The compute data type can still be BF16.

Treat a QLoRA workstation quote as an adapter-tuning system, not as a promise of full-parameter training. Dataset preparation, sequence length, activation checkpointing and rank all affect memory and runtime.

NVFP4 and emerging 4-bit floating point

Newer NVIDIA software and Blackwell-generation hardware introduce 4-bit floating-point paths such as NVFP4. These can change the throughput and memory case for supported models. They also make software qualification more important, because older GPUs or engine versions cannot be assumed to execute the same format efficiently.

Before specifying a platform around NVFP4, require a supported model path, reproducible conversion method and quality result. A future-looking data type in a hardware table is not yet an accepted application design.

Why a smaller model may not be faster

Quantisation reduces bytes moved, but end-to-end serving includes more than weight reads.

  • Kernel availability: A mature BF16 or FP8 kernel can beat a poorly supported 4-bit path.
  • Dequantisation overhead: Weight-only formats may unpack into a higher compute type during execution.
  • Batch and sequence shape: Prefill and decode stress different resources. A result at batch one may reverse at production concurrency.
  • Tensor parallel communication: Fewer bytes per weight do not remove collective communication between GPUs.
  • CPU and request handling: Tokenisation, scheduling and API overhead can dominate small-model latency.
  • Memory headroom: A model that barely fits can suffer from limited cache, low batch size or offload.
  • Model architecture: Dense and mixture-of-experts models have different compute and memory behaviour.

Benchmark on the final serving engine. A PyTorch notebook result is not a substitute for the intended vLLM, TensorRT-LLM, TGI or other production runtime.

Hardware implications beyond GPU memory

GPU topology

If quantisation lets the model fit on one GPU, it can remove tensor-parallel communication and simplify operations. If the service still uses several GPUs, inspect the scale-up fabric and PCIe layout. Aggregate memory is not automatically a single fast pool.

Choose the GPU count from accepted throughput and cache capacity, then check whether the server's PCIe lanes, NVLink or NVSwitch topology matches the parallelism strategy. The GPUMachines configurator is a useful starting point once the model, precision and traffic target are defined.

Host RAM

Provide enough system memory to stage the source checkpoint, build or load the engine, run the serving stack and retain operational headroom without swapping. Conversion workflows may briefly need both the source and quantized representations. The correct RAM figure therefore depends on where conversion occurs and whether model loading is sharded.

Local storage

Use enterprise NVMe for checkpoints, engine artefacts, container layers, logs and rapid restart. Capacity planning should include the original model, one or more quantized variants, temporary conversion output and rollback images. RAID protects service continuity, not model quality.

Network

Single-node quantisation can reduce network complexity by avoiding multi-node tensor parallelism. A distributed deployment still needs a fabric sized for its collective pattern, plus separate management and storage paths where appropriate. Quantisation does not repair an oversubscribed east-west network.

A production acceptance test

Do not approve a lower-precision model from a single generic benchmark. Build a repeatable comparison against the BF16 or FP16 baseline.

1. Freeze the artefacts. Record model repository, revision, file hashes, quantisation method, group size, calibration data and runtime version. 2. Use business prompts. Include representative retrieval, coding, reasoning, multilingual, structured-output and tool-calling tasks. 3. Test difficult cases. Add long documents, rare terminology, ambiguous instructions, refusal cases and prompts known to expose hallucination. 4. Hold decoding constant. Compare the same temperature, top-p, seed policy and maximum output length. 5. Measure service behaviour. Capture time to first token, inter-token latency, tokens per second, throughput and GPU memory at target concurrency. 6. Run long-context checks. Repeat at the planned context lengths and KV-cache precision. 7. Score quality blindly. Use deterministic checks where possible and human review where judgement is required. 8. Soak the system. Test sustained load, restarts, model reloads, memory fragmentation and failure recovery. 9. Define a rejection threshold. Decide the acceptable quality delta before seeing the low-precision result.

Keep the report with the engine artefact. Re-running the same model name after a framework upgrade may produce a different result.

Deployment decision matrix

| Workload | Sensible first test | Why | | --- | --- | --- | | Regulated or high-consequence inference | BF16 baseline, then validated FP8 or INT8 | Keeps a clear fidelity reference and measurable approval path | | High-throughput production serving | BF16 and FP8 side by side | Exposes whether lower precision improves real throughput on the target GPU | | Local 70B inference | 4-bit AWQ or GPTQ plus an 8-bit/BF16 reference | Four-bit may change the workstation class, but quality needs comparison | | Adapter fine-tuning | BF16 LoRA and 4-bit QLoRA tests | Separates training-memory savings from inference format | | Long-context service | Weight format plus independent KV-cache test | Cache can dominate memory at concurrency | | Rapid research iteration | BF16 where it fits | Broad support and fewer conversion variables |

How GPUMachines scopes a quantized LLM system

A useful infrastructure brief includes:

  • exact model and revision;
  • dense or mixture-of-experts architecture;
  • BF16 reference plus candidate quantisation method;
  • serving runtime and pinned software versions;
  • input and output token distributions;
  • simultaneous requests and latency objective;
  • accepted quality delta and evaluation set;
  • context window and KV-cache format;
  • redundancy, restart and growth requirements.

From that brief, GPU capacity, topology, host memory, NVMe, networking, power and cooling can be calculated as one system. Buying the minimum VRAM needed to load a compressed checkpoint leaves too much of the design unspecified.

FAQs

Does 4-bit quantisation reduce GPU memory by exactly 75 percent?

It reduces the theoretical weight bytes to one quarter of BF16. Total GPU memory also includes scales, metadata, higher-precision layers, runtime workspace and KV cache, so end-to-end savings are smaller and vary by engine.

Is FP8 always better than INT8?

No. FP8 has a floating-point exponent and is well suited to supported modern GPU paths, while INT8 can be mature and efficient in other stacks. Compare the exact model, GPU and runtime rather than ranking the labels in isolation.

Are AWQ and GPTQ interchangeable?

They are both post-training weight-quantisation approaches, but they use different methods and can have different kernels, packing formats and model support. Test the actual artefact that will be deployed.

Can a quantized model be fine-tuned?

Adapter methods such as QLoRA can train additional low-rank parameters while the base model is stored at 4-bit. That is not the same as full-parameter fine-tuning. Confirm what remains trainable and which compute precision is used.

Should the KV cache use the same precision as the weights?

Not necessarily. Weight and KV-cache quantisation solve different memory problems. Choose and validate them independently against context length, concurrency and quality.

How much spare VRAM should a production server have?

There is no universal percentage. Reserve measured capacity for the runtime, cache, peak request mix, fragmentation and operational events. Approve the system from a load test, not a weight-memory table.

Does quantisation reduce power consumption?

It can reduce GPU count or improve throughput per watt, but a running GPU may still operate near its configured power limit. Measure complete-system energy at the accepted throughput and quality point.

Sources

Source material was checked on 22 September 2026. Software support matrices change, so confirm the selected GPU, runtime and container release before purchase or deployment.

← Back to blog