GPUmachines

Llama 70B Production Cost: A Transparent Sizing Model

Calculate Llama 70B production cost from model memory, KV cache, measured throughput, availability, power and productive service hours.

Llama 70B Production Cost: A Transparent Sizing Model

The cost of running Llama 70B in production isn't a GPU rental rate. It is the all-in hourly cost of a service divided by the useful tokens delivered inside its latency target. A server that looks cheap but produces 60 acceptable output tokens per second can cost more per million tokens than a dearer system producing 300.

This guide uses Meta Llama 3.3 70B Instruct as the reference model. Meta lists a 128K context window and Grouped-Query Attention. The method also works for other 70B-class dense models, but the KV-cache calculation, licence and supported precisions must be checked against the exact checkpoint.

The production answer in one paragraph

A quantised Llama 70B service can start on one 96 GB professional GPU when context and concurrency are controlled. BF16 weights need roughly 140 GB before runtime overhead, which points to a larger-memory accelerator or multiple GPUs. Production adds at least one spare or second replica, monitoring, storage, support and a deployment route. Measure prompt processing, decode throughput and P95 latency with real traffic, then calculate cost per million input and output tokens separately.

First calculate model memory

The simplest weight estimate is:

weight memory = parameter count x bits per weight / 8

For an approximate 70-billion-parameter model:

| Weight format | Theoretical weight floor | Practical meaning | | --- | ---: | --- | | BF16 or FP16 | about 140 GB | Needs more than a 141 GB card once runtime and cache are included | | FP8 or INT8 | about 70 GB | Can fit a 96 GB GPU with a useful but workload-dependent cache margin | | 4-bit | about 35 GB | Fits 48 GB by weight; context and concurrency can still make it tight |

These are floors, not recommended capacities. Quantisation scales, temporary buffers, CUDA graphs, allocator fragmentation and the serving engine consume more memory. Some formats also trade accuracy or kernel speed for capacity. Benchmark the exact quantised checkpoint rather than buying from the bit count.

KV cache can become larger than expected

Llama 3.3 70B supports long context through Grouped-Query Attention. A worked BF16 KV-cache estimate for the published 80-layer, eight-KV-head configuration is:

2 x layers x KV heads x head dimension x bytes x tokens

That works out to roughly 0.3125 MiB per token per active sequence before implementation differences. The resulting order of magnitude is:

| Context retained for one sequence | Approximate BF16 KV cache | | ---: | ---: | | 8,192 tokens | 2.5 GiB | | 32,768 tokens | 10 GiB | | 65,536 tokens | 20 GiB | | 131,072 tokens | 40 GiB |

Ten simultaneous 32K conversations would therefore imply about 100 GiB of BF16 KV capacity before paging, cache quantisation or engine optimisations. That does not mean every request holds the full limit, but it shows why a 70 GB weight file on a 96 GB card cannot promise unrestricted 128K context and high concurrency.

Measure the real prompt-length distribution. A support assistant may receive 2K prompts most of the day with occasional 20K documents; a code agent may keep much longer histories. One universal context limit wastes capacity or causes unpredictable failures.

Input and output tokens use the GPU differently

Prompt processing, often called prefill, handles many input tokens in parallel. Decode generates output one token at a time and is sensitive to memory bandwidth, batching and latency policy. Quoting one "tokens per second" number without naming the phase makes the cost model unreliable.

Track at least these measures:

  • input tokens processed per second at representative prompt lengths;
  • output tokens per second for one request and for the whole replica;
  • time to first token at P50 and P95;
  • request completion time at the agreed concurrency;
  • failed, cancelled and retried work.

An engine can maximise aggregate throughput by holding requests for a larger batch, yet make interactive users wait too long. Only tokens delivered inside the service objective belong in the economic numerator.

A hardware starting matrix

One 96 GB PCIe GPU

This is a practical entry point for FP8, INT8 or other tested quantised formats. It keeps the model on one GPU, avoids tensor-parallel communication and can suit a private API with moderate context. NVIDIA RTX PRO 6000 Blackwell Server Edition supplies 96 GB and fits supported rack servers; its workstation variants serve local development.

The constraint is cache room. A roughly 70 GB 8-bit checkpoint leaves much less than 26 GB after engine overhead, while a 4-bit checkpoint leaves more space but may change output quality and kernel behaviour.

Two 96 GB GPUs

Two cards can hold BF16 weights or give a quantised service more cache and batching room, provided the engine splits the model efficiently. The GPUs remain separate memory domains, and RTX PRO 6000 Blackwell does not turn them into one transparent pool. PCIe topology and tensor-parallel traffic must be tested.

This layout can also run two independent quantised replicas, one per GPU, which may give better availability and simpler performance than one model split across both cards. The correct choice depends on per-replica memory and the traffic pattern.

H200, B200 or B300

NVIDIA lists 141 GB for H200, 180 GB for B200 and 288 GB for B300. One B200 or B300 has enough nominal capacity for BF16 Llama 70B weights, although runtime and context still need a measured margin. H200's 141 GB is too close to the theoretical BF16 floor for a safe assumption, so use lower precision or more than one GPU.

Eight-GPU HGX systems make sense when one service needs high throughput, large batches or several models, and when the buyer also expects training or larger-model work. They are rarely the cheapest first platform for one lightly used Llama 70B endpoint.

Availability doubles more than the GPU line

A production service needs an answer for driver maintenance, model rollout and hardware failure. One server has no capacity while it reboots or reloads a model. Two replicas across separate hosts allow rolling maintenance and can preserve service after one failure.

That doesn't always mean a strict 100 per cent cost increase. The second replica can carry normal traffic, and both may run at useful utilisation. Yet a design that needs N+1 capacity must include the spare headroom in its cost per token. Quoting the fully loaded throughput of both hosts leaves no failover capacity.

Define the failure state:

  • one GPU process lost;
  • one complete server lost;
  • a rack power or top-of-rack switch failure;
  • a bad model release requiring rollback.

Then calculate the service objective while that event is active. High availability is a capacity reservation, not a checkbox.

Build the hourly cost

For owned hardware, spread capital cost over productive service hours rather than calendar time:

capital cost per productive hour = (purchase + deployment - residual value) / productive service hours

Productive hours should account for planned maintenance, commissioning and the utilisation the organisation can actually sustain. A three-year period contains 26,280 clock hours, but a server used only 30 per cent of the time has 7,884 productive hours before outages.

Add the other hourly costs:

all-in hourly cost = capital + support + hosting/facility + energy + operations + software

Include CPUs, RAM, NVMe, NICs, chassis, spare parts and deployment work in capital. Facility cost may include rack space, power reservation, cooling and remote hands. Operations covers people and tooling used to patch, monitor, secure and support the service.

For rented GPUs, use the effective bill rather than the headline GPU rate. Add attached storage, network egress, idle replicas, reserved capacity, support and any platform fee.

Energy calculation

Measure system draw at the wall or rack PDU during the replay test. GPU board power alone omits CPUs, DIMMs, drives, fans, NICs and PSU losses.

energy cost per hour = measured system kW x PUE x electricity price per kWh

If colocation pricing already includes facility overhead, don't multiply it by PUE again. For a customer site, PUE can represent cooling and electrical overhead when the facility team supplies a defensible figure.

Power often matters less than hardware depreciation for a small service, but it becomes material at high utilisation and large scale. It also affects how many replicas fit in a rack and whether the site can accept the system.

Calculate cost per million tokens

For output tokens:

cost per million output tokens = all-in hourly cost / (accepted output tokens per second x 3,600 / 1,000,000)

Suppose an entire service costs $12 per hour and meets its latency target at the following aggregate decode rates. These figures are hypothetical and are not a GPUMachines quotation:

| Accepted output throughput | Output tokens per hour | Infrastructure cost per million output tokens | | ---: | ---: | ---: | | 100 tokens/s | 360,000 | $33.33 | | 250 tokens/s | 900,000 | $13.33 | | 500 tokens/s | 1,800,000 | $6.67 |

The same server cost changes fivefold across that table. This is why utilisation, batching and the latency target deserve more attention than a small difference in purchase price.

Calculate input-token cost separately because prefill throughput differs from decode. If the service also performs retrieval, reranking, safety checks or tool calls, give those components their own cost lines rather than hiding them inside the LLM figure.

Charge for completed work, not generated waste

Long outputs, retries and abandoned requests consume capacity. A user may cancel after the model has generated thousands of tokens; a failed tool call may restart the reasoning chain. Gross token counters make the service look productive while the customer receives no result.

Record:

  • accepted billable input and output tokens;
  • internal or hidden reasoning tokens where the stack exposes them lawfully;
  • rejected, cancelled and retried requests;
  • cache-hit and cache-miss rates;
  • completed business operations, such as resolved tickets or finished code tasks.

Cost per completed request can be more useful than cost per token for agent workloads. Tokens remain valuable for capacity accounting, but they are not the business outcome.

Storage and model-loading costs

Llama 70B checkpoints occupy tens or hundreds of gigabytes depending on format. Keep a production version, rollback version and conversion working space on enterprise NVMe. Model loading should be measured, especially when replicas autoscale or restart after maintenance.

Shared object storage can hold the source artefact, while local NVMe provides a predictable warm-up path. If ten replicas pull a 70 GB checkpoint at once, the storage network sees a 700 GB burst before serving begins. That event belongs in the availability plan.

Logs, prompts and traces may cost little in capacity but create security and retention duties. Decide what can be stored, for how long, and who can inspect it before enabling verbose production logging.

Software and licence checks

Meta distributes Llama 3.3 under the Llama 3.3 Community License and an Acceptable Use Policy. A commercial deployment should review those terms, the intended languages and any product-specific obligations rather than assuming that "open weights" means unrestricted use.

Pin the serving engine, container, CUDA, driver and model revision. Upgrading one layer can change throughput or memory use. Keep a canary replica for new versions and retain a tested rollback image.

Engine choice also affects continuous batching, prefix caching, KV-cache datatype and tensor parallelism. vLLM, NVIDIA NIM and other stacks can produce different results on the same hardware. Buy from the measured configuration, not a framework name.

Owned, hosted or public cloud

Public cloud suits demand that is brief, uncertain or spread across regions. It also lets a team test several GPU types before committing capital. The cost problem appears when replicas remain online for steady demand and the organisation pays for idle capacity, storage and egress.

Owned hardware works when utilisation is stable and the team can operate it. Customer-site deployment adds facility work; Buy & Host keeps hardware dedicated while placing power, cooling and remote hands in a datacentre.

Compare the same service objective in every case. One public-cloud instance with no spare isn't equivalent to two owned replicas with failover. A bare server quote isn't equivalent to a managed endpoint with monitoring and on-call support.

A 14-day measurement plan

Before signing off a platform, run a controlled pilot:

1. Freeze the model, quantisation, engine and API settings. 2. Replay prompt and output-length distributions from the intended workload. 3. Increase concurrency until P95 time to first token or completion time reaches the agreed limit. 4. Run that load long enough to expose heat, power and memory drift. 5. Restart a replica and measure model load plus traffic recovery. 6. Remove one host or GPU process and verify the reduced-capacity service. 7. Export accepted tokens, failed requests, system draw and operator effort.

The result provides enough evidence to calculate a cost range and choose the next increment. Without it, a per-token figure is a sales assumption.

FAQ

Can Llama 70B run on one GPU?

Yes, when the GPU has enough memory for a tested quantised checkpoint, runtime and KV cache. A 96 GB GPU is a practical single-card target. BF16 weights alone need about 140 GB, so they require a larger GPU or model parallelism.

Is a 48 GB GPU enough?

A 4-bit checkpoint may fit, but context, concurrency and runtime overhead reduce the margin. It can work for evaluation or a restricted service; production needs a replay test and an availability plan.

How much KV cache does 128K context need?

A worked BF16 estimate for one full 128K sequence is about 40 GiB under the standard Llama 3.3 70B configuration. Engine implementation and cache datatype can change the result. Several concurrent long sequences multiply it.

What is a good cost per million tokens?

There is no honest universal number. It depends on input/output mix, latency, region, availability, precision and utilisation. Calculate it from accepted measured throughput and all-in hourly cost.

Should the service use HGX?

HGX fits high-throughput or larger-model estates where eight GPUs and NVLink/NVSwitch stay busy. One modest Llama 70B endpoint often starts more efficiently on a 96 GB PCIe GPU or a small replicated server pair.

Recommendation

Start with a 96 GB single-GPU test for a quantised Llama 70B service, then add a second failure domain before calling it production. Use measured prompt and decode performance to decide whether the next step is another replica, a two-GPU model-parallel server or an HGX node.

GPUMachines can configure a PCIe GPU server, compare HGX systems, or model hosted ownership through Buy & Host. Bring the exact checkpoint, traffic distribution and latency target; those numbers determine the bill.

Sources

← Back to blog