GPUmachines

Enterprise RAG Hardware Requirements: GPU, RAM and Storage

Size enterprise RAG across ingestion, embeddings, reranking, vector search and LLM generation, with practical GPU, RAM, NVMe and network guidance.

Enterprise RAG Hardware Requirements: GPU, RAM and Storage

An enterprise RAG platform is not one LLM running beside a vector database. It is a chain of services with different resource profiles: document extraction, chunking, embedding, indexing, query embedding, retrieval, reranking and answer generation. Multimodal documents can add OCR, table extraction, image understanding and visual reasoning.

The hardware plan should therefore start with two workloads. The ingestion path turns source material into governed, searchable data. The online path retrieves evidence and generates answers under a latency and concurrency target. They may share GPUs during a pilot, but production systems often scale them independently.

NVIDIA's current Enterprise RAG reference architecture documents a four-GPU, 40-vCPU and 200-GiB baseline for its complete blueprint profile. That is a useful worked example, not a universal minimum. A text-only pilot with a small model can use less; a multimodal, highly concurrent service can need far more.

The short answer

  • A pilot can combine ingestion, retrieval and generation on one well-specified GPU workstation or server if jobs are scheduled rather than run at peak load together.
  • A production service should reserve GPU capacity for the LLM and treat extraction, embedding and reranking as measurable services, not free background tasks.
  • Host RAM and NVMe matter because raw documents, parsed output, model artefacts, vector indexes, metadata and replicas exist outside GPU memory.
  • Corpus size does not determine hardware by itself. Document churn, modality, chunk count, embedding dimension, concurrency, context length and latency targets change the answer.
  • Access control must be applied during retrieval. A fast system that retrieves a document the user is not allowed to see is not production ready.

The correct bill of materials comes from a representative corpus and query trace. It cannot be inferred from the phrase "enterprise RAG" alone.

Map the RAG pipeline before choosing GPUs

Document extraction and chunking

Text-native files can often be parsed on CPUs. Scanned pages, tables, charts, diagrams and mixed-layout PDFs may use OCR or vision models and can benefit from GPU acceleration. The ingestion rate is governed by pages per hour, not simply the number of files.

Record the percentage of scanned and multimodal pages, average page count, largest file, daily additions and reprocessing policy. A compliance archive that changes monthly is different from a support knowledge base receiving documents every minute.

NVIDIA's RAG Blueprint supports text, enhanced PDF extraction, OCR and multimodal retrieval. Its Enterprise reference profile assigns GPU resources to NeMo Retriever extraction services rather than hiding extraction inside the LLM allocation.

Embedding

The embedding model turns chunks and queries into vectors. Bulk ingestion needs throughput; online query embedding needs low latency. The model, input length and batch size determine accelerator demand.

Embedding can share a GPU in a small deployment, use a GPU partition where supported, or run as an independent service. Separating it makes utilisation and scaling easier to see. It also prevents a large ingestion run from stealing latency from live users.

Vector database and index

The vector store holds embeddings, source references, metadata and index structures. It may search on CPUs, GPUs or both. Index choice changes memory, build time, recall and latency.

Use this raw-vector calculation as a floor:

vector bytes = vector count x dimensions x bytes per element

One million 2,048-dimensional FP32 vectors occupy about 8.2 GB before index structures, metadata, deleted records, write headroom, replicas or backups. The same corpus may produce very different vector counts under different chunking policies.

NVIDIA's documented one-million-vector Milvus example uses GPU resources for indexing, CPUs for search, 32 GiB of host memory across the vector-database services and substantial object/block storage. It is a reference configuration for a stated schema, not a capacity formula for every database.

Milvus recommends NVMe for production storage and notes that RAM depends on data volume. Its documentation also warns that etcd is sensitive to disk latency. Database metadata and control-plane storage deserve the same care as the vector files.

Reranking

A reranker scores the retrieved candidates before context is sent to the LLM. It often improves relevance, but adds another model invocation and another latency stage. The cost depends on top-K, chunk length, reranker model and query rate.

Treat reranking as an optional, benchmarked service. Compare answer quality and p95 latency with it enabled and disabled. Increasing top-K without evidence can raise both compute cost and noise.

LLM generation

The generator usually has the largest GPU-memory requirement. Its load is driven by model size, precision, input tokens, output tokens, concurrent sequences and the service-level objective.

For dense models, the theoretical weight floor is:

parameter count x bits per weight / 8

An 8B model starts near 16 GB at BF16, 8 GB at FP8/INT8 and 4 GB at 4-bit weights. A 70B model starts near 140 GB, 70 GB and 35 GB respectively. Runtime workspace and KV cache come on top. The quantized versus full-precision guide explains why the smallest checkpoint is not always the fastest deployment.

Inputs needed for an accurate hardware quote

Corpus and ingestion

  • total source data and document count;
  • average and maximum page count;
  • text, scan, table, chart, image and audio proportions;
  • new and changed documents per day;
  • required time from upload to searchable;
  • retention, deletion and re-indexing policy;
  • chunk size, overlap and expected chunks per document;
  • embedding dimensions and element type;
  • replication and backup requirements.

Online retrieval

  • peak and sustained queries per second;
  • concurrent users and concurrent model requests;
  • average and p95 input and output tokens;
  • top-K before and after reranking;
  • number of collections or tenants searched;
  • p50, p95 and p99 latency targets;
  • target retrieval recall and answer-quality threshold.

Governance

  • identity provider and group model;
  • document-level or row-level entitlements;
  • data residency and tenant boundaries;
  • audit-log retention;
  • encryption and key ownership;
  • availability, recovery-point and recovery-time objectives.

Without these fields, GPU quantity is a guess. A product demonstration can tolerate that; a production quotation cannot.

Three practical deployment profiles

Profile 1: text-only pilot

A pilot for a small internal team can run a compact generator, embedding model and vector database on one system. Schedule initial indexing outside user-test windows and keep the unquantized or higher-precision generator available as a quality baseline.

A sensible design has one GPU with enough memory for the chosen generator plus cache, ample system RAM, mirrored enterprise NVMe and a current server CPU. The exact GPU cannot be selected until the model and context are known. A 7B or 8B model is a different class from 70B.

The pilot must still test permissions, citations and document deletion. These are architecture questions, not features to bolt on after hardware approval.

Profile 2: departmental production RAG

For a live service, separate the online path from bursty ingestion. One or more GPUs can host generation, while embedding, reranking and extraction receive dedicated GPU capacity or scheduled partitions. Run the vector database and control-plane services on protected CPU/RAM resources with enterprise NVMe.

Provide redundancy for the generator and database if the service has an uptime objective. Two smaller replicas may be operationally better than one fully utilised server, even when the single server wins on headline throughput.

This profile suits support, engineering, policy or research assistants where hundreds of users share a governed corpus. Size it from a recorded query trace and a full-corpus ingestion run.

Profile 3: multimodal or organisation-wide platform

Large deployments split extraction, embedding, index building, query nodes, rerankers and LLM replicas into independently scalable services. Kubernetes can help operate those services, but it does not reduce their capacity requirements.

NVIDIA's current Enterprise RAG guide uses a complete baseline of four GPUs, 40 vCPUs and 200 GiB RAM, then describes scale-out to much larger clusters. Its cited profile includes a Nemotron generator, embedding and reranking NIMs, multimodal extraction services and Milvus. Changing those components changes the requirement.

At this scale, plan separate storage and management networks, failure domains, observability, model and index versioning, rolling updates and capacity for re-indexing without stopping queries.

CPU, RAM and storage requirements

CPU

CPU work includes API handling, parsing, tokenisation, metadata filtering, database search, compression, orchestration and background jobs. A GPU-rich server with too few CPU cores can leave accelerators waiting.

Measure CPU saturation during concurrent retrieval and ingestion. Check vector-database instruction-set requirements; Milvus lists SIMD extensions for similarity search and index building. Core count should follow service density and replica count rather than a fixed GPU-to-CPU ratio.

System memory

RAM may hold database indexes, caches, parsed documents, queues and model-loading buffers. It also provides recovery headroom when a service restarts or an index is rebuilt.

Plan from measured resident sets at the target corpus size. Include replicas and operating margin. NVIDIA's complete baseline uses 200 GiB, while Milvus recommends 128 GB for a generic cluster deployment. Those figures describe different scopes and should not be mixed into a false universal minimum.

NVMe and durable storage

Use enterprise NVMe for active indexes, database journals, model caches and temporary extraction output. Use durable object storage for originals, parsed artefacts, model versions and backups according to the recovery design.

NVIDIA advises roughly 200 GB of free disk merely for self-hosted RAG Blueprint model downloads and caching. That is software staging space, not corpus capacity. Add source data, extracted text, vectors, index overhead, replicas, logs, snapshots and re-indexing headroom.

Separate capacity from endurance. Continuous ingestion, compaction and index rebuilding can create sustained writes. Mirror or protect local storage where a disk failure would interrupt the service.

Network and scale-out design

RAG is often less fabric-intensive than distributed model training, but service composition creates steady east-west traffic. Documents move through extraction and embedding; queries move through the vector database, reranker and generator; telemetry and checkpoints use other paths.

Use at least three logical planes in a serious deployment:

  • service/data traffic between RAG components;
  • storage traffic for documents, models, indexes and backups;
  • management traffic for orchestration, monitoring and administration.

Multi-node tensor parallelism for a large generator can introduce a much more demanding GPU fabric. If the LLM fits within one NVLink/NVSwitch domain, the rest of the RAG platform may use conventional high-speed Ethernet. If generation spans nodes, design that collective network separately from ordinary API traffic.

Security is part of the retrieval path

RAG does not make private data safe by default. The vector store, caches, prompts and logs can all expose source material.

  • Attach source identity, tenant and access metadata at ingestion.
  • Filter candidates using the requesting user's current permissions before context reaches the model.
  • Preserve citations so an answer can be traced to authorised evidence.
  • Test revoked access and document deletion through every cache and replica.
  • Separate tenants where policy requires it; do not rely on prompt instructions for isolation.
  • Protect model endpoints, service credentials, object storage and database backups.
  • Record query, retrieval and policy decisions without logging more sensitive text than necessary.

A private deployment can reduce external data movement, but it still needs identity, secrets, patching, audit and incident-response controls. The private AI cloud guide covers the wider platform boundary.

Benchmark the complete pipeline

GPU utilisation alone cannot tell whether RAG works. Validate each stage and the user-visible result.

Retrieval quality

  • recall at K for a labelled query set;
  • reranker lift over first-stage retrieval;
  • metadata-filter correctness;
  • citation precision and source coverage;
  • performance on tables, scans and diagrams where relevant.

Answer quality

  • groundedness in retrieved evidence;
  • factual correctness and completeness;
  • refusal when evidence is absent or access is denied;
  • consistency across paraphrased questions;
  • task-specific review by domain owners.

System performance

  • pages or documents ingested per hour;
  • embedding and index-build throughput;
  • vector-search and reranking latency;
  • time to first token and inter-token latency;
  • end-to-end p50, p95 and p99 latency;
  • concurrent requests at the accepted quality setting;
  • recovery time after service, node and storage failures.

Keep the model, prompt, chunking, embedding, index and reranker versions with every result. A hardware comparison is meaningless if the retrieval pipeline changes between tests.

How GPUMachines scopes an enterprise RAG platform

GPUMachines can translate a corpus sample, model choice and service objective into a server or cluster design. That design should cover GPU memory and topology, CPU cores, system RAM, enterprise NVMe, vector-database placement, network interfaces, rack power and resilience.

Start with a representative document set and query pack. The resulting benchmark tells us whether one GPU server is sufficient, whether ingestion needs separate capacity and where redundancy produces more value than a larger single node. Use the server configurator after the workload profile is defined, not as a substitute for it.

FAQs

How many GPUs does enterprise RAG need?

There is no fixed number. A small text-only pilot may share one GPU, while NVIDIA's current full Enterprise RAG reference baseline allocates four GPUs across generation, embedding, reranking, extraction and indexing services. Model size, modality, ingestion rate and concurrency decide the result.

Does the vector database need a GPU?

Not always. CPU search can be appropriate, while a GPU can accelerate index construction or high-throughput search. Benchmark the chosen database, index and corpus at the required recall and latency.

Is a larger LLM always better for RAG?

No. Retrieval quality, chunking, reranking and source governance can dominate the outcome. Compare candidate generators on a labelled enterprise query set before accepting their extra GPU cost.

How much RAM does the vector database need?

It depends on vector count, dimensions, data type, index, metadata, replicas and cache policy. Calculate raw vectors first, then measure the selected index with production-like data and operating headroom.

Can ingestion and inference share the same GPUs?

They can during a pilot or under a scheduler. Production services often separate them so bulk re-indexing cannot damage query latency. GPU partitioning can help where the hardware and software stack support it.

What is the biggest sizing mistake?

Choosing hardware from model weights alone. RAG also needs cache, extraction, embeddings, reranking, vector search, system RAM, durable storage, permissions, observability and enough spare capacity to rebuild or recover.

Sources

Source material was checked on 22 September 2026. Blueprint versions, model profiles and support matrices change, so verify the intended software release before final hardware approval.

← Back to blog