An AI data pipeline is the path from source data to a running model and back again. It includes ingestion, validation, transformation, training reads, checkpoints, model artefacts, inference inputs, vector indexes and retained outputs. Each stage creates a different storage and network pattern.
That distinction matters because a system can deliver high sequential bandwidth yet struggle with millions of small files, metadata lookups or simultaneous checkpoint writes. It can also feed one GPU quickly but collapse when 64 workers request shuffled batches at once.
The short answer
A useful AI data pipeline separates five questions:
1. How quickly can data be ingested and verified? 2. Can preprocessing prepare batches before the GPUs need them? 3. Can storage sustain the read pattern of every training or inference worker? 4. Can the system write and restore checkpoints inside the recovery target? 5. Can model artefacts, vector indexes, logs and retained datasets be governed without slowing the hot path?
Capacity is only one line in the design. Measure throughput, IOPS, metadata rate, tail latency, client concurrency and recovery time with the real dataset format.
For infrastructure options, review scale-out storage, storage servers and the GPU cluster configurator. The storage design should be tied to a workload acceptance test, not a generic GB/s claim.
The six data paths to map
Treating the pipeline as six paths makes sizing easier.
1. Ingestion
Source systems may deliver objects, database exports, event streams, images, video, documents or scientific files. Ingestion needs enough sustained write throughput for the arrival rate, but it also needs validation, checksums, schema checks and a quarantine path for bad data.
The raw copy should normally be immutable. Keep source identity, collection time, licence or consent state and checksum with the data. Reprocessing becomes safer when a transformed dataset can be traced to an exact raw version.
2. Curation and preprocessing
Filtering, decoding, tokenisation, augmentation and feature extraction can be more expensive than reading bytes. A CPU-bound decoder can leave GPUs idle even when the storage array is fast.
NVIDIA DALI represents preprocessing as a graph with CPU, mixed and GPU stages. It can overlap work, prefetch batches and move suitable image, video or audio operations to the GPU. It is not a universal fix: data format, operator support and available GPU headroom determine whether offload helps.
3. Training reads
Training usually reads shuffled samples through many workers. The pattern may be large sequential records, random object reads, small-file metadata traffic or a mixture. Distributed training multiplies concurrency because every rank needs its next batch on time.
MLPerf Storage treats accelerator utilisation as the outcome: storage succeeds when it keeps simulated training clients busy. That is a better procurement question than asking for the array's peak benchmark number.
4. Checkpoint writes and restores
Checkpoints are bursty and collective. Many ranks may write shards at once, then continue computing. A failed run later reads those shards, sometimes from different clients. The design must cover both directions.
Measure checkpoint pause time, background-write interference, restore time and behaviour during a node or storage failure. A platform that feeds training well can still miss its recovery objective.
5. Model and experiment artefacts
Weights, optimiser states, container images, datasets, logs, evaluation results and lineage records do not all belong on the fastest tier. They need versioning, retention, access control and a promotion process from experiment to production.
Object storage is often suitable for durable artefacts and datasets. Shared POSIX storage may be needed by frameworks and legacy tools. Local NVMe can stage hot data. These tiers should be explicit rather than accidental copies maintained by individual users.
6. Inference, vector and cache data
Production inference introduces prompt data, retrieved documents, embeddings, vector indexes, response logs and, in some architectures, externalised KV cache. Latency and random access may matter more than the sequential throughput used for training.
MLPerf Storage v3.0 includes separate vector database and KV cache workloads, which is a useful reminder that inference storage cannot be sized from a training-read test alone.
Start with workload measurements
Collect the following before choosing hardware:
- Total raw, curated and retained capacity.
- Daily ingest and transformation volume.
- Sample count and size distribution, including the median and 99th percentile.
- Compression and decode cost.
- Training workers, GPUs per worker and expected batch rate.
- Read bandwidth and IOPS per worker at steady state.
- Metadata operations per second for the actual namespace.
- Checkpoint size, interval, write window and required restore time.
- Number of concurrent experiments and inference services.
- Vector index and inference-cache read/write pattern.
- Dataset versioning, replication, backup and retention rules.
Average throughput hides stalls. Record the 95th and 99th percentile batch-fetch time and correlate it with GPU utilisation. A short storage pause can hold every rank at a collective barrier.
Training-read bandwidth arithmetic
A first estimate is:
required payload bandwidth = active workers x payload consumed per worker
Then add protocol, filesystem, re-read and concurrency headroom. Do not treat the headroom factor as universal; derive it from tests.
For example, eight workers that each consume 0.8 GB/s require 6.4 GB/s of payload. Applying 30 percent planning headroom gives 8.32 GB/s. This is an illustrative calculation, not a recommendation. Compression, augmentation, caching and uneven sample sizes can move the real number in either direction.
Run the test for long enough to exceed client RAM and page cache. MLPerf Storage rules require datasets larger than aggregate client memory specifically to avoid measuring a cache instead of the storage system.
Checkpoint arithmetic
Checkpoint sizing starts with the write window:
checkpoint bandwidth = checkpoint size / allowed write time
A 2 TB distributed checkpoint that must complete in 60 seconds requires about 33.3 GB/s of aggregate payload throughput. That excludes filesystem overhead, competing training reads, replication and metadata work.
Restores need their own target. If the same 2 TB checkpoint must be available to replacement workers in 90 seconds, the read path needs about 22.2 GB/s of payload. Test remapping after a client failure rather than restoring only to the original writers.
Also measure the application pause. An asynchronous checkpoint that reports a fast API return but saturates storage for the next ten minutes may still reduce useful training throughput.
Small files versus record formats
Millions of small files stress directory traversal, inode or object metadata, open/close operations and random reads. A storage system can advertise tens of GB/s while workers wait on namespace operations.
Record formats such as WebDataset tar archives, TFRecord or application-specific containers can reduce file-count pressure and enable larger reads. NVIDIA DALI includes readers for WebDataset, TFRecord, LMDB, NumPy and other formats.
Packaging is not free. It changes random access, repair, shuffling and update workflows. Benchmark archive size and index strategy with the framework used in production. Do not convert a dataset merely because larger files look better in a synthetic bandwidth test.
Local NVMe, shared storage and object storage
Most production designs use more than one tier.
Local NVMe
Local drives provide low latency and high bandwidth close to the GPU node. They work well for dataset staging, scratch space, temporary caches and checkpoint buffering. The trade-off is duplication and loss of locality when a job moves to another node.
Staging only helps if the scheduler knows where data is located, the copy completes before the job starts, and eviction does not remove another job's working set. Include staging time in job wait-time measurements.
Shared file storage
A scale-out filesystem gives many nodes a common namespace. It can simplify distributed training and checkpoint recovery, but metadata, client tuning and network paths become part of performance.
Test client counts and failure modes at the intended scale. One fast client does not prove that 32 clients can sustain shuffled reads while eight of them checkpoint.
Object storage
Object storage suits durable datasets, model artefacts and large-scale retention. It provides a different consistency, namespace and access model from POSIX files. Applications may read objects directly or stage them into a file or local tier.
MLPerf Storage v3.0 supports an S3 access layer for several workloads. That makes object-based tests more comparable, but a submission still needs to match the access pattern you plan to run.
CPU preprocessing and GPU preprocessing
PyTorch DataLoader can use multiple worker processes, pinned memory, prefetching and persistent workers. These settings need profiling. More workers consume CPU, memory, file descriptors and shared memory, and can increase contention.
NVIDIA DALI can place supported decode and augmentation stages on CPU, mixed or GPU execution paths. This may reduce a CPU bottleneck or overlap preparation with model compute. It can also consume GPU cycles and memory that the model needs.
Compare at least three states:
1. Baseline framework data loader with measured worker settings. 2. Optimised CPU loading with pinned memory and transfer overlap. 3. GPU-assisted preprocessing where the data type and operators are supported.
Report end-to-end samples or tokens per second, not preprocessing speed in isolation.
Where GPUDirect Storage fits
NVIDIA GPUDirect Storage provides a direct DMA path between compatible storage and GPU memory, avoiding a bounce buffer through CPU system memory. This can reduce CPU load, latency and duplicate copies.
It is an application and platform feature, not a label to add to any array. NVIDIA's design guide states that I/O must be a material bottleneck, transfers must target GPU memory, buffers must be pinned, and the application must use CUDA and cuFile APIs. The filesystem, driver, GPU, NIC or NVMe device and PCIe topology must also be supported.
GPUDirect Storage is most persuasive when profiling shows CPU-mediated I/O is the constraint. Validate the direct path, fallback behaviour and container configuration with NVIDIA's supplied tools before using it in a capacity claim.
Network sizing for storage clients
Network line rate is not storage payload rate. Encoding, protocol, congestion control and filesystem overhead reduce usable throughput. Paths may also be limited by PCIe placement or shared switch uplinks.
Convert link rate carefully:
400 Gb/s / 8 = 50 GB/s raw line-rate equivalent
That does not mean one 400GbE port will deliver 50 GB/s of application reads. The server, NIC, switch, optics, storage targets and software must sustain the same flow. Redundancy may reserve links for failover instead of adding payload bandwidth.
Separate storage, compute fabric and management traffic where contention or failure isolation requires it. If networks are converged, test checkpoint bursts, training collectives and storage reads together.
Three reference profiles
Single server or small workgroup
- Durable datasets and model artefacts on shared file or object storage.
- Local NVMe for active data, cache and temporary checkpoints.
- 25, 100 or 200GbE selected from measured demand, not server price.
- Framework-level loading tests before introducing a more complex direct-I/O path.
Multi-node training cluster
- Scale-out shared or object storage with enough clients and targets for aggregate demand.
- Local NVMe staging where repeated epochs justify the copy.
- A dedicated or engineered storage network sized alongside the compute fabric.
- Checkpoint tests during active training, including restore to replacement nodes.
- Namespace, scheduler and data-locality integration.
RAG and inference platform
- Object or file tier for source documents and model artefacts.
- Vector database storage sized for index build, updates and query latency.
- Fast local or shared cache where repeated model and embedding access warrants it.
- Retention controls for prompts, responses and retrieved content.
- Separate tests for online tail latency and offline ingestion.
Acceptance tests before purchase
A storage proof should use the intended clients, network and software. Include:
- Cold-cache and warm-cache training reads.
- The real file-size distribution and shuffle pattern.
- Concurrent readers at the planned GPU count.
- Checkpoint writes while training reads continue.
- Full and partial checkpoint restore after a client failure.
- Dataset staging and eviction at scheduler boundaries.
- Metadata-heavy directory or object listing.
- Network or storage-target failure and recovery.
- Capacity at a realistic fill level rather than an empty array.
- Monitoring that attributes stalls to client, network, storage or preprocessing.
Record GPU utilisation, batch-fetch latency, payload throughput, IOPS, metadata rate, CPU usage and checkpoint pause. A pass condition should be tied to completed training or inference work.
Common design mistakes
- Sizing only from usable capacity.
- Using a sequential benchmark for a small-file dataset.
- Testing one client and multiplying the result by node count.
- Allowing page cache to hide the storage path.
- Ignoring checkpoint bursts and restore time.
- Moving preprocessing to the GPU without accounting for model contention.
- Claiming GPUDirect Storage support without validating the full software and PCIe path.
- Treating a redundant pair of links as twice the guaranteed payload bandwidth.
- Keeping every dataset version on the most expensive tier.
- Omitting lineage, access control and deletion policy from the design.
Frequently asked questions
How much storage bandwidth does an AI cluster need?
Multiply measured payload consumption per active worker by concurrent workers, then add tested headroom. Checkpoint and restore requirements must be calculated separately because they create different, burstier paths.
Is local NVMe enough for training?
It can feed one node very well, but local capacity, staging time, duplication and job movement become operational constraints. Larger clusters usually combine local scratch with durable shared or object storage.
Does GPUDirect Storage remove the CPU completely?
No. It provides a direct data path between compatible storage and GPU memory, but the CPU still prepares I/O and the application must use the supported APIs. Platform qualification remains necessary.
Should datasets be packed into large files?
Packing can reduce metadata pressure and improve read efficiency. It may also complicate random access, shuffling and updates. Test a representative packed format with the actual framework before converting the full corpus.
How should checkpoint storage be sized?
Use total checkpoint size divided by the allowed write window, then test simultaneous writers and restore to replacement clients. Include replication, protocol overhead and competing traffic.
Can the compute and storage network be shared?
Yes, if the converged design is sized and tested for both traffic classes. Collective communication and checkpoint bursts can interfere, so quality of service, topology and failure isolation need validation.
A practical architecture decision
Build the pipeline from measurements rather than a storage product category. Keep raw and durable data on a governed tier, place hot working sets close to compute where reuse justifies staging, and size shared paths for concurrent readers plus checkpoint recovery.
GPUMachines can map those requirements to storage servers, scale-out storage, high-speed networking and GPU nodes. A useful proposal should state the dataset pattern, client count, checkpoint target and acceptance test alongside the hardware.
.jpg)