GPUmachines

DDN vs WEKA for AI Storage: EXAScaler or NeuralMesh?

Compare DDN AI400X3 with EXAScaler and WEKA NeuralMesh across file architecture, GPU data paths, deployment, operations and proof-of-concept criteria.

DDN vs WEKA for AI Storage: EXAScaler or NeuralMesh?

A useful DDN versus WEKA comparison must name the products. DDN AI400X3 uses the EXAScaler parallel filesystem, DDN's enhanced Lustre distribution, for large AI and HPC file workloads. WEKA NeuralMesh is a software-defined, fully distributed parallel filesystem that can run on qualified dedicated servers, in supported clouds or in a converged mode alongside GPU compute.

DDN also sells Infinia, a separate data-intelligence platform aimed at multi-protocol, metadata-heavy, inference and RAG workflows. Infinia should not be treated as another name for EXAScaler. A buyer comparing large-scale training filesystems will usually evaluate AI400X3/EXAScaler against NeuralMesh. A buyer focused on object data, semantic metadata or inference memory tiers may need a different shortlist that includes Infinia.

For the file-system decision, DDN makes its strongest case through an integrated, NVIDIA-certified appliance and a mature Lustre operating model. WEKA makes its strongest case through a portable software architecture, a high-performance client path and deployment flexibility. The better platform depends on data shape, client model, operations and the system that will actually be quoted.

DDN EXAScaler vs WEKA NeuralMesh at a glance

| Decision area | DDN AI400X3 with EXAScaler | WEKA NeuralMesh | What to verify | | --- | --- | --- | --- | | Core design | DDN-integrated appliance running an enhanced Lustre parallel filesystem | Software-defined, distributed parallel filesystem built around containerised processes and NVMe | Whether the team values an appliance boundary or software portability | | Primary access | Parallel POSIX file access through Lustre clients; optional data services depend on the subscription | Native POSIX client plus documented NFS, SMB, S3 and Kubernetes CSI paths | Exact client OS, protocol and feature requirements | | Metadata model | Lustre metadata servers and targets, with DDN management and tuning | Distributed compute processes manage filesystem metadata without dedicated metadata servers | Small-file concurrency, directory hot spots and failover behaviour | | GPU data path | EXAScaler supports NVIDIA GPUDirect Storage in validated designs | NeuralMesh documents RDMA and GPUDirect Storage in supported environments | Driver, client, NIC and PCIe compatibility on the proposed GPU nodes | | Deployment | Dedicated DDN appliance, cloud-managed Lustre options and supported reference architectures | Dedicated, standard converged, NeuralMesh Axon and supported public-cloud deployments | Resource isolation, minimum size and support boundary | | Capacity media | AI400X3 all-flash family; EXAScaler also has hybrid product routes | NVMe performance tier with optional object tiering | Active working set, cold tier and recall behaviour | | Operations | EXAScaler Management Framework, DDN support and Lustre concepts | WEKA management, distributed services and selected infrastructure from qualified partners | Which team owns software, hardware, monitoring and upgrades | | Strong starting fit | Large training, HPC, checkpoints and established Lustre environments | AI pipelines needing fast POSIX, deployment choice and direct GPU data paths | Results from a matched workload proof of concept |

This comparison does not include every DDN or WEKA feature. Licensing, qualified hardware and release support must be checked against the current proposal.

DDN is two different AI storage conversations

DDN's current portfolio includes at least two architectures relevant to AI buyers.

AI400X3 and EXAScaler

AI400X3 is DDN's integrated storage platform for AI and HPC. It uses EXAScaler, which is based on the Lustre parallel filesystem and adds DDN management, appliance integration and product support. NVIDIA's current certified-storage list includes DDN AI400X3 and AI400X3i for file-based AI storage, including Enterprise and NVIDIA Cloud Partner validation levels.

This route is aimed at high-throughput shared files, large client counts, training datasets and checkpoint workloads. The system has the familiar Lustre separation between metadata services and object storage services. DDN supplies a defined hardware and software stack rather than asking the buyer to assemble a community Lustre environment.

Infinia

Infinia is a newer DDN software-defined platform built around a distributed key-value architecture and metadata-rich data services. DDN positions it for object, file and block access, RAG, inference, data discovery and distributed data workflows.

Those claims describe a different product direction from AI400X3/EXAScaler. If the requirement is a parallel training filesystem, compare EXAScaler. If the requirement centres on semantic metadata, object access or KV-cache movement, ask DDN to explain whether Infinia, EXAScaler or a combined design is being proposed. Do not let a benchmark from one product justify the other.

How EXAScaler handles the file path

EXAScaler inherits the central Lustre architecture. Metadata Servers work with Metadata Targets that hold names, permissions and file layouts. Object Storage Servers work with Object Storage Targets that contain file data. Clients combine those services into one POSIX namespace and can stripe large files across several OSTs.

The benefit is a file path designed for parallel access. A large checkpoint or dataset shard can use several targets at once, and more OSS capacity can increase aggregate throughput. Metadata and bulk file data have separate scaling paths.

The buyer still has to design:

  • Metadata target performance and high availability.
  • Object storage target layout and protection.
  • File stripe count and stripe size.
  • LNet interfaces and routes.
  • Client-kernel and module compatibility.
  • Recovery, target failover and upgrade procedures.

DDN's appliance and management framework can reduce integration work, but they do not remove workload-specific layout choices. A wide stripe suitable for a multi-terabyte checkpoint may be wasteful for small images or source files.

How NeuralMesh handles the file path

WEKA documents NeuralMesh as a software-only, container-native parallel filesystem. Its architecture divides work among front-end, compute, drive, management and telemetry processes. Compute processes handle data distribution, protection, clustering, metadata and tiering. Drive processes turn SSDs into networked storage resources.

The system does not use dedicated metadata servers in the Lustre sense. Data and metadata responsibilities are distributed across participating resources. WEKA's native client can use kernel-bypass modes, RDMA and NVIDIA GPUDirect Storage in supported environments.

NeuralMesh can run as a dedicated storage cluster or in converged forms. NeuralMesh Axon places storage services on GPU servers and uses their CPU, memory, local NVMe and network resources. That can reduce separate storage hardware and bring data close to compute. It also means the GPU node's non-GPU resources are shared deliberately with storage.

Convergence is not a free capacity pool. Reserve the required CPU cores, RAM, NVMe endurance and network bandwidth, then test the GPU workload while the storage layer rebuilds or rebalances.

Dedicated appliance or software-defined deployment

The operating boundary is a major difference.

A DDN AI400X3 design is supplied as an integrated product with a defined support path. That can suit organisations that want one vendor to own the storage appliance, EXAScaler software and validated scale-out configuration. Hardware expansion follows DDN's supported building blocks.

WEKA separates the software from a wider choice of qualified infrastructure. The deployment can use WEKApod appliances, partner servers, supported cloud instances or converged GPU-node resources. That gives the buyer more choices, but it also makes the exact server, drive, NIC and support arrangement important.

Ask both vendors the same questions:

1. Who replaces a failed drive, NIC or server? 2. Who owns the operating-system image and firmware matrix? 3. Which changes require vendor approval? 4. Can hardware generations be mixed during expansion? 5. What resources are reserved on every client or converged node? 6. What is the tested minimum configuration? 7. What happens to support when third-party switches or servers are used?

The answer may matter more than a small benchmark difference.

Metadata and small-file workloads

AI storage is not always a large-file problem. Image collections, source repositories, feature stores and unsharded datasets can generate millions of opens, stats and directory operations.

EXAScaler uses Lustre metadata services. DDN can size metadata targets and servers independently from bulk capacity. Distributed Namespace Environment can spread namespaces across several MDTs where the design and directory layout make use of it.

NeuralMesh distributes metadata work across its compute processes and uses hash-based data structures according to WEKA's architecture documentation. There is no separate pair of metadata servers to size, but metadata still consumes CPU, memory, network and storage resources across the cluster.

Test the actual namespace. Create, list, open and remove files at production concurrency. Measure the 95th and 99th percentile response time and observe what happens when a directory becomes hot. Vendor aggregate throughput figures do not answer this question.

Large files, checkpoints and training

DDN AI400X3/EXAScaler has a clear heritage in large AI and HPC file workloads. Striping allows a client or job to use several storage targets for one file, and the appliance route is designed around sustained parallel throughput. DDN's NVIDIA certifications make it a credible starting point for HGX, DGX and cloud-provider storage designs.

WEKA NeuralMesh also targets training and checkpoint workloads. The native client can distribute file I/O across the cluster, and supported RDMA and GDS paths can reduce host-copy overhead. Dedicated and converged deployments allow different ratios between storage and GPU nodes.

For both platforms, a fair training test includes:

  • Dataset reads from the planned number of workers.
  • Synchronous checkpoint bursts.
  • Background ingest or transformation.
  • Metadata scans and model-file opens.
  • A drive or node recovery event.

Record training step time and GPU utilisation as well as storage throughput. The goal is completed model work, not a storage-only headline.

GPUDirect Storage and RDMA

Both product families can participate in NVIDIA GPUDirect Storage designs. That does not mean every configuration uses the same path.

GDS depends on the GPU, NIC, storage client, PCIe topology, driver, CUDA release and application. RDMA capability on a data sheet is necessary in many designs but not sufficient. Confirm the exact compatibility matrix and capture the final node topology.

Run the same workload with and without the GDS path where possible. Record CPU utilisation, throughput, latency and GPU duty cycle. If the application uses small cached reads or spends little time moving data, GDS may not be the deciding factor.

Protocols and data movement

EXAScaler is first a parallel filesystem. Current DDN subscription material also describes S3, SMB and NFS data services, but the implementation and performance path differ from the native Lustre client. Confirm which services are included and whether they are production requirements or convenience gateways.

NeuralMesh documents native client POSIX, NFS, SMB and S3 access in its standard platform. WEKA's Axon converged mode has a more specialised feature set, so do not assume every protocol is available in every deployment model.

If the AI pipeline begins in object storage and trains through POSIX, test the full transition. Measure listing, metadata, ingest, recall and namespace consistency. A protocol checklist does not prove that one copy of data serves every workflow efficiently.

Capacity tiering and economics

An all-NVMe training tier can be expensive when only a small fraction of the dataset is active.

DDN offers all-flash AI400X3 configurations and separate hybrid EXAScaler routes. The proposed system should state which data lives on flash, which lives on capacity media, and how movement is managed.

WEKA can tier colder data to an object store while presenting it through the filesystem namespace. The active flash tier, object capacity and recall policy must be sized together. A recalled object still has the latency and throughput characteristics of the capacity tier and network path.

Compare usable protected capacity, active working-set size, media endurance and the cost of keeping GPUs waiting. Data-reduction or tiering estimates should be measured on representative data rather than copied from another customer.

Recovery and degraded-state behaviour

Normal-state bandwidth is only half of the service.

In an EXAScaler system, test metadata-service failover, OSS or target faults and backend storage recovery according to the supported design. Record client interruption, checkpoint behaviour and performance while protection is being restored.

In NeuralMesh, test a drive process, server failure domain or network path under vendor supervision. Observe data rebuild, process relocation and client behaviour. In converged mode, include the effect on the GPU workload because storage and compute share the same physical nodes.

Use the same business requirement for both products: maximum interruption, minimum throughput during recovery and time to restore protection. Vendor-specific fault names should map to that requirement.

Network design

DDN EXAScaler traffic follows the Lustre client, metadata and object-storage paths over LNet. WEKA native clients communicate with distributed storage processes and may use DPDK, UDP, RDMA or GDS paths depending on the deployment.

For either design, calculate:

  • Client-facing bandwidth per GPU node.
  • Storage-node uplinks and oversubscription.
  • Switch ports, optics and cable quantities.
  • Recovery and rebalance traffic.
  • Management and high-availability paths.
  • PCIe locality between GPUs, NICs and CPU roots.

A storage appliance with four fast ports cannot feed a cluster if the switch fabric or client adapters are undersized. Test one client, one rack and full planned concurrency.

Operations and support

DDN operators work with EXAScaler Management Framework, Lustre concepts, appliance health and DDN support. Existing HPC teams may already understand targets, LNet, layouts and service failover.

WEKA operators work with NeuralMesh processes, filesystems, resource allocation, tiering and the chosen server or cloud infrastructure. A converged deployment also requires coordination with the GPU platform team because storage consumes node resources.

Commercial support is not the same as operational outsourcing. The internal team still needs alert ownership, capacity forecasting, patch windows, data restore procedures and a recovery runbook.

When DDN EXAScaler is the stronger shortlist

Start with DDN AI400X3/EXAScaler when:

  • The primary requirement is a large, parallel POSIX filesystem.
  • Training, HPC and checkpoint throughput dominate.
  • The organisation prefers an integrated storage appliance and support boundary.
  • Lustre expertise already exists.
  • An NVIDIA-certified AI400X3 reference path matches the compute estate.

Do not select it from brand familiarity alone. Confirm metadata behaviour, protocol needs, target protection and expansion increments.

When WEKA NeuralMesh is the stronger shortlist

Start with WEKA NeuralMesh when:

  • A high-performance POSIX path is required across on-premises and cloud environments.
  • The team values software deployment choice or a converged GPU-node option.
  • RDMA and GPUDirect Storage are central to the tested client path.
  • NFS, SMB or S3 access is part of the same namespace requirement.
  • NVMe performance with object tiering matches the active and cold data model.

Do not choose convergence only to remove a storage rack. Measure the CPU, RAM, NVMe and network resources taken from each GPU node.

When DDN Infinia belongs in the discussion

Add Infinia when the requirement centres on RAG, semantic metadata, object access, distributed data discovery or inference data services rather than only a training filesystem. Ask DDN for a current architecture, supported protocols, production references and a proof of concept for the exact release.

Do not use AI400X3 benchmark results as evidence for Infinia, or Infinia metadata claims as evidence for EXAScaler. They are separate products with different data paths.

Proof-of-concept checklist

Use the same GPU clients, network, dataset and acceptance thresholds:

1. Measure one-client and full-cluster read/write throughput. 2. Test the production file-size distribution. 3. Run metadata create, open, stat, list and delete workloads. 4. Launch simultaneous checkpoints from the planned worker count. 5. Measure GPU utilisation and completed training steps. 6. Test object or gateway protocols if they are required. 7. Trigger a supported component failure and rerun the workload during recovery. 8. Add capacity or a node using the documented expansion process. 9. Restore files and validate snapshots or protection workflows. 10. Record complete rack power, switch ports, licences and staff effort.

Keep vendor engineers involved, but require the future operations team to perform monitoring, restore and fault exercises.

Common mistakes

  • Comparing DDN as one product with WEKA as one product.
  • Treating EXAScaler results as Infinia results.
  • Testing only large sequential files.
  • Ignoring client CPU and memory reserved by the storage stack.
  • Assuming GDS works because the NIC supports RDMA.
  • Comparing raw capacity instead of usable protected capacity.
  • Omitting rebuild traffic and failure tests.
  • Choosing a converged layout without measuring interference with GPU jobs.
  • Buying an appliance without checking expansion and lifecycle rules.

GPUMachines storage design

Start with the scale-out storage solution, compare current storage server platforms, and place the storage requirement beside the GPU cluster configurator.

GPUMachines can review the storage path with GPU nodes, east-west fabric, north-south traffic, local NVMe, rack power and deployment model. The deliverable should include a workload model, failure-domain map, network-port calculation, capacity plan and acceptance test.

FAQ

Is DDN faster than WEKA?

There is no universal result. DDN AI400X3/EXAScaler and WEKA NeuralMesh use different architectures and deployment models. Test the proposed systems with the same clients, network, dataset, protection and recovery state.

Is DDN EXAScaler based on Lustre?

Yes. EXAScaler is DDN's enhanced and supported Lustre-based parallel filesystem product. AI400X3 packages it in an integrated AI storage platform.

Is WEKA only an appliance?

No. NeuralMesh is software-defined and can run on qualified partner hardware, WEKApod systems, supported cloud instances or in converged modes. The selected route changes the support boundary.

Which is better for small files?

Both require workload testing. EXAScaler scales Lustre metadata services and targets; NeuralMesh distributes metadata work across its processes. Directory shape, client count, metadata media and recovery state matter.

Which is better for RAG and inference?

WEKA has specific inference and memory-tier features, while DDN positions Infinia for metadata-heavy RAG and inference workflows. Compare those current products rather than assuming the training filesystem comparison answers the question.

Are both NVIDIA-certified?

NVIDIA's current certified-storage list includes DDN AI400X3/AI400X3i and WEKA systems or qualified designs at several validation levels. Certification narrows deployment risk; it does not select a platform for a particular workload.

Technical sources

Vendor performance, efficiency and scale statements should be treated as claims until reproduced on the quoted configuration with representative data.

← Back to blog