One exabyte in a rack is a useful headline only after the buyer asks how much is raw flash, how much depends on data reduction, which workload produced the throughput and what remains available when a chassis is out of service.
WEKA announced its third-generation WEKApod appliances and NeuralMesh 6 on 21 July 2026. The hardware range now comprises WEKApod Nitro, Prime and Prime Max on WEKA-designed chassis. The software release adds native virtual multi-tenancy, a file-and-object protocol stack on the same NVMe-backed data, metadata-first replication, remote caching, Kubernetes operations and shared observability.
WEKA also reports a fully populated 56U rack with up to 1.1 EB of effective capacity, 10.2 TB/s and 210 million IOPS. The company's own note says the effective-capacity figure assumes 3:1 data reduction. These are vendor-published rack figures, not GPUMachines test results, and they do not describe every WEKApod model or every workload.
The launch is still relevant. It shifts WEKApod from software on selected third-party servers towards an appliance whose chassis, thermals, drive attachment and supply chain are under one vendor's control. For buyers, that can simplify accountability. It also makes it more important to understand raw capacity, hardware service, upgrade rights and the boundary between appliance support and the surrounding network.
Executive Summary
- What WEKA launched: third-generation Nitro, Prime and Prime Max appliances plus NeuralMesh 6 software.
- What matters in the hardware: WEKA says its own chassis can hold up to 70 NVMe drives or four servers in 2U, with designs for drive density, thermals and service access.
- What matters in the software: virtual and physical tenant models, shared file and S3 access to the same data, replication with on-demand hydration, data reduction, Kubernetes operations and multi-cluster monitoring.
- What the large numbers mean: 1.1 EB is an effective figure based on 3:1 reduction in a fully populated 56U configuration. Procurement should also compare raw usable capacity after protection and reserve space.
- Who should consider it: AI cloud providers, model builders and enterprises with sustained multi-GPU training or inference, many tenants, large checkpoint flows or a need to serve file and object workloads from one fast tier.
- When it is too much: small research teams, low-duty-cycle GPU estates and buyers without 400GbE-class storage networking may get a better result from fewer nodes, qualified software-defined storage or a lower-cost capacity tier.
GPUMachines can size compute, storage and network together for an AI cluster, using the buyer's model, checkpoint, dataset and inference traces.
What Was Announced
The launch has two linked parts. WEKApod 3 is the appliance family; NeuralMesh 6 is the software platform running on it. Buying one without examining the other would miss the point.
WEKA positions Nitro for performance-heavy AI factories and inference-as-a-service, Prime for enterprise AI with a balance of capacity and speed, and Prime Max for very high capacity. Public launch material says all three are available to order, with delivery beginning in autumn 2026, and that NeuralMesh 6 will ship with the first systems. Availability, region, component configuration and lead time still need confirmation in a quote.
NeuralMesh 6 adds several capabilities that matter beyond benchmark speed:
- virtual tenants with separate networks, identity, encryption and quality-of-service policies;
- composable clusters that reserve hardware resources for stronger physical separation;
- file and native S3 access to the same blocks;
- asynchronous replication, metadata-first browsing and remote caching;
- always-on data reduction with contractual performance terms;
- Kubernetes-native deployment and operations;
- observability across deployments.
That is an ambitious scope. A buyer should test each feature required for production rather than assume one platform release makes every data workflow simpler.
Read the Rack Numbers Carefully
The 1.1 EB, 10.2 TB/s and 210 million IOPS figures come from a fully populated 56U rack. WEKA states that the capacity calculation uses 3:1 data reduction. Effective capacity is therefore not the same as installed flash, usable capacity after protection or guaranteed capacity for data that does not reduce well.
The distinction is easy to miss:
| Capacity term | What it should mean in a proposal | Question to ask | |---|---|---| | Raw flash | Sum of installed drive nameplate capacity | Which drive models and endurance classes are included? | | Usable capacity | Space available after protection, metadata, system reserve and formatting | What remains after the stated fault policy? | | Effective capacity | Usable capacity multiplied by an assumed reduction ratio | Is the ratio measured on our data and contractually defined? | | Operational headroom | Free capacity kept for rebuild, performance and growth | At what fill level do performance and support guidance change? |
Compressed model archives, encrypted objects and pre-compressed media may reduce poorly. Text corpora, repeated checkpoints or some structured datasets may reduce more. The buyer's data sample decides whether 3:1 is conservative or unrealistic.
IOPS and throughput need similar context. A large-block sequential read test can show bandwidth but say little about millions of small files. A cached metadata test can show IOPS but not sustained checkpoint writes. Useful figures require block size, read/write ratio, queue depth, client count, dataset size, cache state, network topology, protection mode and fill level.
Ask WEKA or the supplying partner for test conditions and then run a workload trace. GPUMachines would not compare 10.2 TB/s with a competitor's number unless both results use compatible methods.
Hardware Ownership Changes the Support Question
Earlier WEKA deployments could run on qualified server platforms from system partners. WEKApod 3 moves the appliance line to custom WEKA-engineered hardware. WEKA says it designed chassis architecture, drive interconnect, thermal management and serviceability, and now controls component procurement for the appliance.
That can have real value. Storage software is sensitive to PCIe lane layout, NVMe thermals, NUMA placement, NIC attachment and firmware. A vendor that controls those variables can qualify a narrower matrix and hold one support boundary.
There is a counterweight. The appliance may reduce freedom to substitute components or upgrade outside the approved plan. Procurement should ask:
- Are drives, NICs, memory and power supplies standard field-replaceable units?
- Which parts can the customer replace without affecting support?
- Are capacity upgrades performed in place or by adding nodes?
- How are firmware bundles qualified and scheduled?
- What happens if a selected NAND generation becomes unavailable?
- Which support party owns a fault that crosses storage nodes, switches and GPU clients?
- Can an existing NeuralMesh software deployment absorb WEKApod 3 nodes, and under which version rules?
Supply-chain control is a benefit only when the commercial terms explain component continuity, substitutions and support life.
The Workload Determines Whether Density Helps
AI storage is not one workload. Training, inference, retrieval, checkpointing and model distribution put different pressure on a system.
Distributed training
Training clusters often stream large datasets, read many small samples and write large checkpoints from many ranks at once. Aggregate write bandwidth and metadata behaviour matter, but so do recovery and namespace operations. The acceptance test should include a checkpoint while data loaders are active, then restore that checkpoint after a node or path failure.
Model loading and replica launch
Inference fleets may need to start many replicas at once. That creates a read storm even when steady-state inference is light on storage. Measure time from scheduler request to ready model, not only storage throughput. Include container images, model shards and adapter files.
Retrieval and agent data
RAG and agent workloads can mix object ingestion, vector-index updates, small metadata changes and reads with unpredictable locality. File and S3 access to the same blocks may remove copies, but application consistency and namespace semantics need testing. A file becoming visible through S3 does not guarantee that every application expects the same update and listing behaviour.
KV-cache extension
WEKA discusses NVMe as an extended tier for persistent KV cache. This is latency-sensitive and different from durable model storage. Measure token throughput, time to first token, tail latency, cache hit rate, network use and write endurance. A high storage bandwidth result does not prove that a serving engine can use remote NVMe without decode stalls.
Research and HPC
Research groups often have mixed jobs, irregular schedules and users who create many small files. A high-density appliance can serve them well, but staff skill and cost per useful terabyte may matter more than maximum rack throughput. Open-source Lustre or Ceph, a smaller commercial system or local NVMe plus an archive tier may fit better.
The GPUMachines guide to choosing a parallel file system for AI clusters explains why workload shape should come before a vendor shortlist.
Native File and Object Access Needs Application Tests
NeuralMesh 6 presents the same physical data through POSIX or file protocols and S3. WEKA says this is a native implementation rather than a gateway and gives an example where data written through NFS or POSIX is immediately readable through S3.
Removing full copies can save capacity and shorten pipelines. A training process might write model output as files while a serving or distribution service reads it as objects. A data engineering team could ingest through S3 without staging a second dataset for POSIX clients.
Shared blocks do not remove questions about semantics:
- How are file renames and object keys mapped?
- What consistency does each protocol expose?
- How are permissions translated across POSIX, NFS, SMB and S3 identities?
- Which protocol owns retention, immutability and versioning policy?
- How do snapshots appear through each access method?
- What happens when two protocols update related data?
- Which client and SDK versions are qualified?
A proof of concept should run the exact writer and reader applications together. It should test failure and recovery during an update, not only a clean hand-off.
Multi-Tenancy Is More Than a Quota
WEKA says NeuralMesh 6 supports more than 1,000 virtual tenants per cluster, with tenants as small as 1 TB and provisioning in under 30 minutes. Its technical guide describes each tenant as an isolation and policy domain with capacity limits, protocol resources, encryption, identity and quality of service.
Virtualized RDMA Data Fabric network spaces can give tenants separate VLANs, IP ranges and even overlapping address spaces. Independent LDAP or Active Directory and key-management settings can separate identity and encryption. Composable clusters can assign dedicated CPU, memory and drives to anchor tenants that need physical resource separation.
These are WEKA design claims, and the distinction between logical and physical isolation should be clear in a service contract. An AI cloud buyer should test:
1. noisy-neighbour behaviour during read, write and metadata saturation; 2. per-tenant bandwidth and IOPS limits; 3. control-plane behaviour with hundreds of tenant objects; 4. address overlap and network-space enforcement; 5. identity failure, key rotation and tenant deletion; 6. monitoring and billing data per tenant; 7. upgrade impact across shared and composable clusters; 8. administrator boundaries and audit trails.
Provisioning time is useful, but day-two operations usually decide whether the model scales. The team needs repeatable API workflows for creation, policy changes, capacity growth, incident isolation and secure offboarding.
Metadata-First Replication Changes When Data Moves
NeuralMesh 6 introduces asynchronous replication and remote caching where metadata can arrive before all data blocks. WEKA says a destination can become browsable while content is hydrated on demand. That can help a workload follow available GPU capacity without waiting for a complete dataset copy.
The idea is attractive for cloud bursting and multi-site scheduling. It also changes failure and cost behaviour. A job may start quickly but read remote blocks across a WAN. Performance then depends on cache warmth, link capacity, latency, egress charges and the working set.
Before using the feature in production, define:
- which metadata is available before data hydration;
- the consistency point between source and destination;
- recovery point and recovery time objectives;
- cache eviction and pinning policy;
- WAN encryption and bandwidth controls;
- behaviour when the source or link is unavailable;
- whether a partially hydrated dataset can be promoted to an independent copy;
- how egress and inter-site traffic are measured.
Do not treat metadata-first browsing as completed disaster recovery. A recovery test must prove that required blocks, permissions, keys and application state are available when the source site cannot answer.
Network Sizing Can Make or Break the Appliance
A dense NVMe rack can only deliver its published aggregate rates when clients and switches can move the data. At 10.2 TB/s, the theoretical network equivalent is 81.6 Tb/s before protocol overhead and resilience. That does not mean every deployment needs that rate, but it shows why a pair of ordinary top-of-rack switches cannot expose a fully populated rack's headline bandwidth.
The design should count storage-node NICs, switch ports, leaf-spine uplinks, GPU client paths and failure capacity. Check whether the compute and storage fabrics share switches, and whether a checkpoint surge can affect collective traffic. A separate storage fabric can give cleaner fault and congestion boundaries; shared high-speed Ethernet can reduce hardware when QoS and capacity are well engineered.
WEKA's previous qualified designs have used high-speed NVIDIA ConnectX adapters and 400GbE or NDR connectivity, but a WEKApod 3 quote should state the current NIC and network options. Do not carry an older datasheet specification into a new appliance order.
GPUMachines can combine the storage design with Ethernet cluster networking, GPU rail mapping and management networks. The complete path from client memory to storage media should be reviewed, including NUMA placement, RDMA, MTU, routing and cable reach.
Failure Domains, Rebuilds and Service Access
Density saves rack space, but it places more capacity and bandwidth behind each chassis boundary. WEKA's product page says a 2U chassis can contain up to 70 NVMe drives or four servers. Buyers should learn exactly which components share power, cooling, backplane or management dependencies.
Request a failure-domain diagram covering drives, servers, chassis, power supplies, NICs, switches and racks. Ask how protection stripes are placed and how many simultaneous failures are tolerated. Measure rebuild time at realistic fill level while foreground I/O continues.
Serviceability deserves a physical review. Can a failed drive be identified and replaced without disturbing adjacent cables? Does a server sled leave enough space for fibre bends? Are power and network paths labelled consistently? What lifting or rail procedure is required? How much capacity is lost while a chassis is removed?
The most useful availability figure is not an abstract percentage. It is the expected workload effect and repair process for each component failure.
Power, Cooling and Rack Density
WEKA says the new chassis includes thermal-aware power control and is designed for sustained data-centre conditions. Exact rack power and cooling values depend on the selected Nitro, Prime or Prime Max configuration and should come from the current site-planning guide.
High storage density may free rack positions for GPU compute, but the room still has to reject the heat. Check rack input power, feed redundancy, airflow direction, inlet-temperature range, floor loading and hot-aisle capacity. A 56U figure also assumes a rack format and service arrangement that may differ from the site's standard 42U or 48U design.
Power efficiency should be compared per delivered workload and usable capacity, not per chassis. Data reduction can improve effective capacity per watt when it works on the data, while heavy reduction processing could affect CPU use. Both need measurement.
Appliance, Software-Defined or Open Source
WEKApod offers one-vendor hardware and software accountability. NeuralMesh can also run in other deployment forms, subject to qualification. Buyers should compare these with DDN, VAST, Lustre, Ceph and other platforms using the same test pack.
An appliance suits teams that value a defined support boundary and predictable node design. Software on qualified hardware may give more purchasing flexibility. Open-source storage can lower licence cost and offer control, but the organisation accepts integration, tuning, upgrade and incident ownership.
Our WEKA versus Ceph comparison covers the difference between a commercial parallel platform and an operator-built object/file estate. The Lustre versus WEKA guide looks at training and HPC patterns. Neither choice is settled by a vendor IOPS figure.
Who Should Consider WEKApod 3
AI cloud operators with many tenants are the clearest audience. Native tenant networks, policy domains and fast provisioning could reduce the number of storage clusters they manage, provided isolation and billing tests pass.
Large model builders may value dense checkpoint bandwidth, fast model distribution and a shared file/object namespace. Enterprises with several GPU teams can also benefit when one platform replaces separate training, S3 and collaboration copies.
The system makes the most sense where GPU demand is steady, the network is fast enough and the operations team can use the density. Buyers should have enough data to test reduction and enough clients to exercise the intended scale.
Who Should Not Buy It
A small laboratory with four or eight GPUs may not keep a WEKApod cluster busy. Local NVMe plus backed-up shared storage, a smaller NAS or a modest Ceph system may be easier to justify.
Teams with mostly cold datasets should not pay flash prices for every byte. A fast working tier backed by object storage or archive media can cost less. Buyers whose data is encrypted or already compressed should be cautious about effective-capacity claims.
Organisations without high-speed network skills should budget for integration or choose a supported hosted design. A fast storage rack connected through an undersized fabric will disappoint. Buyers that require delivery before autumn 2026 should also verify availability rather than treating the announcement date as a shipping date.
A Better Proof of Concept
Do not begin with a single benchmark file. Build a test from production traces:
1. Inventory dataset size, file and object counts, size distribution and daily change rate. 2. Capture concurrent reads, checkpoint writes, model launches and metadata operations. 3. Test at expected and near-full capacity, including realistic reduction ratios. 4. Remove a drive, server, link and switch path while workloads continue. 5. Create enough tenants to exercise control-plane and monitoring workflows. 6. Test file-to-S3 hand-offs with the real applications and security model. 7. Replicate to a remote site, interrupt the WAN and prove recovery. 8. Record storage, network, GPU idle time, application completion and operator effort.
Commercial acceptance terms can then refer to observed workload metrics rather than generic peak rates. Keep raw results and configuration files so future software or firmware releases can be compared with the same baseline.
Our Technical View
WEKApod 3 is most interesting because WEKA has taken ownership of the hardware variables around NeuralMesh. Dense NVMe systems are sensitive to lane layout, thermals and service design, so a chassis built for the software can make qualification and fault ownership clearer.
NeuralMesh 6 also addresses real operating problems: tenant separation, duplicate file/object copies and moving data towards available GPUs. The strongest feature for one buyer may be multi-tenancy; for another it may be model distribution or checkpoint recovery. Treating every feature as equally valuable would inflate the design.
The weak point in the launch material is the ease with which effective capacity and aggregate rack performance can be mistaken for guaranteed application behaviour. The 3:1 reduction assumption must be tested, and 10.2 TB/s needs a network and client estate capable of driving it.
GPUMachines would compare WEKApod 3 with commercial and open-source alternatives using raw usable capacity, measured workload completion, failure recovery, staffing and five-year expansion. We would also price the switches, optics and rack power needed to reach the selected performance target. The appliance is a strong candidate for dense shared AI infrastructure, not a default answer for every GPU server.
FAQ
Is 1.1 EB the raw capacity of one WEKApod rack?
No. WEKA describes it as effective capacity for a fully populated 56U configuration and states that the calculation assumes 3:1 data reduction. Ask for raw and usable figures for the quoted configuration.
What is the difference between WEKApod 3 and NeuralMesh 6?
WEKApod 3 is the appliance hardware family. NeuralMesh 6 is the software platform that supplies the distributed storage, protocols, tenant controls, replication, reduction and operations features.
Can the same data be accessed through files and S3?
WEKA says NeuralMesh 6 exposes the same physical blocks through file protocols and native S3. Application consistency, permissions, naming and update behaviour should be tested with the intended clients.
Does every deployment support more than 1,000 tenants?
WEKA publishes support for more than 1,000 virtual tenants per cluster. Practical density depends on capacity, workload, policies and control-plane use. A service provider should validate its tenant profile.
Is WEKApod 3 only for inference?
No. WEKA positions NeuralMesh 6 for training, inference and accelerated computing. The new appliance messaging focuses heavily on inference economics, but suitability still follows the workload trace.
How much network bandwidth do we need?
Enough for the required client throughput after allowing for protocol overhead and failures. Size from measured workload demand, node NICs and switch topology rather than the maximum rack figure alone.
Can WEKApod replace a capacity object store?
It can provide S3 access and dense flash capacity, but an inexpensive object or archive tier may still be appropriate for cold data. Compare durability, retention, replication, cost and recovery requirements.
Can GPUMachines source and integrate WEKApod-class storage?
GPUMachines can review commercial AI storage options, source supported systems subject to vendor and regional availability, and design the surrounding GPU, network, rack, power and hosting plan. Exact WEKApod 3 configuration and delivery require a current quote.
Verdict
WEKApod 3 and NeuralMesh 6 combine a denser appliance with software aimed at shared production AI. The package deserves attention from AI clouds and large GPU estates that need tenant controls, high checkpoint or model-loading rates and fewer copies between file and object workflows.
The ideal buyer has production traces, reducibility samples and a high-speed network plan. The strongest reason to choose the platform is one accountable system spanning hardware, distributed storage and tenant operations. A smaller, colder or lightly used estate should compare lower-cost commercial and open-source designs before committing to an exabyte-class flash story.
Ask GPUMachines to compare WEKApod, DDN, VAST, Lustre, Ceph and other AI storage paths for your cluster, using the same capacity, workload, failure and network assumptions.
Sources and Further Reading
- WEKA: WEKApod 3 announcement (21 July 2026). Primary source for the appliance family, availability statement and hardware positioning.
- WEKA: NeuralMesh 6 announcement (21 July 2026). Primary source for multi-tenancy, file and object access, replication, reduction and operations claims.
- WEKA: The capability stack behind exabyte-scale AI. Primary source for the 56U rack figures and the stated 3:1 data-reduction assumption.
- WEKA: NeuralMesh multitenancy technical guide. Primary technical source for tenant, network, identity and policy design.
- WEKA: WEKApod product page. Primary source for the Nitro, Prime and Prime Max positioning and chassis-density statements.
Capacity, performance, tenant-scale and application results cited from WEKA are vendor-reported. GPUMachines has not independently benchmarked WEKApod 3 or NeuralMesh 6 for this article.
