Weka launches NeuralMesh 6 to cache 100% of tokens and solve GPU memory bottlenecks in AI inference

2 Sources

Share

Weka unveiled NeuralMesh 6, its biggest software overhaul, alongside its first custom-designed hardware line WEKApod. The platform introduces augmented memory grid technology that caches 100% of pre-calculated tokens, preventing GPUs from wasting compute on redundant calculations. Benchmark tests show 10x higher token throughput and 10x more concurrent users compared to standard memory systems.

Weka Addresses Critical GPU Memory Bottlenecks with NeuralMesh 6

Weka has launched NeuralMesh 6, marking what the company describes as its most significant product refresh to date, alongside its first custom-designed hardware line targeting AI inference workloads

1

2

. The AI data storage platform arrives as enterprises shift focus from training AI models to running them in production environments, where GPU memory has become the most expensive and fastest-depleting resource. Long context windows and multi-turn conversations force AI models to repeatedly recompute information they've already processed, consuming GPU memory and compute that could otherwise serve additional users or generate new responses

1

.

Source: VentureBeat

Source: VentureBeat

Augmented Memory Grid Technology Eliminates Redundant Calculations

At the core of NeuralMesh 6 sits the augmented memory grid, an approach that aggregates NAND flash storage to behave like GPU memory at a fraction of the cost

1

. The technology extends GPU memory by ensuring that the critical key-value cache resides within NVme storage, bypassing the recalculation overheads that silently drain GPUs running AI inference workloads

2

. Every prompt triggers two stages: prefill, which calculates attention and is computationally expensive, and decode, which converts that calculation into output. The cost shows up hardest in multi-turn sessions like chat or coding, where each new turn re-triggers prefill for everything that came before it, unless that work has been cached. "If you have 10 turns, you may overcalculate 100 times because you're redoing all of them. If you have 20, you'll overcalculate 400 times," Weka co-founder and CEO Liran Zvibel told VentureBeat. "You can put two orders of magnitude more NAND than you could afford in shared memory, and we can cache 100% of the pre-calculated tokens, so you never need to redo it"

1

.

Benchmark Results Show Dramatic Performance Gains

Weka pointed to benchmark results demonstrating that NeuralMesh 6 with NVMe delivered 10 times higher token throughput and served 10 times more concurrent users on Oracle Cloud Infrastructure compared to standard dynamic random-access memory

2

. The potential payoff is straightforward: better GPU utilization of existing investments, lower inference costs, and faster deployment of new AI workloads without waiting months for additional GPU capacity

1

. "What we're seeing now with customers is they're chasing availability of compute, and once they get new allocation from anyone, they want to be able to grab it and start running right away," Zvibel said

1

.

Source: SiliconANGLE

Source: SiliconANGLE

Four Key Capabilities Target Enterprise AI Production Needs

NeuralMesh 6 adds four capabilities aimed directly at addressing gaps in competitive evaluations. Composable and virtual multi-tenancy gives anchor tenants full hardware-level isolation with dedicated CPU, memory, and storage, while virtual multi-tenancy runs through Weka's RDMA fabric, delivering network-level isolation that scales past 1,000 tenants per cluster with provisioning in under 30 minutes. A single cluster running 50 composable clusters can support up to 50,000 tenants

1

. Unified file and object storage eliminates the typical dual-path architecture where data effectively exists twice. The same physical data on disk is directly readable through either path at once, with no translation layer or second copy. Zvibel is targeting non-AWS GPU clouds specifically, naming Lambda, Nebius, G42, and CoreWeave, with what he described as roughly two orders of magnitude higher performance than conventional S3 and a capacity-based pricing model instead of per-API charges

1

. Metadata-first replication makes destination environments browsable before a full data copy arrives, with data hydrating only when accessed. "They had to wait for all of that to make it to the other side, and this takes days or weeks, in extreme cases a month," Zvibel said. "We now allow our customers to grab some allocation of new GPUs and get up and running within an hour"

1

. AlloyFlash mixes TLC and QLC NAND flash within a single cluster, automatically routing latency-sensitive work to TLC while running bulk-capacity workloads on QLC, cutting cost per terabyte without a performance penalty. Data reduction now runs by default rather than as an option

1

.

Custom-Designed WEKApod Hardware Maximizes AI Data Storage Efficiency

Weka is moving away from general-purpose servers and engineering its own platforms from the ground up with the WEKApod Nitro, WEKApod Prime, and WEKApod Prime Max AI appliances

2

. The new WEKApod systems are built with a PCIe Gen 6 internal fabric and feature a software-management thermal architecture to enhance both performance and capacity. The WEKApod Prime and WEKApod Prime Max arrays come with up to 245 terabytes of ultra-dense solid-state drives and utilize NeuralMesh's proprietary data reduction techniques to create as much as 1.1 exabytes of effective capacity in a single 56-unit rack. WEKApod Nitro is aimed at high-concurrency environments where storage bandwidth determines GPU utilization

2

. "Building our own hardware was not the original plan," Zvibel said. "Our customers' inference economics made it a necessity. We built WEKApod because the alternative was letting someone else's hardware define the limits of what our software could do"

2

.

Timing Aligns with Enterprise Shift to Agentic AI Workflows

According to Weka, the timing could not be better as enterprise AI reaches a critical inflection point with organizations shifting away from training AI models to running them in production inference environments

2

. The startup believes that AI training-focused storage infrastructures are unsuitable for the long-context reasoning, agentic AI workflows, and retrieval-augmented generation workloads that now dominate most organizations' AI efforts . Organizations also face major physical constraints as available space in physical data centers shrinks and energy grids struggle to keep up with power demands. Chief Product Officer Ajay Singh noted that most data centers running production AI workloads today were built chaotically, mixing and matching different platforms, chips, and networking technologies from multiple vendors. "NeuralMesh 6 delivers what they've actually needed all along: a single platform that handles the high-performance file layer and the high-capacity object layer on the same blocks, with native multitenancy, intelligent data mobility and always-on data efficiency built in from the start," Singh said

2

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved