GPU memory is the most expensive resource in AI infrastructure, and recomputation is destroying it. Weka's NeuralMesh 6 platform, launching alongside its first self-designed hardware line Wekapod 3, attacks this directly with Augmented Memory Grid: NAND flash aggregated to behave like GPU memory, caching 100% of pre-calculated tokens so models never reprocess prior context. In multi-turn sessions, that math compounds fast. Weka CEO Liran Zvibel puts it plainly: 20 conversational turns without caching means recalculating prefill attention 400 times over.

NeuralMesh 6 adds four concrete capabilities beyond KV caching. Composable and virtual multi-tenancy supports up to 50,000 tenants per cluster with provisioning under 30 minutes. Unified file and object storage eliminates the translation layer between file-based and S3-based paths, serving the same physical data through either interface simultaneously, a claim Weka is pitching to GPU cloud providers including Lambda, CoreWeave, and Nebius at roughly two orders of magnitude better performance than conventional S3. Metadata-first replication lets destination environments become browsable before data fully transfers, cutting migration time from weeks to under an hour. AlloyFlash mixes TLC and QLC NAND within a single cluster, routing latency-sensitive work to faster TLC and bulk storage to cheaper QLC automatically.

The competitive picture is blunt: Dell, NetApp, Pure Storage, and VAST have all repositioned toward AI infrastructure in the past 18 months. NAND Research analyst Steve McDowell draws a line between repositioned legacy vendors and companies like Weka and VAST that were built for this from the start, calling Augmented Memory Grid the most technically capable KV cache implementation currently on the market. He also flagged something most buyers are missing: Weka backs its data reduction claims with contractual guarantees. Read the full piece for McDowell's specific advice on how to evaluate competing vendor claims, and why the inability to cite real-world customers at scale should be a hard stop.

[READ ORIGINAL →]