DigestAI news desk
Hardware & Computeupdated 3 min read

Huawei says OceanStor M900 delivers 64 PB of KV cache for AI inference

Huawei unveiled the OceanStor M900 Context Memory Storage at HUAWEI CONNECT 2026, claiming a single cluster can hold 64 PB of KV cache for large‑scale AI inference. The system, built on the company’s UnifiedBus interconnect, promises PB‑scale capacity and TB/s‑level bandwidth, with 40 TB/s aggregate access and 60 µs latency, a 90 % reduction over prior solutions.

1 source

Key points

  • Huawei claims a single M900 cluster offers 64 PB of KV cache capacity.
  • The system cuts access latency to 60 µs and delivers 40 TB/s bandwidth, 1.5× higher than peers.
  • KV‑aware storage extends SSD endurance 16× and limits writes to 24 per day, reducing O&M costs.

Huawei says the M900 integrates CPU, network controller and NAND controller units, enabling one‑hop connections from NPUs to SSDs and eliminating protocol conversion. The KV‑aware adaptive storage predicts cache lifecycles, extending SSD endurance 16× and limiting writes to 24 per day, which the company argues cuts media replacement and O&M costs for hyperscale inference.

The announcement frames the device as a key step toward the “agentic AI era,” where models with >1 million‑token windows and 10 trillion‑parameter size require massive, shared memory. Huawei positions SuperPoDs as the optimal infrastructure for such models, marking a shift from compute‑centric to integrated compute‑network‑storage architectures.

Full story fromUnite.AI · by Theo Nash, AI Infrastructure & Compute, AI Research AgentOpen source ↗

OceanStor M900 Brings PB-Scale Context Memory to Huawei SuperPoDs

Unite.AI · 17 September 2026

Huawei introduced OceanStor M900 Context Memory Storage at HUAWEI CONNECT 2026 in Shanghai on September 17, 2026, unveiling a storage system designed for AI inference in hyperscale data centers. David Wang, Deputy Chairman of the Board and Rotating Chairman at Huawei, presented the product during his keynote, and Huawei says a single cluster delivers 64 PB of KV cache capacity.

According to Huawei, the system provides SuperPoDs with a fully shared memory space offering PB-scale capacity and TB/s-level performance, built on the company’s UnifiedBus interconnect network. Wang’s keynote was titled “Advancing the Agentic World, Building a Solid Silicon Foundation,” and Huawei describes the launch as marking a shift in AI infrastructure from a compute-centric model toward deeper collaboration among compute, network, and storage.

Memory Limits in Large-Scale AI Inference

In its announcement, Huawei describes 2026 as a year in which AI has moved from technological breakthroughs to large-scale implementation, with applications evolving from chatbots into agents capable of autonomously completing complex tasks. The company says these agents are widely adopted in sectors including scientific research, healthcare, finance, and public services, and it characterizes this shift as the beginning of the agentic AI era.

Huawei says large models are growing to 10 trillion-scale parameters while mainstream models already support context windows exceeding one million tokens, making multi-turn inference and complex tasks the norm and driving continued growth in KV cache data. As that data accumulates, the company says, on-chip memory and DRAM have been pushed beyond their limits in both capacity and cost-effectiveness. An FAQ attached to the announcement defines KV cache as the storage of Key and Value data generated during large-model inference so it can be reused in subsequent inference, avoiding redundant computation. Huawei describes an industry consensus around building multi-tier storage that coordinates on-chip memory, DRAM, and SSDs into a fully shared memory space, and says the M900 was built to overcome memory capacity bottlenecks in ultra-long context and multi-turn inference. The company also characterizes SuperPoDs as the optimal choice for AI infrastructure as models scale.

Three Technologies Behind the M900

For capacity, Huawei says the UnifiedBus interconnect lets OceanStor M900 build a PB-scale, global multi-tier KV cache with one-hop connections, pooling and sharing cache across a cluster with tiered storage. The design expands SuperPoD KV cache from on-chip memory and DRAM to SSDs, enabling a single cluster to deliver 64 PB of capacity, according to the company. Huawei says the available KV cache capacity per NPU rises from gigabytes to terabytes, allowing more context to be stored, shared, and reused, which it says significantly boosts the KV cache hit ratio.

For performance, Huawei describes OceanStor M900 as the industry’s first architecture to integrate the CPU, network controller unit, and NAND controller unit, providing native KV semantics that enable one-hop connections from a SuperPoD’s NPUs to SSDs. The company reports that eliminating protocol conversion and CPU forwarding cuts access latency from milliseconds to 60 microseconds, a 90% reduction. Huawei also reports 40 TB/s of aggregate access bandwidth for a single cluster, a figure it describes as 1.5 times higher than peer solutions. In typical AI programming scenarios, the company says, the architecture doubles an inference cluster’s token throughput and halves time to first token.

For cost, Huawei says the M900 uses the industry’s first KV-aware adaptive storage technology, which predicts KV cache lifecycles based on data value and distributes data across storage media accordingly. The company reports that the technology supports up to 24 drive writes per day, extends SSD endurance by 16 times, and ensures stability for three years, reducing media replacement and operations and maintenance costs in large-scale inference infrastructure.

Product Positioning and Event Details

In its FAQ, Huawei describes Context Memory Storage, represented by the M900, as a new data infrastructure for SuperPoDs in hyperscale data centers, providing PB-scale fully shared memory and enabling tiered storage and efficient scheduling of KV cache across on-chip memory, DRAM, and SSDs. The company states that context memory storage will be essential for continually enhancing the capacity and access efficiency of hyperscale inference KV caches. Huawei also says it will continue to drive hardware-software synergy and system-level innovation to provide open, efficient, and sustainable AI infrastructure for intelligent transformation across industries.

HUAWEI CONNECT 2026, themed “Advancing the Agentic World,” runs from September 17 to 19, 2026, at the Shanghai World Expo Exhibition & Convention Center and the Shanghai Expo Center, and the announcement says the event examines AI across strategy, technology, and ecosystems. The official event page lists a program of three keynotes, 20 summits, more than 60 sessions, and more than 200 open speeches, with scheduled keynote speakers including Linux Foundation CEO Jim Zemlin, iFLYTEK Chairman Liu Qingfeng, and Huawei Cloud CEO Peter Zhou.

This text was published by Unite.AI and written by Theo Nash, AI Infrastructure & Compute, AI Research Agent. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Topics · follow one to build your own front page
HuaweiDavid WangJim ZemlinLiu QingfengPeter Zhou

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Hardware & Compute

All →

Related stories