Marvell launches AI "memory decoupling" architecture to address the bandwidth bottleneck of Agentic AI inference
According to official news, Marvell Technology announced the launch of a new generation of memory solution portfolio for AI infrastructure, covering server-level AI storage, rack-level CXL memory expansion and pooling, as well as multi-rack optical interconnect shared memory, aimed at addressing the growing memory capacity and bandwidth bottlenecks in Agentic AI inference processes.
Marvell stated that as AI model sizes increase, context windows extend, and KV Cache demand grows, traditional tightly coupled architectures of computing and memory are limiting AI inference efficiency. Through memory disaggregation, memory resources can be made more independent of computing resource expansion, improving GPU utilization and reducing data movement latency. The products released include:
Bravera SC6 PCIe 6.0 SSD controller: Designed for AI inference storage scenarios, it helps cloud service providers migrate more KV Cache to high-performance SSDs, enhancing infrastructure efficiency. This product features an architecture compatible with multi-vendor NAND and is expected to begin sampling in the fourth quarter of 2026.
Marvell Structera X memory expansion solution: Based on CXL technology, it supports rack-level memory expansion and resource pooling, helping data centers share and allocate memory resources more flexibly, reducing AI infrastructure costs.
Marvell Photonic Fabric optical interconnect memory solution: Constructs a shared memory architecture across multiple racks using optical interconnect technology, supporting up to 32TB warm KV Cache offloading and helping AI inference clusters enhance throughput capacity.
Marvell stated that the Photonic Fabric solution can achieve a 2 to 3 times increase in Token throughput under existing data center space and power consumption constraints, supporting larger scale models and longer context AI applications.
Marvell executive Will Chu stated that AI infrastructure is transitioning from a single server architecture to a system where computing, memory, and connectivity operate in synergy, and in the future, memory needs to expand more independently to enhance resource utilization and Token efficiency.
As the demand for AI Agents and large model inference continues to grow, memory capacity, bandwidth, and data transfer efficiency are becoming new focal points of competition in AI infrastructure, following computing power.






