Arm has introduced the Neoverse CSS N4, a compute subsystem built for agentic AI infrastructure across cloud and data center environments. Fabricated on a 3nm node, the platform delivers up to 128 cores per die, clock speeds reaching 3.8 GHz, and the first integration of LPDDR6 memory alongside PCIe Gen 7 connectivity. Early silicon partners including Hongjun Microelectronics have already begun adopting the architecture for custom server chip designs.
The announcement took place during the annual Arm Everywhere China conference. Arm positioned CSS N4 as its most configurable compute subsystem to date, aiming to shorten the path from initial design to final production silicon. Each die can scale from 8 up to 128 cores, with native hardware support for FP8 and MMLA operations. This setup gives data center operators the flexibility to build high density chips tailored specifically for AI model execution, data routing, and system control.
Memory bandwidth and interconnect speeds see major upgrades across the board. The platform supports LPDDR6, DDR5, and MRDIMM configurations running at transfer rates between 8000 and 12000 MT/s. For high throughput communications between compute tiles and accelerators, the design incorporates 128 lanes of PCIe Gen 7. It provides the bandwidth required to keep AI processing clusters fed with data while avoiding local bottlenecks.
Performance metrics show major generational gains over the previous Neoverse CSS N3 architecture. Socket level performance jumps by up to 100%. Efficiency also improves, with performance per watt increasing by 1.25x and memory bandwidth expanding by 1.75x. These improvements help cloud providers handle heavier data movement and multi tenant processing loads without inflating facility power budgets.

