Huawei has unveiled the Ascend 960 SuperNode, marking the first commercial deployment of Near Package Optics technology in a supercomputing node. The system replaces thousands of conventional optical modules with integrated optical engines to train models scaling past 10 trillion parameters. But the real shift lies in how the company redesigned its physical interconnects to eliminate datacenter bottlenecks.
Modern datacenter networking is hitting a hard physical wall. Moving data across copper traces generates massive heat and drains efficiency. Huawei is countering this with its proprietary Hi ONE optical engine, the first mass produced NPO unit featuring an internal light source and a total throughput capacity of 7.2T. The design combines dense optical connections with full liquid cooling across an orthogonal physical chassis.
The numbers behind the Ascend 960 SuperNode show a radical reduction in physical complexity. A single unit integrates 5500 Hi ONE engines, which eliminates the need for 48000 discrete 800G optical modules. This hardware reduction cuts overall power draw by more than 550 kW. At the same time, system availability climbs to 99.8% while doubling mean time between failures. For raw computational throughput, a single 4096 card configuration delivers 8 EFLOPS of FP8 precision and 16 EFLOPS of FP4 precision.
Huawei focuses on building strong AI infrastructure and an open compute ecosystem, supporting native training for leading frontier models to meet the demands of an automated world.
Scaling extends far beyond single racks. By linking multiple Ascend 960 units through a dual layer CLOS four plane network using Lingqu or RoCE fabrics, clusters can scale up to 512000 cards. Adding multi rail network topology pushes the maximum theoretical ceiling to 1,000,000 interconnected accelerators. To support this compute volume, Huawei also upgraded its general computing Kunpeng SuperNode to link up to 4096 nodes into a shared 256TB memory pool. Complementing this is the OceanStor M900 storage cluster, which provides direct access caching to speed up vector searches and agent sandboxes.
Hardware power means nothing without software adoption. Huawei confirmed that its CANN compute architecture has moved into full open source community governance to simplify model porting. The goal is straightforward: make custom silicon as easy to program as standard commodity chips.
Beyond raw compute clusters, Huawei introduced its DIMAK engineering framework, which focuses on data, infrastructure, models, agents, and knowledge. The system is designed to help industrial clients move AI out of isolated pilot tests and into core production environments. Case studies presented at the conference showed real world implementations across major sectors. Migu Music integrated automated maintenance agents to cut root cause diagnostic times to under 5 minutes while reducing incident recovery time by 60%. Manufacturing partnerships, including the HongQi 100 initiative, are expanding the deployment of open operating systems across industrial automation lines, electric power grids, and hardware fabrication plants worldwide.
