Arm Announces CSS for Mobile 2 with Mali G2 Ultra NX GPU

Arm Announces CSS for Mobile 2 with Mali G2 Ultra NX GPU

Arm has announced CSS for Mobile 2, a complete compute design built to run autonomous AI agents and complex visual rendering on smartphones. The platform combines the new Mali G2 Ultra NX graphics processor with an updated C2 CPU cluster powered by 2 SME2 units. Silicon designers gain access to an integrated blueprint that delivers up to 4x better efficiency in neural graphics alongside a 70% speedup for small language models.

Mobile workloads are shifting rapidly away from basic image filters toward autonomous agents that manage apps in the background. At the same time, mobile games demand cinematic lighting that traditional rendering methods struggle to provide within phone thermal limits. Arm built the Mali G2 Ultra NX to address this exact bottleneck. It is the first Mali graphics chip to include dedicated neural accelerators right inside the rendering pipeline.

Putting neural math and standard graphics in the same place changes how mobile games run. The hardware uses machine learning to reconstruct pixels and sharpen textures without draining the battery. Arm reports the chip offers up to 4x higher performance per watt for neural rendering tasks, alongside a 14% uplift in standard game processing compared to the prior generation. A fresh Ray Tracing Unit is also included to handle realistic lighting in modern game engines.

Gaming studios are already writing software for the new hardware. Sumo Digital built a visual showcase called Neural Dawn to test the setup. Tencent Games is linking its MagicDawn tools into the pipeline, while Unity China integrated support into its Tuanjie Engine. Players will see early implementations in titles like Where Winds Meet by NetEase, Arena Breakout Infinite by Tencent, and Infinity Nikki from Infold Games.

Running an AI agent requires heavy coordination across the entire phone, and the main processor handles that workload. The C2 CPU cluster combines the high performance C2 Ultra with efficient C2 Pro cores. By doubling the SME2 capability across the cluster, the design cuts latency when handling text processing and user commands.

The numbers show clear gains over older silicon. The C2 Ultra hits up to 1.7x higher AI performance and 15% faster single thread execution compared to the C1 Ultra. It achieves these speeds while drawing up to 38% less power under matching workloads. Ecosystem partners like Google AI Edge Gallery, Alipay, vivo, and OPPO are already targeting this SME2 setup for mobile applications.

Hardware improvements only matter if app developers can write code for them without friction. Arm is pairing the architecture with software libraries like KleidiAI and the Neural Graphics Development Kit. Programmers can also access the Arm AI Portal to download prebuilt models, testing scripts, and optimization tools tailored for this silicon architecture.

About the author

Majid T.
Majid T.
Owner of Technetbook | 10+ Years of Expertise in Technology | Seasoned Writer, Designer, and Programmer | Specialist in In-Depth Tech Reviews and Industry Insights | Passionate about Driving Innovation and Educating the Tech Community Technetbook

Join the conversation

Newsletter Subscription