AMD and Cerebras Partner on Disaggregated AI Inference Platform

AMD and Cerebras Partner on Disaggregated AI Inference Platform

AMD and Cerebras Systems have formed a technical partnership to build a disaggregated artificial intelligence inference system. The hardware merges AMD Helios rack scale infrastructure with the Cerebras Wafer Scale Engine to tackle speed bottlenecks in advanced software. The joint platform is scheduled to launch on the Cerebras Cloud in the second half of the year.

Processing modern artificial intelligence tasks requires different types of computing power depending on the workload. Simple text generation needs massive data capacity, while real time agents and coding assistants require instant feedback. This partnership splits these tasks between 2 separate chips. AMD Helios handles the initial prompt processing and large data contexts. The Cerebras Wafer Scale Engine then takes over the memory heavy task of generating individual words.

This cooperative division of labor is expected to yield 5x higher efficiency measured in tokens per second per watt. During the Advancing AI conference, AMD chief executive officer Lisa Su discussed the need for more adaptable computing architectures.

AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach.

Andrew Feldman, the chief executive officer at Cerebras, also highlighted the immediate market need for fast processing during the joint announcement.

The demand for ultra fast inference is growing at an unprecedented pace.

The speed of text generation is becoming a key factor in software development, robotics, and scientific research where sluggish response times ruin the user experience. To support this launch, Cerebras plans to install the Helios hardware inside its own data centers. The combined service will roll out as a cloud option, giving developers access to high throughput processing without sacrificing system response times. This setup gives businesses a way to run complex AI workflows without purchasing expensive physical infrastructure.

About the author

Majid T.
Owner of Technetbook | 10+ Years of Expertise in Technology | Seasoned Writer, Designer, and Programmer | Specialist in In-Depth Tech Reviews and Industry Insights | Passionate about Driving Innovation and Educating the Tech Community Technetbook

Join the conversation

Newsletter Subscription