AMD and Liqid have announced a partnership to build next generation artificial intelligence infrastructure. This joint effort combines the new AMD Instinct MI350P processors with Liqid software to let businesses pool multiple accelerators without rebuilding their server rooms. But the actual shipping dates for the hardware are still unannounced.
The centerpiece of this hardware launch is the Liqid UltraStack 30. This system packs a dual socket AMD EPYC 9005 series processor and supports up to 30 PCIe based AMD Instinct MI350P GPUs inside 1 server. By clustering these processors together, businesses can access 69 PFLOPS of FP8 calculation performance. The entire setup shares an enormous pool of 4.3 TB of high bandwidth memory. Powering this massive server setup requires about 22 kW of electricity, which translates to a performance metric of roughly 3.14 PFLOPS per kW.
IT buyers are constantly hunting for ways to lower the cost of running large language models. This new system targets those exact financial concerns. According to the official joint announcement, pooling these PCIe processors leads to 3.7x more tokens per second compared to conventional configurations. Businesses could also see 2.1x more tokens per dollar spent and 1.8x better power efficiency per watt. These numbers mean that companies can deploy frontier models on premises while keeping operating costs low.
This clustered architecture is designed specifically to handle large scale artificial intelligence inference. Workloads like retrieval augmented generation and long context model deployments benefit immediately from having over 4TB of high bandwidth memory in a single machine. The setup also supports mixed fleets of small and large models served from the same shared pool, which increases server density. This allows marketing, legal, and coding assistants to run locally where data privacy is easier to manage.
Looking forward, the hardware is built to support CXL memory pooling. This provides a clear path for future memory expansion across multiple servers. As this technology matures, customers will be able to allocate terabytes of shared memory for complex caching tasks. Combined with large scale GPU pooling, the platform aims to deliver better infrastructure economics for research institutions and cloud service providers.


