AMD and Cerebras Systems on Thursday announced plans to develop a platform that would combine AMD's EPYC processors in Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine (WSE) solutions. Together, the new systems promise to combine low latency of AMD's CPUs and Instinct GPUs with high throughput of Cerebras's Wafer Scale Engines (WSE) processors.
AMD and Cerebras expect the new inter-rack-scale platform — based on AMD Helios rack with EPYC CPUs and Instinct MI400-series accelerators inside — to be responsible for prompt processing and large context windows, whereas Cerebras' WSE will take care of the memory-bandwidth-intensive token-generation stage.
AMD and Cerebras expect their disaggregated inference platform to deliver up to 5X higher tokens per second per watt (T/s/W) by assigning different portions of an inference workload to architectures optimized for them. Therefore, AMD Helios provides rack-scale compute capacity and large volumes of complex requests, whereas the Cerebras WSE handles latency-sensitive token generation. The two compute platforms will operate within a single inference workflow, although the companies have not disclosed additional performance data or explained how the systems will be interconnected.
The underlying idea of the AMD + Cerebras platform is essentially the same as Nvidia's CPX concept, but AMD and Cerebras assign the specialized hardware to the opposite inference stage.
Nvidia's disaggregated design separates inference into context/prefill and generation/decode. The cancelled Rubin CPX GPU with GDDR7 was optimized specifically for the compute-heavy context/prefill stage, while the regular HBM-equipped Rubin GPUs handle the memory-bandwidth-bound generation stage.
By contrast, the AMD and Cerebras platform follows the same disaggregation principle, but the specialization is inverted: AMD's Helios platform with Instinct GPUs handles the prefill stage and processes prompts and large context windows, while the Cerebras WSE takes over the decode stage and handles latency-sensitive token generation.
Cerebras plans to install AMD Helios systems in its own data centers and integrate them with its WSE racks. The combined offering is scheduled to become available initially through Cerebras Cloud in the second half of 2026, according to the two companies.
Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.

3 hours ago
5








English (US) ·