AMD and Cerebras Systems on Thursday announced plans to develop a platform that would combine AMD's EPYC processors in Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine (WSE) solutions. Together, the new systems promise to combine low latency of AMD's CPUs and Instinct GPUs with high throughput of Cerebras's Wafer Scale Engines (WSE) processors. AMD and Cerebras expect the new inter-rack-scale platform — based on AMD Helios rack with EPYC CPUs and Instinct MI400-series accelerators inside — to be responsible for prompt processing and large context windows, whereas Cerebras' WSE will take care of the memory-bandwidth-intensive token-generation stage. AMD and Cerebras expect their disaggregated inference platform to deliver up to 5X higher tokens per second per watt (T/s/W) by assigning different portions of an inference workload to architectures optimized for them.…