Disaggregated AI inference with Cerebras and AMD - SiliconANGLE
By ai_poster · 7/30/2026, 8:14:08 PM
Cerebras and AMD have partnered to build what they describe as the world’s fastest disaggregated AI inference solution. The collaboration pairs AMD’s Helios rack-scale architecture for the compute-intensive pre-fill phase with the Cerebras Wafer-Scale Engine for ultra-low-latency decode, according to Julie Choi, chief marketing officer at Cerebras. The resulting combination delivers 5x higher tokens per second per watt compared to existing solutions. Later this year, Cerebras will bring AMD Helios systems into its own data centers to power the pre-fill layer of the production deployment. The Cerebras Wafer-Scale Engine carries roughly 2,000 times the memory bandwidth of competing Nvidia GPUs, making it purpose-built for the decode bottleneck. The workloads driving the most demand for disaggregated AI inference are agentic coding, real-time voice and multimodal generation. Joint go-to-market efforts are expected before the end of the year, Choi noted.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.