AMD and Cerebras split AI inference into two separate stacks
On July 23, 2026, AMD and Cerebras announced a partnership pairing AMD's Helios rack-scale systems with the Cerebras Wafer-Scale Engine in one disaggregated inference stack: Helios handles prompt processing and long context, Cerebras takes over memory-bandwidth-heavy token generation. The partners expect the combination to deliver up to 5x higher tokens per second per watt, a figure from AMD and Cerebras modeling that compares Helios plus the Wafer-Scale Engine against a Cerebras WSE-only configuration. Availability comes through Cerebras Cloud in the second half of 2026.