Timeline
Cerebras achieved 750 token/s on GPT-5.6 Sol
Cerebras expanded CS-3 production 7× at Milpitas facility
AMD and Cerebras launched disaggregated inference platform at Advancing AI 2026
AMD and Anthropic announce strategic partnership to deploy up to 2GW of MI450 GPUs
Cerebras CS-3 launched achieving 500+ token/s for Llama 2 70B
Cerebras and Flex announce 7x production expansion of CS-3 at Milpitas facility
Disclosed three mechanical innovations for wafer-scale chip cooling: vertical power delivery, flexible interposers, direct-impingement cooling
Announced CS4 wafer-scale chip staying on 5nm due to SRAM scaling flattening.
AMD provides $3.6M MI355X cluster access to vLLM and SGLang maintainers
Launched PCIe-based GPU for AI workloads targeting existing servers