
Cerebras Unveils CS-4, OpenAI and AMD Partnerships to Accelerate AI Inference
MarketBeat
Published: Aug 19, 2026, 02:02 AM
Sentiment Analysis
Cerebras unveiled its CS-4 AI system, which is expected to deliver up to twice the token-generation speed of CS-3, six times higher system-level performance and up to 10 times more tokens per watt in select applications.
The company highlighted partnerships with OpenAI and AMD to accelerate inference, including an AMD-GPU/Cerebras split in which GPUs handle model prefill and Cerebras handles token decoding.
Cerebras is expanding its data-center footprint, with 600 megawatts of power online or under contract by the end of next year, while targeting AI-agent, design, coding and cybersecurity applications that benefit from lower latency.
Cerebras Systems NASDAQ: CBRS outlined a product roadmap centered on faster AI inference, expanded data-center capacity and new partnerships with OpenAI, Arista Networks and Advanced Micro Devices during a company event led by CEO and Co-Founder Andrew Feldman.
Feldman said the company recently went public and is seeking to position its wafer-scale computing technology as an infrastructure platform for real-time AI applications.
He argued that inference speed has become a product-level consideration rather than solely a technical benchmark, particularly as AI shifts toward interactive tools and autonomous agents.
“In AI, speed is productivity,” Feldman said, asserting that faster response times allow users to run more workloads and address more complex tasks without the traditional trade-off between model intelligence and latency.
Feldman highlighted OpenAI’s recently announced GPT-5.6 Sol Ultrafast mode, which he said runs OpenAI’s most intelligent models at up to 14 times the speed of its standard offering for select customers.
He said Cerebras hardware powers the service and presented a comparison based on Humanity’s Last Exam, a graduate-level reasoning benchmark.
According to Feldman, the Cerebras-powered system completed the full 2,500-question benchmark in 11 hours, 11 minutes and 26 seconds, while the comparison system required more than three days.
Thibault Sottiaux, a member of the technical staff at OpenAI, said the company’s strategy has been to build leading models and serve them at scale, with increasing focus on agents.
He said ultrafast inference reduces the need to choose between a smaller, lower-latency model and a larger model that takes longer to respond.
Sottiaux said OpenAI would like Ultrafast to become the default experience over time, though he stressed that the company is still early in deploying it.
He said OpenAI reserves some Ultrafast capacity for incidents, security matters, major internal projects and research efforts, while also allocating capacity for customers through its API.
He added that ChatGPT has surpassed 1 billion users, while the company’s more sophisticated agentic workflows remain used by a smaller but rapidly growing segment.
Feldman said OpenAI has 15 million weekly users of Codex and ChatGPT agents, a figure Sottiaux referenced as recently published by OpenAI.
Feldman said Cerebras has data centers operating or being developed across North America and Europe, naming locations including Santa Clara, Toronto, Dallas, Minneapolis, Montreal, Oklahoma City, Alabama, Lyon, France, Norway and Mikkeli, Finland.
He said the company has brought on 600 megawatts of power that is online or under contract for delivery by the end of next year.
Jayshree Ullal, ...
Source: MarketBeat
This content is not intended as investment advice or a recommendation. Any opinions expressed are solely the personal views of each article.