OpenAI CEO Altman: Cerebras Remains a Close Partner, Both Companies Have Deep Cooperation at the Speed Frontier
OpenAI CEO Sam Altman stated that there is speculation about its collaboration with Cerebras, but Cerebras remains a close partner, and both companies have deep cooperation at the speed frontier.
The speculation arose after Cerebras' stock price fell about 17%. The decline was related to reports that OpenAI had moved part of the ultra-fast version of its new model to NVIDIA graphics processors, raising market concerns that the wafer-scale chip company would lose its most important client's inference share. Paul Meeks of Freedom Capital subsequently upgraded his rating from hold to buy, with a target price of $209, stating that the sell-off was an overreaction.
The collaboration is not just a statement. In January, OpenAI agreed to access 750 megawatts of Cerebras computing power, to be rolled out in phases by 2028, specifically for low-latency inference. According to Reuters, the contract is worth over $10 billion during its term. Cerebras sells its own chip cloud services, with data centers built or leased by it, and OpenAI pays based on usage.
In August, OpenAI connected the Ultrafast tier of GPT-5.6 Sol to Cerebras, achieving an output of about 750 tokens per second, which is up to 14 times faster than the standard tier, initially available for limited preview via an interface. OpenAI's head of computing, Sachin Katti, stated at the time that the combination of computing power should match the workload, with Cerebras providing dedicated low-latency inference, not replacing all graphics processors.
Cerebras CEO Andrew Feldman mentioned that negotiations began with a demonstration last August, when OpenAI's open-source model was more efficient on its chips than traditional graphics processors, focusing on inference and models that need to "think" before answering. The wafer-scale engine integrates computing, storage, and bandwidth on a single large chip, avoiding memory transfers during conventional hardware decoding.
This is a position correction driven by events, not new orders. Sellers are Cerebras shareholders concerned about customer concentration, while buyers believe that inference supply remains tight, and a single client switching to NVIDIA does not equate to contract cancellation. The 750 megawatt order and the already launched Ultrafast still keep revenue expectations with Cerebras, while NVIDIA is receiving another tier of workload. Beneficiaries are chip vendors that can simultaneously enter OpenAI's procurement mix, while those under pressure are wafer chip narratives that treat single client speed orders as their entire valuation.
After the rating upgrade, Cerebras' stock price rebounded about 2% on Friday. The company also signed a deployment of about 100 megawatts of inference hardware with cloud startup Gimlet Labs this week, indicating that low-latency inference clients are not limited to OpenAI.
Source: Public Information
ABAB AI Insight
Cerebras, founded by Andrew Feldman, follows a wafer-scale engine route and does not compete with NVIDIA in general-purpose training cards. After attempting to go public in 2021 and retracting, its path to listing in 2024 is again stalled due to Middle Eastern funding and regulatory scrutiny, with customer concentration being a core risk in its prospectus. OpenAI is being highlighted to reassure stakeholders because revenue visibility is tied to a few large inference contracts.
OpenAI's funding is not solely dependent on one client. The 750 megawatts in January 2026, worth over $10 billion, is for low-latency decoding, with capacity spread out until 2028; training and large-scale pre-filling still utilize NVIDIA and cloud partners' graphics processor clusters. Sachin Katti publicly articulated the strategy as "load-matching systems," allowing Ultrafast to run on Cerebras while another speed tier of new models can still utilize NVIDIA, meaning the two purchases are not mutually exclusive.
In contrast, Google uses its self-developed TPU for internal inference while continuing to purchase from NVIDIA. The difference is that Cerebras lacks internal demand from search or cloud to support it, while OpenAI serves as both a customer and a narrative. The industry position is at the eve of specialized inference scaling: speed can achieve 14 times and 750 tokens per second, but orders are still confirmed by megawatts and multi-year leases, not by demonstrations.
Structurally, this represents a split in the decoding layer within the reconstruction of the industry chain. Training, pre-filling, and decoding no longer need to be on the same chip. After the split, pricing power shifts from "who has the most training cards" to "who can keep weights on the chip and reduce latency." Altman's statement of "close partner" addresses the risk premium of customer concentration transactions without changing OpenAI's continued multi-vendor procurement structure.
ABAB News · Cognitive Laws
- A word from a major client can adjust the premium but cannot change the contract.
- Once the load is split, suppliers shift from substitutes to equals.
- Speed is the product; megawatts are the revenue.