The AI hardware race has mostly been a fight over scarce parts. Cerebras Systems is trying to win it by not needing most of them.
The company’s wafer-scale chips avoid high-bandwidth memory (HBM), advanced CoWoS packaging and 3nm manufacturing. Those happen to be three of the tightest chokepoints in the industry right now.
A chip built around what it leaves out
Cerebras took a different route. Its WSE-3 and the derivative WSE-3T are built on TSMC’s 5nm process node. All of their memory lives on the chip itself as SRAM, rather than on separate HBM stacks.
The WSE-3 was announced in March 2024. Its specs read more like a city plan than a chip sheet: 900,000 AI cores, 44 GB of on-chip SRAM and approximately 4 trillion transistors.
All of that sits within a 46,225 mm² area.
CEO Andrew Feldman has argued the design pays off where it counts most for customers: speed when AI models generate responses. Speaking in June 2026, he described the architecture this way:
“the fastest inference in the world by an order of magnitude”
The research findings describe Feldman’s claim as boosting token-generation speed while removing the slowdowns typical of GPU-based designs.
The backlog behind the bet
As of June 2026, the company reported a backlog of $25.4 billion. More than $20 billion of that comes from a multi-year deal with OpenAI.
Cerebras also says it has over 600 MW of data-center capacity either live or contracted. Its Q2 revenue nearly doubled year-over-year, according to the research findings.
In September 2026, Cerebras agreed to supply CS-4 systems for approximately 100 MW of capacity to Gimlet Labs.
The company also struck a long-term agreement with General Compute, with deployment planned for Q1 2027.
All of this follows Cerebras’ Nasdaq IPO in May 2026. Since listing, the company has started expanding US manufacturing and building out partnerships to increase operational capacity.
Why skipping the queue matters
The research findings note that AI infrastructure demand increasingly hinges on data-center space, power and construction, not just on whether chips are available.
That helps explain why Cerebras now talks about capacity in megawatts rather than chip counts. The 600 MW figure and the roughly 100 MW Gimlet Labs commitment are measures of power and floor space.
What to watch
For Cerebras investors, the core question is execution. A $25.4 billion backlog is a promise, not revenue, and converting it depends on getting data centers built, powered and running on schedule.
The research findings point to data-center ramp costs as a challenge the company is working to overcome.
Customer concentration is the second thing to track. With more than $20 billion of the backlog tied to OpenAI, the health of that single relationship carries outsized weight. Deals like Gimlet Labs and General Compute help diversify the mix, though they are much smaller by comparison.
On performance, Feldman’s order-of-magnitude inference claim will face scrutiny from customers running real workloads. On-chip SRAM is extremely fast, but 44 GB per chip is a fixed budget, and how systems scale for the largest models will shape how broadly the design gets adopted.
Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.

3 hours ago
16







English (US) ·