SAN FRANCISCO, August 18, 2026, 20:13 PDT — Cerebras Systems (NASDAQ:CBRS) unveiled the CS-4, a three-wafer AI rack delivering 750 sparse-FP16 petaflops. First shipments are scheduled this quarter.
- CS-4 provides six times the listed compute of one CS-3 system.
- Each WSE-3 Turbo doubles per-wafer compute and memory bandwidth.
- Pricing and independent production benchmarks were not disclosed.
The sixfold rack gain comes from two changes. Each wafer is twice as fast, while the new Nexus design packages three wafer engines into one active power rack.
That matters for AI agents, where every reasoning step adds delay. Cerebras says CS-4 can spend more computation on verification and tool use without making users wait longer.
| System specification | CS-3 | CS-4 | Change |
|---|---|---|---|
| Wafer engines per system | 1 WSE-3 | 3 WSE-3 Turbo | Threefold count |
| AI compute | 125 PFLOPS | 750 PFLOPS | 6x |
| Memory bandwidth | 21.6 PB/s | 129.6 PB/s | 6x |
| On-chip fabric bandwidth | 26.7 PB/s | 160.5 PB/s | About 6x |
| System I/O bandwidth | 1.2 Tbit/s | 7.2 Tbit/s | 6x |
| I/O latency | 5 microseconds | As low as 2 microseconds | 60% lower |
The WSE-3 Turbo keeps the prior wafer’s four trillion transistors, 900,000 cores and 44GB of SRAM. Cerebras instead raises operating speed through closer power delivery and a redesigned cooling, I/O and control assembly.
| Per-wafer specification | WSE-3 | WSE-3 Turbo | Change |
|---|---|---|---|
| Transistors | 4 trillion | 4 trillion | Unchanged |
| AI-optimised cores | 900,000 | 900,000 | Unchanged |
| On-chip SRAM | 44GB | 44GB | Unchanged |
| AI compute | 125 PFLOPS | 250 PFLOPS | 2x |
| Memory bandwidth | 21.6 PB/s | 43.2 PB/s | 2x |
| Off-chip I/O | 1.2 Tbit/s | 2.4 Tbit/s | 2x |
In Cerebras’ GPT-OSS-120B test, CS-4 produced more than 4,400 tokens per second for each user. The company claimed speeds up to 30 times those of GPU systems. Actual results vary with model, context, precision and serving setup.
“In AI, speed is productivity,” Chief Executive Andrew Feldman said. He argued that fast inference no longer requires a smaller, less capable model. Those claims still need broad independent testing.
The rack’s three rear-mounted “backpacks” combine each wafer with direct liquid cooling, power conversion and I/O. Cerebras says the design has 50% fewer components and uses 60% more automated manufacturing than its predecessor.
| Deployment measure | Prior CS-3 generation | CS-4 Nexus platform |
|---|---|---|
| Compute packaging | Single-wafer system | Three pluggable wafer-scale backpacks |
| Installation time | Days | Hours |
| Relative component count | Baseline | 50% fewer |
| Supported model scale | Up to 24 trillion parameters | More than 50 trillion parameters |
| Availability | Shipping | First shipments this quarter |
| Public list price | Not disclosed | Not disclosed |
Direct Wafer Links connect racks without an intervening switch. Wafer-to-wafer latency falls to two microseconds, supporting models above 50 trillion parameters. Standard RoCE v2 Ethernet remains available.
The programmable I/O also supports disaggregated inference. A separate engine handles prompt processing, then passes token generation to CS-4. Cerebras named AMD (NASDAQ:AMD) Helios and Amazon (NASDAQ:AMZN) Web Services Trainium as planned ecosystem partners.
CS-4 arrives as Cerebras shifts toward cloud revenue. Second-quarter cloud sales reached about $126 million, while hardware revenue fell to $54.1 million from $70.3 million a year earlier. Adjusted gross margin declined to 40.6%.
The new rack could revive hardware demand and lower cloud operating costs. It also raises execution pressure. Cerebras must manufacture, cool and deploy systems quickly enough to support contracted capacity, including OpenAI workloads.
Competitive claims remain vendor-defined. Nvidia (NASDAQ:NVDA) and AMD publish performance using different precisions, models and rack configurations, making headline petaflop comparisons unreliable without matched tests.
Risks: Shipment timing is forward-looking, and pricing remains unknown. The 30-fold GPU advantage comes from Cerebras testing, not a neutral production benchmark.
CS-4’s clearest verified gain is architectural. It combines three faster wafers, modular servicing and lower-latency links. Whether that turns speed into durable margins now depends on deployments.