Cerebras CS-4 Packs Six Times CS-3 Compute Into a New AI Rack

Cerebras CS-4 Packs Six Times CS-3 Compute Into a New AI Rack

August 19, 2026

SAN FRANCISCO, August 18, 2026, 20:13 PDT — Cerebras Systems (NASDAQ:CBRS) unveiled the CS-4, a three-wafer AI rack delivering 750 sparse-FP16 petaflops. First shipments are scheduled this quarter.

  • CS-4 provides six times the listed compute of one CS-3 system.
  • Each WSE-3 Turbo doubles per-wafer compute and memory bandwidth.
  • Pricing and independent production benchmarks were not disclosed.

The sixfold rack gain comes from two changes. Each wafer is twice as fast, while the new Nexus design packages three wafer engines into one active power rack.

That matters for AI agents, where every reasoning step adds delay. Cerebras says CS-4 can spend more computation on verification and tool use without making users wait longer.

System specificationCS-3CS-4Change
Wafer engines per system1 WSE-33 WSE-3 TurboThreefold count
AI compute125 PFLOPS750 PFLOPS6x
Memory bandwidth21.6 PB/s129.6 PB/s6x
On-chip fabric bandwidth26.7 PB/s160.5 PB/sAbout 6x
System I/O bandwidth1.2 Tbit/s7.2 Tbit/s6x
I/O latency5 microsecondsAs low as 2 microseconds60% lower

The WSE-3 Turbo keeps the prior wafer’s four trillion transistors, 900,000 cores and 44GB of SRAM. Cerebras instead raises operating speed through closer power delivery and a redesigned cooling, I/O and control assembly.

Per-wafer specificationWSE-3WSE-3 TurboChange
Transistors4 trillion4 trillionUnchanged
AI-optimised cores900,000900,000Unchanged
On-chip SRAM44GB44GBUnchanged
AI compute125 PFLOPS250 PFLOPS2x
Memory bandwidth21.6 PB/s43.2 PB/s2x
Off-chip I/O1.2 Tbit/s2.4 Tbit/s2x

In Cerebras’ GPT-OSS-120B test, CS-4 produced more than 4,400 tokens per second for each user. The company claimed speeds up to 30 times those of GPU systems. Actual results vary with model, context, precision and serving setup.

“In AI, speed is productivity,” Chief Executive Andrew Feldman said. He argued that fast inference no longer requires a smaller, less capable model. Those claims still need broad independent testing.

The rack’s three rear-mounted “backpacks” combine each wafer with direct liquid cooling, power conversion and I/O. Cerebras says the design has 50% fewer components and uses 60% more automated manufacturing than its predecessor.

Deployment measurePrior CS-3 generationCS-4 Nexus platform
Compute packagingSingle-wafer systemThree pluggable wafer-scale backpacks
Installation timeDaysHours
Relative component countBaseline50% fewer
Supported model scaleUp to 24 trillion parametersMore than 50 trillion parameters
AvailabilityShippingFirst shipments this quarter
Public list priceNot disclosedNot disclosed

Direct Wafer Links connect racks without an intervening switch. Wafer-to-wafer latency falls to two microseconds, supporting models above 50 trillion parameters. Standard RoCE v2 Ethernet remains available.

The programmable I/O also supports disaggregated inference. A separate engine handles prompt processing, then passes token generation to CS-4. Cerebras named AMD (NASDAQ:AMD) Helios and Amazon (NASDAQ:AMZN) Web Services Trainium as planned ecosystem partners.

CS-4 arrives as Cerebras shifts toward cloud revenue. Second-quarter cloud sales reached about $126 million, while hardware revenue fell to $54.1 million from $70.3 million a year earlier. Adjusted gross margin declined to 40.6%.

The new rack could revive hardware demand and lower cloud operating costs. It also raises execution pressure. Cerebras must manufacture, cool and deploy systems quickly enough to support contracted capacity, including OpenAI workloads.

Competitive claims remain vendor-defined. Nvidia (NASDAQ:NVDA) and AMD publish performance using different precisions, models and rack configurations, making headline petaflop comparisons unreliable without matched tests.

Risks: Shipment timing is forward-looking, and pricing remains unknown. The 30-fold GPU advantage comes from Cerebras testing, not a neutral production benchmark.

CS-4’s clearest verified gain is architectural. It combines three faster wafers, modular servicing and lower-latency links. Whether that turns speed into durable margins now depends on deployments.

BEZ KABLI • EXTENDED COVERAGE

Further analysis

What did Cerebras launch?
Cerebras unveiled CS-4, a rack-scale AI accelerator built around three WSE-3 Turbo wafer engines. It lists 750 sparse-FP16 petaflops, 129.6 PB/s of memory bandwidth and 7.2 Tbit/s of I/O.
How much faster is CS-4 than CS-3?
Listed rack compute rises sixfold, from 125 to 750 PFLOPS. Each Turbo wafer doubles compute, while CS-4 uses three wafers instead of one.
Does CS-4 really beat GPUs by 30 times?
Cerebras claims up to a 30-fold speed advantage. Its GPT-OSS-120B test exceeded 4,400 tokens per second per user. Independent matched production tests are not yet available, and results vary by configuration.
When will CS-4 ship, and what will it cost?
First shipments are scheduled for this quarter. Cerebras has not published a list price, so buyers cannot yet compare acquisition costs with GPU racks.
Why does CS-4 matter to AI users?
Faster token generation can reduce delays in coding tools and AI agents. The system also supports models above 50 trillion parameters, but real benefits depend on software, model and workload.

Mateusz Ługowik

Mateusz Ługowik is a senior markets reporter at Bez-kabli.pl, specializing in technology stocks, artificial intelligence and global financial markets. A graduate of the University of Gdańsk, he previously worked in investment research and market analysis. His coverage helps readers understand the key trends, companies and innovations influencing investors worldwide.