Cerebras announced CS-4 on August 18. The rack contains three WSE-3T processors, but the headline compute figure is only meaningful with its numerical format and the scope of each specification.

Start with the unit being described
The CS-4 datasheet lists 250 PFLOPS per wafer and 750 for the three-wafer system. Its footnote specifies sparse FP16. That qualification matters: comparing the number with a dense-compute specification without explaining sparsity does not establish a performance advantage.
| Published specification | One WSE-3T wafer | CS-4 system |
|---|---|---|
| On-chip SRAM | 44 GB | 132 GB across three wafers |
| Memory bandwidth | 43.2 PB/s | 129.6 PB/s |
| I/O bandwidth | 2.4 Tbit/s | 7.2 Tbit/s |
The system totals do not describe one universally accessible memory pool with zero communication cost. Nor does the continued use of 5 nm, four trillion transistors and 900,000 cores prove that the newly announced WSE-3T is identical to its predecessor.
A bandwidth conversion that prevents a costly mistake
Using decimal units, 7.2 Tbit/s divided by eight is 0.9 TB/s. It is not 129.6 PB/s: the latter describes internal memory bandwidth, a different part of the machine. A procurement spreadsheet that places these values in one “network speed” column would compare unrelated capabilities.
The official product view shows the rack and its exposed modules. It helps locate the physical system being discussed; it does not illustrate the path taken by every byte. Direct Wafer Links connect wafers within and across racks, while the announcement separately describes programmable RoCEv2 I/O. Fabric design and required network settings must come from the actual deployment specification.
Evaluate the model you will actually serve
The launch announcement reports 4,400 tokens per second per user for GPT-OSS-120B and an up-to-30× comparison with GPU-based solutions. These are Cerebras results, not our measurements or a guarantee across models. First shipments were scheduled to begin in the announcement quarter; a schedule is not evidence that every customer can already take delivery.
For a useful comparison, hold model version, precision, input and output lengths, concurrent requests and quality requirements constant. Request time to first token and sustained output rate separately. Then compare delivered throughput, power and price for that configuration. Peak arithmetic and one fast generation result cannot replace that workload-specific assessment.
Distinguish wafer and rack figures, sparse FP16 and external I/O; qualify vendor benchmark and shipment schedule; use official CS-4 press image.