Google TPU, AWS Trainium, and Microsoft Maia cover three hyperscaler ASIC bets. Cerebras answers a different question: what if the accelerator is not a reticle-limited die at all, but nearly an entire 300 mm wafer kept intact?
This guide covers the Cerebras Wafer-Scale Engine (WSE-3) and the CS-3 system it powers. You will get architecture numbers from Cerebras's WSE-3 datasheet and March 2024 launch materials; how MemoryX and SwarmX sit around the wafer; the software path (CSoft, PyTorch, Graph Compiler, SDK/CSL); Condor Galaxy and cloud access as documented; how wafer-scale compares with GPU clusters and with Google TPU, AWS Trainium, and Microsoft Maia; and an honest limits section. Numbers below are vendor-published as of September 2026 context, not invented datasheets.
Key Takeaways
- WSE-3: 5 nm, 46,225 mm², 4 trillion transistors, 900,000 AI cores, 44 GB on-chip SRAM.
- Vendor peaks: 125 PFLOPS AI, 21 PB/s SRAM bandwidth, 214 Pb/s on-wafer fabric.
- CS-3 pairs the wafer with MemoryX (to 1.2 PB) and SwarmX clustering (to 2,048 systems).
- Software centers on CSoft, PyTorch via XLA, Graph Compiler, and CSL kernels, not CUDA.
- Wafer-scale wins on single-device simplicity for huge models; it loses when CUDA portability or commodity GPU rental dominate.
What wafer-scale actually means for AI silicon
Every GPU and most ASICs stop at the reticle limit: the largest die a scanner can expose in one shot, then dice, package, and wire across a board or NVLink domain. Cerebras's bet since the first Wafer-Scale Engine (2019) is to keep a giant square of a 300 mm wafer intact, stitch dies into one processor, and put SRAM next to every core so the machine looks like one device to the programmer.
That is not a packaging curiosity. It is a different interconnect and memory story. Instead of shipping activations across copper between dozens of HBM stacks and GPUs, the WSE moves data across an on-wafer fabric between hundreds of thousands of small cores. External weight memory (MemoryX) and cluster fabric (SwarmX) sit outside that wafer when models or clusters grow past one chip.
As of September 2026, the production story that matters for buyers is WSE-3 inside the CS-3 system (announced March 2024). Earlier WSE-1 / CS-1 and WSE-2 / CS-2 generations still appear in labs and Condor Galaxy history, but WSE-3 is the current flagship architecture Cerebras publishes against.
WSE-3 architecture: cores, SRAM, and on-wafer fabric
The primary hardware source is Cerebras's WSE-3 datasheet, backed by the March 13, 2024 press release and the March 12, 2024 CS-3 blog.
Per the datasheet, WSE-3 is fabricated on a 5 nm process with 46,225 mm² of silicon, 4 trillion transistors, and 900,000 AI-optimized cores. On-chip memory is 44 GB of distributed SRAM with 21 PB/s memory bandwidth. The on-wafer interconnect is quoted at 214 Pb/s fabric bandwidth between cores.
Launch materials put peak AI performance at 125 petaFLOPS for the CS-3 / WSE-3 combination. Cerebras positions each core as independently programmable for tensor and sparse linear-algebra work. Versus WSE-2 / CS-2, the CS-3 blog states each core moved to 8-wide FP16 SIMD (2× the prior 4-wide design), with measured up-to-2× tokens per second on Llama 2, Falcon 40B, MPT-30B, and multi-modal models in Cerebras's own testing.
Here is why the SRAM figure matters more than the transistor marketing. HBM on a GPU is fast but sits off-die through an interposer. WSE puts tens of gigabytes of SRAM across the wafer so cores see single-cycle local memory at extreme aggregate bandwidth. Inside Deep Tech's HBM full guide remains the right companion when you compare "stacked DRAM next to a reticle die" with "SRAM tiled across a wafer." They are not interchangeable BOM columns.
| Spec | WSE-3 (Cerebras datasheet / Mar 2024 launch) |
|---|---|
| Process | 5 nm |
| Silicon area | 46,225 mm² |
| Transistors | 4 trillion |
| AI-optimized cores | 900,000 |
| On-chip memory | 44 GB SRAM |
| Memory bandwidth | 21 PB/s |
| On-wafer fabric | 214 Pb/s |
| Peak AI compute (vendor) | 125 PFLOPS |
| Core math note | 8-wide FP16 SIMD (per CS-3 blog vs CS-2) |
Source: Cerebras WSE-3 datasheet; Cerebras press release "Cerebras Systems Unveils World's Fastest AI Chip" (March 13, 2024); Cerebras blog "Cerebras CS-3" (March 12, 2024). PFLOPS and "× vs GPU" multipliers on the datasheet are vendor comparisons, not third-party lab measurements.
CS-3 system: MemoryX, SwarmX, and Weight Streaming
The wafer is only half the product. CS-3 is the system chassis around WSE-3: power, cooling, host attach, and the external memory and cluster fabrics Cerebras calls MemoryX and SwarmX.
Unlike a GPU where HBM capacity is fused to the package, Cerebras decouples compute from weight memory. CS-2 clusters offered MemoryX SKUs around 1.5 TB and 12 TB. For CS-3, Cerebras documented enterprise options (24 TB and 36 TB) and hyperscaler options (120 TB and 1,200 TB / 1.2 PB). The 1.2 PB configuration is the one Cerebras cites for storing models up to 24 trillion parameters in a single logical memory space without partitioning.
Clustering scales through SwarmX. CS-2 supported clusters up to 192 systems; CS-3 marketing raises that to 2,048 CS-3 systems for a vendor claim of 256 exaFLOPS of AI compute. The programming pitch is Weight Streaming: the cluster presents as a single logical device so the model is not hand-sharded the way a multi-thousand-GPU job usually is.
Treat those cluster and training-time claims as capacity planning inputs from the vendor, then measure. A claim that Llama 2 70B trains "from scratch in less than a day" on a full 2,048-node CS-3 cluster (versus roughly a month on Meta's published GPU baseline in Cerebras's comparison) depends on Cerebras's full stack, sparsity features, and definition of the run. It is not a portable FLOPS conversion you can drop onto an NVIDIA quote spreadsheet.
| Layer | Role on CS-3 | Buyer note |
|---|---|---|
| WSE-3 wafer | Compute + on-chip SRAM + fabric | Reticle-limit alternative; unique supply |
| MemoryX | External weight / parameter memory | Scales independently of wafer count |
| SwarmX | Multi-CS-3 interconnect | Up to 2,048 systems (vendor) |
| Weight Streaming | Single-device programming model | Simplifies large-model placement vs GPU sharding |
| Chassis | <½ rack CS-3 system (vendor) | Liquid cooling / power co-designed |
Software stack: CSoft, PyTorch, Graph Compiler, SDK
Hardware without a software path is a science project. Cerebras's documented stack is the Cerebras Software Platform (CSoft), aimed at letting teams train multi-billion-parameter models on a single logical device without writing classic distributed training glue.
The ML path most teams will evaluate first is PyTorch. Cerebras documents a lightweight wrapper that integrates through XLA's lazy tensor backend, with the cerebras.pytorch API configuring a CSX backend and ClusterConfig (including num_csx). The Cerebras Graph Compiler (CGC) turns the network into an executable that allocates cores and schedules communication across the wafer.
For kernels that need closer control, Cerebras ships an SDK and Cerebras Software Language (CSL): a C-like interface to the WSE microarchitecture. Launch materials also claim native support for PyTorch 2.0-era models and techniques such as multi-modal models, vision transformers, mixture of experts, and diffusion, plus hardware acceleration for dynamic and unstructured sparsity (vendor claim: up to 8× training speedup when sparsity applies).
Inference has a separate, productized face: an OpenAI-compatible API at Cerebras Cloud (api.cerebras.ai / cloud.cerebras.ai), with documented rate limits and model endpoints. That path is how many developers meet Cerebras today, even if their training still lives on GPUs. Do not confuse "fast hosted Llama completions" with "your custom research kernel will compile on day one."
CUDA remains the gravity well. Libraries, Triton recipes, and hiring markets still orbit NVIDIA. Cerebras's answer is fewer lines of distributed code and a single-device abstraction, not a drop-in cuDNN replacement. Budget bring-up time the same way you would for Neuron or Maia SDK work.
Condor Galaxy, AI Model Studio, and how you actually get access
On-prem CS-3 systems exist for enterprise and lab buyers. Many teams will never take delivery of a wafer chassis. They will rent capacity.
Condor Galaxy is the Cerebras + G42 supercomputer network. Condor Galaxy 1 and 2 were CS-2 based installations in California. Condor Galaxy 3 is documented as 64× CS-3 for 8 exaFLOPS in Dallas, doubling the prior Condor Galaxy compute capacity in Cerebras's framing. The pitch to users is cloud-style access to a single logical AI supercomputer rather than operating 64 separate boxes.
For training jobs without owning metal, Cerebras documents the AI Model Studio: a pay-per-model service on dedicated CS-3 clusters hosted with Cirrascale Cloud. Institutional programs (for example Neocortex / ACCESS-style research access) have also exposed CS-3 cloud paths to academic users; those programs change over time, so verify current eligibility rather than assuming a permanent public queue.
On the inference side, Cerebras's public materials through 2025 described expanding CS-3-powered datacenters across the U.S. and Europe and a target of serving over 40 million tokens per second by end of 2025 across that footprint. Treat capacity and token-rate targets as vendor roadmap language. Your SLA is whatever your cloud console and contract show in September 2026.
Wafer-scale vs GPU clusters vs hyperscaler ASICs
Readers who already finished Inside Deep Tech's hyperscaler trilogy need a clean comparison, not another slogan.
| Dimension | Cerebras CS-3 / WSE-3 | GPU cluster (typical) | Hyperscaler ASIC (TPU / Trainium / Maia) |
|---|---|---|---|
| Basic unit | Wafer-scale chip in CS-3 | Reticle die + HBM package | Cloud-custom die + fabric |
| On-chip memory story | 44 GB distributed SRAM | Caches + HBM on package | HBM + vendor SRAM/cache mix |
| Scale-up fabric | On-wafer + SwarmX | NVLink / Infinity Fabric / UALink bets | ICI / NeuronLink / Ethernet (vendor) |
| Programming model | Single logical device + Weight Streaming | Multi-GPU sharding (FSDP, Megatron, etc.) | Cloud SDK + XLA/Neuron/Triton paths |
| Where you buy it | Cerebras cloud / Condor Galaxy / on-prem | Hyperscaler GPU SKUs + OEM | Inside one cloud primarily |
| Software gravity | CSoft / PyTorch-XLA / CSL | CUDA ecosystem | Cloud-native SDKs |
Versus GPUs: Cerebras wins when model size and sharding pain dominate, when unstructured sparsity maps to the hardware, and when you want one logical device instead of a mesh of NCCL domains. GPUs win when you need commodity rental tomorrow, CUDA libraries, multi-vendor tooling, or a hiring market that already knows the stack. Open scale-up bets such as UALink matter for GPU fabrics; they do not make a WSE look like an NVLink domain.
Versus Google TPU, AWS Trainium / Inferentia, and Microsoft Maia: all four are non-GPU AI silicon, but the access models differ. TPU, Trainium, and Maia are hyperscaler-first ASICs co-designed with one cloud's racks and compilers. Cerebras sells a wafer-scale product and cloud/on-prem systems that are not locked inside a single hyperscaler's SKU list. Pick the constraint you actually have: cloud tenancy, wafer-scale simplicity, or CUDA portability.
Who this is for (and not for)
This guide is for infra leads, lab directors, and investors evaluating wafer-scale AI systems alongside NVIDIA GPU clusters and hyperscaler ASICs. If you are sizing Condor Galaxy or Cerebras Cloud capacity for large LLM / multi-modal training or ultra-low-latency inference, you are the reader.
This guide is not for buyers who need a public EC2-style GPU SKU this week with CUDA unchanged. If your constraint is "eight H100-class GPUs on a credit card by Friday," start with hyperscaler GPU instances (or Trainium / TPU when those SKUs fit), and treat Cerebras as a parallel evaluation once your model graph and commercial path are clear.
Honest limits
Name the downside or the guide is marketing with footnotes.
- Supply and packaging are unique. You are not buying a commodity GPU SKU from three OEMs; wafer-scale yield, system integration, and Cerebras's delivery schedule gate capacity.
- Software is not CUDA. PyTorch + XLA + CGC + CSL cover a lot, but exotic kernels, research ops, and hiring depth still lag the NVIDIA ecosystem.
- Workload fit is real. Huge dense or sparsifiable models that benefit from Weight Streaming shine; tiny batches, highly dynamic control flow, or graphs that thrash the compiler may not.
- Cost and availability are commercial, not list-price transparent like every GPU VM family. Cloud quotas, Model Studio pricing, and on-prem deals vary; verify current contracts.
- Vendor peak FLOPs, "× vs GPU," and cluster training-time claims need your own measurement on your model, precision, and sparsity settings.
- Hyperscaler ASICs may still win inside a single cloud's first-party inference path even when wafer-scale looks better on a slide.
FAQ
What is the Cerebras Wafer-Scale Engine?
The Wafer-Scale Engine is Cerebras's AI processor built from a giant intact region of a 300 mm wafer rather than a reticle-limited die. WSE-3 (5 nm, 4 trillion transistors, 900,000 cores, 44 GB on-chip SRAM) powers the CS-3 system.
What are the key WSE-3 specs?
Per Cerebras's WSE-3 datasheet and March 2024 launch: 5 nm, 46,225 mm², 4 trillion transistors, 900,000 AI cores, 44 GB on-chip SRAM, 21 PB/s memory bandwidth, 214 Pb/s fabric bandwidth, and 125 PFLOPS peak AI performance (vendor).
How does CS-3 differ from a GPU server?
CS-3 wraps one WSE-3 wafer with MemoryX external weight memory and SwarmX clustering. Compute and parameter memory scale separately, and Weight Streaming presents large clusters as a single logical device instead of a classic multi-GPU shard map.
What software runs on Cerebras systems?
Cerebras documents CSoft, PyTorch integration via an XLA lazy-tensor path (cerebras.pytorch), the Graph Compiler (CGC), and an SDK with CSL for custom kernels. Hosted inference uses an OpenAI-compatible API. CUDA does not run natively.
What is Condor Galaxy?
Condor Galaxy is the Cerebras and G42 AI supercomputer network. Condor Galaxy 3 is documented as 64 CS-3 systems delivering 8 exaFLOPS in Dallas, offered as a cloud-style single logical supercomputer rather than 64 separately programmed boxes.
How does Cerebras compare to Google TPU, AWS Trainium, and Microsoft Maia?
All are non-GPU AI silicon. TPU, Trainium, and Maia are hyperscaler-first ASICs inside one cloud. Cerebras sells wafer-scale CS-3 systems and cloud/on-prem access that are not locked to a single hyperscaler SKU list. Choose based on tenancy, software path, and workload fit.
When should teams stay on NVIDIA GPUs instead?
Stay on GPUs when CUDA libraries, multi-cloud portability, immediate self-serve rental, or the existing GPU hiring market dominate. Evaluate CS-3 when model scale, sharding complexity, sparsity, or single-device programming outweigh ecosystem lock-in.
Are the "24 trillion parameter" and "256 exaFLOPS" figures independent benchmarks?
No. Those are Cerebras vendor claims tied to MemoryX configurations (up to 1.2 PB) and full SwarmX clusters (up to 2,048 CS-3 systems). Use them as architecture intent, then measure your workload.
Monday's move is simple. If GPU sharding pain or inference latency is already on the invoice, pick one production graph, ask whether it fits Cerebras Cloud inference or Model Studio / Condor Galaxy training, and measure tokens per dollar and engineering hours against your current GPU baseline. Datasheet PB/s can wait until that experiment finishes.



