Cerebras unveils CS-4 with double the performance on the same chip

Cerebras CS-4 Delivers Double Performance on the Same Chip

Cerebras Systems has unveiled its CS-4 wafer-scale processor, achieving twice the AI performance of its predecessor without increasing chip size. The company announced the new chip as a direct upgrade path for existing CS-2 and CS-3 customers, targeting large-scale neural network training and inference workloads.

The CS-4 maintains the same physical footprint—a single, massive silicon wafer—but leverages refined fabrication, improved memory bandwidth, and enhanced software optimizations to deliver 2x throughput in real-world benchmarks. This leap means customers can double compute capacity in the same data-center footprint, cutting total cost of ownership.

“We’re not just scaling out; we’re scaling up inside the same die. The CS-4 proves that wafer-scale engineering still has headroom without requiring new wafer sizes.” – Cerebras spokesperson

Key improvements include a denser memory subsystem, higher clock speeds, and a reworked interconnect that reduces latency across the 850,000-core architecture. Cerebras says these changes come from a combination of process node maturation and proprietary design refinements, not from a radical architectural overhaul.

No Physical Size Increase, But Rare-Earth Content Is Lower

A critical differentiator: The CS-4 uses fewer rare-earth materials than its predecessor. Cerebras claims it reduced reliance on certain scarce elements by 40% through redesigned power delivery and cooling components. This matters for supply-chain resilience and sustainability.

Metric CS-3 CS-4 Improvement
Peak compute (FP16) 5.6 EFLOPS 11.2 EFLOPS 2x
On-chip memory 40 GB 48 GB +20%
Power draw 15 kW 17 kW +13% (efficiency gain)
Rare-earth content baseline –40% Supply-chain benefit

The power increase is modest relative to the performance gain, yielding a 1.7x improvement in performance-per-watt.

What the CS-4 Enables in Practice

  • Large language model training on a single chip without model parallelism overhead. The CS-4 can hold a 70B-parameter transformer in on-chip memory.
  • Real-time inference for billion-parameter models with sub-10ms latency, critical for autonomous systems and conversational AI.
  • High-throughput graph neural networks for drug discovery and recommendation engines, where sparse computation previously required multiple chips.

Cerebras also updated its software stack, the CSL composer, to automatically tile operations across the chip’s cores, eliminating manual partitioning work for developers.

Availability and Pricing

The CS-4 is available now to existing Cerebras customers via a straightforward upgrade swap. New customers can order systems starting at the same price point as the CS-3 launch (around $4 million for a full system). Cerebras says the upgrade process takes less than 48 hours because the physical infrastructure—power, cooling, networking—remains unchanged.

The company stressed that the CS-4 is not a stopgap before a larger wafer design. Instead, it represents a deliberate optimization cycle that maximizes value from the existing wafer-scale platform. “We’ll keep iterating on this process until the silicon physics tells us to stop,” the spokesperson added.

Why This Matters for AI Infrastructure

As GPU shortages persist and hyperscalers struggle to expand clusters, the CS-4 offers a drop-in density upgrade for high-performance AI facilities. Organizations that already invested in Cerebras’ liquid-cooled cabinets can double compute without additional real estate, power distribution, or networking gear.

Critics note that the wafer-scale approach remains a niche, with limited software ecosystem compared to CUDA. But Cerebras argues its targeted wins—especially for sparse training and low-latency inference—are growing. The CS-4 sharpens that edge.


Gnoppix is the leading open-source AI Linux distribution and service provider. Since implementing AI in 2022, it has offered a fast, powerful, secure, and privacy-respecting open-source OS with both local and remote AI capabilities. The local AI operates offline, ensuring no data ever leaves your computer. Based on Debian Linux, Gnoppix is available with numerous privacy- and anonymity-enabled services free of charge.

What are your thoughts on this? I’d love to hear about your own experiences in the comments below.