Home / News

Cerebras Unveils CS-4 AI System, Claims 30x Speed Boost Over Nvidia GPUs

• Cerebras Systems has launched the CS-4, a rack-scale AI system it asserts delivers 30 times more tokens per second per user than leading GPUs. • The system is powered by three massive WSE-3 Turbo processors, each containing 4 trillion transistors, making them the largest AI chips ever built. • Cerebras specializes in AI inference using fast SRAM memory, an approach enabled by its wafer-scale design where data travels shorter distances. • Despite the launch, Cerebras shares have fallen over 35% from their debut peak, with recent earnings showing a significant per-share loss.

In a bold challenge to Nvidia's market dominance, AI chip specialist Cerebras Systems has unveiled its latest high-performance computing system, the Cerebras CS-4. The company claims the new rack-scale offering represents the fastest AI accelerator in the industry, specifically engineered for running—or inferencing—massive AI models. According to Cerebras, the CS-4 delivers a staggering 30 times the tokens processed per second per user compared to systems built on traditional graphics processing units (GPUs). The core of the CS-4's performance lies in its unique hardware architecture. The system integrates three of Cerebras’s proprietary Wafer-Scale Engine 3 Turbo (WSE-3 Turbo) processors. Each of these single-chip processors is the size of a dinner plate and packs an unprecedented 4 trillion transistors, cementing their status as the largest semiconductors ever built for artificial intelligence. Crucially, this wafer-scale design allows Cerebras to utilize Static Random-Access Memory (SRAM), which is significantly faster than the Dynamic RAM (DRAM) common in conventional systems. While SRAM is more complex and costly, the large physical space of the WSE chip makes its implementation feasible. Furthermore, the monolithic processor design means data travels minimal on-chip distances, eliminating bottlenecks associated with multi-chip systems that must shuttle information between discrete components. "Historically, fast inference meant using smaller and less capable models. Cerebras CS-4 delivers industry-leading speeds on the largest frontier models, fundamentally changing the paradigm," said Cerebras CEO Andrew Feldman. He added that every design aspect was optimized for maximum speed and throughput, claiming "AI is so fast that it fundamentally reshapes product experiences." Beyond selling hardware, Cerebras also offers cloud-based AI services powered by its chips on a rental basis. However, this significant technological announcement contrasts sharply with the company's recent financial performance. Since going public in May at $185 per share and briefly soaring to $350, Cerebras stock has retreated sharply, trading around $218 recently—a decline of more than 35%. Its second-quarter results revealed a loss per share of $2.98, a stark reversal from a profit of $1.91 in the prior-year period, and its subsequent third-quarter guidance failed to galvanize investor sentiment. The launch of the CS-4 now positions Cerebras' cutting-edge engineering against the formidable commercial headwinds it faces in the competitive AI hardware landscape.