BTC $85,519.77 -1.28%
ETH $2,701.78 -1.15%
BNB $781.22 -2.78%
XRP $1.50 -1.56%
SOL $120.30 -0.76%
TRX $0.3356 -0.05%
DOGE $0.0945 -1.81%
ADA $0.2671 -0.97%
BCH $315.48 -1.66%
LINK $13.78 -2.78%
HYPE $93.89 +3.59%
AAVE $182.36 +0.71%
SUI $1.19 -4.17%
XLM $0.2143 -4.92%
ZEC $1,341.50 -0.05%
AAPL $332.94 -0.19%
AMZN $252.02 -0.14%
GOOGL $346.98 +0.76%
MSFT $525.69 +1.86%
META $743.34 +2.02%
NVDA $239.34 +1.58%
TSLA $379.51 +1.86%
SNDK $1,693.29 -2.19%
INTC $115.93 -1.73%
SPCX $171.78 +7.04%
MU $1,058.33 -1.86%
AMD $631.27 -1.30%
BTC $85,519.77 -1.28%
ETH $2,701.78 -1.15%
BNB $781.22 -2.78%
XRP $1.50 -1.56%
SOL $120.30 -0.76%
TRX $0.3356 -0.05%
DOGE $0.0945 -1.81%
ADA $0.2671 -0.97%
BCH $315.48 -1.66%
LINK $13.78 -2.78%
HYPE $93.89 +3.59%
AAVE $182.36 +0.71%
SUI $1.19 -4.17%
XLM $0.2143 -4.92%
ZEC $1,341.50 -0.05%
AAPL $332.94 -0.19%
AMZN $252.02 -0.14%
GOOGL $346.98 +0.76%
MSFT $525.69 +1.86%
META $743.34 +2.02%
NVDA $239.34 +1.58%
TSLA $379.51 +1.86%
SNDK $1,693.29 -2.19%
INTC $115.93 -1.73%
SPCX $171.78 +7.04%
MU $1,058.33 -1.86%
AMD $631.27 -1.30%
hot_img

Cerebras releases the fourth generation AI inference system CS-4: performance doubled, power consumption doubled, more flexible deployment

2026-08-19 10:38:29

Cerebras released its fourth-generation AI inference system CS-4 this week, based on the same 5nm WSE-3 wafer, achieving double the performance by doubling the clock frequency and power consumption. A single CS-4 cabinet accommodates 3 wafers (CS-3 has 2), featuring a modular "backpack" design that simplifies manufacturing and deployment, with a TDP of approximately 125 to 135kW. The CS-4 can provide an inference speed of nearly 4000 tokens/second/user, about twice that of the CS-3, and supports decomposed inference with heterogeneous systems such as AMD and AWS Trainium.

Cerebras claims that the CS-4 offers about 2000 times the on-chip memory bandwidth of NVIDIA's Rubin (43PB/s), but the 44GB SRAM capacity remains unchanged, and long-context inference still requires multi-wafer stacking. For example, with the DeepSeek V4 Pro (1.6T parameters), approximately 20 systems are needed for a 1M context window, and about 40 systems are required for 256 concurrent users, corresponding to a CAPEX exceeding 20 million USD. Cerebras is collaborating with clients such as OpenAI and plans to achieve approximately double performance improvements each year, aiming for a 20-fold throughput increase by 2027. The "backpack" cabinet design of the CS-4 will continue into the next-generation "Nexus" platform.

app_icon
ChainCatcher Building the Web3 world with innovations.