BTC $79,840.27 +1.26%
ETH $2,490.54 -0.11%
BNB $708.57 +0.79%
XRP $1.43 +1.49%
SOL $107.29 +5.73%
TRX $0.3385 +1.24%
DOGE $0.0877 +0.94%
ADA $0.2101 +0.07%
BCH $264.17 -0.88%
LINK $11.74 +1.60%
HYPE $83.30 +2.75%
AAVE $126.38 +0.84%
SUI $0.7608 +1.38%
XLM $0.1845 +0.64%
ZEC $790.82 +1.24%
BTC $79,840.27 +1.26%
ETH $2,490.54 -0.11%
BNB $708.57 +0.79%
XRP $1.43 +1.49%
SOL $107.29 +5.73%
TRX $0.3385 +1.24%
DOGE $0.0877 +0.94%
ADA $0.2101 +0.07%
BCH $264.17 -0.88%
LINK $11.74 +1.60%
HYPE $83.30 +2.75%
AAVE $126.38 +0.84%
SUI $0.7608 +1.38%
XLM $0.1845 +0.64%
ZEC $790.82 +1.24%

deployment

All
Article
Flash

hot_img Cerebras releases the fourth generation AI inference system CS-4: performance doubled, power consumption doubled, more flexible deployment

Cerebras released its fourth-generation AI inference system CS-4 this week, based on the same 5nm WSE-3 wafer, achieving double the performance by doubling the clock frequency and power consumption. A single CS-4 cabinet accommodates 3 wafers (CS-3 has 2), featuring a modular "backpack" design that simplifies manufacturing and deployment, with a TDP of approximately 125 to 135kW. The CS-4 can provide an inference speed of nearly 4000 tokens/second/user, about twice that of the CS-3, and supports decomposed inference with heterogeneous systems such as AMD and AWS Trainium.Cerebras claims that the CS-4 offers about 2000 times the on-chip memory bandwidth of NVIDIA's Rubin (43PB/s), but the 44GB SRAM capacity remains unchanged, and long-context inference still requires multi-wafer stacking. For example, with the DeepSeek V4 Pro (1.6T parameters), approximately 20 systems are needed for a 1M context window, and about 40 systems are required for 256 concurrent users, corresponding to a CAPEX exceeding 20 million USD. Cerebras is collaborating with clients such as OpenAI and plans to achieve approximately double performance improvements each year, aiming for a 20-fold throughput increase by 2027. The "backpack" cabinet design of the CS-4 will continue into the next-generation "Nexus" platform.

KPMG's research shows that nearly 30% of corporate executives find it difficult to understand the cost of AI on a pay-per-use basis, and nearly half have delayed deployment

According to KPMG's latest survey report involving 2,145 executives from 20 countries, as technology companies like Anthropic, OpenAI, and GitHub recently shifted some of their AI services from fixed subscription models to usage-based billing, businesses are facing challenges in cost forecasting and management during the scaling of AI deployment.The report indicates that 29% of corporate executives find it difficult to understand and control operational costs when scaling AI deployment, and one-third of executives believe that insufficient understanding of AI economics hinders the deployment of AI entities. Due to costs exceeding expected value, nearly half (about 49%) of corporate organizations have chosen to delay or readjust their AI deployment plans; meanwhile, low-cost, high-fidelity large models are accelerating their impact on corporate AI strategies.In addition, tech giants are increasing capital expenditures to build AI capacity. Amazon plans to spend about $200 billion on capital expenditures this year and is investing $1 billion in its AWS frontline engineering organization to assist customers in adopting AI entities; Microsoft's total capital expenditure is expected to reach $190 billion this year, with $2.5 billion allocated to the new entity Microsoft Frontier Company. KPMG emphasizes that, in addition to cost pressures, accountability in AI governance, employee engagement rules, and the prevention of system "hallucinations" remain core challenges faced by businesses today.

NVIDIA plans to market the Vera AI CPU to Chinese customers, and some cloud providers intend to start testing deployments

Sources say that Nvidia has begun marketing its first standalone Central Processing Unit (CPU) product, Vera, to Chinese customers. This chip is designed for Agentic AI systems and has now entered mass production, marking Nvidia's attempt to further expand its presence in the Chinese market through CPU products.Insiders indicate that some Chinese customers have shown interest in Vera. One large Chinese cloud computing company plans to purchase over 300 servers equipped with dual Vera CPUs for testing and will decide whether to scale up purchases after the testing is completed.Vera is built on the Arm Holdings architecture and is Nvidia's first standalone CPU product. Nvidia previously stated that Vera's performance in AI agent-related computing tasks can reach 1.8 times that of competing products, and it is expected that this product will contribute approximately $20 billion in revenue before the end of the current fiscal year (by the end of January next year).Reports point out that as the focus of the AI industry gradually shifts from model training to inference computing, CPUs and custom chips are gaining more attention. Vera also puts Nvidia in direct competition with Intel and Advanced Micro Devices (AMD), which have long dominated the server CPU market.Insiders have noted that due to strict restrictions imposed by the U.S. on high-end GPU exports, CPUs face relatively fewer regulatory hurdles in the Chinese market compared to GPU products. Currently, some Chinese customers plan to first deploy Vera chips in overseas data centers for testing. Meanwhile, software ecosystem compatibility and the existing domestic AI chip deployment system may still affect the subsequent large-scale adoption of Vera.
app_icon
ChainCatcher Building the Web3 world with innovations.