BTC $77,570.53 -1.44%
ETH $2,420.80 -2.14%
BNB $687.99 -0.25%
XRP $1.35 -2.63%
SOL $100.02 -3.13%
TRX $0.3227 -2.64%
DOGE $0.0816 -2.20%
ADA $0.1971 -1.96%
BCH $247.62 -1.27%
LINK $11.25 -2.04%
HYPE $82.94 -1.22%
AAVE $131.84 +4.06%
SUI $0.7260 -0.92%
XLM $0.1751 -1.75%
ZEC $835.07 -2.42%
BTC $77,570.53 -1.44%
ETH $2,420.80 -2.14%
BNB $687.99 -0.25%
XRP $1.35 -2.63%
SOL $100.02 -3.13%
TRX $0.3227 -2.64%
DOGE $0.0816 -2.20%
ADA $0.1971 -1.96%
BCH $247.62 -1.27%
LINK $11.25 -2.04%
HYPE $82.94 -1.22%
AAVE $131.84 +4.06%
SUI $0.7260 -0.92%
XLM $0.1751 -1.75%
ZEC $835.07 -2.42%

server

All
Article
Flash

first_img Google claims that the cost of AI server memory has exceeded 75%, promoting a dual-track strategy for software and hardware

The SEMICON Taiwan 2026 Memory Summit took place on the 1st, where Nikhil Cherian, Senior Director of Supply Chain Infrastructure at Google Cloud under Alphabet, pointed out that with the popularity of multimodal and mixed expert architectures, AI computation has shifted from being power-limited to memory-limited, with high-performance memory accounting for over 75% of the cost of AI server hardware bill of materials. In the face of capacity, bandwidth, and power consumption bottlenecks, Google is breaking through the AI memory bottleneck through a dual-track strategy of hardware offloading for inference and training, and lossless quantization software algorithms.Google adopts an offloading strategy in hardware architecture, launching TPU 8i for low-latency inference and TPU 8t specialized for large-scale training. The TPU 8i is equipped with 288 GB of high-bandwidth memory, with SRAM capacity on the chip increased threefold to 384 MiB, placing dynamic conversation states and key-value caches on the chip itself to achieve zero chip-off latency. The TPU 8t forms a super-large computing cluster with 9600 chips, achieving a shared pool of HBM at a scale of 2 PB, eliminating chip-off data transfer bottlenecks, along with TPU Direct Storage technology.Google has developed the training-free TurboQuant lossless quantization algorithm, compressing the key-value cache of large models from 32 bits to 3 bits, reducing memory usage by six times without loss of accuracy, resulting in an eightfold acceleration in attention computation, and integrating old-generation DRAM technology to extend the lifecycle of components.

hot_img Intel and AMD are seeking to sign long-term CPU supply agreements with Chinese server customers, with some product prices increasing by over 40% within the year

According to Reuters, citing informed sources, due to the surge in demand for AI data centers leading to a tight supply of server CPUs, Intel and AMD are negotiating long-term procurement commitments with Chinese server customers. Agreements typically lock in a year's supply, with some discussions extending to two years or longer. Driven by the construction of AI computing power, CPU demand has expanded from AI accelerators to mainstream processors, with some server CPU products in the Chinese market experiencing a cumulative price increase of over 40% this year, with month-on-month increases exceeding 10% at times.Intel CEO Lip-Bu Tan stated in April that Xeon server CPU demand "continues to exceed supply," and the company has signed multiple long-term contracts in the first quarter, including a multi-year agreement with Google. AMD has raised its forecast for the server CPU market to exceed $120 billion by 2030. The report notes that China, as one of the largest server markets in the world, is intensifying competition for Intel and AMD processors due to the rapid construction of data centers and AI computing clusters, even as Chinese buyers face U.S. export restrictions on advanced AI GPUs. Intel is set to announce its quarterly results on Thursday, with the CPU shortage expected to become a market focus.
app_icon
ChainCatcher Building the Web3 world with innovations.