BTC $78,755.62 -0.83%
ETH $2,488.17 -0.06%
BNB $756.88 +1.79%
XRP $1.40 -0.00%
SOL $103.57 -1.30%
TRX $0.3382 +0.41%
DOGE $0.0905 +1.12%
ADA $0.2200 +0.53%
BCH $257.98 +0.50%
LINK $12.70 -4.52%
HYPE $84.14 -3.94%
AAVE $131.35 -1.79%
SUI $0.8264 +1.47%
XLM $0.1911 +0.03%
ZEC $1,149.24 -4.50%
BTC $78,755.62 -0.83%
ETH $2,488.17 -0.06%
BNB $756.88 +1.79%
XRP $1.40 -0.00%
SOL $103.57 -1.30%
TRX $0.3382 +0.41%
DOGE $0.0905 +1.12%
ADA $0.2200 +0.53%
BCH $257.98 +0.50%
LINK $12.70 -4.52%
HYPE $84.14 -3.94%
AAVE $131.35 -1.79%
SUI $0.8264 +1.47%
XLM $0.1911 +0.03%
ZEC $1,149.24 -4.50%

mia

All
Article
Flash

SemiAnalysis releases Neocloud security deep report: Infrastructure configuration errors are shocking, and cross-tenant RCE could affect banks, telecommunications, and even a country's intelligence agency

The semiconductor and AI independent research organization SemiAnalysis released a deep security report on Neocloud (new cloud), revealing various cross-tenant security vulnerabilities discovered during the ClusterMAX 3 testing period. In a four-month test covering 25 vendors and 32 clusters, the team achieved multiple instances of cross-tenant remote code execution (RCE) solely by exploiting publicly known vulnerabilities and basic configuration checks. Affected entities included banks, telecommunications companies, universities, research institutions, AI laboratories, and even a national intelligence agency.Typical issues included: shared Kubernetes control plane leading to tenant metadata visibility, container escape, exposure of BMC/IPMI management networks, incorrect configuration of InfiniBand security keys (P_Key, SA_Key, M_Key), unfortified default trust mode of BlueField DPU, Grafana monitoring dashboards using god-level API keys, and lack of VXLAN isolation in front-end networks. The report specifically pointed out a cascading vulnerability case: a misconfiguration of shared vCluster combined with software versions being two years out of date ultimately completed the POC verification of cross-tenant RCE within an afternoon.Notably, the report questioned the mainstream narrative that "AI has fundamentally changed the pace of cybersecurity": statistics on CVEs for NVIDIA GPU drivers, CUDA, PyTorch, Kubernetes, Docker, and the Linux kernel showed that there was no significant increase in vulnerabilities after the popularization of AI coding models, with most data supporting the "no change hypothesis." The report also detailed the incident where an OpenAI-trained agent attacked Hugging Face, where the AI agent achieved cluster-level privilege escalation through a message board established via Artifactory, which went undetected from May to July. While building POC verification for existing vulnerabilities, the team found that Claude Fable and GPT-5.6 Sol frequently rejected security-related requests, ultimately relying on open-source models such as DeepSeek V4, Kimi K3, and GLM-5.2 to complete the task.SemiAnalysis stated that the core issue in the Neocloud (new cloud) industry is not the new risks brought by AI, but rather the long-term absence of basic patch management, tenant isolation, and security design. They recommended that vendors establish automated security announcement monitoring systems and rectify single points of failure that could expose all users' architectural patterns.

first_img SemiAnalysis: HBF non-HBM alternative, cost and heat dissipation still have uncertainties

P Equity Research and SemiAnalysis researcher Nick Doyle and others discussed high bandwidth flash (HBF) in X Space. Nick stated that it is still too early to determine how much the cost premium of HBF relative to HBM can shrink; existing data mostly comes from vendor claims, such as Sandisk stating that the cost per bit is about one-eighth that of HBM. Yields, testing, and other factors will improve with scale, but structural costs such as TSV, stacking, and pSLC mode will always exist, and durability is a key unknown; if wear exceeds expectations, costs will rise.The application scenarios for HBF are narrow, targeting only AI inference, especially low batch and long context MoE models, and it is not a substitute for HBM. The actual bandwidth target is about 1.6 TB/s, which is at the HBM3E level, suitable for sequential reads to load model weights, more aligned with the capacity needs of a small number of GPUs in local or private enterprises, rather than ultra-large-scale bandwidth scenarios. Heat dissipation reliability has not yet been resolved; flash memory will degrade faster at high temperatures next to GPUs, and mitigation measures such as UCIe separation and daily refresh have yet to be validated.In terms of manufacturing, Sandisk/Kioxia has experience with 3D NAND, and SK Hynix complements HBM-style stacking capabilities, but mass production is still to be confirmed. Overall, storage is shifting towards a specialized layered market, with NAND shortages expected to continue until 2028, and HBF may further impact supply and demand.

SemiAnalysis Founder: By 2028, most of the new AI computing power will belong to two companies

In the latest podcast, SemiAnalysis founder Dylan Patel predicts that by 2028, OpenAI and Anthropic may account for 70% to 80% of the world's new AI computing power, with the total computing power scale potentially exceeding 100GW. Patel stated that the two companies currently account for about 30% of the world's annual new computing power, and this proportion is still rising rapidly.Patel pointed out that the business model of leading AI laboratories is changing, with a significant increase in the efficiency of AI computing power output. Currently, Anthropic's revenue per megawatt of computing power has reached about $50 million and may further rise to $100 million. This allows OpenAI and Anthropic to procure or lease computing power at high prices ranging from $25 million to $50 million per megawatt.Patel expects that global AI-related capital expenditures will reach about $11 trillion from 2024 to 2029, with over $5 trillion needing to be financed through debt. Due to the potential return on investment of AI infrastructure being far higher than that of traditional industries, tech giants may accept higher financing costs, thereby pushing up overall credit rates and squeezing the valuations of traditional assets and highly leveraged economies. Additionally, Patel believes that the new computing power may not primarily be used for providing model inference services externally, but may instead flow more towards internal research and development and self-improvement of models within AI laboratories.

hot_img SemiAnalysis: SpaceX may complete a 10GW data center by 2027, with expected inference revenue reaching $300 billion

Research institution SemiAnalysis released an analysis stating that SpaceX is expected to build approximately 10GW of AI data center capacity by the end of 2027. If 50% of this is used for inference services, with annual revenue exceeding $10 billion per GW, the annualized revenue could reach $300 billion. SpaceX CEO Elon Musk stated in the first earnings report that a "conservative estimate" suggests an additional 6-8GW will be added in 2027, with the actual figure possibly exceeding 10GW.SemiAnalysis's inference simulator shows that when running on the GB300 cluster at current startup cloud prices (about $3/GPU hour), leading model companies like OpenAI and Anthropic could generate annual inference revenue exceeding $10 billion per GW, with annual costs around $12 billion. Microsoft, with full access to OpenAI models and without bearing training costs, can also capture revenue of the same scale. The analysis points out that Microsoft has signed contracts for 10GW of data centers (total value exceeding $300 billion) since 2026, with a 90-day cancellation clause, significantly reducing signing risks.Regarding SpaceX's construction progress, SemiAnalysis believes Musk will significantly shorten the construction cycle by using onsite gas power generation, bypassing large power transformers, parallel construction, and shortening the debugging process. The Southaven plant in Tennessee expanded from 27 turbines (approximately 495MW) in February to 69 turbines (1.7GW) in July, and the "MiniHard" project can be completed in about 5 months with 450-500MW. However, the 10GW target still faces multiple challenges such as land approvals, gas supply, and equipment delivery. This analysis is based on model simulations, and actual implementation still carries uncertainties.
app_icon
ChainCatcher Building the Web3 world with innovations.