Reproducing the "DeepSeek Moment"? Wall Street unanimously states: Kimi K3 instead strengthens the demand for computing power
Author: Long Yue
The market panicked at Kimi K3 as the "DeepSeek Moment 2.0," but this time, Wall Street's judgment is completely different.
On the night of July 16, Moonlight released Kimi K3 in Shanghai. This 2.8 trillion parameter open-source model scored 57 on the Artificial Analysis Intelligence Index, ranking third to fourth globally, on par with Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5. More importantly, on the Frontend Code Arena programming leaderboard created by the University of California, Berkeley, K3 topped the list with a score of 1679, surpassing Claude Fable 5 and GPT-5.6 Sol, becoming the first open-source model to surpass all overseas closed-source models on an authoritative programming leaderboard.
On July 17, the U.S. semiconductor sector saw a significant decline. The market's reflex is easy to understand—at the beginning of 2025, the release of DeepSeek R1 triggered a sharp drop in computing power stocks, with the logic being: As Chinese models become stronger, do American AI companies still need to spend so much on computing power? If Chinese models can approach cutting-edge capabilities at a lower cost, will the demand for Nvidia, HBM, servers, and network equipment be reassessed?
However, according to news from the Wind Trading Desk, the latest research reports from investment banks such as UBS, Nomura, Bank of America Merrill Lynch, and Citigroup believe: Kimi K3 is not the end of computing power demand, but an accelerator.
Kimi K3 and DeepSeek R1 are not the same kind of impact. R1 made the market see "efficiency"; K3 emphasizes "scale." With 2.8 trillion parameters, a 1M token context, always-on inference, native multimodal capabilities, and MoE architecture, these features are not a story of light assets. They will raise the pressure on inference, memory, network, and storage together.

How strong is Kimi K3?
Kimi K3 was released by Moonlight on July 16, 2026, with the complete model weights scheduled to be opened on July 27. It is a 2.8 trillion parameter open-source large model, referred to by multiple institutions as the largest open-source weight LLM currently available.
The core configuration includes three points:
First, a 1M token context window. The model can handle longer texts, longer codebases, more complex business profiles, and research tasks.
Second, always-on inference. It is not just simple Q&A but is aimed at long-chain reasoning and agent tasks.
Third, native visual capabilities. K3 not only processes text but also targets multimodal tasks such as video, images, game development, front-end design, and CAD.
In terms of architecture, K3 uses Kimi Delta Attention, Attention Residuals, and Stable LatentMoE. The MoE part activates 16 experts for each token out of 896 experts. Moonlight states that compared to Kimi K2, the overall scaling efficiency has improved by about 2.5 times.
This explains why K3 is not "a cheaper K2." According to pricing compiled by Nomura, the input price for K3 is $3 per million tokens, cache hit input is $0.30, and output is $15 per million tokens; according to Artificial Analysis, the cost per task for K3 is about $0.94. This price is lower than Claude Fable 5 at about $2.75 and Claude Opus 4.8 at about $1.80, close to GPT-5.6 Sol's $1.04, but significantly higher than GLM-5.2's $0.32—$0.47 and far above DeepSeek V4 Pro's $0.04.
Therefore, K3's positioning is not the lowest price, but rather the ability to approach cutting-edge models at a lower price.

Four investment banks intensively set the tone: this is not a demand weakening
In response to market concerns about the "DeepSeek Moment" impact, Nomura's Asia-Pacific technology team analyst Duan Bing wrote in a report: "We believe that competition and innovation in the global large model market will not stop. As we get closer to Artificial General Intelligence (AGI), the application of generative AI on both consumer and enterprise sides will continue to expand. Leading AI labs and hyperscale cloud platform companies may continue to invest during this phase to maintain their competitive positions—we interpret this competition as a positive for the AI infrastructure value chain."
Citigroup semiconductor analyst Peter Lee directly titled his July 19 report "Another Jevons Paradox." What does Jevons Paradox mean? Simply put: the efficiency improvement of coal steam engines leads to greater coal consumption because more people can afford it and more scenarios can use it. The same goes for AI models—when high-quality models become cheaper, developers and enterprises will deploy more applications and process more tokens, ultimately increasing computing power consumption.
Peter Lee believes that even if K3 is widely used, the demand for general memory such as server DDR5 and eSSD will still increase. The reason is that K3's inference efficiency is comparable to other cutting-edge models, but the KV cache usage will expand with the growth of context, increasing the pressure on memory rather than decreasing it.
Bank of America Merrill Lynch semiconductor analyst Vivek Arya expressed more directly in his July 17 report. He believes that the response from leading U.S. AI labs is "not less computing power, but more." If Chinese open-source models continue to approach, OpenAI, Anthropic, and Google must maintain differentiation through larger-scale training, heavier inference, and faster iterations. Arya also mentioned a background that is easily overlooked: media reports indicate that Google's Gemini 3.5 Pro has been delayed for several months compared to the plan, and programming performance has not reached internal targets, making it "increasingly difficult to defend" its leading position.
UBS analyst Timo Arcuri's team pointed out in a July 20 report that there are indeed parallels between K3 and DeepSeek R1, but K3 is more about scale—it is the world's largest open-source model, with 2.8 trillion parameters and a 1 million token context window. Analysts emphasize that open-source models generally consume more memory than closed-source cutting-edge models because the context window is longer, and the demand for KV cache continues to grow in absolute terms even after quantization, making the deployment of open-source models more reliant on HBM and storage.
Who truly benefits in this competition?
Storage: the most direct beneficiary sector. UBS estimates that the cumulative free cash flow (FCF) of the storage and memory sector is expected to reach about 30% of market value by 2028, the highest proportion among all sub-sectors—Micron (MU) alone accounts for as much as 47%. Both Citigroup and Nomura maintain buy ratings on Samsung Electronics, citing that the global memory market is in a state of extreme supply tightness. Citigroup analyst Peter Lee pointed out that Kimi K3's memory demand on the inference side is no less than that of other cutting-edge models, and the expansion of KV cache will directly drive the demand for server DDR5 and enterprise-level solid-state drives (eSSD). He specifically noted that the large-scale deployment of Kimi K3 requires "super-node" cluster configurations with more than 64 GPUs.

Computing power infrastructure: TSMC and Nvidia are the primary beneficiaries. Whether the scaling law on the training side remains effective or the growth in token demand on the inference side, it ultimately points to more demand for advanced process chips. Nomura reiterates buy ratings on TSMC, ASE, MediaTek, and others. Nvidia has publicly stated that the inference performance-to-power ratio of modern MoE (Mixture of Experts) models on the GB300 NVL72 has improved by up to 25 times compared to the previous generation Hopper architecture. Models like K3 are naturally beneficiaries of Nvidia's latest hardware.
Network: the super-node trend creates structural opportunities. Kimi K3 requires super-node clusters, and China's local computing power is limited by high-end chip export controls, making it even more reliant on super-node architectures to bridge the performance gap of single cards, which drives demand for network layer suppliers such as optical modules and optical chips. Nomura is optimistic about Zhongji Xuchuang and Suzhou Xuchuang.
Cloud platforms: benefit from ecological aggregation effects. Cloud platforms that host various cutting-edge open-source models have stronger bargaining power and do not rely on a single closed-source model supplier. Nomura is optimistic about Alibaba (BABA) as the core of China's AI cloud ecosystem, as well as data center operators like GDS and VNET.
How fast is the global penetration of Chinese AI models?
This may be the most easily underestimated data point in the entire narrative.
According to statistics from the open API gateway OpenRouter, the token usage of Chinese AI models accounted for less than 2% of global developer traffic a year ago, and now it has exceeded 45%. Data from Bank of America Merrill Lynch corroborates the acceleration of overall AI penetration: currently, about 55% of U.S. enterprises have subscribed to AI models, platforms, or tools, with Anthropic's enterprise adoption rate reaching 42% and OpenAI at 40%. Top AI consumers (the top 1% of enterprise users) have reached an AI spending of $4,833 per employee per month.
The market is diversifying. On one hand, there are Chinese open-source models represented by DeepSeek and Kimi K3, covering the economy and mid-to-high-end cost-performance markets; on the other hand, top U.S. cutting-edge models are focusing on more complex workloads (such as scientific computing), maintaining technological and pricing premiums. Nomura's judgment is that leading large model players on both the U.S. and Chinese sides will benefit—provided that each can continue to stay at the forefront of the technology curve.

What K3 truly changes is the rhythm of competition, not the story of a single company
After the release of K3, the market's first reaction was still to compare it with DeepSeek R1. This comparison is useful, but it cannot stop at the level of "is it going to crush AI hardware again."
DeepSeek made the market reassess training efficiency. K3 shows the market another thing: open-source models can also push scale, long context, agents, and multimodality to the cutting edge.
This will force U.S. leading labs to continue investing and allow Chinese models to continue expanding in the global developer ecosystem. Closed-source top models retain technological and price premiums, while open-source models cover more price ranges and deployment scenarios, cloud vendors provide model distribution and enterprise landing, and the hardware chain bears the training and inference pressure.
In the short term, trading may fluctuate due to the "DeepSeek memory." In the medium term, as long as token usage continues to grow, long contexts and agents continue to spread, computing power, HBM, storage, network, and IDC will remain unavoidable cost items.
This is also why multiple institutions have given similar conclusions after K3: stronger open-source models are not the endpoint of AI infrastructure demand, but may instead be the entry point for the next round of demand diffusion.
However, Bank of America Merrill Lynch also clearly left a tail risk: "If the speed of efficiency gains exceeds the growth of workloads, we may see some pullback in infrastructure construction." In other words, if models become cheaper and cheaper, but usage does not significantly expand in sync, the logic of computing power demand growth will be discounted.












