Bitget App
Trade smarter
Buy cryptoMarketsTradeFuturesEarnSquareMore
Replicating the "DeepSeek Moment"? Wall Street Agrees: Kimi K3 Actually Strengthens the Demand for Computing Power

Replicating the "DeepSeek Moment"? Wall Street Agrees: Kimi K3 Actually Strengthens the Demand for Computing Power

华尔街见闻华尔街见闻2026/07/21 01:46
Show original
By:华尔街见闻

Kimi K3 has once again triggered concerns similar to “DeepSeek-style” worries, causing US semiconductor stocks to fall. However, several Wall Street institutions believe that computing power demand remains strong. This open-source model, with 2.8 trillion parameters and a context window of one million for persistent inference, is expected to drive up demand for KV cache, HBM, storage, and cloud infrastructure. Kimi K3 acts more as a catalyst for the proliferation of AI usage rather than a signal that hardware demand has peaked.

The market is panicking over Kimi K3 as if it's another "DeepSeek Moment 2.0", but this time, Wall Street’s judgment is fundamentally different.

Late at night on July 16, Moonshot AI launched Kimi K3 in Shanghai. This 2.8 trillion parameter open-source model scored 57 on the Artificial Analysis intelligence index, ranking third or fourth globally, tied with Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5. More importantly, on the Frontend Code Arena programming ranking created by the University of California, Berkeley, K3 topped the list with 1679 points, surpassing Claude Fable 5 and GPT-5.6 Sol, making it the first open-source model to outperform all closed-source overseas models on an authoritative programming leaderboard.

On July 17, US semiconductor stocks saw a significant decline. The market’s knee-jerk reaction is understandable—when DeepSeek R1 was released in early 2025, it triggered a sharp fall in computing power stocks. The logic: Chinese models are getting stronger; is it still necessary for US AI companies to pour so much money into computing power? If Chinese models can approach frontier capabilities at lower costs, will Nvidia, HBM, servers, and network equipment demand need to be reconsidered?

However, according to Tradewind Desk, the latest reports from investment banks like UBS, Nomura, Bank of America Merrill Lynch, and Citigroup believe: Kimi K3 is not the end of compute demand—but an accelerator.

Kimi K3 and DeepSeek R1 are not the same kind of shock. R1 highlighted “efficiency” for the market; K3 emphasizes “scale.” 2.8 trillion parameters, 1M token context window, always-on inference, native multimodality, MoE architecture—these features are not a “light-asset” story. They simultaneously raise the pressure on inference, memory, networks, and storage.

Replicating the

How Powerful Is Kimi K3?

Kimi K3 was released by Moonshot AI on July 16, 2026, with plans to open the full model weights on July 27. As an open-source model with 2.8 trillion parameters, multiple institutions have referred to it as the largest open-weight LLM available to date.

The core configurations come in three points:

First, 1M token context window. The model can process longer texts, larger codebases, and more complex enterprise documents and research tasks.

Second, always-on inference. It is not just simple Q&A, but focused on long-chain reasoning and Agent tasks.

Third, native visual capability. K3 handles not only text, but also video, image, game development, front-end design, CAD, and other multimodal tasks.

In terms of architecture, K3 uses Kimi Delta Attention, Attention Residuals, and Stable LatentMoE. The MoE part has 16 of 896 experts activated per token. According to Moonshot AI, the overall scaling efficiency is about 2.5 times higher compared to Kimi K2.

This explains why K3 is not a “cheaper K2.” Nomura’s compiled data shows that K3 input costs $3 per million tokens, cache hit input costs $0.30, and output is $15 per million tokens; according to Artificial Analysis, the average K3 task cost is $0.94. This price is lower than Claude Fable 5 at $2.75 and Claude Opus 4.8 at $1.80, nearly matching GPT-5.6 Sol at $1.04, but significantly higher than GLM-5.2’s $0.32—$0.47, and much higher than DeepSeek V4 Pro at $0.04.

So, K3 is not positioned as the lowest price, but as a model that approaches cutting-edge capability at a much lower cost.

Replicating the

Four Major Investment Banks Signal: This Is Not Demand Destruction

Regarding concerns in the market about another “DeepSeek moment,” Nomura Securities Asia-Pacific technology analyst Duan Bing wrote in the report: “We believe the competition and innovation in the global large-model market will not stop. As we get closer to AGI, generative AI applications will continue to expand for both consumers and enterprises. Cutting-edge AI labs and hyperscale cloud platform companies are likely to further invest to maintain competitive positions—as the law of scale persists, we view this competition as positive for the AI infrastructure value chain.

Citi semiconductor analyst Peter Lee titled his July 19 report directly as “Another Jevons Paradox.” What does the Jevons Paradox mean? In short: when the efficiency of coal steam engines improved, coal consumption actually increased, because more people could afford and utilize it. The same applies to AI models—when high-quality models become cheaper, developers and enterprises deploy more applications and process more tokens, which ultimately drives up compute consumption.

Peter Lee believes that even if K3 sees widespread adoption, general-purpose memory demand for servers such as DDR5 and eSSD will continue to rise. The reason: K3’s inference efficiency is on par with other advanced models, but KV cache usage expands with context length, increasing pressure on memory rather than reducing it.

Bank of America Securities semiconductor analyst Vivek Arya was even more direct in his July 17 report. He noted, the US’s response among advanced AI labs is “not less compute, but more.” If Chinese open-source models keep closing the gap, companies like OpenAI, Anthropic, and Google must sustain differentiation through even larger-scale training, heavier inference, and faster iteration. Arya also highlighted an easily overlooked backdrop: media reported Google’s Gemini 3.5 Pro had been delayed for months compared to plan, with programming performance missing internal goals. Leadership at the forefront “is becoming increasingly difficult to defend.”

UBS analyst Timo Arcuri’s team pointed out in a July 20 report that, while there are parallels between K3 and DeepSeek R1, K3 is much more about scale—it’s the world’s largest open-source model, with 2.8 trillion parameters and a 1 million token context window. The analysts emphasized that open-source models typically require more memory than closed-source cutting-edge models, because of longer context windows. Even after quantization, KV cache demand keeps growing in absolute terms, making open-source model deployment more reliant on HBM and storage.

Replicating the

Who Truly Benefits from This Competition?

Storage: The most direct beneficiary sector. UBS estimates shows cumulative free cash flow (FCF) for storage and memory sectors could reach about 30% of their market cap by 2028—the highest proportion among all sub-sectors—with Micron (MU) alone at 47%. Both Citi and Nomura maintain a buy rating on Samsung Electronics, citing the global memory market’s severe supply constraints. Citi’s Peter Lee noted that Kimi K3’s inference-side memory demand is no less than other leading models; the growing KV cache size will directly boost demand for server DDR5 and enterprise-level SSDs (eSSD). He especially stressed: large-scale Kimi K3 deployment requires “supernode” cluster configurations with over 64 GPUs.

Replicating the

Compute infrastructure: TSMC and Nvidia as primary beneficiaries. Whether it’s the law of scale for training still holding, or surging inference-side token demand, it all ultimately points to more advanced process chips. Nomura reiterates its buy ratings on TSMC, ASE, MediaTek, etc. Nvidia has publicly indicated that modern MoE (Mixture of Experts) models’ inference on GB300 NVL72 achieves a power efficiency ratio up to 25 times better than the previous Hopper architecture. Models like K3 naturally benefit from Nvidia’s latest hardware.

Network: Supernode trend creates structural opportunities. Kimi K3 requires a supernode cluster. Given Chinese compute power is restricted by high-end chip export controls, it relies more on supernode architectures to compensate for single-card performance gaps, driving demand for optical modules, optical chips, and other network-layer suppliers. Nomura is bullish on Zhongji Innolight and Suzhou Innolight.

Cloud platforms: Benefiting from ecosystem aggregation effects. Cloud platforms hosting multiple leading open-source models have stronger bargaining power and are less dependent on a single closed-source supplier. Nomura favors Alibaba (BABA) as the core of the Chinese AI cloud ecosystem, as well as data center operators like GDS and VNET.

How Fast Is the Global Penetration of Chinese AI Models?

This may be the most underestimated data point in the narrative.

According to OpenRouter, the open API gateway, the share of token usage by Chinese AI models among global developer traffic has grown from less than 2% a year ago to over 45% as of now. Data from Bank of America Merrill Lynch confirms the overall acceleration of AI penetration: currently about 55% of US enterprises have subscribed to AI models, platforms, or tools. Anthropic’s enterprise adoption rate is 42%, OpenAI’s is 40%. For top AI consumers (top 1% of enterprise users), per-employee monthly AI spend has reached $4,833.

The market is diverging. On one hand are Chinese open-source models represented by DeepSeek and Kimi K3, covering both budget and mid/high-value markets; on the other, top US frontier models are focusing on more complex workloads (like scientific computing) to maintain technology and pricing premiums. Nomura’s view is: the leading model players on both sides of the US and China will benefit—as long as each continues to push the technology frontier.

Replicating the

Replicating the

The Real Change Brought by K3 Is Competitive Tempo, Not a Single Company’s Story

After K3 launched, the market’s first reaction was to compare it with DeepSeek R1. This comparison is useful but shouldn’t stop at “are we going to over-invest in AI hardware again.”

DeepSeek made the market reassess training efficiency. K3 showed another angle: open-source models can also push scale, long context, Agent, and multimodality into the cutting edge.

This will force US frontier labs to keep investing and allow Chinese models to further expand in the global developer ecosystem. Closed-source leading models retain technology and price premium, open-source models cover broader price points and deployment scenarios, cloud providers offer model distribution and enterprise landing, while hardware suppliers bear training and inference loads.

In the short term, trading may fluctuate due to “DeepSeek memories.” In the medium term, as long as token usage continues growing, long context and Agent use cases keep spreading, compute, HBM, storage, network, and IDC will remain unavoidable cost items.

This is why many institutions have reached a similar conclusion post-K3: more powerful open-source models are not the end point for AI infrastructure demand—on the contrary, they may be the entry to the next round of demand diffusion.

However, Bank of America Merrill Lynch also clearly flagged a tail risk: “If efficiency gains outpace workload growth, we may see some pullback in infrastructure buildout.” In other words, if models become cheaper but usage doesn’t expand accordingly, the logic for compute demand growth will weaken.

0
0

Disclaimer: The content of this article solely reflects the author's opinion and does not represent the platform in any capacity. This article is not intended to serve as a reference for making investment decisions.

Understand the market, then trade.
Bitget offers one-stop trading for cryptocurrencies, stocks, and gold.
Trade now!

You may also like

Turmoil Intensifies Within OpenAI: Ongoing Restructuring, Employee Exhaustion, Continued Management Resignations, IPO Plans Quietly Delayed

OpenAI has undergone nearly five restructurings this year, with a continued wave of executive departures. Chief Revenue Officer Denise Dresser became the latest to leave this week. The company also disbanded its safety team responsible for assessing the "catastrophic risks" of AI models, raising internal concerns. Although revenue has increased to about $40 billion, OpenAI has been surpassed by Anthropic. Employees generally feel exhausted, and the IPO plans have quietly been postponed until next year.

华尔街见闻2026/08/16 08:31