Bitget App
Trade smarter
MarketsTradeFuturesEarnAISquareMore
Latest SemiAnalysis Interview: Storage Still Has Room to Double, CPU Is Just a Supporting Role, CPO Deployment Delayed Until 2029

Latest SemiAnalysis Interview: Storage Still Has Room to Double, CPU Is Just a Supporting Role, CPO Deployment Delayed Until 2029

追风交易台追风交易台2026/08/21 13:23
Show original
By:追风交易台

Every layer of AI infrastructure is under simultaneous pressure, presenting both opportunities and misjudgments. Dylan Patel, founder of SemiAnalysis, recently gave a podcast interview systematically outlining the core dynamics and investment logic of the current AI infrastructure stack. His insights cover the economics of AI models, the memory supercycle, the repricing of CPUs, CPO timeline risks, and structural opportunities in data center energy supply.

In response to widespread market skepticism regarding AI investment returns (ROI), Dylan revealed that Anthropic achieved positive free cash flow in Q2 this year, with annualized recurring revenue exceeding $50 billion,and a gross margin over 70%. On the enterprise side, the productivity leap brought by the latest AI models far exceeds the increase in compute costs, leading companies to cut other software expenses to maintain their explosive AI budgets.

At the hardware evolution level, the paradigm shift toward inference models is reshaping market demand. Dylan emphasizes thatthere is a structural shortage in memory lasting several years, still leaving 2-3x upside potential;meanwhile, although agents and reinforcement learning have pushed up CPU demand, the market is overpricing CPUs, and much of the CPU growth comes from historical backfill, with their absolute value in AI servers still far lower than GPU.

Dylan believes that the highly anticipated large-scale implementation of Co-Packaged Optics (CPO)has been clearly postponed to late 2028 or 2029,unexpectedly extending the boom period for copper cable connectors. Additionally, the limitations in grid transmission and distribution are forcing data centers to turn to “behind the meter” power (self-built power), creating massive industrial energy and power conversion supply chain investment opportunities beyond traditional chip investment.

Latest SemiAnalysis Interview: Storage Still Has Room to Double, CPU Is Just a Supporting Role, CPO Deployment Delayed Until 2029 image 0

Anthropic Leads in Self-Financing, AI Demand Narrative Begins to Materialize

As for the market’s doubts around the ROI of AI companies, Dylan Patel responded with concrete data.

"Anthropic already achieved positive free cash flow in Q2, made a profit in April, in May, and June looks to be the same." He noted that Anthropic has surpassed $50 billion in annualized recurring revenue, with gross margins above 70%. OpenAI’s revenue is also growing rapidly as Codex adoption increases.

SemiAnalysis’s own spending trajectory corroborates this trend. Last November, the company’s 90-person team had annualized AI expenses of less than $100,000; by the end of January this year, due to large-scale deployment of Claude Code, that number soared to $4 million; currently, it has reached $11 million, with the peak weekly annualized rate once hitting $14 million. "Employee labor costs plus AI costs—AI has already surpassed one-third, and it will likely reach half by year-end."

He also pointed out that newer, stronger models do not necessarily cost more in practice. Older models might require 100,000 tokens and 10 interactions to complete a task, while new models may need only 25,000 tokens and a single interaction. "Each time the model upgrades from 4.6 Opus to 4.7 Opus, our spending drops for a week, then spikes—because people realize things that weren’t possible before are now doable." He believes this is one of Anthropic’s main advantages in competition with OpenAI:higher token efficiency and lower overall user costs.

Memory: A Structural Shortage, Not a Normal Cycle

Of all hardware categories, Dylan Patel is most resolute about memory.

"This is not a short-term shortage; it’s a multi-year structural shortfall." He notes that memory capacity only grows 20-30% per year, while AI demand is doubling and redoubling, with the gap steadily widening.

The driving logic here comes from inference model pressure on KV cache. Traditional conversational inference context length was in the thousands of tokens with limited KV cache usage; but with inference models such as o1, context length has exploded, KV cache is swelling rapidly, and memory is the most direct beneficiary. SemiAnalysis published a report in December 2024 specifically highlighting this trend.

Rigid constraints on the supply side will force downstream markets to reallocate limited memory resources. He predicts that consumer electronics with less price elasticity will come under pressure first—mid- and low-end smartphone shipments are down 40%, and iPhone and MacBook prices will rise next year. "Memory prices will keep rising, consumer electronics will get squeezed down to a new level, and only when AI gets all the memory it needs will it be enough." He adds that while a cyclical downturn will eventually come, "from trough to trough, long-term growth is beyond doubt."

CPU: Limited Backfill Opportunity, Avoid Overextending Projections

The CPU is a new star in this year’s AI infrastructure narrative, but Dylan Patel holds a clear cautionary stance.

The logic for CPU demand recovery is clear: reinforcement learning requires massive CPU usage for environment validation (unit testing, simulation, etc.); agent inference requires models to frequently call tools and interact with the real world, which depend heavily on CPU processing power. At the same time, after years of large-scale AI chip shipments with severely inadequate CPU pairing, the market is now in a concentrated restocking stage. ARM, Intel, and AMD have benefited, and Nvidia’s Vera CPU has provided a $2 billion revenue forecast.

"But I want to give an important warning: there are massive backfill effects here." He states that once the historic gap is filled, only incremental demand remains, and growth returns to normal. In absolute dollar terms, a Blackwell chip costs around $50,000, a CPU about $5,000—even as CPU expansion increases, the dollar amount is still far lower than AI accelerators."Memory and AI accelerator chips are where the real money is; CPUs are being repriced after being undervalued. Pricing is more reasonable now, but they won’t grow faster than AI chips indefinitely."

Optical Interconnect: Optimistic Long Term, Cautious on CPO in the Short to Medium Term

Networking and optical interconnects are another area of high market enthusiasm, but Dylan Patel is cautious about the large-scale rollout of CPO (Co-Packaged Optics).

"I think true mass production of CPO will only happen between late 2028 and 2029." He points out that current manufacturing yields, chip design, and supply chain maturity have not reached mass deployment standards. Nvidia’s Rubin and its successor Feynman will still use all-copper solutions, so CPO on the GPU side needs several more chip generations.

He revealed that last week, SemiAnalysis released a report to institutional subscribers that, in the mid-term, he is more optimistic about copper cables and non-CPO optical solutions and remains cautious on CPO. Some design changes in downstream chips (such as Rubin Ultra’s Kyber removing 800V design) further delay CPO’s arrival. Companies like Amphenol (copper cable connectors) will benefit more than expected because of this.

"CPO will happen in the long run, and copper will be replaced eventually, but the timeline has shifted—short to mid-term, copper still has significant opportunity."

Electricity: Self-Built Power Will Become Mainstream, Diverse Paths for Innovation

Data center power supply is fast becoming the hardest physical constraint on AI growth. According to Dylan Patel, data centers will add 20GW of new power consumption this year, 30GW next year, and 50GW the year after—a nearly explosive rate of growth.

He breaks down energy issues into three dimensions:transmission, generation, and conversion. Transmission is the toughest challenge, involving regulatory policy, local power company monopolies, and cost-sharing mechanisms—unlikely to change anytime soon. Generation and conversion offer much broader opportunities.

He predicts that in the next few years, half of new data center power consumption will come from “behind the meter” power—that is, self-built corporate power, not the public grid. The mainstream solution is Combined Cycle Gas Turbine (CCGT) from GE Vernova, Mitsubishi, Siemens, etc.; there are also reciprocating engines, industrial gas turbines, and even unconventional solutions retrofitted from marine, train, and truck engines. "It may sound rough, but it works, and it’s already being used."

In the longer term, he believes that in about two years, the total cost of solar plus storage will be lower than that of gas power; further out, orbital data centers—computing chips deployed in space where solar panels don’t need to penetrate the atmosphere—could provide much higher energy density than on Earth and require no storage.

Conversion also holds significant investment opportunities: from IGBT, silicon carbide, gallium nitride MOSFETs to solid-state transformers, UPS, and supercapacitors, the entire voltage conversion chain is evolving rapidly. SemiAnalysis’ largest research team is now not semiconductors, but their so-called “DEI” (Data Center, Energy & Industrial) group, tracking deployment dynamics of every data center and power plant worldwide.

Below is the full interview transcript:

Host:Welcome back to The Next Big Thing podcast. I’m the host, joined today by colleague Klay Hyman and our new partner, Dylan Patel, founder of the research institution SemiAnalysis. It’s a special day—we’ll be digging deep into today’s AI infrastructure landscape with Dylan. I’ve been following his content for a long time; he appears frequently on podcasts, and SemiAnalysis’s newsletter has always been must-read material for me. Recently, they published a deep-dive report on orbital data centers—if you’re interested, check it out; it’s incredibly detailed.

Dylan, I’d love to hear about how SemiAnalysis got started. I know that recently everyone on Substack talks about the company’s revenue and success, but I think people often focus on the success in the present and forget the journey and effort behind it.

Origins of SemiAnalysis

Dylan Patel: The origin of SemiAnalysis actually comes from posting online—some would call it "shitposting." Thinking back, my first posts about semiconductors happened in my teens; I started writing about chips, smartphones, smartphone displays and SoCs—even before I had a smartphone, I was obsessed with this stuff. Same with gaming hardware, PC hardware, console hardware—I posted everywhere online.

People on our team—now we’re 90 and even have a full marketing department—always complain: "Dylan, please stop replying to random trolls online; it makes us look unprofessional." But I still have this trigger—whenever someone criticizes me online, I want to reply. Maybe it’s not a good thing, but it’s true.

I managed forums throughout my whole adolescence. When I started earning income, I began investing. Later, I worked as a quant trader twice, then founded my own company—but throughout it all, I kept posting. I used to write anonymous blogs, but by 2020 I was sick of my job—being a quant trader loses its glamour once you realize you can make money but it isn’t as cool as it sounds. So I basically quit and started a business.

I wasn’t sure how far the road would go, but I posted under my real name on my own WordPress site—writing about tech, business, finance, supply chain, and whatever I was interested in. I grew up in a small business family; my parents ran a rural motel in Georgia and lived in it. Later, we ran gas stations too, so I did small business stuff and loved business in general.

Then the WordPress site turned into Substack, I started charging, and wrote about the entire semiconductor and AI supply chain. Over the last four years, I’ve traveled everywhere, attending 40 conferences a year—from top academic AI events to industrial meetings about some chemistry in the chip supply chain, from servers to networking to wafer fabrication, the full stack. Some conferences have just 300 people, mostly in Japanese, and only five of us spoke English—but I still went. Others had tens of thousands. After three times at any given conference, you really master the field’s language, get to know key people, ask deep questions, and the whole ecosystem builds in your mind.

SemiAnalysis’s Growth and Team Building

Dylan Patel:I’ve always focused on technical "inflection points"—whenever I spot a shift in technology or the supply chain at a conference, I analyze what it means for business or financial flows. Some reports are purely technical and ignored by finance; others make it clear this is a bottleneck, that company will win with next-gen tech—and I can articulate that ahead of anyone in the market, including any hedge fund.

The third person was Myin, who used to work at a hedge fund and moved to Japan for his wife—a true “free agent.” He was the first with a hedge fund background, the other two had technical backgrounds. Once he joined, we started building financial models and transitioned from pure newsletters to information services, research report sales, and dataset sales.

After that, the snowball started rolling. From 2023 to 2024, we went from 2 to 7 people; from the end of 2024 to early 2025, 7 to 20; then from 2025 to 2026, 20 to 60; and this year up another 30, now at 90.

The most exciting thing about SemiAnalysis is the density of talent—I don’t know another company with this level of expertise and focus. We have people from ASML, Applied Materials, Lam Research, Intel, TSMC, Nvidia, Microsoft, Amazon, people who’ve worked on models at OpenAI, FSD at Tesla, at Cohere, focused on data centers, even someone who’s built power plants in Kazakhstan—the talent density is insane. Half are engineers from every industry layer, half are ex-hedge fund people, or smart passionate people we found online.

SemiAnalysis now has many business lines: data services, consulting, information services, newsletters, media, and we’re about to hold a large conference—it’s been quite the adventure.

Nvidia GTC Live and the ‘King of Compute’ Belt

Host:Speaking of brave adventures, I’m reminded of a particularly memorable moment. WisdomTree and SemiAnalysis have collaborated for several months now. At Nvidia’s GTC in March, I was watching the live stream from Charlotte, North Carolina—there were reportedly 55,000 watching online, and 20,000 adults packed into the arena.

Dylan Patel:Yes, it really was 20,000 adults.

Host:Jensen Huang directly name-checked you on stage, saying you’d accused him of "hiding" some number, and your chart was projected right on stage. I was—maybe on your behalf—a little excited, watching the world’s largest market cap company’s CEO directly respond to your research and your questioning of some number. Can you talk about what that was like in person? Sounds like you were very much in the middle of the action.

Dylan Patel:That moment was truly surreal. One of the things SemiAnalysis does is we have a team of engineers benchmarking all open source AI models and hardware, which is pretty thankless work. We stay closely connected with the whole industry; hardware-wise, we have received over $5,000 in hardware donations from OpenAI, Microsoft, Amazon, Google, Coreweave, Nebius, Crusoe, and Oracle for running these tests.


0
0

Disclaimer: The content of this article solely reflects the author's opinion and does not represent the platform in any capacity. This article is not intended to serve as a reference for making investment decisions.