The Lead That Narrows
Lisa Su put up a chart at Moscone on Wednesday titled “AMD Helios Delivers Industry Leading Rack Scale Performance.” Throughput on the y-axis, interactivity on the x. The advantage over NVIDIA’s Vera Rubin runs +15% at low interactivity, +12% at medium, +10% at high, on Kimi K2 Thinking, a reasoning model AMD picked. I photographed it from my seat and did not think about it again until I was back in the pictures. Nobody around me reacted either. It went by like every other bar chart.
I have held NVIDIA since 2016 and I do not own AMD. Every figure below comes off one of my photographs.
Su had spent the opening arguing that agents are the shift, that 2026 is the first year the world runs more compute on inference than training, and that inference is roughly 60% of AI compute. High interactivity is where that work sits, and where Helios is least separated from NVIDIA’s next rack.

I went looking for the trick in that 15 and found the opposite. Divide Helios by its 72 GPUs and each one is doing about 40 petaflops of FP4. NVIDIA’s own Rubin documentation puts a Rubin package at 35 petaflops of dense NVFP4, and 72 of those is 2.52 exaflops, exactly what AMD’s +15% credits Vera Rubin with. Packages against packages, dense against dense. Moor Insights notes AMD counts MXFP4 where NVIDIA counts NVFP4 and the block scaling differs, but on the question that would have moved the levels, AMD took the conservative NVIDIA number and did the arithmetic straight.
NVIDIA does not market 35. It markets 50 petaflops per package, 3.6 exaflops per rack, and that delta is not silicon. It is an adaptive compression engine in the Transformer Engine that strips zeros out of the data stream in flight, replacing the 2:4 structured sparsity NVIDIA used to double marketing numbers with. Against that figure, Helios at 2.9 exaflops is 19% behind. I will not treat the 50 as realized performance, because nobody outside NVIDIA has benchmarked what that engine delivers on a real workload. Both of these are vendor numbers on unshipped parts. AMD’s +15% is a specification comparison rather than an independent benchmark, and I am holding it to the same standard I just held NVIDIA’s 50 to. Either way, AMD’s lead is a lead in dense silicon, and NVIDIA’s answer to it is not more silicon.
Which is why the curve decays. Both endpoints sit on AMD’s own slide: compute +15%, HBM bandwidth +6%. Low interactivity is batch-heavy work running close to peak throughput, so the lead sits near the flops ratio. As interactivity rises, the workload becomes increasingly sensitive to decode latency and memory movement rather than to peak batch throughput. Weights and KV cache have to be moved for every token generated, and the number that governs that is bandwidth. The line is walking from AMD’s compute number toward AMD’s bandwidth number, and the bandwidth number is six.
Scale-up bandwidth on the same slide reads Same, at 260 TB/s. Nobody said that out loud.

One discrepancy I owe you: my recording has Su saying “50% more compute” at this point, while the slide and Moor Insights both read 15, so I am treating the 50 as an artifact of my own transcription.
The keynote started at 9:30am with open seats around me. In March I stood outside the SAP Center in a line that had been forming for hours. Jensen Huang is the Taylor Swift of AI. That says nothing about the silicon and something about the demand pull AMD is still building. They did copy the DGX Spark, with the Gorgon Halo I am holding above.
The Number I Am Not Swallowing
The $1.4 trillion accelerator TAM does all the work in the bull case, mine included. Last year at this same event Su put it at about $500 billion by 2028. This year it is $1.4 trillion by 2030, so the horizon moved with the number. The server CPU guide went from $60 billion to over $200 billion, against a $25 billion base.
She flagged it herself, in the sentence before the chart went up: “it’s really hard to put this market data up because I feel like every few months we’re changing the perspective based on what our customers are saying.”

Capex has to land as megawatts where the grid will not interconnect new load for years, and Deloitte has 11% of organizations running agents in production against 38% piloting.
The Ramp
AMD earned a win here. In February a SemiAnalysis report said manufacturing problems would push the Helios ramp to Q2 2027, and AMD denied it flatly. Moor Insights put the timing question to Su directly and reports her answer as “Declaring production means we are ready to ship,” with the ramp running through Q4 into the first half of 2027. Whether AMD can build a rack-scale system is settled.
My call is that meaningful volume lands in the second half of 2027, two quarters past AMD’s own window. Helios puts 72 MI455X GPUs and 18 Venice CPUs in one rack, tied together with UALink over Ethernet on Broadcom silicon. Helios is AMD’s first Blackwell-style rack-scale system, but with more external cabling and integration complexity in the part of the stack NVIDIA has spent years learning to manage. Jensen described the Vera Rubin tray in January as having no cables, no hoses and no fans. Six months later Su held up the Helios tray and told the room it weighs more than 160 pounds.

Nick Dorsey, posting publicly on X, reported a conversation with the ODM Wiwynn:
“He said mi455 has lot of challenges right now, that they’re not sure about how it’s ramping or going to ramp. It’s their first rack scale solution and it will have Broadcom scale up networking so AMD is dealing with a lot of complexity […] He said we will see how they come together 2nd half next year”
The same post carries the other half. Customers are optimistic, and “if AMD can make the rack work well then they are going to buy it.” That channel comment supports my second-half 2027 base case, although it is one datapoint rather than confirmation. What would prove me wrong is material Helios revenue in AMD’s Q1 or Q2 2027 results, since a clean Q4 is already guided and would tell me nothing.
The evidence against me is something I stood in front of. Supermicro, which I own, had liquid-cooled AMD racks on its stand with a large team working them, and their product manager, twenty years there, was optimistic. Build quality looked good. Those were CPU racks and Helios is a GPU rack, so what I saw is a channel building and cooling AMD silicon at volume. It counts against me anyway.
The Validators Are Incentivized. The Demand Is Still Real.

OpenAI and Meta each hold a performance-based warrant for up to 160 million AMD shares, roughly 10% of the company apiece, vesting on deployment milestones against 6 GW commitments. Together that is close to a fifth of AMD. AMD’s investment of up to $5 billion in Anthropic was reported July 22, and Tom Brown committed 2 GW the next day.
Be precise about what that does and does not mean. Nobody was paid to read a script. The warrants vest on deployment, so the incentive is to build, and Su says all three are co-developing Helios, which is a deeper commitment than an endorsement. What it does mean is that AMD’s loudest public validation came from parties holding unusually large equity claims on the outcome they were validating. You weigh that differently than a buyer with nothing at stake.
Which is why the most useful thing I heard all day happened in a hallway. A neocloud operator stopped me. His customers had told him to hold on until Helios shipped, and now he wants to scale into the hundreds of megawatts. Demand queued behind availability, from someone with no equity in the outcome.
The Part of AMD I Would Not Bet Against
Two days earlier I was on the other side of this argument. NVIDIA published its Vera whitepaper naming AMD’s flagship, EPYC Turin 9755, beside its own SPECrate score of 925 against 898, vendor-run and estimated. As I wrote in NVIDIA Vera: When CPU Latency Becomes GPU Economics the morning the embargo lifted: “AMD gets the first word almost immediately: its AI event runs July 22 and 23, and NVIDIA just handed every analyst in that room the question to ask.”
Forty-eight hours later Su answered from a stage. Venice runs at 1.8x Turin on 512 threads, and she claimed 20% higher per-core performance than Vera using NVIDIA’s own published data. Each company benchmarked the other’s unshipped part, in public, by name, in the same week.
She also put Venice in a tier she called sandboxes, dedicated CPU capacity for the agent work that is not inference. Every agent step that is not a forward pass is orchestration and tool calls, so the logic holds. But she described the tier without showing what runs there or what it costs, and the $200 billion server CPU number depends on whether anything does. The evidence I could stand in front of was plumbing: nobody liquid-cools a CPU rack, as Supermicro had, unless the thermal load justifies it. AT&T said it runs more than 45 billion tokens a day on AMD across 122 production pipelines. The CPU franchise is the part of AMD I would not bet against.
Where This Read Breaks
AMD is squeezed from two directions, and it can gain merchant GPU share while both jaws close on the economics.
Above it is custom silicon, and people read that threat as hyperscalers walking away from merchant chips. In The Third Announcement Most People Missed at Google Cloud Next I argued the opposite, that “Google is not trying to beat NVIDIA. Google is making Google Cloud the neutral venue where TPU and NVIDIA both monetize enterprise agents.” Two fleets run side by side there, and AMD is in neither. AMD loses the moment the merchant slot is contested rather than expanded.
Below it is a funded field of inference specialists, roughly $58 billion of private paper by Deedy Das’s July 25 tally, with Cerebras public at about $48 billion. AMD’s fast-inference answer is a Cerebras partnership, and Cerebras answers to its own shareholders. NVIDIA licensed Groq’s technology outright and hired Jonathan Ross with it. What would tell me the squeeze is biting is merchant GPU share falling while accelerator spend keeps rising.
AMD is also missing a layer. The stack runs silicon, interconnect, software, and models, and AMD launched the first three this week. NVIDIA ships Nemotron as a full development stack with no revenue or user thresholds, a distribution channel into every enterprise building its own AI.
The strongest case against me is software. In Top 10 Tech Predictions for 2026 I wrote that “AMD’s ROCm remains 3-5 years behind in tooling maturity.” What AMD showed was an attempt to make the coding agents developers already use fluent in ROCm, with demos through Cursor, Claude and Codex. Nobody has to learn AMD’s tooling if the agent already knows it. If that lands, the software gap closes faster than the hardware gap and my read gets worse.
So What?
The market wants a credible second source to NVIDIA, and AMD is finally offering one. The question is no longer whether AMD can ship. It is what you pay for the shipping.
At 10% of its own $1.4 trillion 2030 accelerator TAM, AMD is doing $140 billion a year, against data center revenue of $5.8 billion in the March quarter. That 10% assumes the squeeze does not bite, and it rests on Su’s claim that GPUs stay the vast majority of the market.
Wall Street ran that arithmetic the next day and got there from the other direction. Baird took its target to a street-high $1,250 from $625, modeling $147 billion of AMD AI GPU revenue by 2030 on an assumed 15% share. That is within $7 billion of the figure above, except Baird calls it 15% and I call it 10%. Fifteen percent of $1.4 trillion is $210 billion, so Baird’s share is being applied to something nearer $980 billion. The most bullish note on the street is implicitly applying AMD’s share to an addressable pool roughly 30% below AMD’s headline TAM, most likely because it does not count custom silicon, or categories it does not consider available to merchant GPUs, as reachable. The top jaw is already embedded inside the bull case.
Jeff Pu’s $640 target is explicitly 45 times 2027 earnings. NVIDIA trades in the high teens to low twenties on forward earnings, depending on whose estimate you use. Same demand, same week, and the market pays more than twice the multiple for the smaller and later-arriving one.

The warrants sit outside that multiple entirely. Those 320 million shares strike at a penny, which makes them a customer acquisition cost that lands in the share count where a P/E cannot see it. A forward estimate will not carry them either, because under ASC 260 contingently issuable shares stay out of diluted EPS until the conditions are met, and these vest on milestones nobody has hit.
Run it as illustration rather than as a target. Against 1.65 billion shares, full vesting takes roughly 16% off per-share earnings, which puts that same 45 times nearer $535 than $640. That is arithmetic on somebody else’s published number, not my price target on AMD, and it only happens if AMD ships six gigawatts, which describes a company that has earned the dilution.
What would move me is an independent inference benchmark that tests NVIDIA’s compression claim against AMD’s dense flops.
AMD no longer needs to prove it can compete. Helios is credible, the customers are real, and the market is large enough for AMD to grow dramatically. But credible competition is not the same thing as attractive relative value. AMD is being valued as though the execution, the market share and the TAM all arrive together. NVIDIA is already delivering into the same demand, without promising customers warrants equal to nearly a fifth of its current share count. AMD may win. I still think NVIDIA is the cheaper way to own the outcome.
Disclosure: Long NVDA, NOW, LITE, CRDO, TSEM, LSCC, ALAB, WOLF, SMCI, BE, and ORCL (2027 LEAPS). This is not investment advice.
Coming Up
-
The Bernstein Societe Generale conference, and what the inference startups are claiming. The other jaw of the squeeze, first-hand.
-
A dinner in San Francisco with AI founders, and what they say about buying compute off-stage.
-
Nobody has benchmarked Helios. Every number on that comparison slide is a spec or an estimate, the microbenchmark footnotes cite one to four GPUs, and the only thing that ran live on stage was a small model on a single GPU. I am going through AMD’s endnotes line by line for subscribers.
-
Venice is the same question wearing a different hat. AMD called it in production, then footnoted the comparison as preliminary estimates based on engineering projections. The SPEC26 subscores that would settle it have not been published.
-
AMD reports Q2 on August 4 after the close, which tests the ramp language. NVIDIA’s FQ2 lands August 26.
Leave a Reply