11 min read

Four Physics Bets, One Constraint

Two weeks ago we sat in front of four founders at Rosewood Sand Hill who had almost nothing in common except the thing they were all trying to get around. The Rosewood is a beautiful hotel, and the room was mostly private-side money and investment bankers. Bernstein and Société Générale had put the founders on a panel called Power & Performance at eleven in the morning, with Stacy Rasgon, who covers the public side of this fight, moderating. Nick Harris of Lightmatter, betting on a photonic interposer. Sid Sheth of d-Matrix, betting on digital compute sitting inside SRAM. Naveen Verma of EnCharge, betting on capacitor-based analog. Taner Ozcelik of Mythic, betting on flash analog after a company reset that could have ended the company instead.

Four incompatible physical bets against the same constraint. We had a question written down for the floor, and time ran out with our hand still up:

“Each of you is a different physics bet against the same constraint, HBM. Reasoning models just made inference capacity-bound, not just bandwidth-bound. Who up here benefits from that shift and who gets hurt?”

It never got asked. They would have split on it, and the direction of that split is the entire investment question.

The session ran under the Chatham House Rule, so we can tell you what the room concluded and not who concluded it. What we took away was that training was conceded to NVIDIA and inference was the battleground. Notable, in a room full of companies that exist to take silicon share from NVIDIA, and it narrows this piece, because a fight nobody is having is not where the money gets made or lost.

Photonics, in-memory digital, analog charge domain and flash analog do not agree about anything at the device level. They agree completely about how you win. Isolate a narrower workload than the incumbent has to support, strip out the hardware that workload does not need, optimize around its actual bottleneck, post a number the incumbent cannot match on that workload, and use the number to get designed in somewhere. It is the classic specialist attack, it is well funded this time, and it is the strategy that lost the last time somebody ran it against this particular company.

3dfx Won Benchmarks. NVIDIA Changed the Product.

The graphics wars are not a story about NVIDIA beating better-funded competitors on frames per second. 3dfx won the benchmarks for about three years, from the first Voodoo in 1996 through Voodoo2, and then stopped winning them: by Voodoo3 it was behind on 32-bit color and large textures. But it did not die of a bad benchmark. It bought STB and turned its board partners into competitors overnight, it missed Rampage, and it kept Glide proprietary while Direct3D and OpenGL became what developers actually wrote to. Every one of those is a decision about the system around the chip. NVIDIA agreed to buy the assets in December 2000. Some of us were gamer kids while that war was running, playing everything on NVIDIA cards, which is part of why the lesson stuck.

What separated NVIDIA was a product family spanning entry through professional, a release cadence nobody else sustained, drivers that worked across a messy installed base, and engineers embedded with game developers before volume availability. Developer support was the moat, and NVIDIA funded it as a moat while most of the field funded it as marketing. Then the definition of the product moved: GeForce 256 pulled transform and lighting onto one processor in 1999, and CUDA turned a graphics component into a general-purpose computer in 2006. A competitor could still build an excellent graphics chip. It was now competing against an installed programming model, a developer base and a roadmap.

Three stacked panels showing the industry's defining metric moving from frames per second in the graphics era, to tokens per second in the current inference era, to cost per successful task, with what each metric measures and what it leaves out listed beneath each rung. — by Ben Pouladian, BEP Research

Each rung is measured further from the chip and closer to the invoice. That is the direction of travel, and it is the direction that favors whoever owns the layers in between.

Tokens Per Second Is Not an Economic Outcome

There is no honest leaderboard called inference performance, and the reason is that the number has too many free parameters. A published throughput figure moves on model size, weight precision, prompt length, output length, batch size, concurrency, speculative decoding, context window, cache strategy, utilization, whether the measurement is one user or the whole system, and how much model-specific optimization went into the software build. Change any one and the ranking changes with it.

People are playing benchmark games with tokens per second and tokens per watt, and tokens are everywhere now anyway; the open-source wave made them nearly free to try. But run agents making multiple tool calls, watch some of them fail, and the number a CFO opens the spreadsheet wanting is cost per task.

So we are not reprinting anybody’s headline token rate. The challenger field publishes real numbers on real silicon, but they are company-reported, on configurations the company chose, and in several cases they are projections rather than measured deployments. Cerebras publishes token generation speeds on supported models, and those are the company’s own figures. d-Matrix publishes total-cost-of-ownership advantages that are projections or preliminary results. Nobody is lying. But you cannot put two of those numbers on the same axis and have the comparison mean anything.

Frames per second had the same gap and nobody minded, because the number was never the whole purchase. Tokens per second does not tell you whether the model finished the customer’s work correctly.

Table of four companies on the Power and Performance panel, listing each one's architectural bet, the part of the memory problem it attacks, and the open question that decides whether it scales. — by Ben Pouladian, BEP Research

Four architectures that disagree about physics and agree about strategy. The right-hand column is where each bet gets settled, and none of those questions is answered by a throughput number.

The Denominator Already Moved, and NVIDIA Moved With It

In July we went through what happened when five model launches in three days all competed on the buyer’s cost per completed task. Databricks had published its own benchmark against a multi-million-line codebase and concluded that token costs are often a poor indicator of overall task costs. Sonnet 5 was roughly 1.7 times cheaper per token than Opus 4.8 and still more expensive per task, because it burned 1.9 times more tokens getting there. The buyers stopped paying for tokens some time ago, and the two numbers do not rank the same.

Then you divide by the pass rate and the spread widens again. As we wrote in Jensen Huang: “Most Companies Will Be Built on Harnesses”: “A cheap task that fails is the most expensive task there is, because a human gets paid to clean it up.” That piece worked the argument at the model layer. Four founders on one stage in July is the same argument arriving at the silicon layer, and it arrives carrying a problem for them, because none of the costs that separate a cheap token from a cheap task are measured at the chip.

We turned that arithmetic into a tool, and it went live today. Octane runs 20 models through the same 32 tasks, grades every attempt the same way, and prices each model per finished task rather than per token. The six models at the top of the board are doing work we cannot tell apart, at a 73x spread in price, and the line on its masthead is this whole section in ten words: price is what you pay for tokens, cost is what you pay for mistakes.

Paired bars for four models showing published cost per task beside cost per successful task after dividing by pass rate, with the percentage inflation labeled above each pair, ranging from 37 percent to 85 percent, and the spread between cheapest and dearest widening from 5.3x to 6.9x. — by Ben Pouladian, BEP Research

Dividing by the pass rate does not reorder this field. It spreads it, from 5.3x between cheapest and dearest to 6.9x, because failure taxes the weakest models hardest. The ranking that flips is the one against a rival on different silicon.

Dynamo separates prompt processing from token generation, routes requests to whichever GPU already holds the relevant context, pins high-value cache, and manages a memory hierarchy that runs from on-chip SRAM out to cloud storage. That is not a faster chip. It is an argument that the unit of the product is the fleet, and the argument has a number attached to it: MLPerf v6.0 in April showed the same GB300 NVL72 rack going from 2,907 to 8,064 tokens per second per GPU on DeepSeek-R1 in six months. Same silicon, same power envelope, software only. An installed base that improves 2.7x without new hardware is a very hard thing to win a socket away from.

The same move showed up in the silicon a year earlier. As we wrote in The Fourth Piece Ships: “The LPU is NVIDIA’s answer to the purpose-built inference chip, and it comes with an ecosystem moat that no startup can replicate.” Whether that ecosystem claim survives contact with a specialist who is materially better is the open question of the next two years. The shape of the response is not in doubt. When a challenger demonstrates an advantage in one layer, NVIDIA expands the stack and makes that layer one optimization inside a system it controls.

The startups need inference to become a separable commodity layer where specialized silicon can win on its own merits. NVIDIA needs inference to be a system problem where hardware, networking, scheduling, cache management and fleet utilization have to be optimized together.

Today Ran the Experiment at Both Ends

Today that fork stopped being hypothetical, twice in one afternoon. SpaceX played its Q2 call live on x.com, and we listened as Musk went all-in on NVIDIA in public. Verbatim: “Going forward, we’ve decided to build exclusively on Nvidia, because we think the Vera Rubin architecture is the best architecture. We think it’s the best AI computer.” SpaceX expects to end this year with over two gigawatts of compute and, in Musk’s framing, to be “closer to 10 GW of compute than 5 GW” by the end of next year, with a stated target of up to 20 gigawatts of power and cooling online ahead of the GPUs to fill it. The word he reached for was computer, and no throughput number appears anywhere in the sentence. A buyer planning gigawatts chose on the architecture. That is the layer, and it is this argument arriving in the customer’s own words. And consider the mouth it arrived from. Elon is regarded as the best hardware operator alive, building the hardest hardware in the world, and he is not trying to make his own AI chip and he is not looking at AMD. When that man says NVIDIA has the best architecture, it is a big, big endorsement, and it landed the same afternoon AMD printed its record quarter. He also told analysts that “our understanding with NVIDIA is that we will receive a very significant percentage of their GPUs next year.” The platform and its largest customer are allocating each other.

Then he took the platform off the planet. Starmind, in his words, is “essentially an optimized Vera Rubin NVL72 computer,” launching next year: a solar-powered rack in sun-synchronous orbit, continuous sunlight for power, vacuum for cooling, results lasered down through Starlink. If the input cost of energy is the sun, it is zero, and energy is most of the bill in this business. It also looks real rather than aspirational; his phrase was “not some sort of far future distant thing.” One more thing we cannot stop noticing: AMD named its rack Helios, the sun, probably with this customer in mind. The customer went and built the sun-powered computer instead. And the same company’s own Starmind page carries the counterweight. It says “We are AI chip vendor agnostic. Our system architecture supports compute modules from any provider,” and it describes a planned in-house AI chip fab with Tesla. Exclusive is a decision that can be unmade, and NVIDIA’s most enthusiastic customer is architected to unmake it.

Below the divider, for paid subscribers: what the AMD print actually repriced and why the earnings multiple is the tell, the two numbers Su was asked to bless and the one she re-anchored, the yield concession and the HBM admission, the ROCm.AI move nobody wrote up, and what SpaceX’s 10% number makes of the other 90%, which is where our most speculative call in months lives. Plus the ranked bear case, led by the one that would make this whole reframe commercially inert.


Originally published on BEP Research on Substack. Subscribe for more.

Posted in

Leave a Reply

Discover more from Ben Pouladian

Subscribe now to keep reading and get access to the full archive.

Continue reading