What Jensen told Dwarkesh that everyone missed
4 min read

The Story Everyone Knows Is Already Over

A lot of people have DM’d me since I posted my thread on the Dwarkesh Patel interview with Jensen Huang on Wednesday. I’m heading to Google Cloud Next this week, and paid subscribers will get my field notes from Vegas. But first, the framework.

For fifty years, progress in computing had a simple shape. Engineers shrank the transistors. Chips got faster. You bought the new one. That was the deal. Moore’s Law was the name people gave to the rhythm, and the rhythm was reliable enough that entire industries planned their roadmaps around it.

That story is over. Transistors have hit the atomic wall. At the leading edge of semiconductor manufacturing, the smallest features on a chip are now roughly ten atoms across. You cannot build a switch smaller than an atom. The physics that powered the digital revolution for half a century has quietly run out of room.

And yet the chips powering artificial intelligence keep getting dramatically faster, generation after generation, at a pace Moore’s Law never delivered even in its best years. NVIDIA’s newest platforms deliver roughly ten times the usable output of the generation they replace. Same factory. Same electrical grid. Same engineers. How?

The answer is that the thing being built is no longer a chip. It’s a factory. And the factory doesn’t just ship once. It keeps getting better after it ships.

From Chip to System to Factory

When Jensen Huang started calling NVIDIA’s approach extreme co-design, most of the industry treated it as marketing. It wasn’t. It was a description of what happens when every part of a computing system, including the chip, the memory, the networking between chips, the networking between racks, the cooling, the power delivery, and the software that schedules every operation, is designed together from day one rather than assembled from parts bought on the open market.

Consider what actually ships when NVIDIA delivers its latest platform. You don’t get a GPU. You get a rack. The rack contains 72 GPUs wired together so tightly that software treats them as one giant processor. Inside that rack is enough copper cable to stretch the length of two city blocks. The power density is the equivalent of two hundred suburban homes running full tilt inside a single seven-foot-tall cabinet. The cooling system circulates water warm enough to shower in, because the chips produce so much heat that traditional air cooling cannot keep up. Every one of those design decisions, from the cable routing to the water temperature to the way the GPUs share memory, was made so that the other decisions would work.

Now zoom out one level. That rack is itself just a unit inside a larger structure. At NVIDIA’s most recent GTC conference, the company unveiled five different rack types that work together: one for training, one for the fast answers a chatbot gives you, one for the memory that keeps long conversations coherent, one for the networking that ties everything together, one for the agents that do work on your behalf. None of these racks function alone. They function as a single building that converts electricity into intelligence. The industry has a name for this building now. It’s called an AI factory.

This is the first shift to understand. The product you are buying, if you are a hyperscaler or a government or a large enterprise, is not silicon. It is a factory that manufactures tokens, the small units of output that make up every answer ChatGPT gives, every image Midjourney renders, every line of code an AI coding assistant writes. The factory has a throughput measured in tokens per second. It has an input measured in megawatts of electricity. And like any factory, what matters is how much product it ships for each unit of electricity it consumes.

In six months, on the same silicon, one rack just got 2.7 times more productive. The Financial Times just reported that Anthropic cannot ship its most advanced model because the factories to run it do not exist yet. And Jensen Huang told Dwarkesh Patel the single line that explains three years of NVIDIA strategy, the line most viewers glossed over. Here is what it all means.


Originally published on BEP Research on Substack. Subscribe for more.

Posted in

Leave a Reply

Discover more from Ben Pouladian

Subscribe now to keep reading and get access to the full archive.

Continue reading