Free Models Have a Compute Bill
NVIDIA is reportedly in talks to buy Hugging Face for around $13 billion. The number is a scoop, not a filing. Insider had talks above $13 billion that can still die. The Information had $12.9 billion agreed. Bloomberg had about $14 billion with a $1 billion retention package, then on Tuesday evening moved it to $12.9 billion plus the retention pool, an agreement possible as soon as this week, no final terms, both companies declining to comment. It is September 3. Nobody has filed.
Days earlier came Poolside: a $6 billion non-exclusive license to its Model Factory, $1 billion of equity, and jobs for more than 100 of its engineers. The two moves are one strategy.
The context for both is a chart. NVIDIA is already the biggest corporate publisher on Hugging Face, and it pulled away from the field this year: by AI World’s count of repositories with real traction, more than 500 added since September 2025, ahead of Alibaba Cloud, Hugging Face itself, and Tencent.
We were on ETN the morning after the print, and put the problem this way, on tape:
“Poolside was making a really good open source coding model, but the challenges of open source is like if you’re spending all this money on making this model, but you’re kind of giving it out for free, how do you pay for all this compute?”
That question is the whole deal, more than the headline number. An open model is a product with billions of dollars of cost of goods and a price of zero. For three years the answer to who pays was venture capital, and venture capital wants the loop to close eventually.
It closes a different way. As we said on the same tape: for a lot of the open source guys the business model didn’t work, and you needed someone like an NVIDIA to come in, fund it with the compute, and take care of the talent.
An open model cannot charge more for its tokens than the cheapest stranger hosting its weights, so the durable owner of open source is whoever gets paid when the tokens run.
That is the Compute Patron. To make all these models you need compute, and compute is expensive. How do you expect to have something open source and not be able to pay for it? What is your end game? You can’t be a charity. The patron’s answer: if you create more demand for compute through open source, you still need to sell more compute, which is NVIDIA compute, and a lot of these models will be optimized for NVIDIA general-purpose compute and CUDA.
The Loop Breaks at Step Four
The business model of open source did not fail on quality. It failed on a missing step. In The Token Dollar we mapped AI compute finance as a loop: borrow dollars, buy GPUs, mint tokens, sell tokens for dollars, service the carry, roll the principal. An open-weights release gives step four away. The tokens still get minted, on the same GPUs, and anyone who downloads the file can sell them. DeepSeek and Mistral run price sheets on their own weights; Together and Fireworks run them on everyone’s. The publisher keeps whatever margin the cheapest host leaves, which rounds to nothing. That is why an open lab’s runway is somebody else’s balance sheet, always.
Jensen said the cost side plainly in July: “Frankly, I think closed models are cheaper.” His list of what open costs the builder ran through fine-tuning, guardrails, evaluation, and building the computers to host it, and he landed on “there’s nothing cheap about doing that.” If open models are not cheap, someone has to pay for them, on purpose, indefinitely.
Run the arithmetic on one lab. Poolside’s training capacity was a 40,000-plus GPU commitment at CoreWeave, estimated at $5 to $6 billion over five years, with no token margin on the other side. They then missed a six-week window to raise about $2 billion for the GB300 cluster due in January, and sold the factory to the company that sells the compute. The letter to investors is the instrument: $6 billion non-exclusive license, $1 billion of equity at $12 billion pre-money, founders stay. Not an acquisition. The structure is the tell.
We tried to get in contact with Eiso before and after the talks and never got through. Laguna was a promising frontier coding model and the team behind it was executing well. Maybe the investors didn’t see a path to revenue or profitability versus the competitors that were out there, and this is the best case for them: the founders stayed, and the employees left. Licensing is a faster path to getting people into the office to work on the product, just like Groq, announced on Christmas Eve with Jonathan Ross and his team going to NVIDIA. We don’t think it could possibly be a rescue, and it also benefits the open source ecosystem. A missed raise and a rescue are different things: the equity came in at $12 billion pre-money, which is not a distress price, and a non-exclusive license leaves Poolside free to sell the same factory again.

Now look at who carries that kind of bill today, model by model. Meta pays for Llama out of the ad business. Alibaba pays for Qwen out of the cloud business. Zhipu and Moonshot pay for GLM and Kimi out of funding rounds and, increasingly, out of Beijing’s interest in them existing. Our theory, and it is only a theory, is that somehow they are being subsidized by the Chinese government; there is no way they can do all this at the cost they are saying. Every open model of consequence is subsidized by a business that is not tokens.

The Patron Gets Paid in CUDA
NVIDIA is the one patron whose subsidy is structural, because the free model is a demand generator for the thing NVIDIA actually sells. If you buy the box, you cannot put frontier OpenAI or Anthropic weights on it. Those stay in their clouds. What you can run on your own iron is Qwen, GLM, Kimi, Llama, and Nemotron. Most of the open-model inference in the world, on-prem, at a neocloud, on a DGX Spark under a desk, clears through NVIDIA silicon today, and Nemotron is the executor tuned to that silicon on purpose. Most, not all, and we have already published the other side. On August 4, in All Roads and Rockets Lead to NVIDIA, we wrote that “inference also sells behind an API, so a provider can swap hardware under an OpenAI-compatible endpoint without telling anyone,” and open weights make the swap easier, because the same file runs on TPU, Trainium, and MI. What the patron structure changes is which iron the free model is tuned for first, and NVIDIA is hoping that more compute in every form factor it sells, Jetson Thor, the DGX Spark, the datacenter GPUs, captures one of these open source models. Whether that holds the volume is the sixth gate at the end of this piece. We would probably see the drift in the neoclouds first, or on TPUs. Up until now it hasn’t happened, and on our read NVIDIA is increasing its share of inference.
The demand side showed up too, and it is personal for a lot of buyers now, including us. We put a DGX Spark under a desk this summer, the machine we called The Edge Token Factory in June, and we had a hard time setting it up. Connect to Wi-Fi, then command prompt, then a bunch of updates, and we had to use ChatGPT to ping it and find it on the network. When you can’t find it on your network and it’s hard to connect, it gets a little scary, and then you’re like, okay, how do I get open-source models? Where do I get them from? Everyone was pointing towards Hugging Face as the place to be. We had to create a Hugging Face account and get a token and do all that stuff, and those are barriers to the process to get open source models. The initial setup took an hour and a half. In the end we had it all connected, with Claude Code basically running the Spark for us on the Qwen models, NVIDIA’s own build of them, and we were able to do some image generation. With NVIDIA owning the shelf, it gets more seamless and easy to set up and less guessing, because it is all coming from one trusted provider. NVIDIA shows up on the screen now instead. On the Blockspace tape we put it the way an owner would: run a good enough model at home and your landlord isn’t Dario or Sam. You own your own tokens.

Open-source share of enterprise LLM spend fell from 19 percent to 11 percent in a year while Decagon still runs about 90 percent of its workloads on open models, and open-weight models went from a third of the tokens routed through OpenRouter in late 2025 to a majority by mid-2026, per the July State of Open Source AI report. The work stayed. The invoice left.
The CUDA Playbook, One Layer Up
Poolside and Hugging Face are the first rows of NVIDIA’s deal ledger that touch the model and its distribution. In January, in NVIDIA’s Inference Stack Depth Strategy, we called the pattern the CUDA playbook replicated at the inference layer. That ledger ran past $30 billion in announced deals, licenses, and reported talks, every row below the model. Add $7 billion for Poolside and a reported $13 billion for Hugging Face and it is past $50 billion, still mixing checks and term sheets.

Hugging Face is the distribution row, and distribution is where this stops being an M&A story. Every executor in the enterprise agent stack lives there, and so does NVIDIA, more than any other company, as the chart at the top shows: Nemotron, Cosmos, GR00T, and optimized versions of other labs’ models. The biggest corporate publisher on the shelf is the one reportedly buying it, and a good share of what it publishes is other people’s weights tuned to its iron. America needs an open source leader, and no one else is really stepping up to the plate. It is funny which company it is: people think, why would you want to give away something for free? This is NVIDIA. Trusted distribution matters, and putting all the pieces together, the full stack story we have written about all year, this is the next full stack solution.
As we wrote in The Open Source War Is Over: “Nemotron was built to be the layer under the frontier, not to beat it, and to make sure every token in the system, orchestrator and executor alike, gets minted on NVIDIA hardware.” If the report is right, the company that ships the American executor would also own the shelf that every executor, American and Chinese, ships from. We said as much on ETN the morning after the print, before any of the numbers firmed: NVIDIA is “building a full stack solution for open source all the way from Hugging Face to the Nemotron models that they’re creating.”
The Bear Case: A Patron Is Not a Neutral
Four ways this read gets hurt. The first three are in rising order of damage. The fourth is the one we lived.
1. The deal is a report. Talks die and sizes shrink, and a $13 billion headline can end as a commercial partnership with a press release. If it dies, the patron economics still exist; Nemotron and Poolside already prove them. What dies with it is the depot, and the depot is what the $13 billion was for.
2. Neutrality was the asset. Hugging Face works because it belongs to no lab and no silicon vendor, and ownership by NVIDIA tests that instantly. Our own answer as an uploader: we’re okay with it, at least it is a trusted source, and they have other models hosted there as well. People might pull weights, but others might keep it; it’s not for everyone. The observable: where new model uploads and download share go in the two quarters after a close. Migration is the miss.
3. The frontier closes the seam. Carried forward unchanged from July: the moment Anthropic or OpenAI sells a first-party small model, fine-tunable on your task, served in your perimeter, the two-tier stack becomes a one-vendor product and the executor tier loses its reason to exist. This is the one that changes the conclusion rather than the slope, and it was the one that changed the conclusion in July too.
4. Open source stays a hobby. What if open source is just like a hobby and doesn’t really take off, and everyone just wants closed source? Based on the experience we had, it wasn’t really easy to set up. For open source to be a real winner it needs an iPhone or Mac type of experience, where you get up and running out of the box simply and don’t have to be a developer or data scientist to put it together. It is the bear the reported deal is built to fix.
We don’t think you can look at revenue numbers on this one. It is all about scarcity, moat, and user value, and the talent behind the team, and possibly other bidders. Stripe buying OpenRouter for more than $7 billion last month is the same type of concept. The one public anchor is the $4.5 billion valuation from its 2023 round; the reported price is roughly three times that, on revenue the company does not disclose. What we can underwrite is the stack, and the stack does not need Hugging Face to earn its keep. It needs open models to keep improving and keep running on NVIDIA.
So What?
Ask who pays for the next open model release, because the answer determines whether it is a durable layer of the stack or a runway with weights attached. Venture patience and Beijing’s interest both run out. A patron who books the subsidy as ecosystem spend, against the highest-margin hardware business on earth, does not run out. People didn’t believe in CUDA, and NVIDIA was just spending, hoping developers would come and use it. Eventually it found its use case. When it’s early, people never see it. We honestly didn’t see it then either. We bought the stock in 2016 because NVIDIA was using its GPUs to solve hard math problems for crypto mining, and our thought was that if you can solve hard math problems you can probably solve other problems, and these systems have a value for creating some sort of intelligence. We see it now.
The instinct runs all the way up the house. On September 2, Open Athena, a nonprofit, announced a 535-billion-parameter Marin training run paid for with what it called “a generous gift of compute from The Jen-Hsun and Lori Huang Foundation.” Jensen’s line in the announcement: researchers “need access to frontier-scale models they can build, study, and understand.” A nonprofit cannot sell tokens either. It is just interesting to see Jensen and his family spending money on a nonprofit to make other open-source models. That is something he believes in. He doesn’t have to do this. He could have taken this money and built a museum, or donated to some university campus to be named after him. The patron paid anyway.
It should deflate all the token prices. Open source tokens should be much cheaper than frontier, but then you are paying upfront hardware and maintenance costs to host them yourself, and that invoice goes to the patron. For the closed labs, the floor just got harder: Anthropic and OpenAI now price against an executor tier maintained by a competitor that never needs token revenue.
For NVDA itself, we read the reported deal as CUDA spend, and we are not resizing a position we have held publicly since 2016 on a press report. Still holding core strong. It deepens the moat, and these are wise investments that pay off in the long run. It is hard to look at a $5 trillion company and think there is still room to grow, but the market isn’t fully penetrated; the TAM is so big that there are still opportunities.
Six gates, all observable from outside, in the table below. The Hugging Face talks confirm near the reported size, which grades the depot. Nemotron stays free with no paid token API, which is the patron gate. The depot stays neutral and its access never gets tied to NVIDIA hardware, both live only after a close. The frontier does not close the seam with a first-party small model in the customer perimeter, which changes the conclusion as it did in July. And open-model inference keeps clearing mostly on NVIDIA silicon, where two straight quarters of non-NVIDIA share rising grades the ticker, not the patron. It would break the call for us, probably in a few years, but not now. Miss the first gate and NVIDIA is still the patron. Crack the second and the patron was just a shopper.
The model is free; the compute never was.
Underwrite the patron, not the model.

Coming Up
-
The paid follow-up: what Hugging Face is actually worth as a business, and what a closed deal does to our coverage universe. Ships when the deal status hardens.
-
GTC Berlin is October 20, and we will be there. If you are going, reply to this email; last GTC produced two on-record interviews that became pieces.
-
One more piece of housekeeping. BEP’s principal runs occasional SPVs in private companies across this stack, offered to verified accredited investors under Rule 506(c). If you want to see those deals when they open, join the list at tools.bepresearch.com/syndicate.
Related BEP Research
-
The Open Source War Is Over (X, July 7): frontier orchestrates, open executes, and NVIDIA armed both sides
-
The Token Dollar (April 4): the loop this piece breaks at step four
-
The Slide in NVIDIA’s RTX Spark Keynote (June 1): The Edge Token Factory
-
All Roads and Rockets Lead to NVIDIA (August 4): the Palantir deployment that closed the July gate
Disclosure: BEP Research’s principal is long NVDA, LITE, CRDO, TSEM, ALAB, WOLF, NOW, SMCI, BE, NBIS, and ORCL (2027 LEAPS), and has been publicly long NVDA since 2016. This is investment research, not investment advice. Do your own work.

Leave a Reply