Why NVIDIA Bought HuggingFace
Happy Thursday.
The current thing in tech and business is WIRED’s profile of AI text detector Pangram and its founder Max Spero, which some are calling a hit piece.
Today’s Lineup
- Pablo Torre Finds Out Host Pablo Torres at 11:30 AM
- SciFin founder Mohit Aron at 11:50 AM
- Pocket cofounder and CEO Akshay Narisetti at 12:00 PM
- Aura founder and CEO Hari Ravichandran at 12:15 PM
- Minnesota Timberwolves CEO Matt Caldwell and Jump cofounder Jordi Leiser at 12:30 PM
- Portal Space Systems CEO Jeff Thornburg at 12:45 PM
- Baseten Co-heads of Model Training Charles O’Neill and Mudith Jayasekara at 12:55 PM
- Snowflake CEO Sridhar Ramaswamy at 1:10 PM
Run of Show
Why Nvidia Bought HuggingFace
The news that Nvidia spent nearly $13 billion to acquire Hugging Face is all over the timeline. It was leaked a few weeks ago, but now we have the final number. It’s $12.9303 billion, which just happens to be the exact decimal code for the Hugging Face emoji. And if you take that number and turn it into a color code, you get a shade of green that hints at Nvidia. Serious Easter eggs in the purchase price, which is completely ridiculous given what a huge outcome this is. Back in the social media boom, a billion-dollar outcome was enough to drive a month-long news cycle, and now it seems like we’re getting decacorn liquidity events every few weeks.
But the deal makes a lot of sense, especially when you consider Jensen’s messaging on the importance of open source and Nvidia’s position in the AI ecosystem. Hugging Face does a lot of things well, but many people just think of it as the GitHub of AI. They don’t own the smartest AI models. They don’t try to build them. But they have created a nexus for people to upload, discover, test, and modify models. And the numbers are good. They have over 18 million developers, 200,000 companies using the product, 3 million models—which is a shocking number—and over half a million datasets.
But Hugging Face started as a completely different idea. The founder, Clément Delangue, worked at a French computer vision startup called Moodstocks that was eventually acquired by Google. In 2016, he teamed up with two co-founders, one a mathematician and the other a scientist who had worked in patent law, and they started building an AI that could basically talk about anything. This was obviously post-Siri and Alexa, but the goal here was to build something a little bit more like a Tamagotchi for AI: a funny, emotional digital friend targeted at teenagers.
So that was the first version of Hugging Face. The app let users name their bot, text it, send selfies, and trade emojis. It was explicitly marketed as an AI best friend for bored teenagers, and they scaled it pretty significantly. By 2018, it was processing a million messages a day and had more than 100 million messages in total.
They were able to raise a series of financing rounds. The big one that grabbed headlines was that Kevin Durant was in the $1.2 million seed round. Betaworks and SV Angel were also in there in 2017, and then they did a proper $4 million seed round led by Ronny Conway’s A Capital in 2018. People used the product. The company had raised money, and the technology worked well enough to feel magical. But this was still 2018. It was pre-GPT-3, and it just didn’t become a durable consumer business.
But in 2018, Google released BERT. People were excited about it, but it arrived via a research paper, implemented in Google’s TensorFlow framework, and was mostly accessible to specialists. So the Hugging Face team converted BERT to PyTorch and released the conversion for free. Developers loved it.
The team added more models, and eventually thousands of others. Eventually, the team stopped trying to build the one application that everyone would use and instead focused on building tools that everyone could use to create their own AI applications. True picks-and-shovels trade. Instead of trying to pick a winner, you just host every model, and the flywheel starts compounding: more models, more developers, more model creators, more companies. The product basically developed the same simple network effect that made GitHub the default home for software.
Over time, Hugging Face grew from a code library to a place where developers could publish models, version them, attach datasets, discuss changes, report problems, and even build demos through Hugging Face Spaces. Companies could maintain private repositories and collaborate internally.
2019 saw Lux Capital put $15 million in. That injection led to a $40 million Series B in 2021, and Lux came back for the Series C in 2022: $100 million at a $2 billion valuation, joined by Sequoia and Coatue. The big step up came in August 2023, when Hugging Face raised $235 million at a $4.5 billion valuation from Salesforce, Google, Amazon, Nvidia, AMD, Intel, Qualcomm, and IBM. Basically a who’s who of potential acquirers.
It was a good move because it looked sort of like a peace treaty among the major AI infrastructure companies. You get everyone around the table, and everyone’s aligned with the mission. Hugging Face became this neutral territory where it supported competing clouds, competing chips, competing frameworks, and competing models. No single company could control the platform.
Even though they did a number of rounds, Hugging Face only raised $395 million, which is pretty small considering the $13 billion outcome here. It was also very capital-efficient because they weren’t actually buying chips and serving models directly. It was reported that they became profitable in 2025 and still had half the money they had raised. Nvidia eventually came to offer $500 million in late 2025 at a $7 billion valuation, but apparently Clem said no.
So it’s a pretty high revenue multiple. Hugging Face is reportedly doing something around $150 million in ARR. You’re looking at close to 80 or 90x revenue, but it’s growing really significantly, and Nvidia gets something really special here, so they’re happy to pay. They’re trying to position Hugging Face as the front door to the Nvidia ecosystem in AI broadly. As open models proliferate, Hugging Face distributes those models and Nvidia sells chips to the people running them. Exciting outcome for all involved. — John
Model Mayhem
It’s model mayhem again. This week Anthropic, Google and Meta all released major upgrades to their LLMs within roughly 24 hours of each other.
- Anthropic launched Claude Fable 5.1 alongside the restricted Claude Mythos 5.1
- Google released Gemini 3.8 Flash + a cybersecurity-focused version
- Meta released Muse Spark 1.3
It might not be that much of a surprise that Anthropic seems to have the strongest model of the three, with Fable 5.1 scoring 66 on the Artificial Analysis Intelligence Index, which is the highest result ever on the index. This score is also ahead of Opus 5 which got a 63 and Fable, which got a 62.
Anthropic says Fable 5.1 is also cheaper and more efficient, made possible by an improved caching system — it should make ordinary workloads ~25% cheaper and longer-horizon agentic jobs 45% cheaper, the company says.
Interestingly, and as Ben Thompson pointed out, Anthropic is also dropping its no-Zero Data Retention policy, which there was a whole news cycle around a few weeks ago, with Alex Karp and Satya Nadella warning against models that don’t allow companies to have full data sovereignty. The policy, which will be replaced by a new system called Enterprise Frontier Safeguards (EFS), may have contributed to lower Fable adoption among enterprises and it sounds like it was a direct response to user feedback.
Google’s Gemini 3.8 Flash is the company’s third Flash release in 6 weeks. It scored 73.7% on DeepSWE — just behind Opus 5 and competitive with models that cost several times more. Independent testing gave it a 59 intelligence score, which isn’t the absolute frontier, but it’s a great result for a model generating roughly 300 tokens per second.
Meta’s Muse Spark 1.3 did very well on benchmarks, scoring 75.4% on DeepSWE (beating both Opus 5 and GPT-5.6 Sol). It didn’t sweep the board — Opus still beats it on several professional-work and computer-use evaluations — but it got 62 on the Intelligence Index, which is behind only the newest Claude models.
Tons of stuff to think through and discuss here, which we’ll be doing on the show today. — Brandon
Clip Spotlight: Palo Alto Networks’ Nikesh Arora tells us what they’re looking for in new hires
Palo Alto Networks CEO Nikesh Arora says roughly 50% of the company’s new hires over the last six months have been early-career.
Three talent profiles PAN is prioritizing in the AI era:
1 – People who can actually use coding agents
“One set of people we want are people who are AI-savvy, who know how to use AI to get their jobs done better, who know how to use coding agents.”
“These are people who are learning, and they’re typically having to run through hackathons to get to Palo Alto. We don’t look at where you went to school. We care about whether you can use coding agents, whether you can use AI.”
“If we keep following that trend, it’s reasonably likely that we will overwhelm the non-AI-savvy people with more AI-savvy people. That’s kind of ground zero. You’ve got to get that done.”
2 - Cybersecurity experts who know AI
“The second part is there are people who are really good at cybersecurity and know AI and they’d like to get involved.”
“If you want to work in cybersecurity, there’s no better place to work than Palo Alto Networks. We’re the largest cybersecurity company in the world. You want to have an impact? You want lots of proprietary training data? We’ve got it.”
3 – Pure AI researchers
“The third part is raw, pure AI researchers who can do wonders with small language models and large language models.”
“There we have people who are AI literate. We have some PhDs in AI who work here who have a cybersecurity bent, but they can go back to their roots in AI.”
Headlines
Eric Lu solves complete factorization of RSA-260
Axios: Widespread AI outage underway
WIRED on Pangram: Meet the AI police who can make or break careers—in publishing and beyond.
Bernie Sanders and Greg Casar announce the Ban Artificial Superintelligence Act
NYC Mayor Zohran Mamdani says city will impose 1-year moratorium on AI for students
Los Angeles Unified School District bans students from accessing AI on school-issued devices
Nvidia closes $12.93B HuggingFace acquisition
Sequoia’s Konstantine Buhler: The Cognitive Revolution
USDOT Secretary Sean Duffy announces “Americas Great Corridors of Commerce”
wordgrammer: Prediction: AI will collapse
WSJ: See the New Building Techniques Turbocharging the Data-Center Boom
Axios: The Information bets bigger on video with new AI show
WSJ: Steve Ballmer and the Clippers Get the NBA’s Doomsday Hammer
Posts of the Day
Special thanks to our sponsors: Ramp, Shopify, CrowdStrike, MongoDB, NYSE, Codex, Public, Console, Railway, Figma, and Cisco.