How big is the open-model threat to AI hyperscalers?

In a speech earlier this year, Vivian Balakrishnan, Singapore’s Foreign Minister, outlined how he increasingly uses AI to organise his life. AI agents draft briefs for him ahead of foreign trips, his speeches and even answers to parliamentary questions.

He wasn’t talking ChatGPT, Claude, Gemini or even Grok. As a self-confessed ‘geek and tinkerer’, Balakrishnan went to GitHub, downloaded code, and spun up a system, then communicated with his bots through WhatsApp.

Presumably this sits humming away in a state-of-the-art Singaporean data centre somewhere?

all this stuff – my most daily-used agent – is running off a Raspberry Pi, which is at least two or three years old. All it has is eight gigabytes of RAM.

all this stuff – my most daily-used agent – is running off a Raspberry Pi, which is at least two or three years old. All it has is eight gigabytes of RAM.

Alphaville is not a tech blog. But given how the fate of half of US economic growth, a good chunk of earnings uplifts and global stock market gains, and a fast-growing portion of private and public credit appear to hang on the AI trade, we increasingly find ourselves having to make an effort.

Known unknown threats to the AI trade

One of the ‘known unknown’ threats that S&P Global Ratings outlined last week hanging over the biggest companies driving the AI juggernaut is the notion that open-weight models close the performance gap with expensive frontier models.

We can break this down further into two component risks. First: that queries put to frontier LLMs — which require centralised cloud infrastructure — can be answered more quickly and more cheaply by the kind of locally hosted small language model setup that Balakrishnan uses to organise his life. And second: that big complex open-weight models become pretty indistinguishable from big complex closed-weight models, can be accessed at a fraction of the cost, and come with riders that are valuable in their own right.

Little ickle AI models

The degree to which SLMs are displacing frontier models is far from clear. After all, there are only so many geeks out there with the gumption to spin up their own AI rigs. But, if such small language models got good and the economics of switching were right, we can see that, as well as all the nerdy types unafraid of code, it might make sense for companies to flip their low-grade queries into locally based models.

So how good are these desktop AI models?

Our first idea was to jump on to Hugging Face and pop in a prompt asking a selection of cloud-based frontier models and bite-sized SLMs to write the first verse in a sonnet about avocados. This proved fun.

Our second idea was to seek out research from people who are actual experts. This proved useful.

A team of Stanford University computer scientists led by Jon Saad-Falcon and Avanika Narayan has been benchmarking the relative performance of SLMs against frontier models for the kind of everyday tasks that dominate chat requests since 2023. For relatively simple tasks, the academics reckon SLMs are pretty good.

In fact, the researchers find that SLMs can find the same — or better — answer to the kind of queries that dominate current chat requests almost every time, and are catching up fast on reasoning tasks:

Admittedly, most of the queries are not going to be too computationally testing. But versus chucking everything into cloud-based LLMs, routing such queries locally cuts compute and energy costs by between 60 and 80 per cent. Moreover, while the equipment they’re using to run these models — an Apple M4 Max — retails at around 15x the price of a Raspberry Pi 4 Model B 8GB starter kit, we’re still talking equipment costs in the Tom Ford gilet ball-park, rather than anything approaching the monthly pay cheque of an XTX Markets intern.

So we can see how SLRs might appeal to management seeking to repair their bottom lines and reputations for cost control after token-maxxing experiments backfired so spectacularly.

In fact, given quite how basic most AI queries appear to be, the researchers’ results sound like the sort of thing that might even challenge the economics of building some of the myriad mahoosive data centres being planned. This is what Joachim Klement, head of equity strategy at Panmure Liberum, reckons. In his excellent blog — through which we first encountered the Stanford paper — he writes that:

[i]f their results are true, then we will hardly need any data centres in the future, and the hyperscalers are wasting hundreds of billions of dollars in investments. . . . [and] . . . one can replace data centres and their expensive cutting-edge semiconductor infrastructure in four out of five use cases.

[i]f their results are true, then we will hardly need any data centres in the future, and the hyperscalers are wasting hundreds of billions of dollars in investments. . . . [and] . . . one can replace data centres and their expensive cutting-edge semiconductor infrastructure in four out of five use cases.

We’re honestly not plugged in enough to the unit economics of AI, or the likely shape and nature of future AI demand, to make this call. And we note that the paper — and indeed the more general research project in which it nests — is predicated on solving the question as to how to scale compute in a world that lacks the data centres, and maybe electricity, to manage exponentially growing demand.

Moreover, we note that one of the co-authors is John Hennessy, a man Marc Andreessen once called “the godfather of Silicon Valley” and current chair of Alphabet. If he thought that SLMs were going to undo the economics of data centre build-out, he’d presumably be pressing Alphabet to maybe not splurge a trillion dollars and change into capex between now and the end of 2029. As far as we know, he hasn’t pressed them to stop.

What about cloud-based open models though?

The second aspect of S&P’s ‘known unknown’ threat to the big AI economic/credit/everything trade comes from the use of big open-weight cloud-hosted models. Cloud-based open-weight models, just like their closed-weight counterparts, gobble up data centre processing capacity. As such, it’s hard to see what direct problem they might pose to the economics of data centre build-out. Although it’s easier to see how they might be a problem for AI labs like Anthropic and OpenAI.

Open-weight models are a lot like closed-weight frontier models — like Anthropic’s Claude, OpenAI’s ChatGPT, Google’s Gemini and SpaceX’s Grok. It’s just that users don’t pay these labs for their use. In fact, Hugging Face — the platform recently acquired by Nvidia — hosts over 3mn open-source and open-weight models that you can just download for free.

Some of these are small models, like IBM’s Granite 4.2 3B model, that you could pop on your laptop to maybe summarise pdfs, classify documents, chat with you or compose odes to avocados. Good luck though trying to download Moonshot’s Kimi K3 on to your PC. Operating across 2.8tn parameters, it’s just a different animal. According to Citi Research, Kimi K3 trounces every closed-weight frontier model ever built prior to [checks calendar] three months ago.

Sadly, working out the break-even usage rate for a firm considering downloading an open model and running its own system — a sort of scaled-up version of Singapore’s Balakrishnan Raspberry Pi — versus taking an enterprise subscription to one of the proprietary systems is too complex a task for Alphaville to calculate easily.

We can see that going down the frontier closed-model route means avoiding the hassle of sourcing compute and storage. And everyone’s going to have to pay for the power and compute used in inference, aka usage, whichever route they choose. But the open-model route means you avoid paying through the nose for shiny closed-model tokens.

One way to see which way the wind is blowing on open-model versus closed-model usage is by looking at data from router firms like OpenRouter, the New York start-up that Stripe agreed to buy last month.

AI routers work a bit like an AI query triage service with a toll booth strapped on. Clients rock up, basically model-indifferent, and rather than tie themselves to any given model, they just send their queries to the router, which then flips them on to the lowest-cost model that will produce a good enough response for whatever the task at hand. Asking really tough closed-end frontier-model-worthy questions? To a frontier model they go.

Increasingly, the share of queries that are truly closed-weight frontier-model-worthy is declining. Back at the start of the year, three-fifths of its queries were routed through to closed-weight proprietary models. The latest share is just a quarter.

The decline has been accompanied by a steady trickle of news stories featuring companies switching away from closed-weight models. MainFT ran a great piece talking about how DoorDash, Siemens and Airbnb have been making the switch in pursuit of cost savings back in July. And the New York Times reported last week that AT&T has been switching higher and higher proportions of its AI usage towards open models. As they wrote:

By May, open models accounted for 20 percent of AT&T’s A.I. use. That has since risen to 40 percent and may jump to 60 percent in the coming months, Mr. Markus [the company’s chief data and AI officer] said in an interview.“We believe it could go much, much higher,” he said, adding that AT&T was saving up to 80 percent on A.I. costs compared with earlier this year.

By May, open models accounted for 20 percent of AT&T’s A.I. use. That has since risen to 40 percent and may jump to 60 percent in the coming months, Mr. Markus [the company’s chief data and AI officer] said in an interview.

“We believe it could go much, much higher,” he said, adding that AT&T was saving up to 80 percent on A.I. costs compared with earlier this year.

Moreover, open models have other appeals.

You’re not being paranoid if everyone is out to get you

MainFT reported on Wednesday that the world of mathematics has been thrown into a complete tiz this week. OpenAI claimed to have found a “singularity” in the Navier-Stokes equations in three dimensions, one of the six unsolved “Millennium Prize Problems” set by the Clay Mathematics Institute. In a statement posted to X.com the lab wrote, “The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.”

A few hours earlier, Tristan Buckmaster, a professor of mathematics at NYU who has been working on closely related problems for some time, posted his own statement in which he went to great lengths to NOT EXACTLY formally accuse OpenAI of stealing an almost complete proof of Navier-Stokes from his sessions on Codex, GPT-5.6 Sol, and Astra. But he did not sound at all happy, and it’s worth a read in full if you’ve got five minutes.

OpenAI’s response was to issue a QT of its main post on X.com congratulating Buckmaster on his findings, saying that OpenAI’s researchers hadn’t seen his work, but also that “[w]hile unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.

😬😬😬

It doesn’t take an Oxford University AI research fellow to wonder what this means for companies already paranoid about information security.

The heaviest prospective users of the most advanced frontier models tend to be those in the most knowledge-intensive industries, whose existence rests on the security of their data — be they quant hedge funds and prop trading firms building proprietary algorithms, biotech and pharmaceutical companies, defence and engineering companies, law firms, whatever. And while most companies have been pretty happy to sign off proprietary models, any company protective of its IP will want to keep its data out of training algorithms that might be made available to competitors. And this is an idea that looks like it is gaining commercial traction.

A bunch of law firms have signalled that they are going down the sovereign AI route, and we guess they’ll be tweaking open models to get there. For example, MainFT reported on Thursday that the US’s second-largest law firm Latham & Watkins was buying up Nvidia hardware on which it would host Nvidia Nemotron 3 open-weight models that it plans to fine-tune:

“Sometimes we may have information that is so sensitive, client information that we really want to protect, we don’t want to put it to any cloud vendor,” said Rene Mendoza, the firm’s chief information officer.

“Sometimes we may have information that is so sensitive, client information that we really want to protect, we don’t want to put it to any cloud vendor,” said Rene Mendoza, the firm’s chief information officer.

And Tareq Islam, a strategic adviser to ApexE3, a capital markets AI infrastructure firm, tells Alphaville that some London-based asset managers, as well as not wanting to build dependency on any single proprietary model, are reluctant to post their most confidential data and intellectual property to closed-model providers. Clients of ApexE3 include Vanguard, the world’s second-largest asset manager.

All in all, we can see why Deloitte launched a new ‘Open Model Engineering practice’ to help companies around the world scale open models earlier this month.

But what has this got to do with data centres and the AI trade?

The whole world won’t be running off open-weight models any time soon. But even if they did — as long as these models were the big cloud-based models — this still sounds like quite the bull case not only for the construction boom that has sustained US economic growth and the chip companies that supply all the data centre hardware, but also for the productivity miracle that’s yet to arrive. However, it’s not clear whether this is really a bull case for all of the hyperscalers.

OpenAI and Anthropic are private companies yet to list or issue rated debt. While totemically and technologically significant, companies at this stage of their lifecycle are really not supposed to be financially important. Certainly not financially important enough to impact the global economy.

But their centrality to the industry that they dominate — and will perhaps continue to dominate — and the degree of circular financing in place means that this doesn’t look nailed on.

While it’s now over five months old, it’s worth reminding ourselves of this amazeballs wiring diagram that the Global Valuation, Accounting and Tax team over at Morgan Stanley, led by Todd Castagno, put together — and which Louis turned into an interactive wonder:

Since the chart was put together, a lot of things have happened, not least OpenAI shutting down its AI video app Sora, which killed the Disney/OpenAI deal. But also the renegotiation of the OpenAI/Microsoft relationship in April, the $100bn backstop Nvidia provided in August, and a bunch of other things. Please mentally adjust the chart as necessary.

Hyperscalers’ core businesses are generally hugely profitable and cash-generative. The data centre capacity that they own, or to which they have rights, would, in a compute-constrained world, be valuable regardless of which models are making them purr. Bascially, they have sufficient financial strength to be able to survive some big bets that turn bad.

But the leading AI labs are cash-burning monsters, kept alive by the faith that their hyperscaler investors have in them. This is faith not only in the evident brilliance of their inventions, but also in the promise that their genius can be monetised at a pace that eclipses payments due to service debt, lease commitments, forward power agreements, the works.

There are hurdles to clear in this path to monetisation, not least in fending off the competitive threat that open models present. And in a scenario where a leading AI lab fails to clear all the hurdles, it’s frankly crazy-hard to establish quite where all the receipts are held.

The Bank for International Settlements, the so-called “central banks’ central bank”, warned in June that worse than expected returns for hyperscalers could lead to an investment bust that sinks global stock markets. This could be a case of “they would say that, wouldn’t they?” Because no one thanks the BIS when things go right.

So sure, in the best of all possible worlds, AI boosts economic productivity, eliminates world poverty, fills government coffers with revenue to repay its debts, gives us all more time to play, and steers clear of wiping out all human life.

But we can see that this best of all possible worlds needn’t be one that is problem-free for anyone considering the potential risks and rewards of piling into hyperscaler debt, or indeed frontier lab stocks.

Additional reporting by Oliver Hawkins.

Further reading:
Companies turn to Chinese AI models to cut costs (MainFT)

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论