What’s really going on with AI token price deflation?

Chances are you’ve seen this chart a few times, but what does it mean?

Silicon Data’s Token Expenditure Index is a blended dollar-cost benchmark for one million tokens. From a series high of $2.0651 in late May, it’s down 52.4 per cent.

Token price deflation, it’s often argued, is bad for companies that make tokens. The market for units of synthetic intelligence is heading towards commoditisation, because anything that can be mass-produced will be. Most enterprises are still in the exploring stage with AI and they’re already on sale at less than half price.

What the token expenditure index doesn’t show is total token expenditure. It’s a measure of average price that’s weighted by estimated usage and normalised for whether they’re for input or output. (LLM labs routinely charge about five times more for output tokens than input tokens because they want to encourage users to ask a lot of complicated questions, then make the money back by delivering the answers.)

Charles-Henry Monchau, CIO of Syz Group, has a LinkedIn-ish explainer on LinkedIn about why the recent collapse in observed pricing is because token mills are getting more efficient, and because enterprises have been routing lower-priority requests to cheaper open-weight LLMs, causing a bit of a price war:

Technical deflation expands margins and is unambiguously good for the ecosystem. Mix deflation is neutral for aggregate spend but corrosive for the frontier labs’ pricing power. Competitive price deflation is a margin war. The index blends all three into one number and tells you nothing about the composition. Anyone using this chart as a directional signal without decomposing it is trading on noise.

Technical deflation expands margins and is unambiguously good for the ecosystem. Mix deflation is neutral for aggregate spend but corrosive for the frontier labs’ pricing power. Competitive price deflation is a margin war. The index blends all three into one number and tells you nothing about the composition. Anyone using this chart as a directional signal without decomposing it is trading on noise.

Seaport Research Partners’ tech analyst Jay Goldberg — best known for maintaining a heterodox “sell” recommendation on Nvidia — has tried to build a better mousetrap.

The data he collects shows a similar pattern to the Silicon Data index, with token prices plunging for frontier LLMs from $50 per million tokens in late 2023 to about $6 now. More interestingly, he finds the median token price has held fairly steady at just below $1 per million since mid-2024 as the LLM marketplace became more competitive. A floor price of about $0.11 (basically just the cost of electricity) has barely budged.

The same is true for the ceiling price. Narrowing the sample to the most expensive flagship models, there’s been very little movement in pricing since late 2023. Deflation at the frontier end is mostly just a function of more models launching and dragging down the average.

Pricing data is only so useful without also estimating the cost of milling a token. Goldberg estimates, on what look like fairly typical inputs, that the cost of production is falling much faster than the market price:

© Seaport Research

(We’re not going to show all the working for this chart, sorry. Stuff like assumed uptime, rental rates and depreciation creates a lot of heated argument more suited to the trade press than a markets blog. Goldberg says his cost assumptions are reasonable and we’ve seen no reason to argue. Also note that the cost inflation forecasts for the 2027 and 2028 hardware are mechanical estimates using what can be extrapolated from pre-release spec sheets, which might not reflect real-world performance.)

What’s revealed is a two-tier market. Gross margins at the (all-American) frontier are about 70 per cent. Meanwhile, the (mostly Chinese) open-weight LLM shops are operating at gross margins of approximately 20 per cent. Profit margin differentials between operators could be disrupted, of course, but appear to have held pretty steady for the past three years.

Training is the elephant in the DCF. Tokens relate to inference alone. Capital requirements for model training are a whole other flaming dumpster:

All of which leads us back to two fairly simple observations.

The S-1 flotation docs for Anthropic and OpenAI, should we ever be allowed to see them, are going to show some decent operating metrics. There’s a lot to like about year after year of steady gross margins in the 70 per cent ballpark when businesses are doing gangbusters growth.

But to make sense as investments, they’ll need to keep defending competitive leads while somehow calling a ceasefire to the open-model performance arms race. Maybe scaremongering around safety and protectionist cartel formation within national boundaries will do the job. Or maybe it won’t.

This is probably where a majority of prospective investors are already, let’s be honest. It was just nice to see some new data that will confirm a lot of priors.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论