How One Startup Is Bringing Wall Street–Style Contracts for AI Tokens

If the number of times I called someone this week and heard their cellphone make the familiar ring of a European phone is any indication, many of you are on vacation. Still, there are few breaks these days in the world of financing AI.
With trillions of dollars due to pour into chips and data centers and companies spending heavily on their own use of the technology, nearly everyone involved is looking for ways to make their spending and returns more predictable.
That’s already spurred the creation of marketplaces and contracts for graphics processing units that let buyers and sellers lock in compute rental pricing. And on Tuesday, CME Group announced an October launch date and details for two exchange-traded GPU futures contracts, pending regulatory review.
Now those efforts are turning to tokens, the units that companies actually pay for when they use AI models.
Compute Exchange, a startup that matches chip owners with short-term renters, is moving into a similar matchmaking role for tokens, The Information has learned. It will offer token forwards, privately negotiated contracts that give enterprises a way to lock in their token prices over a timeframe of as long as six months, hedging against future swings in AI costs.
The token forwards offer a different pricing model than GPU contracts, which typically lock in the hourly cost of computing capacity used to run AI models. Token contracts would instead provide consistent pricing over a much longer time period, and for an item that enterprises actually use in their day-to-day.
Customers can select from a variety of mostly open-weight models offered by one of the six inference providers that have signed up with Compute Exchange. Forwards aren’t quite the same as standardized, exchange-traded futures like the ones CME is planning, but they mark another step in bringing financial market-style hedging to more kinds of AI costs.
Tokens are the snippets of words or phrases that get sent to AI models and come back as its response in a process called inference. It is this action rather than the training of new models that accounts for the majority of the world’s interaction with AI. Inference is set to explode as agents begin to handle more and more of the world’s AI traffic.
Deloitte estimates inference will account for about two-thirds of AI workloads by the end of 2026, up from one-third in 2023.
The early days of AI experimentation saw companies offer employees unfettered access to the technology. Meta Platforms and others defined the era of tokenmaxxing, setting up internal competitive league tables that encouraged employees to experiment with AI early and often, consuming massive quantities of tokens. Then companies including Uber and ServiceNow disclosed that they had blown through their annual AI budgets in a matter of months, kicking off new restrictions on employee spending.
Any opportunity to lock in the price of tokens should help companies, startups and others get a better handle on their spending, according to Val Bercovici, chief AI officer for WEKA, an AI data and memory infrastructure company. WEKA helps AI systems use GPU memory more efficiently.
“Everyone is scrambling to understand their cloud bills,” Bercovici said. “Tokenmaxxing quickly turned from a badge of honor [to be] at the top of the leaderboard to a badge of shame.”
Recent developments including larger and more complex models, larger context windows and agents that can run for days and weeks on their own, as well as emerging cybersecurity threats, have caused an explosion in token consumption, he said. That makes it difficult for companies to get a handle on everything they need to know to make informed choices about how to best manage costs.
“This market is far too complex and too opaque, and you need clearinghouses to add transparency,” he said.
In parallel with launching token forwards, Compute Exchange is creating something called standardized token units, a methodology that benchmarks pricing proposals and model performance.
According to Carmen Li, CEO of Compute Exchange, the company’s pitch is that token forwards offer an alternative to the predominant method of consuming AI today, which is accessing models through broad-based application programming interfaces.
Relying on APIs can make it difficult for users to manage costs or measure performance because the model companies charge based on token usage, while also retaining the ability to change how quickly tokens are processed—for instance, they can change a user’s place in the line or otherwise throttle usage without disclosing it, she said. Users can be relegated to paying on-demand or spot prices rather than a locked-in price, which Bercovici described as the reserve price.
Bercovici said spot prices for tokens have become increasingly volatile in recent months, something that’s apparent on OpenRouter, where prices for new models used to be stable over hours, days and even weeks. Now those prices can change hourly, he said.
The inference providers Compute Exchange has signed up predominantly run models developed by companies like DeepSeek, Moonshot AI and Alibaba that are significantly cheaper than the frontier models from Anthropic and OpenAI. The models are also increasingly competitive in terms of performance, Li said.
Companies in need of tokens will visit the Compute Exchange website and share a snapshot of precisely what they need—which model they want to use, ideal pricing, throughput rates and other specifications. The providers, whom Li declined to name, will compete to win that bid.
Once a match is made, the consumer of tokens will know who the provider is and will be obligated to use the number of tokens they have purchased for the duration of the contract. Compute Exchange will collect a 4% fee for the service.
Many companies may not want to commit to specific models over many months, especially with new and better models released seemingly every week. In response, Compute Exchange and its providers may allow companies to convert to different models within the confines of an existing contract, Li said.
Of course, as with any nascent financial market, token forwards will only be useful if there are enough buyers and sellers. Forwards also introduce the chance of counterparty risks, leaving one side exposed if the other can’t pay up or can’t provide the promised inference capacity.
But for enterprises looking to harness AI, the ability to get consistent token pricing may become a way to survive in a world that is changing every day, according to Bercovici. “If you are trying to plan for this or do any type of capacity planning, you are chasing your tail all the time,” he said.