Spending on AI Is Becoming Almost Impossible for Businesses to Budget

First, there was #tokenmaxxing, whereby American businesses encouraged their workers to use as much AI as possible. Then came the bill.

Companies started to realize they need to be more careful about counting their tokens, the units that measure AI use. And that’s hard to do: A recent study found that only 11% of nearly 400 businesses surveyed were able to accurately forecast AI spending.

Unlike traditional software, AI behaves more like a human worker: It takes action, makes decisions, sometimes even makes mistakes—all on the clock.

While more-advanced models tend to cost more per token, they can sometimes perform tasks more efficiently, leading to lower overall costs. Likewise, asking a “cheap” model to do something it isn’t suited to handle could cause a token run-up.

Researchers from Stanford University, Carnegie Mellon University, and the University of California, Berkeley, as well as Microsoft Research put this to the test earlier this year, running models through more than 6,800 tasks spanning math, programming, science and other areas. In 32% of cases, lower-priced models actually cost more than higher-priced models.

“The practical takeaway is clear,” said Lingjiao Chen, one of the researchers. “Price alone should not be used to infer which model is actually cheaper.”

Here’s an example of the same prompt presented to two Google models, the Gemini 3.1 Pro—generally used for tasks that require reasoning and judgment—and the less expensive, lighter, speedier Gemini 3 Flash.

The pricier Pro model finished in 85 steps, whereas the less expensive Flash model went through nearly 1,000 steps…then failed.

The less expensive model failed after running up $14 worth of token use, while the pricier one succeeded for just $1. While this result could have been an anomaly, it is something that can happen in regular model use.

Even the same model might use a different number of tokens each time it completes a task.

Suppose you pay a lawn service $20 an hour to mow. The first week, it takes two hours and costs $40. A week later, the job takes eight hours and costs $160. On week three, the job takes five hours, but only half your lawn is trimmed.

Here’s an example of that, from the researchers’ data. We selected these two prompts to demonstrate the variability of results.

Google, Anthropic and OpenAI—whose models researchers used for this testing—have since released new models that perform better in industry benchmarks.

“Some prompt-level fluctuation is inherent to AI, and our testing shows this averages out across a high volume of real-world, diverse workloads,” said a Google spokeswoman in a statement. “Total costs depend on many factors for a given task,” she added, “which can make it hard to forecast new and evolving technology with precision.” She said the company offers customers more control via spending caps and flexible pricing.

As more companies add AI to their daily workflows, they will also need strategies to manage their use, such as training employees on which model to use for a particular task. Otherwise, they could end up with IT bills that come out of a black box.

News Corp, owner of The Wall Street Journal, has a content-licensing partnership with OpenAI.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论