How to Price an AI Agent Feature When Usage Swings Wildly
Flat per-seat pricing on an AI agent feature works until one customer's workflow burns through thousands of tokens in a weekend, and then your margin on that account goes negative.
Quick Summary - TLDR:
- A flat per-seat fee on an AI agent feature has no ceiling on your cost side, so one heavy user can erase the margin on an entire account
- Zapier and Intercom both moved toward usage-based or hybrid pricing for their AI features after flat bundling made costs unpredictable
- A base fee plus metered overage only works if you set the included allotment using your actual cost data, not a guess
- Credit pools that roll unused capacity into a shared balance reduce the chance that a single spiky user blows your margin
- A cost-plus floor, where you never charge less than token cost times a fixed multiple, is the backstop that keeps the whole model from going underwater
Here's the problem nobody puts in the pitch deck: an AI agent feature doesn't cost you a predictable amount per user the way a SaaS seat does. A seat costs you database rows and some compute, more or less flat no matter how the person uses the product. An agent costs you tokens, and tokens scale with how much the agent reasons, how many tool calls it makes, and how long the task runs. Two customers paying the same $49 a month can cost you $2 and $200 respectively, and you won't know which is which until the invoice from OpenAI or Anthropic lands.
Founders keep bolting flat per-seat pricing onto agent features anyway, because it's familiar and it's what their existing billing system already does. It's also how you end up subsidizing your heaviest users with revenue from your lightest ones, and heavy users don't churn, they stay and keep burning tokens while your gross margin on that account slides toward zero.
A traditional SaaS feature has a cost curve that's basically flat against usage. An AI agent feature has a cost curve that moves in direct proportion to usage, often non-linearly, because agentic workflows chain multiple model calls together and each step adds tokens. A support agent that resolves a simple password reset might burn 2,000 tokens. The same agent debugging a multi-step integration issue for an enterprise customer might burn 80,000 tokens in one conversation because it's calling tools, re-reading context, and retrying failed steps.
That variance is the whole problem. You can't price for the average case because the average hides a tail of power users who will find every edge of what your agent can do and use it there. Startups that ignore this discover it the hard way, usually about 60 days after launch when the first finance review shows COGS eating into what looked like a healthy 80% gross margin.
Intercom ran into exactly this with Fin, its AI customer service agent. Fin originally launched priced per resolution, a flat $0.99 per successful resolution, which is itself a form of usage-based pricing rather than a seat fee, precisely because Intercom's own cost to run each resolution varies by conversation complexity. Charging a flat seat price for unlimited resolutions would have meant eating the cost difference between a one-turn FAQ answer and a long multi-tool troubleshooting session. Zapier made a similar move, shifting its AI-powered Zaps to consume task credits based on actual usage rather than bundling unlimited AI actions into a flat tier, after finding that automation-heavy users were running far more AI steps than the flat pricing accounted for.
The three mechanics that actually fix this
Base fee plus metered overage is the first piece, and it's the one most founders half-implement. You charge a flat monthly fee that includes a set allotment of usage, say 500,000 tokens or 1,000 agent actions, and anything beyond that gets billed per unit. The part people get wrong is setting the included allotment by guessing at a round number instead of pulling it from real cost data. Before you set that number, run your actual token costs per customer segment for at least 30 days of real usage, then set the included allotment at roughly the 60th to 70th percentile of usage, not the median and not the max. That leaves your typical customer feeling like they got a generous deal while your heavy users pay for what they actually consume.
Credit pools solve a narrower but real problem: the customer who's light in week one and heavy in week three. If you reset usage monthly with no rollover, that customer either gets throttled awkwardly mid-cycle or you eat the overage as a courtesy and never get paid for it. A credit pool lets unused capacity carry forward, typically capped at one or two months of rollover, so the billing smooths out without you having to manually comp anyone. This also removes a support headache: customers hate being told they hit a wall on day 18 of a 30 day cycle, and a pool gives you room to let that slide without changing your unit economics.
The third piece is the one that actually protects your business if the first two are set wrong: a cost-plus floor. Decide on a multiple, say 3x your raw inference cost, and make sure no pricing tier, overage rate, or enterprise discount ever prices below that floor. This matters most when a sales team negotiates a custom deal for a big account and shaves the overage rate to close it. Without a hard floor written into your pricing logic, that one deal can turn a flagship customer into a money loser, and you often won't notice until the next quarterly cost review.
Building the model before you get burned
Start by instrumenting cost per action before you touch pricing at all. You need per-customer, per-feature token cost data, not an aggregate average across your whole user base, because the aggregate hides exactly the variance that's going to hurt you. Pull at least 30 days, ideally 90, segmented by customer tier and use case.
From there, set your base fee to cover your median-cost customer with healthy margin, then set your included allotment near the 60th to 70th percentile as described above. Price the overage rate at your cost-plus floor multiple, not a discount off it. And build an internal dashboard, even a basic one, that flags any account whose actual margin drops below your floor in real time, because the first sign of a broken pricing tier is usually one account, not your whole book.
Frankly, most startups skip the instrumentation step because it's unglamorous and doesn't feel like shipping, and that's exactly why they get burned. The pricing model isn't the hard part. Knowing your real cost per customer is the hard part, and the pricing model only works once you have it.
One more thing worth saying plainly: don't try to make this pricing model simple for the sake of simplicity. A base fee, a metered overage, a rollover pool, and a cost floor is four moving parts, and that's more than a one-line pricing page can explain cleanly. Customers tolerate complexity in pricing when it's transparent and tied to something they understand, usage, the same way they tolerate a cell phone bill with data overage charges. What they won't tolerate is a surprise invoice with no visibility into why, so give customers a usage dashboard showing where they sit against their allotment in real time. That single feature prevents more churn than any pricing tweak will.
Also read: How Much Salary Should a Startup Founder Take After Funding • How to Vet a VC Before Raising by Running Diligence on Them • How Does Milestone Based Venture Debt Tranching Actually Work
This article is posted in Startup News, check it out for more related stories.