# You pay your AI agent per resolution. Who decides what resolved means?

More AI support agents are now billed per resolution, and the vendor does the counting. How to write your own definition of resolved, check it, and put it in the contract before the first invoice.

By Precision Code Studios. Published October 11, 2026. Category: AI. https://www.precisioncodestudios.com/insights/ai-agent-per-resolution-pricing-define-resolved

The first monthly report from a new AI support agent usually looks wonderful. Thousands of conversations, a large share of them marked resolved, and a cost per resolution that is a fraction of what a human contact costs. Then someone on the support team does something nobody asked them to do. They open fifty of the conversations marked resolved and read them.

Some are exactly what was promised: a customer asks to change a delivery address, the agent changes it, the customer says thanks. Others end with the customer typing "never mind" and closing the window. A few were handed a link to the returns policy, and the same email address turns up in the phone queue forty minutes later. Every one of those conversations carries the same label and, if the agent is priced per resolution, the same charge.

If you are about to buy or build a chat or voice agent for support, lead qualification or scheduling, this is a conversation to have before the contract is signed, not after the third invoice.

## One word now does two jobs

Several of the large support platforms have moved their AI agents away from per-seat licences towards outcome pricing, where you pay for each conversation the agent resolves. As pricing models go, it is a fair one: you pay when the thing works. The catch is that someone has to decide when the thing worked. By default that someone is the vendor, applying a rule you did not write, often with a language model grading conversations that a language model also conducted.

The same number then travels into a second document. The business case for an agent nearly always contains a line like this: if the agent resolves this share of contacts, we need this many fewer hours on the support desk next year. Inflate the resolution rate and the staffing plan inherits the error, which surfaces months later as a queue nobody budgeted for. Gartner predicted in June 2025 that by 2027, half of the organisations that expected to cut their customer service workforce significantly because of AI would abandon those plans. Plans like that fail for many reasons. Generous counting is one of the cheapest to prevent.

> **The rule:** If the vendor's label is also your invoice, the definition of resolved is a contract term. Write it yourself, and do not let the party being paid be the only one checking it.

## Resolved, from weakest evidence to strongest

Think of it as a ladder, and notice that most dashboards stand on the bottom two rungs. A conversation ending proves only that it ended. No handoff to a human proves the agent kept the customer, which is containment: useful for planning volume, silent on whether anyone was helped. Customers do not arrive neutral, either. In a Gartner survey of 5,728 customers carried out in December 2023, 64 percent said they would prefer companies did not use AI in their customer service at all. When a customer who never wanted to talk to a bot goes quiet, you cannot read that silence as satisfaction.

*Figure: A five-rung ladder of evidence for calling a conversation resolved. From weakest: conversation ended, no handoff to a human, quiet on every channel, customer confirmed, action in the system of record. A line above rung two marks where billing should start. Rungs one and two measure containment. Only rungs three to five say anything about whether the customer was helped.*

The useful rungs sit higher. Silence across every channel for a sensible period is good evidence. A customer saying in their own words that it worked is better, and rare. Best of all is a change in a system of record, meaning the database or business application that holds the truth for that process: the refund appears in billing, the appointment moves in the calendar, the address changes on the order.

That top rung has a consequence buyers rarely notice. An agent that only answers questions can never reach it, because there is nothing for it to change. An agent wired into your systems, one that actually rebooks and refunds, can be measured with the same evidence your finance team would accept. When you compare options, how the agent will be measured is as much a design decision as which model it runs on.

## Write the definition before the pilot

A workable definition has five parts. None is exotic, and all of them are far easier to agree before anyone has a number to defend.

- **Evidence for each intent.** For each kind of request, name what proves it was done. A booking change is proved by the booking record. A question about opening hours leaves no record, so it can only ever be proved by silence.
- **A window that fits the problem.** No further contact about the same issue for a set period. Seven days suits most questions, but a billing problem has to survive the next invoice and a delivery problem the delivery date, so set the window per intent.
- **Every channel, joined.** A customer who leaves the chat and rings you has not been helped, and you only see it if chat, phone and email are matched to the same person. With anonymous web chat this is the hardest engineering in the whole exercise. Do it anyway.
- **Honest exclusions.** A customer who asked for a person and did not get one is a failure, whatever happened next. So is a conversation the agent closed because it ran out of things to say.
- **A blind audit.** Every month, someone who cannot see the agent's label reads a random sample and marks each conversation proven, unproven or failed. If a model does the grading, compare its labels with the human ones: that agreement rate tells you how far to trust the dashboard between audits.

*Figure: Three swimlanes for chat, phone and email. A chat about returning an item is marked resolved by the AI, the same customer phones 39 minutes later because the label never arrived, then emails the next day asking for a refund. Per-channel reporting counts a success, a handled call and an open ticket. The customer has had one problem the whole time.*

Pay attention to the middle category. Plenty of conversations will be unproven rather than failed: the customer got an answer, left, and never came back, which might mean the answer was perfect. Decide in the contract how those are billed. Some buyers pay for unproven conversations at a reduced rate; others pay only for proven ones and accept a higher unit price. Either position is defensible. Leaving it undecided is how the argument ends up happening in month four, with the renewal on the table.

## A worked example, with the arithmetic shown

Take an illustrative month: 10,000 conversations, of which the platform labels 6,000 resolved. You draw 200 of the resolved ones at random and check them against phone and email records. Suppose 30 of those customers contacted you again about the same issue within a week, and another 20 transcripts end with the customer asking for a person and not getting one, or abandoning halfway through an answer. That is 50 out of 200, or a quarter, that fail your definition.

Scaled up, about 1,500 of the 6,000 billed resolutions were nothing of the kind, and the true rate is nearer 45 percent than 60. The invoice is 1,500 units too high, which is irritating. The staffing plan has a different hole. The repeat contacts, 15 percent of the sample, scale to roughly 900 customers who came back to your people with the problem the plan assumed was gone, and a customer on a second attempt is rarely calmer or quicker to help. The other 600 or so are invisible. Some found the answer elsewhere, some will surface next month, and some may have taken their business to a competitor. No dashboard will ever count that last group for you.

Two notes on that arithmetic. A sample of 200 is enough to tell a false-claim rate of about a quarter from one of about a tenth, but not a quarter from a fifth, so do not chase small movements from month to month. And run the same audit on your human team before the agent goes live, because people close tickets that come back too. Without that baseline you are comparing the agent either with perfection, where it will look worse than it is, or with nothing, where it will look better.

## Questions to put to any vendor, or to whoever builds it

Whether you license a platform or have an agent built for you, the same questions apply. A good supplier will have the answers ready and will not be offended by any of them.

- **Who decides a conversation is resolved, and can I read the rule?** You want a written definition, not a description of a model's judgement.
- **Is the grading done by the same model that answered?** If so, ask how often it is checked against human readers, and ask to see the most recent check.
- **Can the agent see my phone and email history?** If not, it cannot know about the customer who rang afterwards, and its figures are containment under a better name.
- **Can I export every transcript with its label?** Your audit depends on raw data in a format you can open without the vendor's dashboard.
- **What happens when a customer asks for a person?** The handover should be quick, and those conversations should never count as resolved.
- **If my audit disagrees with your count, what changes?** A credit, a re-run, a revised definition. Any of those is better than a shrug.

## Sometimes the answer is not an agent

The transcripts can tell you something more useful than a resolution rate. If a large share of your contacts are customers asking where their order is, the cheapest fix may be a better tracking page and a proactive email, not an agent that answers the same question politely ten thousand times a month. Build the measurement before you choose the tool and it will show you that before you have paid for anything.

Where an agent is the right answer, the definition tends to improve the agent too. Once the team can see which requests end in proof and which end in a customer phoning back, they know exactly which workflow to wire into which system next. We would rather ship an agent that claims fewer resolutions and can prove every one than one with a handsome dashboard and a busy phone line.

A customer who gives up is the cheapest conversation you will ever have, right up until they stop being a customer.

## The short film: transcript

Paying per resolution? Define resolved first.: the short version (1:13).

Plenty of support teams now buy AI agents the way they buy electricity, by the unit. The unit is a resolution, and almost nobody asks who decides what one is.

That single number does two jobs. It sets the bill every month, and it's the figure the business case quietly uses to decide how many people you'll need on the support desk next year.

Picture someone who asks the chat about a return, gets a link, closes the window, and rings your phone line forty minutes later. One dashboard calls that a win. Your phone team calls it Monday.

Write the rule yourself before the pilot, one kind of request at a time. An address change showing up on the order is proof. A billing query that stays quiet past the next invoice is good evidence. A closed window is neither.

Two hundred transcripts is enough to tell a quarter of false claims from one in ten. Run the same check on your human team first, so you know what good looks like before the agent arrives.

Agree the meaning of one word before you sign, and most of the later arguments never happen.

Read the full article at Precision Code Studios.
