AI support: 90% deflected, 40% actually resolved
A support system can report 90% deflection on a 40% resolution rate, because a customer who gets an unhelpful answer and gives up counts as deflected. The most-cited deployment reversed.
TL;DR. Customer service is the largest AI deployment by volume and the worst-measured in this series, for one specific reason. Deflection counts a ticket that ended without a human, which includes a customer who received an unhelpful answer and gave up. A platform can report 90% deflection while true resolution sits around 40%, and vendors report whichever number is highest. The most-cited deployment automated two thirds of chats, cut resolution time from eleven minutes to under two, reduced headcount by roughly 40%, and then reversed course in 2025 after satisfaction fell on complex tickets. The fastest read on any vendor is not their case study but their pricing: a supplier that bills per resolution cannot profit from suppressing a handoff, and one that bills per containment can.
---
Three metrics get used interchangeably in this field and they measure different events.
Deflection is the share of conversations a human agent never touched. Containment is the share that stayed in the AI channel without transfer. Resolution is the share where the customer's problem was actually solved.
The first two count the same thing as a success: a customer who asked a question, got an answer that did not help, and closed the window.
A platform can therefore report 90% deflection on a 40% resolution rate, and both numbers are accurate. One measures whether a human was involved. The other measures whether the customer got what they came for.
Vendors report whichever is highest. Independent benchmarks measure resolution or re-contact, which are harder to game. And almost every published comparison in this industry sets a vendor's deflection figure against another party's resolution figure without noticing they are not the same quantity.
What the honest numbers look like
Once you insist on resolution, the picture is more modest and more useful.
Cross-customer averages published by one vendor that bills on resolution report 67% in 2025 rising to 76% in 2026. Synthesis across vendor disclosures and independent aggregates puts roughly two thirds as a support-case median, 70 to 75% as a strong deployment, and above 80% as best-in-class, all on a favourable mix of well-structured intents.
An enterprise aggregate from a major platform reports 41.2% tier-1 deflection, which is a different metric measuring a different thing and appears in the same comparison tables anyway.
Two contextual points make these numbers readable.
Rule-based systems run 15 to 25 percentage points lower on containment and deflection than retrieval-based ones. That gap is the actual size of the recent improvement, and it is substantial.
And a low rate is not a failure in every sector. Professional services and healthcare show lower containment because the questions are truly harder and regulatory caution routes more conversations to people. A 35% containment rate in a legal context can be excellent performance, and comparing it against a retail benchmark is comparing different problems.
The deployment everyone cites, including its reversal
One case dominates discussion of this field and it is usually cited at the halfway point.
A fintech moved to AI-first support in 2023. Within weeks the system handled 2.3 million conversations, around 66% of chats. Average resolution time fell from eleven minutes to under two. Repeat inquiries dropped about 25%. The company said the system did the work of 700 full-time agents, froze hiring, and let attrition reduce global headcount from roughly 6,500 to 3,800.
Those numbers are real and the efficiency was real.
Then in early 2025 the chief executive reversed course, announcing hiring of remote service agents so customers would always have the option of a person. Internally, executives had acknowledged that AI-only support produced lower quality. Satisfaction had fallen on complex and emotionally charged tickets, and confidently wrong answers had surfaced on a minority of edge cases.
The reversal is more informative than the original result, and the lesson is not that the deployment failed. Two thirds automation with a tenfold reduction in handling time is a large success. The lesson is that the last third was not a smaller version of the first two thirds. It was a different problem, containing the cases where being wrong is expensive and where the customer wants acknowledgement rather than an answer.
A system optimised on aggregate volume will be tuned by the two thirds and evaluated by the third, which is the shape of the disappointment.
The pricing test
The single most useful practical finding in this domain, and it requires no benchmark at all.
Watch how a vendor bills.
A supplier charging per contained conversation has a structural interest in containment being high, and containment goes up when handoffs go down, whether or not the customer was helped. A supplier charging only for verified resolution earns nothing on an escalation or a failed conversation, and therefore has no reason to suppress a handoff to protect a number.
At least one major vendor now prices this way explicitly, charging per resolution with no charge for escalations or failures. Another defines resolution as full end-to-end autonomous handling with accountable ownership through to close, and bills only on tickets meeting that standard, which only works commercially if the definition is strict enough to exclude deflection.
The pricing model is a faster read on vendor incentives than any case study, because a case study is chosen and a pricing model is committed to. It is the closest thing to a hard signal available to a buyer in this market.
The diagnostic that separates resolution from containment
A clean test, usable on your own data without any vendor cooperation.
Track resolution and satisfaction together over time.
Rising resolution with stable satisfaction is truly effective automation. The system is solving more and customers are no less happy.
Rising resolution with falling satisfaction is containment wearing resolution's name. The number went up because fewer conversations escaped, not because more problems were solved, and the people who did not escape were unhappy about it.
Two supporting measures make it harder to fool.
Re-contact rate. A customer who returns within a week about the same issue was not resolved, whatever the ticket status says. This is the single best correction to an inflated resolution number and most dashboards do not show it.
Abandonment. A conversation that ends without a stated outcome is counted as deflected by most systems and should be counted as unknown. Splitting abandonment out of deflection typically moves the headline figure substantially.
What actually determines whether it works
Three conditions, and the first is not about the model.
Authenticated context. A system that knows who is typing, what they bought and what state their account is in can resolve things. A system without that can only answer general questions, and general questions were already answered by the help centre. The successful fintech deployment worked because the assistant had purchase history before the first message. Without exposed account state you do not have a deployment plan, you have a demo.
The ability to take an action. Resolution frequently means issuing a refund, changing a booking or cancelling a subscription. A system that can describe how to do those things but not do them is producing a better help article, not resolving a ticket. The gap between generating a response and executing a change is where most of the difference between vendors sits.
And a deliberate human boundary, set before launch. The reversal above happened because the boundary was not deliberate enough at the start. Deciding in advance which intents, tiers and customer segments always reach a person, and instrumenting the routing from day one, is cheaper than discovering the boundary through complaints.
Why this domain measures worst of the six
Six domains in, customer service is the one with the weakest measurement, and the reason is structural rather than negligent.
Nobody in the transaction wants an accurate number.
The vendor is paid on a metric they help define. The buyer approved a budget against a projected saving and now reports against it. The support director's headcount was reduced on the strength of a figure. And the customer, the only party with direct knowledge of whether their problem was solved, is not asked, or is asked at a moment selected by the system.
Compare that with the other five.
Law has an adversary who is paid to find the error. Finance has an examiner with statutory authority and no commercial interest. Medicine has a payer deciding whether to reimburse. Science has replication, weakly and slowly. Education has trials, of low certainty but at least externally designed.
Customer service has a dashboard, and the people who read it chose what it counts.
This is the cleanest illustration of the detection asymmetry the series has been building toward. It is not that support automation is worse than the others. It is that support automation is the only one of the six where every party to the measurement benefits from the same answer, and where the one party with contrary information has no channel.
The practical implication is a preference ordering for evidence. A metric produced by someone who loses money when it is high beats a metric produced by someone who gains. That is why the pricing test above outperforms every benchmark in this domain, and it is a rule worth carrying into any field where the only available numbers come from the parties to the deal.
What is unresolved
Whether resolution rates are comparable across vendors at all. Every supplier defines resolution slightly differently, and the definitions are chosen. Until an independent body specifies the measurement, cross-vendor comparison is comparing self-selected definitions rather than performance.
What the long-run satisfaction effect is. The available data covers deployment and immediate response. Whether customers who learn that a company's front line is automated behave differently over years, in retention or in willingness to contact at all, is not measured anywhere.
Whether the projection is a forecast or an aspiration. A large industry survey reports practitioners expecting AI to handle 50% of cases by 2027, up from 30% in 2025. That is a survey of expectations rather than a model, and expectation surveys in this field have historically run ahead of outcomes.
And what happens to the humans on the remaining tier. If automation absorbs the routine work, the remaining queue is entirely difficult, emotional and escalated. That is a harder job than the one that existed before, staffed by fewer people, and nobody has published anything on the attrition consequences.
The counter-argument
Deflection is a legitimate metric for a legitimate question. A support organisation truly needs to know its human workload, and deflection measures that correctly. The problem is not the metric; it is using a cost metric to answer a quality question. Criticising deflection for not measuring resolution is criticising a thermometer for not measuring pressure.
The reversal has been over-read. A company that automated two thirds of its support, cut handling time by more than 80%, and then chose to reinvest some savings in human coverage for complex cases has run a successful programme and tuned it. Describing that as a failure requires ignoring that the automation was retained.
And the sector's measurement problems predate AI. Human support has always been measured by handle time, first-contact resolution and abandon rate, all of which have known perverse incentives. An agent closing a ticket to protect their average is the same failure as a bot suppressing a handoff. The metric problem is inherited, not introduced.
Two thirds is also truly a lot. Retail support volumes are enormous and the majority of contacts are simple, repetitive and well-specified. A technology that handles the common case reliably and hands over the rest is doing what good automation does, and holding it to a standard of total replacement is a standard nobody set except marketing.
The short version
Three metrics get used interchangeably and measure different events. Deflection counts conversations no human touched, containment counts conversations that did not transfer, resolution counts problems actually solved. The first two record a customer who got an unhelpful answer and gave up as a success, which is why a platform can report 90% deflection on a 40% resolution rate with both numbers accurate.
On honest numbers the picture is more modest. Cross-customer resolution averages of 67% in 2025 rising to 76% in 2026, with roughly two thirds as a support-case median and above 80% as best-in-class on a favourable intent mix. Retrieval-based systems run 15 to 25 points above rule-based ones, which is the real size of the recent improvement.
The most-cited deployment automated 66% of chats, cut resolution from eleven minutes to under two, and reduced headcount from about 6,500 to 3,800. In early 2025 it reversed, restoring a guaranteed human option after satisfaction fell on complex tickets and confidently wrong answers appeared on edge cases. The lesson is that the last third was not a smaller version of the first two thirds. It was a different problem, and a system tuned on aggregate volume gets tuned by the easy cases and judged by the hard ones.
The most useful practical finding requires no benchmark: watch how a vendor bills. A supplier charging per contained conversation profits from suppressing handoffs. One charging only for verified resolution earns nothing on an escalation and has no reason to. A pricing model is committed to; a case study is chosen.
And the diagnostic that catches the substitution: track resolution and satisfaction together. Rising resolution with stable satisfaction is real automation. Rising resolution with falling satisfaction is containment wearing resolution's name. Add re-contact rate, because a customer who returns within a week about the same issue was not resolved whatever the ticket says.
Common questions
What is the difference between deflection and resolution in AI customer service? Deflection is the share of conversations no human agent touched. Resolution is the share where the customer's problem was actually solved. The difference matters because deflection counts a customer who received an unhelpful answer and abandoned the chat as a success. A platform can report 90% deflection while true resolution sits around 40%, and both figures are accurate measurements of different things.
What is a good AI resolution rate for customer support? Roughly two thirds is a reasonable support-case median, 70 to 75% is a strong deployment, and above 80% is best-in-class, all on a favourable mix of well-structured intents. One vendor billing on resolution reports cross-customer averages of 67% in 2025 rising to 76% in 2026. Sectors like healthcare, insurance and professional services benchmark lower because the questions are truly harder, and 35% in a legal context can be excellent.
Why did Klarna reverse its AI customer service decision? The deployment succeeded on volume and struggled on the remainder. It handled around 66% of chats, cut average resolution from eleven minutes to under two, and contributed to headcount falling from about 6,500 to 3,800. In early 2025 the company restored a guaranteed option to reach a person, after satisfaction fell on complex and emotionally charged tickets and confidently wrong answers appeared on a minority of edge cases. The automation was retained; the boundary was moved.
How can I tell if a support AI vendor's numbers are real? Look at how they charge. A vendor billing per contained conversation has a structural interest in containment being high, which improves when handoffs are suppressed whether or not customers were helped. A vendor billing only for verified resolution earns nothing on an escalation and has no reason to suppress one. The pricing model is a faster read on incentives than any case study, because a pricing model is committed to and a case study is selected.
How do I measure whether AI support is actually working? Track resolution and satisfaction together over time. Rising resolution with stable satisfaction indicates truly effective automation. Rising resolution with falling satisfaction indicates containment being reported as resolution: the number rose because fewer conversations escaped, not because more problems were solved. Add re-contact rate, since a customer returning within a week about the same issue was not resolved regardless of ticket status.
What makes an AI support deployment succeed? Three things, and only one concerns the model. Authenticated context, so the system knows who is typing and what state their account is in, since without that it can only answer general questions the help centre already answered. The ability to take an action such as issuing a refund or changing a booking, since describing how to do something is not resolving a ticket. And a human boundary decided before launch rather than discovered through complaints.
Is AI going to replace customer service agents? The evidence points to a change in composition rather than replacement. Automation absorbs the high-volume, well-specified tier, and the remaining queue becomes entirely difficult, escalated and emotional. One large survey reports practitioners expecting AI to handle 50% of cases by 2027, up from 30% in 2025, though that is a survey of expectations rather than a forecast model. What nobody has published is what happens to attrition among the people left staffing a queue with the easy work removed.
Why do vendor benchmarks disagree so much? Partly because they measure different events and partly because each vendor defines resolution in its own way. Comparing an 80% deflection figure against a 41.2% tier-1 deflection figure against a 76% resolution figure produces a table of numbers that cannot be ranked, since no two describe the same quantity. Until an independent body specifies the measurement, cross-vendor comparison is comparing self-selected definitions rather than performance.
Related articles
- Testing a system that answers differently every timeForty percent of organisations hit significant quality regressions within ninety days of deploying an LLM application. Exact-match assertions are the cause: they reject valid answers and occasionally accept wrong ones.
- AI in journalism: 45% of news answers had a flawTwenty-two broadcasters in eighteen countries evaluated 3,000 AI answers about the news. Forty-five per cent carried a significant issue, and the worst performer failed on 76%.
- AI in education: students feel twice the gain they getA 2026 review of 66 randomised trials found AI tutoring improved satisfaction by 0.93 and confidence by 0.91, against 0.53 for knowledge. The authors rated all three very low certainty.
- AI in government: 126 use cases, 65 not made publicOne agency reported 126 active AI use cases and auditors found the inventory still incomplete, with tools contracted to build criminal cases missing from it entirely.