Across the largest public dataset on this question, 195 deployments spanning 38 vendors, the median AI resolution rate is 70%, the middle half of the field sits between 56% and 80%, and the full published range runs from 15% to 98.3%. That is the honest answer as asked, and close to useless on its own, because one deployment can truthfully report 70%, 51% or 42% depending only on which denominator the vendor chose.
This article is about how the number is built rather than what it is. For the benchmark listing itself, our 2026 resolution rate benchmarks for AI customer support is the companion piece. What follows decides whether any of those ranges mean anything: the four metrics reported as one, how far apart they sit, and a test for any vendor claim.
The numbers, in one block
Metric | Figure | Source |
|---|---|---|
Median AI resolution rate | 70% | My AskAI, 195 deployments across 38 vendors, May 2026 snapshot |
Middle half of the field (P25 to P75) | 56% to 80% | Same dataset |
Full published range | 15% to 98.3% | Same dataset |
Rates labelled "Resolution" vs rates labelled "Automation" | 72.5% vs 61%, a 12-point gap | Same dataset, median by metric family |
AI chatbot resolution across 220 million+ live chat interactions | 44.8% | Comm100 2026 AI Live Chat Benchmark Report |
Agentic platforms with backend integration | 70% to 85% end-to-end resolution | Notch, 2026 resolution rate benchmarks |
Healthy human handoff rate | 15% to 30% | Notch |
Human agent resolution rate | 65% | Intercom, stated in the terms of its Fin Guarantee Success Program |
Two of those lines sit 25 points apart. Both are correct. They count different things over different populations, the subject of this article.
Four numbers that get reported as one
Before any percentage means anything you need to know what sits in the numerator and the denominator. Four metrics circulate under the loose heading of "resolution rate", and they measure four different events.
Metric | Numerator | Denominator | What it actually proves |
|---|---|---|---|
Resolution rate | Conversations the AI resolved | Conversations the AI was involved in | How well the AI performs on the work it was given |
Automation rate | Conversations the AI resolved | All conversations | How much of your total volume the AI took off the queue |
Deflection rate | Contacts that ended without reaching a human | All contacts | That a person was not needed, not that the issue was fixed |
Containment rate | Conversations that stayed in the AI channel | All conversations | That the customer did not escalate, which may mean they gave up |
Resolution and automation share a numerator and differ only in the denominator, which is the largest source of confusion in the category. If your AI is involved in 60% of conversations and resolves 70% of those, your resolution rate is 70% and your automation rate is 42%. Both are true. Only one tells your CFO how much work left the queue.
Intercom publishes the relationship openly in its Fin documentation: Automation Rate = Involvement Rate x Resolution Rate. Any vendor that cannot draw that relationship for you on request is either not measuring it or not telling you.
Deflection and containment are weaker still. Both are satisfied by a customer who was handed a help centre link and went away. My AskAI's dataset puts median deflection at 70% and median containment at 58.2%, on samples of 17 and 6, so read those as directional.
A denominator change moved one resolution rate 17 points, with no change in behaviour
The clearest published evidence that the definition drives the number comes from Intercom, in a support article explaining a 2026 change to how Fin's metrics are calculated. Conversations where Fin was active but never had the opportunity to answer, because escalation guidance fired or a workflow triggered first, were removed from the "Fin involved" population. Intercom's own worked example:
Metric | Previous definition | Updated definition | Change |
|---|---|---|---|
Involvement rate | 750 / 1,000 = 75% | 500 / 1,000 = 50% | Down 25 points |
Resolution rate | 250 / 750 = 33% | 250 / 500 = 50% | Up 17 points |
Automation rate | 250 / 1,000 = 25% | 250 / 1,000 = 25% | No change |
Same month, same product, same 250 resolutions. The resolution rate rose 17 points, and Intercom states plainly that the number of resolutions did not change and billing was not affected. The change is defensible and Intercom documented it publicly. The point for a buyer is narrower: a 17-point movement in a headline resolution rate can be produced by an accounting decision alone. If a vendor's published rate improved between two case studies, ask which changed, the agent or the denominator.
The 12-point swing created by the word on the label
My AskAI built its dataset from 278 stat rows, 250 published competitor figures across 38 vendors plus 28 of its own deployments, of which 195 carried a clean, comparable AI-handling rate. Each row was tagged with the metric family the vendor used:
Metric family | n | Min | P25 | Median | Mean | P75 | Max |
|---|---|---|---|---|---|---|---|
Resolution | 108 | 20 | 63.2 | 72.5 | 70.5 | 81.8 | 95 |
One-touch / FCR | 3 | 60 | 60 | 75 | 75.7 | 92 | 92 |
Deflection | 17 | 15 | 62.5 | 70 | 67.8 | 82.5 | 92 |
Self-serve | 9 | 30 | 35 | 70 | 61.6 | 83.8 | 98.3 |
Automation | 52 | 15 | 45 | 61 | 59.8 | 78 | 90 |
Containment | 6 | 30 | 41.2 | 58.2 | 55.2 | 70 | 70 |
Figures published under the word "Resolution" carry a median of 72.5%. Figures published under "Automation" carry 61%. That is an 11.5-point gap produced by the label, on samples of 108 and 52, before any difference in capability. Two vendors doing identical work can post numbers 12 points apart and both be telling the truth.
A shortlist built from vendor homepages is therefore not a comparison. It is a list of numbers computed six different ways, and normalising them is a twenty-minute job almost nobody does before the second demo.
The distribution, once you stop reading only the top of it
Industry is the second-strongest predictor after setup quality, and the spread tracks how repetitive and low-stakes the typical ticket is, not how sophisticated the sector is.
Industry | n | P25 | Median | P75 |
|---|---|---|---|---|
Education | 8 | 71.2 | 81.5 | 93.8 |
Media / Entertainment | 9 | 64.5 | 80 | 88 |
Travel / Transport / Logistics | 10 | 72.2 | 78.5 | 86.8 |
Financial services / Fintech | 30 | 59.8 | 70 | 84 |
SaaS / Software / Tech | 41 | 58 | 68 | 80 |
Retail / eCommerce / DTC | 43 | 50 | 66 | 77 |
Health / Wellness | 19 | 56 | 65 | 75 |
Manufacturing / Hardware | 15 | 43 | 64 | 75 |
Telecom | 2 | 57.5 | 58.2 | 59 |
Benchmark against your own row, not the global median. A 66% rate is below par for education and above par for retail. Small cells are directional.
Company size, the variable most buyers assume matters, is close to noise. Median rates stay inside a flat 65% to 72% band across every revenue tier from under $1M to over $1B, and every employee tier from under 50 staff to over 5,000.
Now the counterweight. The My AskAI figures come largely from published vendor material, and vendors publish their wins. The Comm100 2026 AI Live Chat Benchmark Report analysed more than 220 million live chat interactions and found AI chatbots resolved 44.8% of them. That population is not self-selected, which makes it closer to an average deployment than any vendor homepage, this one included. My AskAI says the same of its own data: the true field average is probably below the 70% median it reports.
Hold both. 70% is what the field publishes. Something nearer 45% is what the field does across an unfiltered population of live chat.
Around 70% is the centre of gravity, and the least useful sentence in this article
The 70% line is unusually stable. The competitor-only slice lands on a 70% median too, and My AskAI's own customer base averages 72%. A figure repeating across hundreds of independent deployments is a real centre of gravity.
It is also the sentence least worth acting on. It tells you nothing about which metric produced it, nothing about the ticket mix underneath it, and nothing about what a comparable deployment in your industry would return. The largest source of variance is not the vendor but the setup: knowledge coverage, connected APIs and tools, and tuned escalation rules. My AskAI reports one account moving from 24% to 80% on the same product through configuration alone. Within a single product, the gap between the best and worst deployment is wider than the gap between products.
So the useful question in an evaluation is not "what is your resolution rate". It is "what does a deployment like mine look like at month three, six and twelve, and what does that need from my side".
What to expect, by platform tier
Platform architecture sets a ceiling before anyone touches configuration. The four-tier ladder below comes from Notch's 2026 benchmark work. Read it as a capability guide, not a league table.
1. Legacy chatbots: 10% to 25% resolution
Decision-tree bots and keyword matchers were never designed to complete anything. They function as intake and routing layers. If your current tool sits here, almost any move upward looks dramatic, which is why year-one improvement percentages from this baseline should be discounted.
Expect the headline metric on offer to be deflection or containment.
Check repeat contact rate before crediting the tool with anything.
Treat the gap to 60% as a knowledge and integration project.
2. Standard AI assistants: 40% to 60% resolution
These carry real business logic and answer well from a knowledge base. Their limit is backend connectivity: they can tell a customer what the refund policy is and cannot issue the refund. That caps them wherever your ticket mix requires an action rather than an answer.
Audit what share of your volume needs a system write, not a read.
That share is roughly your ceiling loss against an integrated platform.
Ask whether the published figures came from FAQ-heavy queues.
3. AI-native platforms in year one: 55% to 70% first contact resolution
This is the realistic target for a mature year-one deployment on integrated systems and maintained knowledge, as opposed to a curated 90-day pilot on a hand-picked queue. Most buyers who believe they are shopping for 85% are shopping for this band.
Ask for the twelve-month curve, not the launch-week number.
Ask what share of the queue the figure covered.
Pair every rate with AI-only CSAT for the same period.
4. Agentic platforms with backend integration: 70% to 85% end-to-end resolution
Platforms here connect to billing, CRM, policy and case systems and execute rather than retrieve. The extra 15 to 25 points over the tier above comes almost entirely from being able to finish a task. This is the band Lorikeet builds for, and at least ten vendors publish figures inside or above it, so the band is not a differentiator.
Ask which external systems the agent writes to, by name.
Ask what happens when a write fails midway through a multi-step task.
Ask for a figure measured on the complex queue, not the whole queue.
Across all four tiers, a healthy handoff rate sits between 15% and 30% depending on complexity. A platform reporting a 3% handoff rate in a regulated queue is not outperforming the field. It is failing to escalate.
A lower rate is often the correct rate
Resolution rate is not a difficulty-adjusted metric, and that is the most under-argued point in the category. Notch puts it precisely: an AI platform resolving healthcare queries at 46% is handling more complex interactions than one reporting 80% on order status lookups for an eCommerce operation. The 46% is harder work.
The distribution supports it. Education sits at an 81.5% median because the questions repeat and nothing much is at stake. Health and wellness sits at 65% and financial services at 70%, because a meaningful share of those tickets should reach a person: a suspected fraud case, a hardship request, a clinical question, a complaint carrying a regulatory obligation. There, escalation is the correct outcome and a lower automation rate is the system working as designed.
The reverse also holds. If an AI never discloses that it is an AI and never offers an easy route to a human, "resolution" quietly becomes "the customer gave up". My AskAI publishes one deployment at 26% by design, because that customer routes most tickets to humans deliberately, and it still posts 85% AI CSAT. A number is only as trustworthy as the escalation path behind it.
Two published examples from regulated queues, in each case study's own wording:
Breeze, a fintech moving fiat to stablecoin and back, reports that in 30 days Lorikeet's agent independently resolved 40% of its complex support volume, including more than 90% independent resolution of the tickets it chooses to solve, across KYC reviews, transaction statuses and declines. Read the two figures together. The 40% is the whole-queue number, the 90%+ is the number over the tickets the agent selected. Same deployment, 50 points apart, both published, neither wrong.
Hnry, an accountancy platform for the self-employed, went live in mid-May 2026 from a baseline of zero and handled more than 17,000 support conversations in the first month. It reports around 70% of conversations automated at the peak week of the Australian end of financial year, up from roughly 58% across the full period. That is an automation figure and the case study calls it automation, not a resolution rate. Hnry deployed on Australian tax first, the hardest jurisdiction it has.
Work out your own four numbers
Take one month of volume and compute all four. It ends most vendor arguments.
Involvement rate = conversations the AI was given / total conversations
Resolution rate = conversations resolved / conversations the AI was given
Automation rate = conversations resolved / total conversations
Deflection rate = contacts that never reached a human / total contacts
A worked example on 10,000 conversations a month, re-runnable with your figures:
Input | Value |
|---|---|
Total conversations in the month | 10,000 |
Routed straight to humans by rule | 2,500 |
AI active but blocked before it could answer | 1,500 |
Conversations the AI was genuinely given | 6,000 |
Conversations the AI closed with no human involved | 4,200 |
Of those, closed with a link or article only | 900 |
Output | Working | Result |
|---|---|---|
Involvement rate | 6,000 / 10,000 | 60% |
Resolution rate | 4,200 / 6,000 | 70% |
Automation rate | 4,200 / 10,000 | 42% |
Resolution rate excluding link-only closes | 3,300 / 6,000 | 55% |
One month, one deployment, four defensible headline numbers between 42% and 70%. Every vendor in this category will quote you the 70%. The number that changes your staffing model is the 42%, and the number that tells you whether customers were helped is the 55%, but only paired with the reopen rate.
Five questions that settle any resolution-rate claim
Each closes one of the standard ways a published rate gets inflated.
What is the denominator? All conversations, or only the ones the AI was given. This single question accounts for most of the 12-point spread in the published data.
What counts as resolved, in writing? Confirmed by the customer, assumed after silence, or contained in channel. Most platforms count a conversation resolved when the customer does not reply inside a window, and Intercom's documentation is explicit that Fin's resolved bucket includes confirmed and assumed resolutions. Ask for the split, in the contract rather than the deck.
What was the ticket mix? A pilot on order status and password resets returns a number with nothing to do with your dispute queue. Ask which intents were in scope, what share of volume they represent, and how many needed an action in an external system rather than an answer.
What is the repeat contact rate on AI-closed tickets at 72 hours? Compare it against the same figure for human-closed tickets. If AI-closed customers come back more often, the rate is inflated by roughly the size of that gap.
What is the CSAT on AI-only conversations? Separately from blended CSAT, for the same period as the rate. A blended score sitting next to an AI-only rate hides what it appears to prove.
A vendor that answers all five in writing has earned the benefit of the doubt. One that answers none has handed you a marketing figure.
What these three vendors actually publish
Three, deliberately: longer lists add names rather than clarity.
1. Intercom Fin: publishes the formula, counts assumed resolutions
Fin is the most transparent vendor in the category on metric construction. It documents the automation, involvement and resolution formulas publicly, published the worked example of its own denominator change, and states in the terms of its Fin Guarantee Success Program that its studies show 65% is the resolution rate of humans, a rare piece of public honesty about the bar being cleared. Its published list price on fin.ai/pricing is $0.99 per outcome with a 50 outcome monthly minimum.
Suits: teams on Intercom, or wanting AI on an existing helpdesk with reporting and quality tooling in one place. Does not suit: volume dominated by multi-step actions across several external systems, where the real question is what happens when step four of six fails.
2. Decagon: strong published figures, no public definition
Decagon publishes customer outcomes prominently and, as of 4 September 2026, publishes neither a pricing page nor a public definition of what it counts as resolved. Its homepage carries three customer tiles side by side quoting 70% chat and voice resolution, an 80% deflection rate, and a 95% cost reduction. Three metric families on one page is the norm rather than an outlier, and it is why the definitional question has to be asked in the room.
Suits: larger enterprises with the procurement capacity to extract definitions during diligence. Does not suit: teams that need to compare unit economics from public information before taking a call.
3. Lorikeet: publishes the price per resolution, not a headline resolution rate
Lorikeet publishes $0.95 per chat, email or SMS resolution on Start ($1,500 a month, paid annually) and $0.80 on Scale ($4,000 a month), one credit equal to one dollar, no per seat charges, implementation and platform fees included on both plans. The commercial promise, quoted from the pricing page: "We only charge for successfully resolved tickets. If you're unhappy with how Lorikeet handled a ticket, you don't pay for that ticket."
What Lorikeet does not publish is a company-wide resolution rate, and this article is not going to invent one. Lorikeet should not be chosen on a headline resolution number, and the company does not lead on that metric. Ten or more vendors publish figures in the 70% to 98% band, several higher than anything Lorikeet has put in public, and as the data above shows, most of that spread is definitional rather than real. Published customer outcomes use each case study's own wording, which is why Breeze reads as independent resolution and Hnry reads as automation.
Suits: Series A and later, digitally native B2C or B2SMB companies in fintech, healthtech and insurance running 2,000 or more monthly tickets, where resolution means calling an API to actually do something, under SOC 2 Type 2 and ISO 27001, with US-based zero-retention inference and a HIPAA BAA on Scale and above. Does not suit: the cases below.
Who this is not for
Stated plainly, because a resolution-rate article is where buyers talk themselves into the wrong tool.
If the lowest cost per ticket is the only goal. Lorikeet publishes $0.95 and $0.80 per resolution, and several deflection tools publish list prices well below that. On simple ticket volume the cheaper tool wins on arithmetic. Cost-only buyers are a documented disqualifier in Lorikeet's own ICP definition.
Below roughly 1,500 to 2,000 tickets a month. Start costs $1,500 a month and includes 18,000 credits a year. At 500 tickets a month and 50% automation you would consume about 3,000 credits a year while paying for 18,000, an effective cost near $6 per resolution. The plan floor dominates the unit rate.
FAQ and password-reset volume with no system actions. A maintained help centre plus a native deflection bot serves this better and costs less. With knowledge base access alone, Lorikeet performs like any other chatbot.
When a helpdesk-native add-on is the better buy. If you are on Zendesk or Intercom, want AI on the same tickets, and need native reporting, knowledge base sync and agent-copilot tooling, the add-on is one vendor and one data model. Lorikeet is behind on reporting and observability, knowledge base management and customer memory, and deliberately does not build an agent-assist copilot.
If you are buying on setup speed. Lorikeet trades plug-and-play for configurability. A team with no engineering support that wants to be live within the hour should pick something else.
Pure B2B SaaS with account-managed, low-volume, high-touch support. The economics do not work and neither does the model.
Key takeaways
Across 195 deployments and 38 vendors the median AI resolution rate is 70%, half the field sits between 56% and 80%, and the full range runs 15% to 98.3%.
Figures published under the word "Resolution" carry a 72.5% median while figures published under "Automation" carry 61%, a 12-point gap created by the label rather than by capability.
Comm100's analysis of more than 220 million live chat interactions found 44.8% AI resolution, an unfiltered population that guides better than any homepage.
Intercom's own documented metric change moved a resolution rate from 33% to 50% with no change in behaviour, resolutions or billing, which is what an accounting decision alone is worth.
A 46% rate in a healthcare queue is harder work than 80% on order status lookups, and a 15% to 30% handoff rate in a regulated queue is correct behaviour rather than underperformance.
Setup drives more variance than vendor choice: the same product has moved from 24% to 80% on one account through configuration alone.
What to do next
Pull one month of conversation data and compute involvement, resolution, automation and deflection for your own queue using the four formulas above. Take those numbers into your next vendor call and ask which one their published figure corresponds to. Whichever platform you buy, that removes the largest single source of disappointment in this category: discovering in month four that the 80% you were sold and the 42% you got were the same measurement described differently.







