Published AI customer support resolution rates in 2026 sit between roughly 50% and 76% of conversations for mature deployments, and between 40% and 60% in the first months after launch. Ada reports an average automated resolution rate of about 52% across more than 550 deployments, Intercom reports a 76% average for Fin across 12,000 customers, and Gartner predicts agentic AI will autonomously resolve 80% of common customer service issues by 2029. The spread comes from three things: what each vendor counts as resolved, how much of your queue needs a backend action rather than an answer, and how complete your knowledge is. Your own rate will land inside that band according to those three variables, and you can predict where before you sign anything.
Key takeaways
Vendor-published platform averages are about 52% automated resolution (Ada, 550+ deployments) and 76% (Intercom Fin, 12,000 customers). Both report best-in-class deployments at 84% or higher.
Containment runs about 20 points above resolution on the same conversations: Ada measures 72% containment against 52% resolution. A rate quoted without a written definition is a containment number until proven otherwise.
Expect 40% to 60% at launch and 60% or more after 6 to 12 months of knowledge and workflow work. Teams that skip the knowledge work stall between 30% and 45%.
Three levers move the number: knowledge coverage, access to backend actions, and channel mix. One Ada customer, Dott, went from 32% to 77% after building 25 API-powered automations.
Decision rule: write the definition of resolved for each top ticket category, test on 100 to 300 of your hardest historical tickets, and treat repeat contact within 24 to 48 hours as the check on the headline rate.
What resolution rate do AI support vendors actually publish?
The published platform averages fall between the low 50s and the mid 70s, with individual top deployments in the mid 80s. Each figure below is under the publisher's own definition, which is why they are listed separately rather than averaged.
Intercom Fin: 76% average. Intercom's 2026 benchmarks page puts Fin's average resolution rate at 76% across 12,000 customers, improving about 1% a month, with top performers at 80% to 84%. The same page gives a market-wide range of 40% to 60% on initial deployment, growing to 60% or more within 6 to 12 months. Intercom's own history is instructive: Fin started at 23% and climbed to 76%. For high-volume enterprises, Intercom guarantees a 65% resolution rate, which is probably the most useful single floor figure in the market because a vendor is willing to pay if it is missed.
Ada: 52% average, 72% containment. Ada's automated resolution guide reports that across more than 550 deployments, containment averages around 72% while automated resolution averages around 52%, measured on the same conversations. Best-in-class deployments reach 84% or higher, and most enterprise deployments launch in the 50% to 65% range.
Zendesk: a definition and a worked example, not a platform average. Zendesk's AI resolution rate explainer defines the metric as issues fully resolved end to end without human intervention, and illustrates it with 150 of 200 requests fully resolved, a 75% rate. It does not publish a customer-base average on that page.
Klarna: two-thirds of chats handled. Klarna's February 2024 announcement reported that its assistant held 2.3 million conversations, two-thirds of its customer service chats, in its first month, with a 25% drop in repeat inquiries. Note the verb: handled, not resolved. The repeat-inquiry figure is the closer proxy for resolution.
Gartner: 80% of common issues by 2029. Gartner's March 2025 prediction is that agentic AI will autonomously resolve 80% of common customer service issues without human intervention by 2029, with a 30% reduction in operational costs. The qualifier "common" matters: it excludes the complex tail.
Two things to notice about this list. Every average is computed across a vendor's own customer base, which skews toward ecommerce and software support where order status and account questions dominate. And every figure is self-reported under a definition the vendor wrote. Neither point makes the numbers wrong. Both mean a regulated lender should not paste them into a forecast.
What is the difference between resolution rate and deflection rate?
Deflection and containment count conversations that never reached a human. Resolution counts conversations where the customer's problem was actually fixed. The gap between the two is where most disappointing AI contracts live, and it is measurable: Ada's 72% containment against 52% resolution is a 20-point difference on identical conversations.
The metrics sit on a ladder from loosest to strictest:
Deflection. The customer was pointed at an article or a portal and did not open a ticket. Intercom's comparison guide puts it plainly: a 70% resolution rate that includes 30 percentage points of article views is really a 40% resolution rate.
Containment. A real conversation happened and it ended without a handoff. It cannot distinguish a satisfied customer from one who gave up. It can be pushed up by hiding the escalation path.
Automated resolution. The conversation was contained, and a grading model judged the answer relevant and accurate. This is a genuine improvement, but in most platforms the grader belongs to the same vendor that handled the conversation, and the grade is often tied to billing.
Verified resolution. The issue was fixed under a rubric the buyer wrote, checked against the customer's subsequent behaviour. Zendesk's definition adds the practical test: no escalation and no repeat contact within 24 to 48 hours.
Billing definitions are a fast way to see which rung a vendor stands on. Intercom's pricing page counts resolutions, procedure handoffs and disqualifications as billable outcomes, and does not charge when a conversation is simply passed to the team without an outcome. That is transparent, and it also means a handoff at the end of a procedure can be an outcome. Read every vendor's equivalent paragraph before comparing two percentages. The longer version of this distinction is in resolution rate versus deflection rate, and the human-team baseline is covered in first contact resolution benchmarks.
Why does the same AI resolve 45% at one company and 80% at another?
Because resolution rate is a property of the queue and the integration, not of the model. Three levers explain most of the variance between deployments running comparable technology.
Knowledge coverage
An agent cannot resolve a question that has no documented, current answer. Intercom's benchmark data shows teams that launch without structured, comprehensive knowledge content stall between 30% and 45%. The organisations that get past that plateau treat knowledge as an operating function: Gartner's October 2025 survey of 321 service leaders found 58% plan to upskill agents into knowledge management specialists. A practical test before launch: list your top 50 intents by volume and check each one for a single, current, unambiguous answer. Every intent without one is a guaranteed escalation.
Access to backend actions
Explaining a refund policy is an answer. Issuing the refund is a resolution. The share of your queue that needs an action (refund, address change, card freeze, dispute filing, rebooking) sets a hard ceiling on what an answer-only agent can resolve. Ada's Dott case study shows the effect: its automated resolution rate went from 32% to 77% after the team built 25 API-powered automations, including one that closes a stuck ride directly. The controls that make action access safe in a regulated environment are covered in how to safely let AI take actions in backend systems.
Channel mix
Chat, email and voice do not produce the same rate. Chat is synchronous and short, so identity checks and clarifying questions are cheap. Email arrives as multi-message threads with days between replies, so one resolution spans several turns. Voice adds identity verification by speech, background noise, and a latency budget under a second before the conversation feels broken. A vendor's chat-heavy average tells you little about your phone queue. Ask for the rate by channel, and set separate targets.
Ticket mix and regulatory share
Some categories must reach a human by policy or by law: formal complaints, customers showing signs of vulnerability or financial hardship, fraud that needs a decision rather than a lookup. A queue where 20% of contacts fall into those buckets has a ceiling near 80% before any model quality enters the picture. Count that share first, then set expectations for the remainder.
How do you set a realistic resolution target for your own queue?
You set it by measuring your own tickets under your own definition before go-live, then staging targets against published ranges. Six steps:
Write the definition per category. One sentence each for your top ten ticket types. A card dispute is resolved when the dispute is filed and the customer knows the timeline, not when the process is explained. If a definition is hard to write, that category is where vendor numbers will mislead you most.
Pull a hostile sample. 100 to 300 real historical tickets weighted toward multi-message threads, backend-action tickets, and angry customers.
Replay them in simulation. Run the configured agent against that sample and score every transcript against your written definitions. Any credible vendor can support pre-launch simulations on your data; a vendor that offers only a curated demo is asking you to accept its average as your forecast.
Stage the targets. Using the published ranges: 40% to 60% in the first quarter, 60% or more by month 12, with the ceiling set by your regulatory share and action coverage.
Instrument the checks. Repeat contact within 24 to 48 hours, CSAT on AI-only conversations reported separately from human CSAT, and QA review of a meaningful share of conversations graded as resolved. A resolution rate rising while recontact rises is a containment rate with better branding.
Put it in the contract. The definition, the measurement method, your audit rights, and the commercial consequence of a resolution you mark as bad. If pricing is per resolution, confirm in writing that escalations are free.
What this looks like in practice: a worked example
Take a regulated consumer lender: a UK car finance provider under FCA supervision, with a queue mixing settlement figures, payment date changes, affordability questions and complaints. Vendor averages transfer badly here: a meaningful share of contacts needs an action and another share must reach a human by rule.
Carmoola runs this queue on Lorikeet and publishes that 60% of inbound support is resolved end to end, with no human involvement, under Carmoola's own definition of resolved. Carmoola set the rubric. The platform pairs deterministic structured workflows for steps that must be exact (identity checks, payment changes) with natural-language workflows for the conversation around them, in one interaction. Post-hoc QA through Coach, Lorikeet's automated QA layer, covers 100% of conversations rather than a sampled fraction. Commercially, pricing is per resolution: escalations to a human are not charged, and if Carmoola reviews a conversation and marks the resolution as bad, it is not charged for it. Unresolved or unsatisfactory tickets cost nothing, which removes the vendor's incentive to grade generously.
Two other published results show how much the metric depends on the queue. easykind, a healthcare provider, reports a 92% reduction in email responses its team had to send. Summ, an Australian tax platform, cut tax-time resolution times by 97% during its seasonal peak. Those are three different measures (an end-to-end resolution share, an email volume reduction, and a peak-period handling share) and they should not be averaged into a platform number. Lorikeet does not publish one, and that is the honest limitation here: if you want a single percentage to put on a slide next to Fin's 76%, Lorikeet will not give you one. What you get instead is a rate produced on your own tickets, under your own rubric, verified by 100% QA, and priced so that a bad resolution costs you nothing. The other cost is on your side: customer-defined resolution means you must write the rubric before launch rather than accept a vendor's grade. Because Lorikeet is built for financial services and other regulated queues, with SOC 2, PII redaction, role-based access and US, UK and Australian data residency, that rubric can include the compliance conditions your own audit team will test. The practical first step is to run a simulation on 200 of your hardest tickets and read every transcript before you talk about a rate.
What still needs a human
A well-run deployment lowers its own resolution rate on purpose in four places, and a vendor that claims none of them is describing containment.
Complaints, vulnerability and hardship. Regulated firms route these to trained people by rule. The agent's job is to recognise the signal early, capture the facts, and hand over with full context.
Decisions rather than lookups. An agent can file a card dispute, compute the regulatory deadline and send the acknowledgement. Whether the customer gets the money back is a decision your disputes team makes. The same split applies to fraud and to any claim: the agent gathers and executes; it never forms an opinion on the outcome.
Silence in the knowledge base. When there is no documented answer, the correct behaviour is to escalate, and every escalation for that reason is a knowledge gap to close rather than a model failure.
Customers who want a person. Gartner's July 2024 survey found 64% of customers would prefer that companies did not use AI in customer service, and 53% would consider switching to a competitor over it. A visible, fast path to a human is part of what keeps the resolved share genuine. Gartner's 2026 survey found nearly 80% of organisations plan to move at least some agents into new roles built around complex and emotionally sensitive interactions, which is where that human capacity goes.
So the answer to "what resolution rate can AI customer support achieve?" is a band, not a point: 40% to 60% of conversations early, 60% to 76% once knowledge and actions are in place, and the mid 80s for the best-instrumented deployments in transactional queues, with about 20 points of daylight between any of those numbers and the containment figure a dashboard will show you by default. Where you land depends on your definition, your action coverage and your channel mix. Measure those three on your own tickets and the benchmark stops being a vendor's claim and becomes your forecast.







