
Thomas Wing-Evans
·
Updated
·
Fact-checked against Gartner & Forrester data
A VP of Support can roll out an AI agent without hurting CSAT by treating it as a staged operational change, not a launch: start with two or three narrow, high-volume intents, prove quality in simulation against real tickets, expose a small share of live traffic, and only widen once AI CSAT on those intents sits within a few points of your human CSAT. Every AI conversation gets scored, every handoff carries full context, and a weekly review decides what to fix, expand or pull back. Platforms such as Lorikeet, Zendesk AI agents and Fin (now part of Salesforce) all support some version of this loop; the discipline around them matters more than the vendor.
The stakes are real. In a Gartner survey of 5,728 customers, 64% said they would prefer companies did not use AI for customer service, and 53% would consider switching to a competitor if they found out a company was going to. The top concern was that AI would make it harder to reach a person. A rollout that protects CSAT is one that answers that fear directly.
Key takeaways
Go narrow first. Launch on 2 to 3 intents that are high volume, low risk and fully resolvable through your systems, and leave everything else with the human team.
Use a CSAT gate, not a date. Widen an intent only when AI CSAT is within about 5 points of human CSAT on the same intent for two consecutive weeks. Magic Eden's first month on its new AI agent landed at 74% against a 78% human baseline.
Make the exit obvious. Gartner found difficulty reaching a person is customers' top AI concern, so every flow needs a clear route to a human that keeps the conversation history.
Score 100% of AI conversations. Survey response rates are too thin to catch bad answers early; automated QA on every conversation is what makes a weekly improvement loop possible.
Plan the team, not just the bot. Gartner predicts half of the organisations planning big agent cuts will abandon those plans by 2027; reskill agents into review, knowledge and escalation roles.
Which tools help you roll out an AI agent without hurting CSAT?
Most AI agent platforms now include some way to test before launch and review quality after it. Three worth comparing:
Lorikeet generates simulations from your real tickets, runs them in bulk and shows side-by-side batch diffs of what a workflow change improved or broke, so you can deploy topic by topic only where the results hold (simulations). It runs inside Zendesk, Intercom, Salesforce and other helpdesks you already use.
Zendesk AI agents run natively in Zendesk and, per Zendesk's AI agents page, let you set policies, review behaviour and evaluate every interaction with built-in QA.
Fin (now part of Salesforce) offers a testing suite for simulations, regression testing and manual inspection, and Fin says it hands off to your human team with full customer context on any helpdesk.
If you are deciding between native helpdesk AI and a layer on top, see our comparison of an AI layer vs helpdesk built-in AI.
How can a VP of Support roll out an AI agent without hurting CSAT?
Roll it out in five phases, each with an exit criterion you agree in advance, so nobody argues about readiness on launch day. The phases below assume a team of roughly 15 to 100 agents on a helpdesk such as Zendesk.
Phase 1: pick the intents and write down the rules (weeks 1 to 2)
Pull the last 60 to 90 days of tickets and rank intents by volume. Shortlist the ones where the right answer depends on data you can reach by API (order status, account changes, document requests) rather than judgement. For each, write the policy the AI must follow and the conditions that force a handoff. SageSure, an insurer, found this was most of the work: its operations lead said they were surprised by how much they had to clean up knowledge articles and define processes, and those documents then doubled as training material for people (SageSure story).
Phase 2: simulate against real tickets (weeks 2 to 4)
Before any customer sees the agent, replay a few hundred historical conversations per intent and grade the outputs against what your best agent would have done. Add adversarial cases: angry customers, false claims of authority, people who switch topic halfway. A practical bar is that the agent gets the correct outcome on at least 90% of replayed tickets for an intent and escalates the rest cleanly. Some teams also run a shadow period, where the AI drafts replies that agents see but customers do not, to calibrate tone.
Phase 3: expose a small share of live traffic (weeks 4 to 6)
Route 5 to 10% of tickets on the launch intents to the AI, ideally on one channel and in business hours so humans can catch problems fast. Compare AI CSAT, reopen rate and escalation rate against the human baseline for the same intents, not against your blended CSAT, because AI is handling the easiest tickets and a blended comparison flatters it.
Phase 4: widen by gate (weeks 6 to 12)
Step to 25%, 50% and then all traffic for an intent only when it clears the CSAT gate. Add new intents one or two at a time. Magic Eden is a useful benchmark: its previous AI agent ran at about 45% CSAT against 78% for humans; within the first month on its new workflows it reached 74%, about 30 points higher than before and 4 points from human level (Magic Eden story).
Phase 5: run the weekly improvement loop (ongoing)
Once live, the work shifts to a weekly cycle: review low-scoring conversations, fix the knowledge gap or workflow step behind them, re-run the simulation suite to confirm the fix did not break something else, then ship. This is where most CSAT gains come from after launch.
Which intents should an AI agent handle first?
Start with intents that are frequent, have a single correct outcome and can be completed end to end through your own systems. Good first candidates are order or application status, document and receipt requests, address or contact detail changes, and simple account actions. Avoid, for now, anything involving complaints, hardship, disputes, cancellations with retention offers, or regulated advice. SageSure began by tagging and categorising incoming email, then routing, then answering common questions, and only then letting the agent act on policies, such as processing alarm certificates and written cancellations within strict rules. Questions about premium increases and refund timing still go to a person.
For a fuller list of ticket types that should stay human, see which customer issues should never be automated.
How do you hand off to a human without losing the customer?
A handoff protects CSAT when the customer never has to repeat themselves and knows a person is coming. Zendesk's 2026 CX Trends report found 81% of consumers want representatives to pick up where they left off, and 74% get frustrated when they have to repeat information. Gartner's guidance in the same 2024 release is that an AI chatbot should tell customers it will connect them to an agent when it cannot solve the problem, and the agent conversation should pick up where the bot left off.
In practice, set four rules:
Always honour a request for a person. No loops, no "are you sure".
Pass a summary, not just a transcript. Intent, what the AI checked, what it tried, and why it stopped.
Escalate on risk signals, such as distress, legal threats, vulnerability or a guardrail firing, even if the customer has not asked.
Route to the right queue with a priority, so escalated tickets do not wait behind new ones.
In Lorikeet, when a guardrail fires the conversation can escalate to your team with full context, and each guardrail event is recorded as a tracked outcome your QA team can review (guardrails).
How do you monitor CSAT and quality once the AI agent is live?
Monitor two things separately: what customers say (CSAT on AI-resolved tickets) and whether the AI was actually right (quality scoring on every conversation). Surveys alone are not enough, because most customers do not answer them and a polite customer can rate a wrong answer highly. Score 100% of AI conversations for correctness against policy, tone, and whether the outcome was really achieved, then sample the lowest scorers for human review. Our guide to automated QA for customer support covers how scoring rubrics work.
Track these weekly, per intent, AI against human:
CSAT on resolved conversations, and the gap to human CSAT on the same intent.
Reopen or recontact rate within 7 days, the best early sign of a false resolution.
Escalation rate and the top reasons for escalation.
Quality score distribution, especially the share below your pass mark.
Guardrail events and any policy breaches.
The trend context matters too. Forrester's 2025 CX Index found US and Canadian perceptions of CX quality fell for a fourth consecutive year, and listed disappointing implementations of technology, including AI, among the causes. A careless rollout adds to that drift; a measured one can reverse it.
How does a support ops manager run an AI agent inside Zendesk day to day?
Day to day, a support ops manager runs the AI agent like a member of the team with its own queue, its own scorecard and a short daily and weekly rhythm, all inside the Zendesk views they already use. The AI works the tickets; the ops manager owns its configuration, its quality and its boundaries.
Daily (about 30 minutes):
Check a Zendesk view of tickets the AI escalated overnight and confirm each one landed in the right queue with a usable summary.
Read the lowest-scored AI conversations from the previous day and tag the cause: missing knowledge, wrong workflow step, tone, or a policy gap.
Watch for spikes in one intent, which often mean a product bug or outage the AI should acknowledge rather than troubleshoot.
Weekly (about 2 hours with the team lead):
Review the per-intent scorecard against the CSAT gate and decide which intents widen, hold or roll back.
Update help center articles and macros that the low scorers exposed, so humans and AI give the same answer.
Change one or two workflows, re-run the simulation suite, and ship only if nothing regressed.
Pick the next intent to automate from the escalation reasons.
With Lorikeet, the agent connects to Zendesk and turns tickets into completed work: classification, action and closure (integrations). Coach quality-scores every conversation, human or AI, and changes go live without an engineering ticket. Carmoola's team uses Coach to see which escalated conversations to automate next, and moved from 40% of inbound conversations resolved end to end on day one to 60% across WhatsApp, chat and email.
What happens to the human support team?
The human team moves up the queue rather than out of the building, and the rollout goes better when you say so early. Gartner polled 163 customer service leaders in March 2025 and found 95% plan to retain human agents to define AI's role; it also predicts that by 2027, 50% of organisations that expected to significantly cut their customer service workforce will abandon those plans.
Change management that works:
Tell agents what the AI will and will not handle before it goes live, and show them the escalation summary format.
Create new roles from the work that appears: AI conversation reviewer, knowledge owner, workflow builder, escalation specialist.
Give senior agents the hard queue. Escalated tickets are harder and more emotional than the average ticket used to be; staff and coach for that.
Design the team structure up front. SageSure's operations lead said he wished he had known the team structure he needed from day one, because he could have moved faster.
If the goal is to absorb growth rather than cut headcount, our guide on scaling customer support without hiring covers the capacity maths.
What still needs a human?
Even a well-run AI agent should leave these decisions with people:
Judgement calls with consequences: goodwill credits above a threshold, hardship arrangements, exceptions to policy.
Emotionally loaded contacts: bereavement, vulnerability, distress, formal complaints.
Regulated or clinical questions: in healthcare, financial services and insurance, the AI should route these to qualified staff, never answer them itself.
Setting the rules: which intents the AI owns, what the CSAT gate is, and when to roll back. Gartner's March 2025 forecast that agentic AI will resolve 80% of common issues by 2029 is a ceiling to plan toward, not a reason to skip the gate.
Sign-off on changes: someone accountable approves new workflows and policy edits before they ship.
Putting it into practice
Export 60 to 90 days of tickets and rank intents by volume and risk.
Write policies and handoff rules for the top 2 to 3 intents.
Simulate against a few hundred real tickets per intent; aim for 90% correct outcomes.
Go live on 5 to 10% of traffic with 100% quality scoring.
Widen only when AI CSAT is within about 5 points of human CSAT for two weeks.
Run the weekly loop and redesign roles as the queue changes.
If your Head of AI is running the vendor evaluation alongside this plan, our guide on how a Head of AI should evaluate AI agents for customer support pairs with it. To test this rollout on your own tickets, Lorikeet offers a 30-day free trial, and if you mark a resolution as bad you are not charged for it.
Frequently asked questions
How can a VP of Support roll out an AI agent without hurting CSAT?
Start with 2 to 3 narrow, high-volume intents, simulate against real tickets, go live on 5 to 10% of traffic, and widen only when AI CSAT is within about 5 points of human CSAT on the same intents. Score every AI conversation and keep a clear route to a human.
How long does a phased AI agent rollout take?
A typical phased rollout takes 8 to 12 weeks to reach full traffic on the first intents: about 2 weeks to choose intents and write rules, 2 weeks of simulation, 2 weeks on a small share of live traffic, then gated expansion. New intents keep being added after that.
What CSAT should an AI agent hit before you expand it?
Compare AI CSAT with human CSAT on the same intent rather than your blended score. A practical gate is staying within about 5 points of the human baseline for two consecutive weeks, alongside a stable reopen rate.
What is shadow mode for an AI support agent?
Shadow mode means the AI drafts responses on live tickets that agents can see but customers cannot. It is a low-risk way to calibrate tone and accuracy before the AI answers customers directly, and it complements simulation on historical tickets.
How do customers feel about AI in customer service?
Gartner's survey of 5,728 customers found 64% would prefer companies did not use AI for customer service, and the top concern was that it would be harder to reach a person. Clear handoffs that keep context address that concern directly.
Will an AI agent replace my support agents?
Most leaders do not expect it to. A Gartner poll of 163 customer service leaders found 95% plan to retain human agents, and agents usually move into escalations, quality review and knowledge ownership as the AI takes routine intents.
Try Lorikeet on your own tickets
Start a 30-day free trial. Coach sets up your first concierge in minutes.

