
Steve Hind
·
Updated
·
Fact-checked against Gartner & Forrester data
A Head of AI should evaluate AI agents for customer support the way they would evaluate any production system that can take actions: decide what to build and what to buy, constrain every tool the agent can call, test it against real and adversarial conversations before launch, and judge it on verified resolutions in a pilot with a holdout group. Benchmarks and demos tell you very little. What matters is how the agent behaves on your tickets, inside your permissions model, with every decision logged.
This framework is written for the AI or ML lead who has been handed the support automation decision. It maps each step to public standards (NIST AI RMF, OWASP Top 10 for LLM Applications, ISO/IEC 42001) so the evaluation also produces the evidence your security, risk and finance colleagues will ask for.
Key takeaways
Score reliability, not a single pass. In the τ-bench research, gpt-4o succeeded on fewer than 50% of tasks and its pass^8 score in retail was below 25%. Run every test scenario several times.
Least privilege is enforced in code, not prompts. OWASP traces excessive agency to excessive functionality, permissions and autonomy. Scope every tool to the workflow that needs it.
Measure resolution, not deflection. A conversation that ends without a human is not the same as a fixed problem. Score 100% of conversations for quality, not a sample.
Expect most agent projects to fail on cost, value or controls. Gartner predicts over 40% of agentic AI projects will be canceled by the end of 2027 for those three reasons.
Pilot with a holdout. Compare AI-handled topics against a control group on the same metrics before you widen scope.
Which platforms support this kind of evaluation?
Several AI support platforms now ship testing tooling a Head of AI can inspect directly. Lorikeet generates simulations from your real tickets, runs them in bulk batches, and lets you author adversarial scenarios such as prompt injection, false authority claims and mid-conversation goal switches as guardrail tests, with side-by-side diffs between runs (simulations page). Fin (now part of Salesforce) says it has a full testing suite to run simulations, regression testing and manual inspection before changes go live. Decagon lists simulations at scale, live A/B testing and always-on QA on its platform page. Sierra's research team published τ-bench, a public benchmark for agents that interact with simulated users and tools under domain policies. Ask each vendor to run its tooling on your tickets, not theirs.
How should a Head of AI evaluate AI agents for customer support?
Run the evaluation as eight gates, each with a pass condition you write down before you see a vendor demo:
Scope. List the top contact reasons by volume and mark which need an action in a backend system (refund, address change, password reset) versus an answer.
Build or buy. Decide which layers you own: models, orchestration, tools, evaluation, QA.
Workflow design. Decide which steps must be deterministic and which can be handled in natural language.
Tool safety. Map every API the agent can call to a permission, an identity context and an approval rule.
Pre-launch testing. Simulations from historical tickets plus red-team scenarios, repeated to measure consistency.
Production metrics. Verified resolution rate, quality score on every conversation, escalation reasons.
Observability and data. Audit trail of every message, tool call and model decision; data retention and training terms.
Pilot. A topic-scoped rollout with a holdout group and a decision date.
The NIST AI RMF Core organises the same work into four functions, govern, map, measure and manage, with governance designed to be cross-cutting. If your organisation already uses that vocabulary, label each gate with its function and you have a risk record as a by-product. NIST released the AI RMF on 26 January 2023 for voluntary use, and added a Generative AI Profile (NIST AI 600-1) on 26 July 2024 to help organisations identify risks unique to generative AI (NIST AI RMF).
Should you build or buy, and how much do model choices matter?
Buy the layer that is not your differentiator, and own the parts that encode your business: policies, SOPs, APIs and the evaluation set. A prototype that answers questions from a knowledge base is the easy part. The production system around it is the expensive part: workflow orchestration, tool permissions, guardrails, regression testing, handoff to humans, voice, audit logs and quality scoring, maintained as models change underneath you. We cover the trade-off in more depth in AI layer vs helpdesk built-in AI.
On models, ask three questions rather than which model a vendor uses. Can the vendor switch models without you rewriting workflows? Is every model decision logged so you can see which model produced which reply? And what are the data terms with each model provider? Lorikeet's trust page, for example, names its core sub-processors (Google Cloud Platform, OpenAI, Anthropic and Baseten) and states that it holds zero-data-retention agreements with all model vendors and does not fine-tune models on customer data (trust and security page). Ask every vendor for the same disclosures in writing.
Deterministic or natural-language workflows?
Use deterministic code for any step where a wrong answer is a compliance event, and natural language for everything around it. Identity verification, payments, refunds above a threshold and regulated disclosures should run as code the agent can call but cannot change. Lorikeet describes this as pockets of determinism: sensitive steps like payments and identity checks run as code the agent can invoke but never alter, inside workflows that are otherwise natural-language (how it works). Whatever you buy, check where that line sits and who can move it.
How do you keep an agent's tools and actions safe?
Assume the model will be manipulated at some point and design so the damage is contained. The OWASP Top 10 for LLM Applications 2025 lists prompt injection as LLM01 and excessive agency as LLM06. OWASP's excessive agency guidance traces the root cause to excessive functionality, excessive permissions or excessive autonomy, and its mitigations read like an evaluation checklist:
Minimise tools and their functions. An agent that needs to read order status should not hold a token that can also issue refunds.
Execute in the customer's context. Actions on behalf of a customer should run with that customer's authorisation and minimum privileges, not a shared admin identity.
Require approval for high-impact actions. OWASP recommends human-in-the-loop approval before high-impact actions are taken.
Complete mediation. Authorisation is enforced in downstream systems rather than by asking the model whether an action is allowed.
Log, monitor and rate-limit. These do not prevent excessive agency but limit the damage.
Then test the hallucination path. OWASP's misinformation entry (LLM09) cites the Air Canada case, where the airline's chatbot gave travellers wrong information and the airline was successfully sued. Ask each vendor what happens when the agent lacks the information to answer: does it check grounding against your sources, and does it escalate rather than guess? Lorikeet's guardrails run in two tiers: hard boundaries in code (customer isolation, server-side identity validation, workflow-scoped tool access, execution caps) and AI-layer checks on incoming and outgoing messages that can block, rewrite or escalate a reply (guardrails page). Its own page is explicit that prompt injection cannot be prevented entirely, which is the honest answer you should expect from any vendor.
How should you test an AI support agent before customers see it?
Test on your own historical tickets, in bulk, several times each, and add adversarial scenarios a red team would try. Three practices separate a real evaluation from a demo.
Use your ticket distribution. Sample a few hundred resolved tickets across your top contact reasons, including the messy ones with partial information and angry customers. Vendor demo scripts are built to pass.
Check end state, not wording. The τ-bench paper evaluates agents by comparing the database state at the end of a conversation with the annotated goal state. That is the right test for support: did the refund post, was the address changed, was the right ticket field set? A fluent reply that leaves the system in the wrong state is a failure.
Measure consistency. τ-bench introduced pass^k, a metric for the reliability of agent behaviour over multiple trials of the same task. Its authors found even gpt-4o succeeded on fewer than 50% of tasks and scored below 25% on pass^8 in the retail domain. Run each scenario repeatedly and report the worst case, not the best.
Red-team scenarios should include prompt injection, a customer claiming to be someone with authority, a customer switching goals mid-conversation, requests for another customer's data, and attempts to talk the agent past a policy limit. Keep these as a regression suite and re-run it after every workflow or model change. For a worked example in a regulated sector, see how to test healthtech AI support with simulation.
Which metrics prove the agent is working in production?
Verified resolution rate is the primary metric; deflection is not. Deflection counts conversations that did not reach a human, which includes customers who gave up and came back on another channel. Resolution counts conversations where the issue was actually fixed, confirmed by system state, customer confirmation or no repeat contact within a set window. The difference is explained in resolve, don't deflect.
Alongside resolution, track:
Quality score on 100% of conversations. Sampled QA misses rare failures, and rare failures are the ones that end up in front of a regulator. Lorikeet's Coach scores every conversation, human or AI, against your own quality criteria (Coach page).
Escalation rate and reason. Escalations are healthy when the agent hits a policy boundary; they are a defect when it simply could not find an answer.
Guardrail events. How often each guardrail fired, and whether it blocked, rewrote or escalated.
Repeat contact. The cheapest check that a resolution was real.
Cost per verified resolution. Your finance partner will want this; see our CFO guide to payback period.
Observability, audit and data handling
You should be able to open any conversation and see every message, every tool call with its inputs and outputs, every guardrail decision and which model produced each reply. If a vendor cannot show that trail in the evaluation, you will not get it after signature. For governance, ISO/IEC 42001:2023 specifies requirements for establishing, implementing, maintaining and continually improving an AI management system, and ISO describes it as the world's first AI management system standard. Map the vendor's controls to whichever of NIST or ISO your organisation uses, and verify certifications on the vendor's public trust centre rather than a slide. Our SOC 2 and HIPAA verification guide lists what to check, and the companion piece on how a CISO approves an AI support agent covers the security review.
How do you run a pilot with a holdout?
Pick two to four contact reasons where the agent passed simulation, route a share of that traffic to the agent and keep a comparable share with humans, then compare both groups on the same metrics before widening scope. Write the pass criteria first: resolution rate, quality score, repeat contact and CSAT for the AI group must meet or beat the control.
Carmoola, a UK car finance lender, ran a proof of concept with Lorikeet before rolling out. Its previous automation, built on a knowledge base and form-driven prompts, resolved about 30% of inbound questions. On day one the Lorikeet agent resolved 40% of inbound conversations end to end, and that now sits at 60% across WhatsApp, chat and email. Its outbound outreach was run as an A/B test and lifted one conversion metric by 60% (Carmoola customer story). That is the shape of evidence to ask for: a baseline, a controlled comparison and a number tied to a business metric.
Commercial terms matter for pilot design too. Under per-resolution pricing the vendor carries part of the risk. Lorikeet's pricing page lists Start at $2,100 a month and Scale at $5,100 a month, paid annually, with chat, email and SMS resolutions at $0.99 and $0.90 respectively, and its product page says you are not charged for a resolution you mark as bad (pricing page). Gartner's warning is worth keeping in view: it predicts over 40% of agentic AI projects will be canceled by the end of 2027 due to escalating costs, unclear business value or inadequate risk controls, and estimates only about 130 of the thousands of agentic AI vendors are real (Gartner). A holdout pilot is how you find out which side of that line you are on.
What still needs a human?
An evaluation framework does not remove judgement; it tells you where to put it.
Policy exceptions and goodwill. Decisions that bend policy for a specific customer should stay with a person, with the agent handing over full context.
Regulated judgement. Hardship, complaints with legal exposure, clinical questions in healthtech and anything a regulator treats as advice should escalate. AI agents should not give clinical advice.
Approval of high-impact actions. Following OWASP, keep human approval on actions above your risk threshold.
Sign-off on the evaluation itself. Security, legal and compliance owners sign off the tool permissions, data terms and escalation rules. The Head of AI owns the evidence; they should not be the only signature.
Reviewing what the scores surface. Scoring every conversation only helps if someone reads the failures and changes a workflow.
If you want to run this framework against your own tickets, you can start a 30-day free trial of Lorikeet and replay historical conversations through simulation before any customer sees the agent.
Frequently asked questions
What is the best way for a Head of AI to evaluate AI agents for customer support?
Treat it as a production system evaluation with written pass criteria: decide build versus buy, scope every tool to least privilege, test against your own historical tickets and adversarial scenarios several times each, measure verified resolution and quality on every conversation, and finish with a topic-scoped pilot against a holdout group.
What is the difference between resolution rate and deflection rate?
Deflection counts conversations that never reached a human, including customers who gave up. Resolution counts conversations where the problem was actually fixed, confirmed by system state, customer confirmation or no repeat contact. Resolution is the metric to hold an AI agent to.
How do you test an AI support agent for prompt injection?
Write adversarial scenarios such as instructions hidden in a message, false claims of authority, mid-conversation goal switches and requests for another customer's data, run them repeatedly before launch, and keep them as a regression suite. Testing reduces risk rather than eliminating it, so pair it with runtime guardrails.
Which frameworks apply to evaluating AI agents for customer support?
The NIST AI Risk Management Framework and its Generative AI Profile (NIST AI 600-1) cover risk management, the OWASP Top 10 for LLM Applications covers security risks such as prompt injection and excessive agency, and ISO/IEC 42001 sets requirements for an AI management system.
Can an AI support agent be made immune to hallucinations or prompt injection?
No. OWASP lists both prompt injection and misinformation among the top risks for LLM applications. The realistic goal is containment: grounding checks, escalation when the agent lacks information, tools scoped to each workflow and authorisation enforced in code.
How should an AI agent pilot be structured?
Pick a few contact reasons that passed simulation, route part of that traffic to the agent and keep a comparable control group with humans, and compare resolution, quality score, repeat contact and CSAT on a fixed decision date before widening scope.
How much does it cost to pilot Lorikeet?
Lorikeet's pricing page lists Start at $2,100 a month and Scale at $5,100 a month, paid annually, with chat, email and SMS resolutions at $0.99 and $0.90 and voice at $1.50 and $1.20. Escalations and unresolved tickets are not charged, and there is a 30-day free trial.
Try Lorikeet on your own tickets
Start a 30-day free trial. Coach sets up your first concierge in minutes.

