A ticket can read beautifully, score 5 out of 5 on tone, and still be wrong. If your quality score only measures how a reply sounds, you are grading the handwriting and ignoring the answer.
AI ticket quality scoring is the practice of evaluating customer support interactions against a rubric to judge whether the resolution was correct - factually accurate, consistent with your standard operating procedures, and complete - rather than only whether it was polite. In 2026, the strongest programs score correct logic and adherence to policy first, run automated checks across 100% of tickets instead of a 2% manual sample, and apply the same rubric to AI agents and human agents so the two are comparable.
Traditional QA samples 1-3% of tickets and scores mostly tone and formatting. That misses the failures that matter: a confident, well-written, factually wrong answer.
A good quality score measures correctness against your SOPs and knowledge base, not just sentiment - did the agent follow the right procedure and give the right answer.
AI-evaluated QA makes 100% coverage practical, so the score becomes a population statistic instead of a sample you have to defend.
One rubric for humans and AI is the only way to compare them honestly and to find systemic gaps in training, knowledge, or workflow.
A score is only useful if it routes to an action: a knowledge-base fix, a workflow edit, a coaching note, or a guardrail change.
Last updated: June 2026
Most support teams have run quality assurance the same way for a decade. A QA analyst pulls a handful of tickets each week, scores them on a form that weights greeting, empathy, grammar, and closing, and reports an average to the team lead. It is a reasonable program for a world where the bottleneck is agent etiquette. It is the wrong program for a world where an AI agent resolves the majority of tickets and the failure mode is not rudeness but a fluent, well-formatted answer that quietly contradicts your refund policy. This guide explains what a quality score should actually measure, how to design a rubric that captures correctness, when to score everything versus sample, how to apply one standard to both human and AI agents, and how to turn scores into changes rather than dashboards nobody reads. Lorikeet Coach and its Ticket Quality Score appear as a worked example throughout, because it was built to do exactly this.
What a Good Quality Score Actually Measures
The single most common mistake in support QA is scoring the surface of a reply instead of its substance. Tone, empathy, and grammar are easy to measure, so they dominate most scorecards. They are also the least likely place a modern agent fails. A large language model writes a polite, well-structured paragraph by default. What it does not do by default is know that your premium tier waives the restocking fee, or that a chargeback over a certain amount requires manager approval, or that this customer's account is flagged and the standard refund flow does not apply.
There is a second reason surface metrics dominate: they are inter-rater reliable. Two reviewers will usually agree on whether a greeting was present. They are far less likely to agree on whether a nuanced policy answer was correct, unless the rubric forces them back to the documented policy. So programs drift toward the dimensions that are easy to agree on, and the score quietly stops measuring the thing that creates risk. The fix is not to abandon agreement; it is to make correctness checkable enough that reviewers can agree on it too.
A quality score worth keeping measures three things in priority order.
Correct logic: Did the agent take the right steps in the right order to resolve the issue. For a failed-transfer ticket, that means diagnosing the cause before promising a fix, checking the account state before issuing a refund, and escalating when a threshold is crossed. Correct logic is about the reasoning path, not the final sentence.
Factual accuracy against SOPs: Did the answer match your documented policy and knowledge base. This is the dimension manual QA misses most often, because verifying it requires the reviewer to know the policy cold and to read the ticket against it. An answer can be confident, fluent, and flatly contradict the SOP it should have followed.
Resolution completeness: Did the interaction actually resolve the customer's issue, or did it close the ticket while leaving the underlying problem open. A ticket marked resolved that generates a repeat contact two days later was not resolved.
Tone and policy compliance (disclosures, PII handling, scripted language) still matter and still belong on the rubric. They belong below correctness, not above it. The order is the point: a beautifully empathetic wrong answer is worse than a terse correct one, because the wrong answer creates the refund, the complaint, or the regulator's question.
Ticket Quality Score (TQS): A composite score that evaluates a resolved interaction against a rubric covering correctness, policy adherence, and resolution, applied uniformly across agents and channels. In Lorikeet, the Coach agent generates a TQS for every ticket and a breakdown of why it scored the way it did.
Designing the Rubric
A rubric is the contract between your quality program and reality. If it is vague, every reviewer interprets it differently and your scores drift. If it is rigid, it punishes good judgment on edge cases. The goal is a rubric specific enough that two reviewers (or a reviewer and an AI evaluator) land on the same score for the same ticket, and grounded enough that the score maps to something you can fix.
Start From Outcomes, Not Etiquette
Write the rubric backward from the question "what would make this ticket a failure we care about." For a fintech, that list looks like: gave a refund the policy did not allow, missed a required disclosure, failed to verify identity before a sensitive action, told the customer the wrong thing about their balance or status, or closed a ticket the customer had to reopen. Each of those becomes a scored dimension. Etiquette dimensions get added after, and weighted lower.
Make Each Dimension Independently Checkable
A dimension like "handled the ticket well" is useless because it bundles five judgments into one number. Split it: "followed the correct procedure," "answer matched the knowledge base," "escalated when required," "resolved the stated issue," "complied with disclosure rules." Each should be answerable yes, no, or partial with reference to a specific artifact - the SOP, the KB article, the policy threshold. When a dimension is checkable against a document, an AI evaluator can score it consistently and a human can audit the AI's score.
Weight Toward the Failures That Cost You
Not all failures are equal. A missing sign-off is a minor deduction. A factually wrong answer about a regulated process is a near-zero, regardless of how good the rest of the ticket looked. Encode that in the weighting so the composite score moves when the dangerous thing happens, not just when the cosmetic thing does. A common pattern is to make correctness dimensions capable of capping the total: if the answer was factually wrong, the ticket cannot score above a floor no matter how polite it was.
Tie Every Dimension to a Source of Truth
Correctness can only be scored if there is a documented right answer to compare against. That means your SOPs and knowledge base are part of your QA infrastructure, not separate from it. A rubric that asks "was this factually accurate" without pointing the reviewer at the policy that defines accuracy is asking for an opinion. Lorikeet Coach scores against the same knowledge the Concierge agent uses to resolve tickets, so accuracy is judged against the live source of truth rather than a reviewer's memory.
Scoring 100% Versus Sampling
Manual QA samples because humans are expensive. A team that handles 50,000 tickets a month and reviews 2% is scoring 1,000 of them, which feels like a lot of work and is statistically almost nothing once you slice by agent, channel, topic, and week. The cells get small, the variance gets large, and a single bad reviewer day moves an agent's average. Worse, sampling is blind by construction: the failures you most need to find are rare, and a random 2% will usually miss them.
AI-evaluated QA changes the economics. When an automated evaluator scores every resolved ticket against the rubric, coverage goes to 100% and the score stops being a sample you defend and becomes a population you analyze. That unlocks three things sampling cannot do.
You catch the rare, expensive failure, because you are not relying on it landing in a 2% draw.
You can segment without running out of data: quality by topic, by channel, by workflow node, by time of day, with enough volume in each cell to trust the trend.
The score becomes a leading indicator. A drift in accuracy on one ticket type shows up in days, not at the next quarterly QA review.
Sampling does not disappear entirely. Humans still review a sample, but the job changes: instead of being the entire QA program, the human sample becomes a calibration check on the AI evaluator. You spot-check whether the machine's scores agree with expert judgment, recalibrate the rubric when they diverge, and spend human attention on the genuinely ambiguous tickets rather than on re-confirming that 980 routine tickets were fine.
There is also a cultural shift hidden in the move to full coverage. When QA samples a tiny fraction, agents and team leads treat a bad score as bad luck - the reviewer happened to pull a hard ticket on a hard day. The number is contested as often as it is acted on. When every ticket is scored against the same rubric, the average is no longer a draw; it is the agent's actual quality, and the conversation moves from "was this sample fair" to "what is driving this pattern." That is the difference between a QA program people argue with and one they use. The score earns trust precisely because it stops being a sample.
The honest limitation: an AI evaluator is only as good as the rubric and the source of truth behind it. If your SOPs are out of date or your knowledge base contradicts itself, 100% scoring will faithfully measure tickets against the wrong standard. Automated QA raises coverage; it does not absolve you of keeping your policies correct. That is why the calibration sample matters.
Scoring Human and AI Agents on One Rubric
Once an AI agent handles a meaningful share of volume, you have two populations to evaluate, and the temptation is to score them differently - a forgiving rubric for the AI because it is new, a stricter one for humans because that is what QA always did. Resist it. The point of a quality score is comparability, and comparability requires one rubric.
A single standard does three useful things. It tells you honestly whether the AI is better or worse than your human baseline on each ticket type, which is the question leadership actually asks. It surfaces systemic problems that are not about who handled the ticket: if both humans and AI score badly on a given topic, the problem is your knowledge base or your policy, not your agents. And it makes the handoff legible - when the AI escalates a ticket to a human, you can see whether the human's resolution scored higher, and whether the escalation was warranted.
The mechanics differ slightly. For an AI agent, the evaluator can also read the full reasoning trace - every tool call and decision - so a low score comes with a precise diagnosis: the agent retrieved the wrong KB article at step three, or skipped the verification step. For a human agent, the evaluator scores the same outcome dimensions from the ticket record. The rubric is identical; the AI simply leaves a richer trail to diagnose. Lorikeet Coach runs standalone for this reason - teams use it to score human-handled tickets even where the Concierge agent is not deployed, at roughly $0.10 per ticket, so the QA program covers the whole operation rather than only the automated part.
Acting on Scores
A quality score that produces a dashboard and nothing else is overhead. The value is in the loop from score to fix to re-measure. Every low score should resolve to one of a small number of actions, and the score's breakdown should make it obvious which.
Knowledge-Base and SOP Fixes
When accuracy scores cluster on a single topic, the usual cause is not the agent but the source. The KB article is ambiguous, outdated, or missing, so every agent (human and AI) gives a slightly wrong answer in the same direction. The fix is upstream: correct the article, and the scores on that topic recover. This is the highest-leverage action because one edit fixes a whole population of tickets, not one.
Workflow and Guardrail Changes
When the logic dimension fails - the agent skipped a verification step, escalated too late, or took actions out of order - the fix is in the workflow, not the knowledge. For an AI agent that means editing the workflow or tightening a guardrail so the bad path is blocked before launch. Lorikeet's approach is to test these changes against historical tickets in simulation before they go live, so a guardrail meant to fix one failure does not quietly break ten other paths.
Coaching and Calibration
When the failure is specific to one human agent and not the population, it is a coaching signal: a concrete ticket, a specific dimension, a documented right answer. That is far more useful to a team lead than a monthly average. And when the AI evaluator's score disagrees with an expert reviewer on a calibration ticket, the action is to adjust the rubric. The loop runs both ways: the rubric scores tickets, and the calibration sample scores the rubric.
If you are evaluating ticket quality at scale, see how Lorikeet Coach scores 100% of your tickets against your SOPs.
Lorikeet's Take on Ticket Quality Scoring
Most quality programs measure the wrong thing well. They produce a precise, defensible number for how polite and well-formatted your tickets are, and they have almost no visibility into whether the answers were right. In a regulated business that is backwards. The expensive failure is not a missing greeting; it is a confident, fluent, wrong answer about a refund, a balance, or a disclosure, and that failure sails through a tone-weighted scorecard untouched.
We built Coach to score correctness first. It evaluates every resolved ticket against your knowledge and your procedures, applies the same Ticket Quality Score to human and AI agents, and returns not just a number but the reasoning behind it - which step failed and why. The honest caveat is the one above: it scores tickets against your source of truth, so the program is only as good as the SOPs behind it. Used well, that is a feature, because it makes your knowledge base accuracy measurable too. If correctness on the tickets that matter is the bar your team uses, that is the bar Coach was built to measure.
Key Takeaways
A good quality score measures correct logic and factual accuracy against your SOPs first, with tone and formatting weighted below correctness.
Design the rubric backward from the failures you actually care about, make each dimension independently checkable against a document, and weight toward the failures that cost you.
AI-evaluated QA makes 100% scoring practical, turning the score from a small defensible sample into a population you can segment and trend.
One rubric for human and AI agents is the only way to compare them honestly and to separate agent problems from knowledge and workflow problems.
A score is only worth producing if it routes to an action: a KB fix, a workflow or guardrail change, or a specific coaching note - and Lorikeet Coach delivers this at roughly $0.10 per ticket, standalone or alongside the Concierge agent.
Conclusion
Ticket quality scoring in 2026 is no longer a weekly ritual of pulling a few tickets and grading their manners. The volume now runs through AI and human agents whose dangerous failure is not rudeness but a well-written wrong answer, and a quality program that cannot see that failure is measuring the wrong thing precisely. The shift is from sampling tone to scoring correctness across the whole population, against a rubric grounded in your real SOPs, applied identically to every agent, and wired to an action every time it flags a problem.
Start by rewriting your rubric around the failures you would be embarrassed to explain to a regulator or a CFO, tie each dimension to a documented source of truth, and decide what action each low score triggers before you turn the scoring on. The score is not the goal. The fixes it drives are.








