Call center quality assurance: how QA scorecards really work

Riellvriany Indriawan
Written by

Riellvriany Indriawan

Katelin Teen
Reviewed by

Katelin Teen

Last edited July 7, 2026

Expert Verified
A quality assurance scorecard reviewing a customer support conversation

What call center quality assurance actually means

Every support team eventually asks some version of "are we actually doing a good job, or does it just feel that way?" Quality assurance is the answer to that question turned into a process. It's the systematic review of support interactions - calls, chats, emails, tickets - against a defined standard, so that "good service" stops being a vibe and starts being a score a manager can point to.

The mechanism underneath almost every QA program is the scorecard: a structured evaluation form with weighted rating categories that a reviewer - human or, increasingly, AI - fills out after reading or listening to a conversation. As Zendesk puts it, scorecards exist for "evaluating agent performance, identifying areas for improvement, and ensuring that your team meets organizational goals." QA sits alongside workforce scheduling and forecasting under the umbrella term Workforce Engagement Management (WEM) - most vendors, including Zendesk, sell it as one of two or three modules in that suite.

It matters for reasons that compound as a team grows. A team of three agents can be QA'd by a manager who just... reads everything. A team of thirty can't. Scorecards plus sampling are what let a QA function produce comparable scores across hundreds of agents without a manager personally reading every ticket. They're also the raw material for coaching - pinpointing that one agent consistently skips the closing line, rather than a vague "be more thorough" note in a 1:1. And in regulated industries (financial services, healthcare, debt collection), QA scorecards often carry legal weight: a missed disclosure isn't a style note, it's a compliance failure.

The business case isn't abstract, either. On Zendesk's own QA product page, the radio station network Audacy reports a 19% increase in agent productivity and a 15% increase in one-touch resolutions after adopting structured QA; Kahoot! saw CSAT climb 5 percentage points alongside a 150% jump in how many tickets actually got reviewed; Liberty reported +2.3% CSAT and +4% on its Internal Quality Score. None of those are QA-for-QA's-sake numbers - they're what happens when a team can finally see its own blind spots.

How a QA scorecard actually works

Strip away the vendor dashboards and a scorecard is just a weighted average. Each category - accuracy, tone, empathy, grammar, resolution, compliance - gets a rating scale and a weight from 0 to 100. Zendesk's documentation is blunt about the math: "To calculate the review score for an interaction, you multiply each category's score by its weight, then divide the total by the sum of the weights." A team that cares more about root-cause understanding than documentation style just weights root cause higher - the formula does the rest.

The scale you pick trades speed for nuance. Zendesk QA offers four:

ScaleWhat it looks likeBest for
BinaryGood / badHigh volume, fast reviews
3-pointGood / satisfactory / badA little more nuance, still quick
4-pointGood / slightly good / slightly bad / badForces a decisive call, no neutral middle
5-pointA–E style gradingMost detailed feedback, slowest to grade

Two mechanics do most of the heavy lifting in a mature QA program. The first is the critical category - mark "disclosed the cancellation fee" as critical, and a failing score there zeroes the whole review to 0%, no matter how warm and well-spelled the rest of the reply was. That's the lever regulated teams pull to make sure a compliance miss can't get averaged away by good tone. The second is root-cause tagging: instead of just scoring "empathy: bad," a reviewer picks a predefined reason from a list, which is what actually makes QA data useful for coaching instead of just being a number an agent resents.

Zendesk QA's own AutoQA system ships with eight default categories that get scored automatically the moment a ticket closes:

CategoryWhat it checks
GreetingDid the agent greet the customer?
EmpathyWas the agent empathetic to the concern?
Spelling and grammarMistakes, misspellings, style errors
ClosingDid the agent close properly, offer further help?
Solution offeredDid the agent propose a fix during the chat?
ToneClassified into 7 root tones (cheerful, supportive, calm, etc.)
ReadabilityWord complexity and sentence length
ComprehensionDid the agent actually understand the issue?

(Zendesk Help Center)

AutoQA scoring a conversation against custom prompt categories like Tone, Solution, and Clear Instructions, landing on a 92% pass score, as taken from Zendesk
AutoQA scoring a conversation against custom prompt categories like Tone, Solution, and Clear Instructions, landing on a 92% pass score, as taken from Zendesk

Here's the math in miniature: say Accuracy is weighted 40, Tone weighted 20, and Compliance weighted 40. An agent scores 90% on accuracy, 100% on tone, and 80% on compliance. Weighted out, that's (90×40 + 100×20 + 80×40) ÷ 100 = 88%. Change the weights and the same underlying performance produces a different headline number - which is exactly why calibration (below) matters so much.

Weighted QA scorecard math: three category cards with weights and scores flow into a single 88% result
Weighted QA scorecard math: three category cards with weights and scores flow into a single 88% result

Sampling: the math that quietly stopped working

For most of QA's history, a human reviewer couldn't listen to or read every interaction, so teams built the whole discipline around random sampling - pull a handful of tickets per agent per week, score those, extrapolate.

The problem is that the math looks fine at the team level and falls apart at the agent level. Leo Gasperin laid this out plainly in a LinkedIn post describing his own former team: they spent 200 hours a month reviewing tickets against strict criteria and still only covered about 5% of total volume. Reviewing 5-10 tickets a week out of one agent's roughly 500 monthly tickets doesn't produce a statistically reliable read on that agent, even though the overall sampling rate looks acceptable on a QA dashboard. His fix, verbatim: "Audit 100% with AI agents configured as evaluators, then only review the problems." A commenter on the same post, Shubhashree S., put the failure mode in one line: "sampling creates confidence, but not clarity."

That's not a fringe complaint. On G2, a MaestroQA reviewer working as an Executive Director described the before-state directly: "we were doing QA on a random sampling of tickets" before switching to a tool that let her team "perform targeted QAs" instead. It's the same story from a different angle at evaluagent, where a Senior Advisor in aviation support named the exact problem AI auto-QA is meant to fix: it "solves the issue of limited visibility from QA sampling, and reduces the time and effort spent on manual reviews."

Manual sampling reviewing 5% of tickets versus AI auto-QA reviewing 100% of tickets, shown as two grids of ticket icons
Manual sampling reviewing 5% of tickets versus AI auto-QA reviewing 100% of tickets, shown as two grids of ticket icons

Calibration: getting reviewers to actually agree

Scorecards only work if two reviewers scoring the same ticket land in the same place. That alignment exercise is called calibration, and Zendesk's own definition is exactly what it sounds like: "the practice of having all your reviewers rate the same batch of conversations and compare their scores and comments... ensures that your reviewers are aligned in their evaluations, providing consistent feedback to agents regardless of who conducts the review."

Mechanically, it usually runs "review first, then discuss": reviewers score independently, a manager marks one review as the baseline, and the group meets to reconcile gaps - often with a QA lead making the final call on any dispute. What actually needs aligning is rarely the big stuff; it's the edge cases. How do you score a category that never came up in this particular conversation? How long should written feedback be? Calibration scores are deliberately kept out of an agent's live Internal Quality Score and live in their own dashboard, precisely so the exercise stays about training reviewers rather than becoming a second, hidden scoring channel. Most teams run this monthly.

The shift to AI reviewing 100% of conversations

Every major QA vendor I looked at is converging on the same pitch, and it's worth naming plainly: AI reviewing every interaction, replacing manual sampling entirely, while keeping the scorecard-and-calibration structure as the backbone underneath.

Zendesk QA (formerly Klaus) leads with AutoQA - scoring 100% of conversations across voice, chat, email, and AI-agent transcripts, with admins able to write custom prompt-based categories in plain language instead of code. Its Spotlight feature automatically flags churn risk and knowledge gaps without a human reading every flagged thread, and Real-time QA surfaces issues while a conversation is still live rather than after the fact.

Spotlight automatically flagging a live conversation for escalation risk and a vulnerable customer
Spotlight automatically flagging a live conversation for escalation risk and a vulnerable customer

The part I find most telling is AI Agent QA - the same scoring standards used on human agents, applied to bots. Zendesk explicitly compares human and AI agent scores side by side to see where the bot needs work, which quietly confirms something worth sitting with: once an AI agent is handling real tickets, it needs a QA process too, not a pass on account of being software.

QA automatically scoring 100% of AI agent conversations using the same categories applied to human agents
QA automatically scoring 100% of AI agent conversations using the same categories applied to human agents

MaestroQA has drifted further from "QA tool" branding and now calls itself a conversation-data platform - its own homepage frames the pivot directly: "We started as a Contact Center QA company... Since the arrival of ChatGPT in 2023, analyzing conversation data has become a C-level priority." Its DraftKings case study is framed as replacing manual QA reviews outright, and Brex reports 20x more at-risk customers flagged after rebuilding its QA process around the platform - a jump large enough that, per the case study, Brex's COO reportedly told the team "no more manual QA" after a single preview.

Playvox, now part of NICE, plays a similar role inside a broader workforce engagement suite - flexible scorecards, calibration, and native Zendesk/Salesforce integrations - though its marketing pages actively block scraping, so treat specifics about it as directional rather than verified.

Zendesk QAMaestroQAPlayvox
PositioningQA add-on to Support/SuiteConversation-data platformQA inside WEM suite (now part of NICE)
Coverage modelAutoQA scores 100% of conversationsAI review replacing manual samplingCustom scorecards, unconfirmed AI coverage %
Scores AI agents tooYes - AI Agent QAYes - bot/chatbot monitoringUnconfirmed
CalibrationBuilt-in, baseline comparisonNot the primary focusBuilt-in
ComplianceSOC 2 Type 2, GDPR, HIPAA-enabledNot publishedNot published
PricingAdd-on, not publicQuote-gatedQuote-gated

Practitioners on the ground back up the coaching framing more than the grading framing. A MaestroQA reviewer working as an independent contractor put it well: "it stops treating quality reviews like a simple 'pass/fail' test and turns them into a way to actually coach people." And a genuine QA analyst - not a manager, the person actually doing the grading day to day - described two and a half years using MaestroQA "first as an agent, monitoring my quality scores, and now as a Quality Analyst, doing all of the grading," calling the process "extremely straightforward."

It's not friction-free. A Founder in telecommunications flagged a real cost gotcha on G2: "AI features require additional purchase which drives the cost up significantly" - budget for AI scoring as a line item on top of the base QA product, not a free upgrade. And an Operations exec at evaluagent surfaced a smaller but real UX complaint from the agent's side: score notification emails that only say a review happened, without the actual result, which he described as "unsettling" and something that "makes me feel anxious, prompting me to open Evaluagent right away." A QA program that's technically rigorous can still land badly if the agent-facing side of it feels like getting summoned to the principal's office.

When the agent being reviewed is AI, not human

This is where I have some real numbers of my own to put on the table, because it's the exact question a QA program has to answer once an AI agent joins the queue: not "does the AI sound right," but "does it hold up against the same scorecard you'd run on a person."

In a cross-validated trial we ran against real Zendesk traffic for an e-commerce team, eesel's AI hit 93% triage accuracy and 100% spam detection with zero false positives, on an inbox where spam made up 22% of total volume. Draft directional accuracy came in at 88% - meaning the substance of the reply was on track almost nine times out of ten. But only 12% of drafts got sent as-is by agents, and the factual error rate sat at 7%. Category performance varied a lot: returns and refunds drafts were rated useful 93.8% of the time, warranty claims 96.4%, and product inquiries and refund-status lookups both hit 100%.

Bar chart: triage accuracy 93%, draft directional accuracy 88%, sent as-is by agents 12%, factual error rate 7%
Bar chart: triage accuracy 93%, draft directional accuracy 88%, sent as-is by agents 12%, factual error rate 7%

The gap between 88% directional accuracy and 12% as-is adoption is the part worth sitting with. It wasn't that agents didn't trust the AI's judgment - the dominant pattern was "glance and rewrite": agents took an 8-15 sentence draft and cut it down to 1-3 sentences before sending. When we broke that rewriting pattern down, roughly 65% of it was pure length-and-tone editing, fixable by training the agent on the team's own past sent replies. Another 20% needed data the AI simply didn't have connected yet, like live ERP or logistics status. Only about 5% of rewrites happened because the draft was flat-out factually wrong. Training on a sample of 200 recent agent replies moved as-is adoption from 12% toward the 30-40% range in follow-up testing.

I bring this up because it's exactly the kind of pattern a real QA scorecard is built to surface - and exactly the kind of pattern that's invisible if you're only sampling 5% of tickets. If your QA program can't tell you whether your AI agent's drafts are wrong or just stylistically different from how your team writes, you're not actually measuring quality; you're measuring a vibe with extra steps.

That distinction shows up before a customer even sees a live reply, too. One prospect evaluating eesel - a support lead at a Belgian company - built his own 67-test evaluation before trusting the tool with real tickets, rating the knowledge answers "solid" across the board (he ultimately walked for an unrelated reason: chat widget speed, not answer quality). That's a QA process in miniature, run by a buyer rather than a vendor, and it's the instinct I'd want every team evaluating an AI agent to have: don't take "it demos well" as your scorecard.

A practical QA scorecard you can build this week

You don't need a WEM platform to start. A working first scorecard needs four things:

  1. Pick 4-6 categories that map to what actually matters to your business - accuracy, tone, resolution, and one compliance-critical category if you're in a regulated space. Fewer, well-chosen categories beat a bloated 15-item form nobody fills out consistently.
  2. Pick binary or 3-point scales to start. Nuanced 5-point scales sound rigorous but slow reviews down and introduce more disagreement between reviewers - exactly the thing calibration exists to fix. Start simple, add nuance once the program is running.
  3. Weight what actually drives outcomes, not what's easiest to score. If root-cause diagnosis matters more to your renewal rate than a perfectly-punctuated reply, weight it that way - the math doesn't care what's convenient to grade.
  4. Calibrate before you coach. Run one calibration session with your reviewers on the same 10 tickets before anyone's real score goes out the door. Disagreement here is normal and expected - it's the whole point of doing it first.
Coaching sessions and follow-up action items scheduled from QA review data
Coaching sessions and follow-up action items scheduled from QA review data

Once that scorecard exists for your human agents, run it against your AI agent too - including any part of your stack handling escalations or call center automation. The same weighted categories, the same critical-compliance rule, the same calibration discipline. An AI agent that's never been held to your actual scorecard is running on vibes, same as an untrained new hire would be.

eesel and quality assurance

I work eesel's support queue, and the honest answer to "should I trust an AI agent's replies without checking" is no - you shouldn't trust anyone's replies without checking, human or AI. That's why eesel builds QA into the rollout itself rather than treating it as a separate program bolted on afterward. Before an agent ever answers a live ticket, simulation mode runs it against your own historical tickets so you can see coverage by theme and catch gaps before a customer ever sees them - the closest thing to a QA scorecard run pre-launch instead of after the fact. Once it's live, confidence-based routing acts as an ongoing QA gate on every single reply: low-confidence answers get held back as a draft for a human to check rather than sent blind, which is the same logic Zendesk's AI Agent QA applies after the fact, just moved earlier in the pipeline.

eesel AI reports dashboard with analytics
eesel AI reports dashboard with analytics

If you're already running a QA program on your human team and weighing whether an AI agent could hold up to the same bar, that's exactly the conversation worth having - eesel is free to try, and pricing is usage-based per ticket rather than a flat seat fee, so testing it against your own scorecard doesn't require a procurement cycle first.

Frequently Asked Questions

What is call center quality assurance?
Call center (or customer support) quality assurance is the practice of reviewing support interactions - calls, chats, emails, tickets - against a structured QA scorecard to check whether the agent handled the interaction correctly, on brand, and compliantly. It's the mechanism that turns 'did we do a good job' into a measured, repeatable process rather than a manager's gut feeling.
What does a QA scorecard actually measure?
A scorecard is a set of weighted categories - accuracy, tone, empathy, compliance, resolution - each scored on a scale (binary, 3-point, 4-point, or 5-point) and multiplied by its weight to produce an overall score. Some categories are marked 'critical', meaning a low score there zeroes out the whole review regardless of everything else. See our breakdown of Zendesk's QA workflow for a worked example.
How often should QA reviewers calibrate?
Most teams run calibration sessions monthly, where reviewers independently score the same batch of conversations, then meet to reconcile disagreements about the rating scale, edge cases, and feedback tone. Calibration scores are kept separate from an agent's live quality score so the exercise stays about aligning reviewers, not judging agents.
Can AI actually do call center quality assurance?
Yes - Zendesk QA's AutoQA, MaestroQA, and Playvox all now score 100% of conversations automatically using the same category logic a human reviewer would apply, replacing the old model of manually sampling a handful of tickets per agent per week. eesel goes a step further by applying similar guardrails to its own AI-drafted replies before they ever reach a customer.
Why does manual QA sampling break down as a team grows?
A team reviewing, say, 5-10 tickets per agent per week can look statistically fine at the team level while covering under 5% of total ticket volume - nowhere near enough to catch a real pattern in any one agent's work. That gap is the main reason teams are shifting toward full-coverage AI review instead of sampling.
How much does call center QA software cost?
Zendesk QA is sold as an add-on to existing Support or Suite plans, and MaestroQA and Playvox are both quote-gated with no public pricing - you fill out a contact-sales form and get a quote sized to your ticket volume and seat count. Budget for AI-scoring features to cost extra on top of the base QA product; several reviewers specifically flag this as a pricing surprise.
What's the difference between QA and CSAT?
CSAT is the customer's own rating of one interaction; QA is a reviewer's structured assessment of how the agent handled it. A ticket can score high on CSAT because the customer got their refund, while still failing a QA scorecard on tone or compliance - the two metrics catch different failure modes.

Share this article

Riellvriany Indriawan

Article by

Riellvriany Indriawan

Riell is a designer and writer at eesel AI with about two years of experience researching CX platforms, AI chatbots, and helpdesk software. She combines her design background with a sharp eye for how these tools actually look and feel in practice — making her comparisons unusually visual and user-focused.

Related Posts

All posts →
Illustration of a support agent surrounded by a repetitive stack of tickets, representing customer service motivation and burnout
Customer Support

Customer service motivation: what actually keeps agents going

Pizza parties and leaderboards don't fix customer service motivation. Here's what Gallup's data, Reddit's agents, and eesel's own customers say actually works.

Riellvriany IndriawanRiellvriany IndriawanJul 6, 2026
Illustration of support tickets flowing through escalation tiers to the right agent
Customer Support

Ticket escalation process: how to design one that works

A practical guide to the ticket escalation process: the two types of escalation, what triggers one, a step-by-step build, and where AI now fits in.

Riellvriany IndriawanRiellvriany IndriawanJul 6, 2026
Illustration of customer behavior analysis lenses and support data charts
customer support

Customer behavior analysis: a support team's guide for 2026

A practical guide to customer behavior analysis for support and CX teams: the five lenses, the metrics that matter, and how to run one without a data team.

Riellvriany IndriawanRiellvriany IndriawanJul 6, 2026
Illustration of a live chat conversation with speech bubbles on an eesel-blue banner
Customer support

Live chat etiquette: 9 rules that make support feel human

The live chat etiquette rules that actually move CSAT: fast acknowledgements, human tone, honest holds, and clean handoffs. From someone who works the queue.

Riellvriany IndriawanRiellvriany IndriawanJul 6, 2026
Banner image for How to build a Zendesk quality assurance workflow that actually works
Guides

How to build a Zendesk quality assurance workflow that actually works

A practical guide to building a quality assurance workflow in Zendesk, covering scorecards, review processes, and AI-powered automation.

Stevia PutriStevia PutriMar 2, 2026
Illustration of a support ticket being graded against a six-criterion QA scorecard
Customer Service

QA feedback examples for customer service teams

Copy-paste QA feedback examples for support teams: sample scorecard comments for accuracy, tone, clarity, policy, and escalation, plus how to write them.

Riellvriany IndriawanRiellvriany IndriawanJul 6, 2026
Illustration of a modern call center dashboard with rising performance charts and support headsets
Guides

13 call center improvement strategies that actually work in 2026

Real call center improvement strategies for 2026: fix the metrics you track, kill repetitive volume, route smarter, and use AI where it actually pays off.

Riellvriany IndriawanRiellvriany IndriawanJul 5, 2026
Banner image for How to use Zendesk QA for agent feedback: A complete guide
Guides

How to use Zendesk QA for agent feedback: A complete guide

A practical guide to using Zendesk QA for agent feedback, including how reviewers grade conversations and how agents can respond to feedback.

Stevia PutriStevia PutriMar 2, 2026
Banner image for How to create a Zendesk QA scorecard: A complete guide for 2026
Guides

How to create a Zendesk QA scorecard: A complete guide for 2026

A practical guide to creating QA scorecards in Zendesk, from initial setup to optimization. Learn rating scales, categories, and best practices for measuring support quality.

Stevia PutriStevia PutriMar 2, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free