Support QA calibration: how to get every reviewer scoring the same (2026)

Riellvriany Indriawan
Written by

Riellvriany Indriawan

Katelin Teen
Reviewed by

Katelin Teen

Last edited October 5, 2026

Expert Verified
Hand-drawn illustration of three support QA reviewers each holding a tablet and placing star-rating cards side by side on a shared line, two cards showing three stars and one showing four, below a customer conversation bubble

What is support QA calibration?

Zendesk puts it plainly in its calibration setup guide: calibration is "the practice of having all your reviewers rate the same batch of conversations and compare their scores and comments." That's the whole idea. Same conversations, same QA scorecard, different people, then a conversation about why the numbers don't match. If you haven't built the scorecard yet, start with Zendesk QA scorecard creation and come back.

I work eesel's support queue every day, and the reason this matters is that QA scores don't stay inside the QA tool. They end up in 1:1s, in coaching plans, sometimes in bonuses. If one reviewer gives a reply 4 out of 5 for tone and another gives the same reply 2, the agent's score depends on whose queue they landed in. Agents notice that fast, and once they decide QA is a lottery, the feedback stops landing.

Reddit

"You could have five QA agents listen to the same call and get five different results. Which isn't fair to the agent as that score could be used to determine if they'll be fired."

Calibration is the fix for that. It's not a training session for agents, and in most tools agents never see it. Zendesk's setup guide says "Agents do not see any conversation-related information in calibration", and calibration scores "don't affect the existing Internal Quality Score (IQS)." It's a check on the graders, run by the graders.

Why QA scores drift apart

Reviewers don't disagree at random. They disagree on the categories where the rubric leaves room for judgment: tone, empathy, "was the reply the right length", "did the agent go above and beyond." Process and accuracy usually line up, because a refund policy is either followed or it isn't. Tone is where two careful people read the same sentence and land two points apart.

The other source of drift is what a "pass" means in the first place. Here's a real example from eesel's own work. I looked at one real-traffic trial on an e-commerce Zendesk inbox, where 284 AI-drafted replies were graded three different ways.

Bar chart of the same 284 replies: 88% pointed the right way, 12% sent without edits, 7% factually wrong
Bar chart of the same 284 replies: 88% pointed the right way, 12% sent without edits, 7% factually wrong

Same replies, three verdicts: 88% pointed the right way, only 12% were sent without edits, and 7% were factually wrong. None of those numbers is wrong. They just answer different questions. A reviewer who grades "would I send this as is?" and a reviewer who grades "is the answer correct?" will score the same queue miles apart, and both will think they're being fair. When I dug into why agents rewrote the drafts, about 65% of rewrites were for length and tone, about 20% needed order data the draft couldn't see, and about 5% were the draft being factually wrong. Length and tone are exactly the categories where human reviewers split too. (I broke down that kind of rewrite analysis in my error analysis guide.)

That's why calibration starts with the bar, not the score. If your scorecard doesn't say what a 3 versus a 4 looks like for tone, no amount of discussion will make reviewers agree for long.

How to run a QA calibration session

Here's the session I'd run, in five steps. It takes about an hour once a month, plus 10 to 15 minutes of scoring per reviewer beforehand.

A five-step calibration loop: pick 5-8 conversations, everyone scores blind, reveal side by side, talk through the biggest gap, write the ruling into the rubric, repeated every 2-4 weeks
A five-step calibration loop: pick 5-8 conversations, everyone scores blind, reveal side by side, talk through the biggest gap, write the ruling into the rubric, repeated every 2-4 weeks

1. Pick 5 to 8 conversations that will cause an argument

Zendesk's calibration guide suggests starting with about five conversations. I'd go up to eight once the team is used to it. More than that and the discussion gets rushed.

The selection matters more than the count. Easy tickets teach you nothing, because everyone scores them the same. MaestroQA's calibration advice is to pull tickets with low QA scores, tricky topics, or a recently changed process or product. evaluagent adds a fair point in its calibration explainer: keep the set balanced, not just bad interactions, or the session turns into a blame review.

My mix: one conversation that touches a policy that just changed, one where the customer was upset from the first message, one long multi-reply thread, one that went fine but felt "off," and one an agent disputed recently. If you support more than one language, add one non-English conversation too; my multilingual support QA guide covers how to calibrate across language teams.

2. Everyone scores blind

This is the step teams skip, and it's the one that makes the rest work. If reviewers can see each other's scores, or score together in the room, the first confident person sets the number and everyone drifts toward it. evaluagent's calibration webinar describes the paper version of this problem: people "can change their scores on their paper, and they can be swayed to another person's point of view before they get an opportunity to speak."

In Zendesk QA, blind or open is a workspace setting you choose per role. Reviewers can be set to see only their own reviews, or to see peers' reviews only after submitting their own. Leads and managers get the same "after submitting" option, or "always."

Zendesk QA workspace calibration settings with the Calibration toggle on, previous reviews set to Not visible, and reviewers set to see all calibration reviews after submitting a review, as taken from Zendesk's help center
Zendesk QA workspace calibration settings with the Calibration toggle on, previous reviews set to Not visible, and reviewers set to see all calibration reviews after submitting a review, as taken from Zendesk's help center

One practical gotcha: Zendesk QA sends no notifications for calibration. Per Zendesk's docs, "it's up to the workspace manager or lead to inform reviewers about the calibration session and its due date." Put the due date in your team channel, or half the reviewers will score it 10 minutes before the meeting.

3. Reveal the scores side by side, against a baseline

Once everyone has scored, put the scores in one grid: reviewers down the side, scorecard categories across the top. You're looking for two things, which I'll come back to in the next section: categories with a wide spread, and reviewers who are off on everything.

You also need a baseline, meaning one agreed "right" score to compare against. Zendesk QA lets a manager mark one review as the Baseline review, which, per Zendesk, "makes it easier to compare scores against it." The calibration dashboard then shows each reviewer's score next to the baseline score.

The Zendesk QA calibration dashboard with session, conversation and review tables, where each review shows its score next to the baseline score and a Yes/No baseline column, as taken from Zendesk's help center
The Zendesk QA calibration dashboard with session, conversation and review tables, where each review shows its score next to the baseline score and a Yes/No baseline column, as taken from Zendesk's help center

MaestroQA does the same thing with a designated final calibrator. Its team alignment guide calls that person "your source of truth to measure alignment against," and each grader gets an alignment score against them.

4. Talk through the biggest gap, not every gap

With 5 conversations and 6 categories you have 30 places to disagree. You won't get through them in an hour, and you shouldn't try. Sort by spread and spend the meeting on the two or three widest gaps.

The aim is a decision, not a consensus. One Stitch Fix leader quoted on MaestroQA's calibrations page put it well: "We've normalized the expectation that we're driving toward commitment, not consensus." evaluagent makes the same point from the other side: "It's not a democracy, if you like, where majority rules." If three reviewers say 4 and one says 2, the answer isn't automatically 4. The person who knows the policy best makes the call, and everyone commits to it.

5. Write the ruling into the rubric

This is where most calibration sessions quietly fail. The team agrees, everyone nods, and three weeks later the same disagreement shows up again because nothing was written down.

Every resolved gap should turn into one line in your scorecard guidance, with the conversation as the worked example. "Tone: a reply that apologizes twice for the same issue scores 3, not 4. Example: ticket #48213." That's the artifact that makes the next session shorter. It also feeds your support SOPs and the examples you use when onboarding new agents, since a new hire needs the same rulings your reviewers just agreed on. Good rulings read like good QA feedback examples: specific, tied to a line in the conversation, and short.

A QA lead on r/callcentres described the strict version of this:

Reddit

"Calibrations by quality need to be strictly enforced and maintained. My last team I ran was quality and every month we'd pick a random call, all score separately, then get together and decide the correct score for each quality point. Whatever we all decided on, we had to be within 3% of that score (and certain elements had to be 100% on point)."

Rubric problem or reviewer problem?

This is the part of calibration that I think most guides skip, and it's the most useful. When scores don't line up, the pattern of the disagreement tells you what to fix.

Two score tables: on the left every reviewer agrees on process and accuracy but splits from 2 to 5 on tone, so fix the rubric; on the right one reviewer, Cy, scores low on every category, so coach the reviewer
Two score tables: on the left every reviewer agrees on process and accuracy but splits from 2 to 5 on tone, so fix the rubric; on the right one reviewer, Cy, scores low on every category, so coach the reviewer

Everyone splits on one category. If process and accuracy line up but tone ranges from 2 to 5 across four reviewers, nobody is "wrong." The criterion is open to interpretation, so fix the rubric. Rewrite the tone guidance with concrete examples (my customer service tone guide has a template), and re-score the same conversation next session to check the gap closed.

One reviewer is off on everything. If one person scores two points lower than everyone else on every category, the rubric is fine. That reviewer is reading it differently, so coach the reviewer. Sit down with them, the same way you'd coach an agent, walk through one conversation category by category, and have them score one more on their own.

The whole team is off from the baseline in the same direction. Everyone gives 4s and the lead gave 2s. That usually means the lead knows something the team doesn't, like a policy change that never made it into the knowledge base. That's a communication fix, not a scoring one.

Try it with your own numbers. Type in each reviewer's scores for one conversation and the checker tells you which of these you're looking at.

Zendesk QA also tracks the reviewer side over time with a separate Reviewer QA dashboard. It runs randomized assignments where someone evaluates the reviewers' own reviews, and Zendesk's dashboard docs say it's there to "track reviewer accuracy, evaluation history, and coaching trends." Evaluated reviewers see their own scores but can't see who evaluated them.

The Zendesk QA Reviewer QA dashboard showing a 96% Reviewer QA score, 43 evaluations, 12 evaluated reviewers and a Reviewer QA scores over time line chart, as taken from Zendesk's help center
The Zendesk QA Reviewer QA dashboard showing a 96% Reviewer QA score, 43 evaluations, 12 evaluated reviewers and a Reviewer QA scores over time line chart, as taken from Zendesk's help center

What counts as "calibrated"?

No QA vendor I checked publishes a hard pass mark. Here's what the people who do this for a living use instead.

SourceCalibration targetWhat it's measured against
Zendesk's calibration guideGap "usually falls around 5 percent"Baseline set during the session
u/mattemer on r/callcentresWithin 3%, some elements 100%Score agreed by the group
MaestroQAAlignment score per grader, no fixed thresholdFinal calibrator's score
Scorebuddy% of answers matching, plus Avg VariationThe original score being re-scored
Zendesk QA AutoQAAbove 75% accuracy is "high agreement"Human reviewers' scores

My take: aim for every reviewer within about 5 points of the baseline on the total score, and no category where reviewers are 2 or more points apart on a 5-point scale. Then add the rule from the r/callcentres quote above: some items have to be 100%. A missed identity check or a wrong refund amount isn't a matter of taste. Mark those as critical or auto-fail items on the scorecard (my scorecard criteria guide shows how in Zendesk QA) so they never get "calibrated" into a partial score.

Watch the total, too. Two reviewers can land on the same 80% overall while disagreeing on three categories underneath, one high where the other is low. Always compare per category, not just the headline score.

How often should you calibrate?

The honest answer is "more often than you're doing it now." Here's what vendors and practitioners suggest:

  • New QA program or new scorecard: MaestroQA's guide says it sees teams calibrate "as frequently as once a week," then move to monthly once alignment is high, and step back up when a new rubric launches.
  • Steady state: Zendesk's guide says monthly sessions suit most teams, and "the more reviewers you have, the more often you should calibrate."
  • Large QA teams: a former QA manager on r/callcentres ran "a weekly calibration call so that everyone was on the same page" alongside 5 audits per rep per month.

A lone QA analyst on r/talesfromcallcenters described the flip side, where the reviewer is the one being graded:

Reddit

"I was the lone QA agent for a call center of 55-70 people. We had a calibration session each week which involved the call center manager, the management staff, and myself. They chose ten calls that I had evaluated to listen to and score... I was held to a 95% accuracy rating which compared to the average of their call scores."

I'd add one trigger that isn't on any calendar: calibrate after anything that changes what "right" means. A pricing change, a new returns policy, a new escalation path. That's when reviewers drift fastest, because some of them know about the change and some don't. It's also worth a session whenever a new reviewer joins, the same way you'd shorten an agent's ramp time with worked examples, before their scores count toward anyone's support KPIs.

QA calibration features compared

Most dedicated QA tools have some kind of calibration feature. They differ in whether scoring is blind, what the "right answer" is measured against, and what you see afterward.

ToolFeatureBlind scoringMeasured againstAgents see it?
Zendesk QACalibration sessions, Reviewer QAConfigurable per roleBaseline review set by a managerNo
MaestroQA (now Rippit)Team Calibration, Grade-the-GraderYes, graders see only their own score and the finalFinal calibrator, alignment scoreNot documented
ScorebuddyCalibration listsNot documentedOriginal score being re-scoredOptional toggle per list
evaluagentCalibration sessions, Check-the-CheckerYes, "blind evaluate"Agreed "calibrated outcome"Not documented
Playvox (NICE)CalibrationNot documented"Expert opinions"Not documented
Level AICalibrationNot documentedAuditor benchmarkNot documented

A few details worth knowing before you pick:

Zendesk QA needs the QA or Workforce Engagement Management add-on. Every Suite plan says Quality Assurance "Requires QA add-on," and the Workforce Engagement Bundle is listed at $50 per agent per month, billed yearly. Calibration isn't gated to a higher tier on top of that. One limit to know about: the calibration dashboard lists conversations and their reviews, and a customer comment on the setup article notes the data "can only be downloaded conversation by conversation." If you want a reviewer's trend across many sessions, plan to pull it into a spreadsheet, or lean on the Reviewer QA dashboard.

MaestroQA, now rebranded as Rippit, separates two jobs: live team calibration and Grade-the-Grader, where a benchmark grader blindly re-grades tickets someone else already scored. Its calibration view shows how each calibrator answered each question, with a "Criteria Alignment" percentage.

The MaestroQA calibration view with a Thorough Solution question, a Calibrator Responses panel showing 89% criteria alignment, three Great responses at 75% marked as the Final Calibration Selection and one Good, could improve response at 25%, as taken from MaestroQA's calibrations page
The MaestroQA calibration view with a Thorough Solution question, a Calibrator Responses panel showing 89% criteria alignment, three Great responses at 75% marked as the Final Calibration Selection and one Good, could improve response at 25%, as taken from MaestroQA's calibrations page

Scorebuddy builds calibration from existing scored results: you add them to a calibration list and assign it to users, including external users, which is handy if you outsource part of your support to a BPO partner. Its glossary defines Avg Variation as "the average of all differences between a set of Calibration's numerical scores, and the numerical scores of the original Scores." The calibration module is in its entry Foundation plan, though prices are "Request a price."

If you run QA inside your helpdesk instead, check what it actually covers. I keep a longer list of AI QA tools if you're shopping. Gorgias offers Auto QA, which scores conversations automatically, but I didn't find a human-reviewer calibration workflow in its docs. Same for Freshdesk and Front in the docs I searched. If your helpdesk doesn't have it, a shared spreadsheet with the grid from step 3 does the job fine for a team of three to five reviewers.

Calibrating AI graders and AI agents

More teams now let AI score 100% of conversations, with humans reviewing a sample. That doesn't remove calibration. It adds one more reviewer to calibrate, and it's the one that reviews the most tickets.

Zendesk QA measures this directly. Its AutoQA dashboard has an Accuracy score that "measures the consistency and level of agreement between ratings generated by AutoQA and those provided by human reviewers," with above 75% counting as high agreement. It also tracks Modified AutoQA conversations, where a human changed the AI's score in at least one category. That list is your calibration agenda for the AI, and it's the fastest way to do support QA with AI without trusting it blindly: if humans keep overriding the same category, the AI grader and your rubric disagree, and one of them needs fixing. AutoQA scores also stay out of the agent's IQS, so the AI can't quietly move an agent's score on its own.

Agents feel it when nobody does:

Reddit

"My company uses AI to score calls, too. You can have a great call, but if it doesn't pick up on certain keywords you get a lousy score for it"

Treat the AI grader like a new reviewer joining the team. Put a few AutoQA-scored conversations into your calibration session, score them blind alongside it, and see where it lands. My guide to AI support QA covers the rest of that setup, and can AI do support QA is honest about where it still falls short.

The same logic applies when the agent is AI. If an AI agent answers tickets, your reviewers need to agree on how to grade its replies before they start, or you'll get the 88% versus 12% problem from earlier, the same gap you see when teams argue about AI resolution rate: one reviewer marks a reply as a pass because the answer is right, another marks it as a fail because it's three paragraphs too long. Agree on the bar first, then grade. A written brand voice for the AI makes the tone category much easier to calibrate. My post on AI agents in Zendesk QA walks through scoring bot conversations.

Common QA calibration mistakes

  • Scoring together in the room. You lose the one thing you came to measure, the original disagreement. Score blind first, always.
  • Only calibrating easy tickets. If every session ends in agreement, you're picking the wrong conversations.
  • Treating the average as the answer. Three people saying 4 doesn't make 4 right. Someone owns the ruling.
  • No written follow-up. If the ruling doesn't make it into the scorecard guidance, you'll have the same argument next month.
  • Comparing only total scores. Two reviewers can both give 80% while disagreeing on half the categories.
  • Never calibrating the AI grader. It reviews more tickets than anyone else on the team. It needs a seat in the session.
  • Running it without agents' disputes. Disputed reviews are free calibration material. Zendesk QA's disputes dashboard even has a "Disputed reviewers" table that "shows which reviewers may need additional help with consistency and accuracy." Pull from it when you pick conversations, and see my Zendesk QA agent feedback post for how disputes flow back to agents.

Try eesel for AI replies your QA team can trust

If you're adding AI to your support queue, the calibration question comes up on day one: how will your reviewers grade what it writes? eesel's AI helpdesk teammate joins your existing Zendesk, Freshdesk, Gorgias or Help Scout queue and lets your team settle that before it goes live. It replays real past tickets so reviewers can score its answers against what your agents actually sent, it starts as internal-note drafts, and every run shows the sources and reasoning behind the reply. When a reviewer corrects it, the correction becomes a standing instruction, so the same ruling applies to every future ticket. It's free to try with 100 credits, and paid plans start at $299 a month on the eesel pricing page.

eesel Activity page listing recent tasks with Approved, Rejected and Pending filters and linked Zendesk ticket numbers
eesel Activity page listing recent tasks with Approved, Rejected and Pending filters and linked Zendesk ticket numbers

Try eesel

Frequently Asked Questions

What is support QA calibration?
Support QA calibration is when every QA reviewer scores the same set of support conversations on their own, then compares scores to agree on one standard. The goal is that an agent gets the same score whoever reviews them. It sits on top of your normal support QA process and usually doesn't count toward agent scores.
How do you run a QA calibration session?
Pick 5 to 8 tricky conversations, have every reviewer score them blind, reveal the scores side by side against a baseline, discuss the two or three biggest gaps, and write each ruling into the scorecard guidance. My Zendesk QA workflow guide shows where calibration fits in a weekly QA rhythm.
How often should support teams do QA calibration?
Monthly suits most teams, weekly while a new scorecard or QA program beds in. Add an extra session after any policy change and whenever a new reviewer joins. Teams with more reviewers should calibrate more often, since there are more ways for support QA calibration to drift.
What is a good calibration score in QA?
A common target is every reviewer within about 5 points of the baseline on the total score, with no category where reviewers are 2 or more points apart on a 5-point scale. Zendesk's own guide puts the usual tolerance around 5 percent. Items like compliance checks should be 100%, set as auto-fail on your QA scorecard.
Should QA calibration be blind?
Yes. Scoring blind first is how you see the real disagreement before anyone is swayed by a louder voice. In Zendesk QA you can set reviewers to see peers' calibration reviews only after submitting their own.
Which QA tools have calibration features?
Zendesk QA, MaestroQA (now Rippit), Scorebuddy, evaluagent, Playvox and Level AI all list calibration features, with different ways of setting the right answer. Helpdesk-native QA tends to skip it. My roundup of customer support QA tools compares them in more depth.
How do you calibrate AI QA scores with human reviewers?
Treat the AI grader as one more reviewer: put AI-scored conversations into your calibration session, score them blind, and track how often humans override it. Zendesk QA's AutoQA accuracy score counts above 75% agreement as high. If an AI agent writes the replies, eesel's AI helpdesk teammate replays past tickets so reviewers can agree on how to grade it before launch.

Share this article

Riellvriany Indriawan

Article by

Riellvriany Indriawan

Riell is a designer and writer at eesel AI with about two years of experience researching CX platforms, AI chatbots, and helpdesk software. She combines her design background with a sharp eye for how these tools actually look and feel in practice — making her comparisons unusually visual and user-focused.

Related Posts

All posts →
Hand-drawn illustration of a QA reviewer with a magnifying glass and clipboard checking a set of support reply bubbles, four marked with check marks and one flagged
Guides

Multilingual support QA: how to review tickets in every language (2026)

Multilingual support QA fails where grammar scores pass. How to sample by language, who should review what, what Zendesk QA scores per language, and how to calibrate.

Kurnia KharismaKurnia KharismaOct 5, 2026
Hand-drawn illustration of a support lead with a magnifying glass sorting wrong ticket replies into bins for knowledge, sources, policy and handoff while two teammates watch
Guides

Customer support error analysis: how to find and fix why support replies go wrong

Customer support error analysis traces each wrong reply, human or AI, to its first failure. Label, count, fix the biggest bucket, then replay past tickets.

KiraKiraOct 5, 2026
Hand-drawn illustration of a support agent at a tablet marking up a reply with a red pen, circling a typo and a question mark, with a messy speech bubble turning into a clean one
Guides

Customer support language quality: how to measure and fix it in 2026

Customer support language quality is more than grammar. How to score spelling, clarity and tone, which helpdesk tools fix replies before send, and what AI changes.

Riellvriany IndriawanRiellvriany IndriawanOct 5, 2026
Hand-drawn illustration of a team lead and a support agent reviewing a scored conversation on a tablet, with a scorecard, coaching notes and a headset on the desk
Guides

Customer support coaching software: 10 best tools for 2026

Customer support coaching software compared: 10 tools sorted by when they coach (before, during or after the ticket), with real pricing, G2 scores and honest limits.

Kurnia KharismaKurnia KharismaOct 5, 2026
Hand-drawn illustration of a team lead pointing at a laptop while a support agent with a headset reviews a ticket, with note bubbles and a thumbs-up above them
Guides

Support agent feedback: how to give it so agents actually use it

Support agent feedback fails when it's late, sampled and scored without context. A weekly rhythm, a fair dispute path, and where feedback lives in your helpdesk.

Riellvriany IndriawanRiellvriany IndriawanOct 5, 2026
Hand-drawn illustration of a new support agent walking up a teal ramp of ticket cards toward a helpdesk screen, with a checked progress line above
Guides

Support agent ramp time: how to measure it and cut it in 2026

Support agent ramp time is three finish lines, not one. How long it really takes, how to track it in your helpdesk, and what actually shortens it.

Riellvriany IndriawanRiellvriany IndriawanOct 5, 2026
Hand-drawn illustration of a support agent with a headset writing in an open glossary notebook of pinned term cards, with arrows carrying speech bubbles to a globe and customers around it while a friendly AI helper points at the notebook
Guides

Support translation glossary: how to build one that every tool follows

A support translation glossary keeps brand names, technical terms and formality consistent across languages. What each helpdesk supports and how to build one.

KiraKiraOct 5, 2026
Hand-drawn illustration of a support lead marking a stack of help center articles with a red pen, using a checklist with keep, update and archive icons, while a colleague looks on
Guides

Support knowledge base audit: how to keep, fix, merge or archive every article

A support knowledge base audit sorts every help center article into keep, update, merge or archive. What to pull, how to triage, and which helpdesks help.

Riellvriany IndriawanRiellvriany IndriawanOct 5, 2026
Hand-drawn illustration of three support agents on a globe passing a box of tickets to each other as the sun and moon move across the sky
Guides

Follow-the-sun support: how to run a 24-hour queue without losing tickets at every handoff

Follow-the-sun support gets you 24-hour coverage without night shifts, but every shift change is a place tickets leak. Here's how to set it up, what it costs, and where AI fits.

Riellvriany IndriawanRiellvriany IndriawanOct 5, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free