Case Study 05Scoring fairness · Agent performance

What We Refuse to Score: The Fairness Gap That Ends QA Programmes

The fastest way to lose a contact centre floor is to score agents for things they did not control. Agents notice within one cycle and the programme never recovers its legitimacy. This case study is about the machinery whose entire purpose is to not score things — and why that is the hardest kind of machinery to sell.

Excluded
not scored zero, when unjudgeable
Context
what customer anger counts as
100/100
calls the gate agreed on in validation
1 cycle
for a floor to notice unfairness
A column of scored calls beside a set held back behind a gate — excluded rather than scored zero, each carrying the reason it was not judged.

Key findings

  1. A QA programme does not die from being wrong. It dies from being unfair — agents notice within one cycle, and the scores keep getting produced long after anyone acts on them.
  2. Customer frustration is context, not the agent's fault. It is scored as the agent's response to it, never as their conduct.
  3. A call with no conversation in it cannot be judged. Those are excluded and recorded as excluded, not scored zero.
  4. Fairness costs recall, and recall is what demos well. This is machinery whose entire purpose is to not score things, which is hard to sell and easy to skip.

The failure that ends programmes

QA programmes rarely die from inaccuracy. They die from unfairness.

The sequence is always the same. A score lands that an agent knows is wrong. They can explain exactly why it is wrong, in one sentence, to anyone who will listen. The team lead escalates, or the union does. And from that week onward the scores are still produced, still distributed, and quietly ignored by everyone including the people producing them.

The programme does not get cancelled. It gets disbelieved, which is worse, because the cost stays and the benefit leaves.

Three unfairness patterns recur across the category:

  1. The customer was furious, so the agent’s score dropped — even though the agent handled it perfectly. Sentiment tools measure the call, not the agent’s contribution to it.
  2. The call was never really answered, or the customer’s channel was lost — and the agent scored zero on every skill, because there was no conversation in which to follow a script.
  3. The agent used a phrasing the keyword list had never seen — so a skill they genuinely performed scored as absent.

Each one produces a defensible-looking number. Each is wrong in a direction that is expensive for the agent to dispute and cheap for the vendor to ignore.

The agent does not need to prove the score is unfair. They only need to believe it, and tell four colleagues.

Why the category skips this

Because fairness costs recall, and recall is what demos well.

Every gate described below reduces the number of scored calls, the number of flagged behaviours, and the apparent thoroughness of the product. In a competitive evaluation, the vendor who scores everything looks more complete than the vendor who explains what they declined to score.

And it requires building machinery whose entire purpose is to not score things. That is difficult to specify, invisible in a demo, and it is the first thing cut when a roadmap gets compressed.

What we changed: four gates, each closing one pattern

1. Customer frustration is context, not fault

This is a top-level rule on the analysis, stated in exactly these terms:

Customer frustration is context, not the agent’s fault. Never let it lower an agent’s score; score only the agent’s response.

It is enforced structurally as well as by instruction. The weekly report gives customer frustration its own page — “Customer Frustration — And How The Team Handled It” — which separates the event from the handling at the level of the document, not just the score. A supervisor reading it sees an upset customer and, next to it, what the agent did about it.

An angry customer is a fact about the week. It is not a fact about the agent.

2. Unjudgeable calls are excluded, not scored zero

Some things look like calls and are not. The clearest example: sixty-eight seconds of “hello… hello ma’am… hello hello” with zero customer words. There is no conversation in it. There is nothing to score.

Scoring it anyway produces a zero on every skill — and that zero enters the agent’s average, their trend, and their ranking against colleagues, for a call in which nothing happened that they controlled.

The engagement gate catches two shapes:

  • Unanswered or dropped — one side never really speaks.
  • Customer channel missing — a recording or diarization defect. Classified as a data-quality issue and explicitly marked not agent behaviour.

Two design details make it trustworthy rather than merely present:

It is role-agnostic by design. The gate tests the quietest channel, never “the customer channel”. So it behaves identically whether it runs before or after speaker roles are worked out — which means an attribution mistake (see Which Voice Is the Agent) cannot leak into the fairness decision. On a real run it agreed with the role-based test on 100 of 100 calls at every cutoff.

The same code is imported by both the filter step and the independent audit step, specifically so the rule cannot drift between them. One definition, two callers, no possibility of the audit using a different standard than the filter.

And excluded calls are recorded as excluded rather than dropped. In a recent live workspace run, 88 calls were scored and four were carried in the skip list — visible, named, and available to anyone who asks why the denominator is 88.

3. Detection keeps learning, so novel phrasings do not silently zero a skill

An agent who performs a skill in words the rule set has never seen scores as though they did not perform it. Over time this penalises exactly the agents who do not read from the script — including the good ones.

So when a genuine phrasing is missed, a candidate rule is generated with positive examples, at least one negative example — the “when this must be FALSE” guard — and a positional zone. It is then self-tested for precision against a rolling corpus before it is appended. Never hand-edited into the live rules.

The stated principle for this subsystem is one line:

Fairness over recall.

A rule that catches a new phrasing but also fires on something innocent is rejected, even though it would raise the detection rate.

4. Coaching cannot be dodged by re-labelling

There is an obvious way to satisfy a rule that says “every negative agent line must carry coaching”: stop calling lines negative.

So the rule distinguishes two cases, and refuses the batch if they are confused:

Case Rule
Stance-negative — the agent is being unpleasant Keep it negative and coach it
Topic-bleed — the agent is discussing something unpleasant Fix the sentiment to neutral, no coaching

You may not downgrade a sentiment to dodge the coaching requirement.

This one matters more than it looks, because it is the gate that keeps the other three honest. Without it, every fairness rule becomes an incentive to quietly reclassify.

What this costs, stated plainly

Every gate here lowers a number that a competitor will show you.

Fewer calls are scored. Fewer behaviours are flagged. The denominator on an agent’s weekly score moves week to week, which is mildly annoying to explain and occasionally has to be explained.

We think that is the correct trade, because a score that an agent accepts is worth more than a score that covers more calls. But it is a trade, and you should evaluate it as one rather than take our word for which side of it is better.

What we do not claim

We do not claim the gates catch every unfair score. They close three specific patterns we have watched destroy programmes; they are not a general fairness guarantee, and no vendor can honestly offer one.

We do not claim agents will like their scores. Fair and welcome are different things. The claim is narrower and more useful: when an agent disputes a score, there is a specific rule, a specific segment and a specific set of words to have the argument about.

Where to start

If you are already running QA, the fastest diagnostic costs nothing: pull the five lowest agent scores from last month and read the calls. Count how many are low because of something the agent did, and how many are low because of something that happened to them.

That ratio is your fairness gap, and it is usually the reason the programme is not changing behaviour.

Book a demo and we will walk you through what the gates exclude and why each exclusion is recorded rather than silent.


Related: Four Hundred Flags on the verification step ahead of these gates, and Two Supervisors, Two Scores on the consistency problem underneath them.

Curious what is in the 95% you never hear?

Book a demo and we will walk you through the platform — how the reviews work, what the reports contain, and how the evidence trail is built.

Book a Demo