Case Study 08Coaching · Agent development

A Score Is Not an Instruction: The Coaching Gap in Every QA Programme

A QA score tells an agent they failed. It does not tell them what to have said. This case study is about the translation step between a number and a behaviour — why rule engines cannot do it, why generic advice is worth nothing, and what a usable alternative looks like.

Say-able
the bar for every suggestion
Agents
the only party ever coached
Same
language as the flagged line
By person
how the weekly report ends
A flagged line of agent speech with the exact words isolated, and a replacement line offered directly beneath it in the same language.

Key findings

  1. A score tells an agent they failed. Producing the sentence they should have said instead is a different and much harder problem.
  2. Generic advice — show more empathy — is what a rule engine can produce, and it is worth nothing on a Monday morning.
  3. Coaching is mandatory on any negative or brand-risky agent line, and customers are never coached. Both are enforced mechanically, not by convention.
  4. The obvious way to satisfy a coaching requirement is to stop calling lines negative. An anti-gaming rule closes that door explicitly.

The gap between a number and a behaviour

A QA score tells an agent they failed. It does not tell them what to have said.

Empathy: 2/5. Opening score 67%. Rule adherence: small concern. Every one of those is a verdict, and none of them is an instruction. The agent reads it, agrees or disagrees, and goes back to the floor doing exactly what they did before — because nothing in the score describes an alternative action.

The translation from score to behaviour is done by a team leader, in a one-to-one, from memory, at the end of a shift. Which means it happens late, inconsistently, and only for the agents whose scores were bad enough to make the list.

The agents in the middle — where most of the recoverable performance actually sits — get nothing at all. Not because anyone decided they should not be coached, but because there are eleven of them and one team leader with forty minutes.

The score was never the deliverable. The changed sentence on the next call was.

Why rule engines cannot close it

Producing a specific, say-able alternative for a specific bad line requires understanding the line.

A rule engine knows that a phrase matched a pattern. It does not know what the agent was trying to do, what the customer had just said, or what would have worked better. So what it can produce is the category of advice that matched — show more empathy, follow the closing script — and that advice is worth nothing, because the agent already knows they were supposed to show empathy.

The gap is not information. It is wording. The agent needs the sentence.

What we changed: emit the replacement line

For any agent line that is negative or brand-risky, the platform produces a concrete professional rephrasing the agent could have said instead — isolated to the exact flagged words, in the same language, delivered as a drop-in replacement.

{ "flagged":    ["exact phrases from the segment"],
  "issue":      "<short reason>",
  "suggestion": "<drop-in replacement line, same language>" }

The bar set for the suggestion is one line long and does most of the work:

Concrete and say-able out loud. Not generic advice.

“Be more professional” fails that bar. A sentence the agent could have used, in the language they were speaking, does not.

Four decisions that raise this above a suggestions feature

1. Coaching is mandatory, not optional. Any negative or brand-risky agent segment must carry it. There is no configuration that turns this into a nice-to-have that gets skipped when a batch is large.

2. Customers are never coached — only agents. Enforced mechanically. It sounds obvious until you have seen a report advising a customer on their tone, which is what happens when the coaching step does not know who it is talking about. (This depends entirely on getting speaker attribution right — see Which Voice Is the Agent.)

3. The anti-gaming rule. The obvious way to satisfy “coach every negative line” is to stop calling lines negative. So the rule distinguishes:

Case Rule
Stance-negative — the agent is being unpleasant Keep it negative and coach it
Topic-bleed — the agent is discussing something unpleasant Fix the sentiment to neutral, no coaching

You may not downgrade a sentiment to dodge the coaching requirement.

An agent calmly explaining a repossession process is discussing something unpleasant. That is not negative conduct, and marking it as such is its own unfairness. But the distinction has to be drawn deliberately, or it becomes the escape hatch.

4. A mechanical gate enforces all of it. The batch is refused before load if a negative agent segment has no coaching, if coaching landed on a customer line, or if a brand-risk segment was not forced negative.

Where it surfaces

In the application, flagged spans are highlighted inline in the transcript with a “Better phrasing” view showing Issue / Flagged / Say instead. On the waveform, flagged agent moments are marked so a supervisor can jump straight to them.

And the weekly report closes the loop at floor level: the last page is “What To Do This Week, By Person” — owner, behaviour, due date. Not a list of scores to interpret. A list of work.

What we do not claim

This one needs stating precisely, because the capability and the automation are at different stages.

What is true: every flagged agent line comes back with a specific alternative — the exact words, in the same language, that the agent could have used instead. That is the deliverable, it is real, and it is what appears in the documentation and the report.

What is not yet true: automatic per-call coaching stored in the application for every agent, on every call, as a permanent record. The generation and the interface exist; the persistence layer does not yet. We would rather say that here than let you discover it in month two.

And an open question we have not answered. Should an agent see the machine’s rephrasing of their own words directly, without their supervisor in between? There is a reasonable case each way — immediacy and dignity pull in opposite directions — and we have not decided. We are raising it publicly because a product that pretends this question has an obvious answer has probably not thought about it.

Where to start

Take last month’s lowest-scoring agent and read three of their flagged calls. For each flagged line, write down what they should have said instead.

If you can do it in under a minute per line, your coaching gap is a capacity problem. If you cannot, it is a wording problem — and that is the one worth buying help with.

Book a demo and we will show you the flagged-line view, with the alternative alongside the original.


Related: Two Supervisors, Two Scores on the consistent score this depends on, and The Dashboard Nobody Opens on why the coaching has to arrive as assigned work.

Curious what is in the 95% you never hear?

Book a demo and we will walk you through the platform — how the reviews work, what the reports contain, and how the evidence trail is built.

Book a Demo