Case Study 06Speaker attribution · Technical

Which Voice Is the Agent: The Silent Failure Under Every Scorecard

Every agent score depends on one prior question — which voice is the agent? Get it wrong and the scorecard inverts: the customer's rudeness becomes the agent's conduct violation. This case study is about why the standard heuristic fails on exactly the calls you care about, and what replaces it.

+100
weight on a self-introduction
+5
weight on talking the most
45s
window the decision is made in
Recorded
every swap, auditable after the fact
Two channel lanes with the opening forty-five seconds bracketed, weighted cues scoring one lane as the agent, and a recorded swap decision.

Key findings

  1. Speaker attribution is the highest-consequence silent failure in call QA. Get it wrong and every downstream number inverts rather than degrades.
  2. The industry heuristic — the agent is whoever talked most — fails precisely on the calls worth reviewing: the quiet agent, the dominant customer, the transfer.
  3. Content decides the question, not talk-time. Talk-time is kept as a weak last-resort tiebreak, weighted twenty times lower than a self-introduction.
  4. Every attribution decision is recorded on the analysis record, so a disputed score can be audited back to the moment the channel was chosen.

One question sits underneath every other one

Before a call can be scored, something has to answer a question so basic that most systems never surface it: which voice is the agent?

Get it right and nothing happens — which is why it is invisible. Get it wrong and the entire scorecard does not degrade, it inverts:

  • The customer’s rudeness is recorded as the agent’s conduct violation.
  • The agent’s empathy is read as customer sentiment.
  • The greeting the agent actually gave scores as a missed opening.
  • The coaching note goes to the person who did nothing wrong.

This is the single highest-consequence silent failure in call QA. It produces a scorecard that is internally consistent, confidently presented, and about the wrong person.

A wrong answer here does not make the score noisy. It makes the score belong to somebody else.

Why it happens more than vendors admit

The standard heuristic is the agent is whoever talked most.

It is a reasonable first guess. Agents usually do talk more, and on a clean stereo recording with a cooperative customer it is right most of the time. Published accuracy for the approach hovers around the 80% mark in practice.

The problem is the shape of the 20%. It is not randomly distributed across your call volume — it concentrates in exactly the calls a QA programme exists to find:

  • The agent went quiet. They were unsure, or being shouted at, or had lost control of the call. These are the coaching calls.
  • The customer dominated. Long complaint, long escalation. These are the compliance calls.
  • The call was transferred, so a third voice appears partway through.
  • The recording is mono, so “who said that” is an inference rather than a fact from the outset.

An 80%-accurate heuristic that fails on the interesting 20% is not 80% accurate for your purposes. It is close to useless on the population you actually care about, and perfectly reliable on the calls you would never have reviewed anyway.

What we changed: content decides, talk-time is a tiebreak

A layered detector works out which channel is the agent from what is said, not from how much, scoring the opening forty-five seconds of the call.

Signal Weight
Self-introduction with affiliation — “calling from X” +100
“X speaking” +100
Service greeting +40
Opens the call +15
Talks the most +5weak last-resort tiebreak only

The ratio is the design. A self-introduction is worth twenty times what talking the most is worth, because one of those is evidence and the other is a correlation. Talk-time is not removed — it still breaks a genuine tie — but it can never outvote somebody saying who they are.

Four layers on top of the score

1. An explicit override. If the ingestion supplied an agent name, the first segment containing that name defines the agent’s channel outright. No scoring, no inference. Known beats guessed.

2. Multilingual cue scoring. The cues are transliterated and Devanagari as well as English — namaste, vanakkam, नमस्ते, mera naam, bol raha, की तरफ से — plus collections vocabulary such as loan, emi, बकाया and रिकवरी. This replaced English-only pattern matching, which on a Hindi call scores nothing at all and therefore falls through to talk-time every time. (More on that in The Language Tax.)

3. Deterministic normalisation before anything else runs. Channel 0 is normalised to Agent as a first layer, before any language model sees the transcript. That step carries a hard guarantee worth quoting, because it is the kind of thing that is easy to promise and easy to violate:

Words are reproduced VERBATIM. The only thing this script changes is the channel/role LABEL. It never rewords, corrects spelling, merges, or invents text.

A step that relabels speakers must not touch a single word, because the words are what an agent will later be quoted on.

4. Regression tests, including named ones. The detector ships with self-tests, two of which encode specific failures that occurred and must not recur. A speaker-attribution component without regression tests is a component that will silently regress, because nothing downstream will tell you it did.

The part that makes it auditable

Every analysis record carries whether a channel swap was applied and a normalisation block naming the source of the decision.

That is what turns a disputed score into a question with an answer. When an agent says that wasn’t me, there is a record of which channel was chosen, on what basis, and whether it was swapped — rather than a shrug and a re-listen.

What we do not claim

We do not claim perfect attribution. Mono recordings with two similar voices and no self-introduction are genuinely hard, and any vendor claiming certainty on that population is describing a demo rather than a floor.

We do not claim the weights are optimal. They are deliberate and they are documented, which is a different and more checkable claim. You can look at them and disagree.

What we do claim is narrower: the decision is made from content, the reasoning is recorded, and when it goes wrong you can find out that it went wrong. Almost no competitor documents their speaker-attribution logic at all — which is worth noticing, given that every number they sell you depends on it.

Where to start

If you are evaluating any call-QA product, ask the question directly: how do you decide which speaker is the agent, and where is that decision recorded?

The answers sort vendors quickly. “Our diarization handles it” means talk-time. “It comes from the dialler” means it is right until a call is transferred. A real answer names the signals and their order.

Book a demo and we will show you the attribution decision on a call where the agent barely spoke.


Related: Mono, Stereo and the Honest Answer on the recording format underneath this, and What We Refuse to Score on the gate that stays role-agnostic so attribution mistakes cannot leak into it.

Curious what is in the 95% you never hear?

Book a demo and we will walk you through the platform — how the reviews work, what the reports contain, and how the evidence trail is built.

Book a Demo