Case Study 11Recording quality · Telephony
Mono, Stereo and the Honest Answer: A Trade-off the Category Hides
Whether your dialler records in stereo or mono is usually a default nobody chose, and it has an outsized effect on QA quality. Most vendors either deliver worse analysis quietly or make you re-plumb your telephony first. This case study is about publishing the trade-off instead.
- Both
- mono and stereo accepted
- Stereo
- the format that analyses better
- Auto
- downgrade when a stream lies
- Said
- out loud, rather than buried
Key findings
- Mono versus stereo is usually a telephony default nobody chose, and it changes how much of your QA output is fact rather than inference.
- Admitting the difference costs deals. Requiring stereo also costs deals. So the industry norm is to say nothing.
- Mono is handled by duplicating the file and muting the other speaker's intervals in each copy, so per-channel analysis has real audio to work with either way.
- A file that declares itself stereo but carries a single-channel stream is downgraded before submission, rather than producing confident garbage.
A decision nobody made, with consequences nobody sees
Somewhere in your telephony configuration there is a setting that determines whether calls are recorded in stereo — agent on one channel, customer on the other — or mono, with both mixed into a single track.
Almost nobody chose it. It is a default from a configuration set years ago, by someone who has since left, for reasons that were about storage cost rather than analysis quality.
It has an outsized effect on everything a QA programme produces. In stereo, “who said that” is a fact — it is a property of the file. In mono it becomes a machine-learning inference, and every metric built on top of speaker identity degrades at once:
- talk ratio,
- cross-talk and interruption counts,
- dead air attribution — whose silence was it,
- per-speaker sentiment,
- per-speaker emotion,
- and the agent scorecard itself, which depends entirely on knowing which voice to score.
None of that degradation is visible in the output. The dashboard renders identically. The numbers have the same number of decimal places.
The mono call and the stereo call produce reports that look exactly alike. Only one of them is mostly measurement.
Why nobody tells you
Because both honest options cost money.
Admitting the difference costs deals. “Our analysis is meaningfully better on stereo” invites the follow-up question, and the follow-up question stalls the evaluation while someone checks.
Requiring stereo also costs deals — and worse, it puts the telephony team in the critical path of a QA purchase, which is how a two-week onboarding becomes a two-quarter project. (The same dynamic described in Twenty Minutes, Not Two Quarters.)
So the industry norm is to say nothing. Take the mono file, run it, produce a report, and let the customer assume the numbers mean what the numbers usually mean.
What we changed: handle both, and publish the difference
Stereo. The real left and right channels are split and analysed independently. Speaker identity is read from the file rather than inferred.
Mono. The file is duplicated, and in each copy the other speaker’s intervals are muted. Both copies then carry real audio for one speaker and silence for the other.
That second one is non-obvious and worth explaining, because the naive alternative is to run per-channel analysis on the mixed track twice and get the same answer twice. Duplicate-and-mute means a mono call still produces a genuine per-speaker emotion profile rather than a blank field or a duplicated one — the acoustic model is looking at one voice at a time, as it was designed to.
Automatic downgrade. If a file declares itself stereo but any stream carries fewer than two channels, the stereo flag is dropped before submission. Files lie about this more often than you would expect, usually after a transcoding step somewhere upstream. Trusting the declaration would produce a channel split on a track that has no second channel — which does not fail loudly, it produces a confident empty result for one speaker.
Speaker attribution on mono is then handled by inference plus a content-based role detector rather than by assuming channel order — described in Which Voice Is the Agent.
The position, stated plainly
Recordings may be mono or stereo, though stereo with separate agent and customer channels provides superior analysis quality.
And the counterpart, which matters just as much:
You should not have to change your dialler to get a good audit either way.
Both halves are load-bearing. The first is the honest technical answer. The second is the commitment that the honest answer will not be used as a reason to send you away and make you fix your telephony first.
If you are already stereo, you get the better version. If you are mono, you get a version that works, built deliberately for mono rather than degraded into it — and you get told which one you are getting.
What we do not claim
We do not claim mono and stereo are equivalent. They are not, we have said which is better, and we would rather you hear it from us than discover it when a talk-ratio number does not match your own logs.
We do not claim the mono handling recovers everything. Speaker attribution on a mono call with two similar voices and no self-introduction is genuinely hard, and there is no processing trick that makes it as reliable as two separate channels.
We do not claim you should re-plumb your telephony. If it is easy, stereo is better. If it is a project, it is not worth blocking a QA programme on — and that is a judgement about your organisation, not about audio.
Where to start
Pull one recording from your storage and check whether it has one channel or two. It takes about a minute and most QA leads have never done it.
Whatever you find is worth knowing before you evaluate anything in this category, because it determines which half of every vendor’s accuracy claim applies to you.
Book a demo and we will show you how a mono call and a stereo call differ in what the analysis can say.
Related: Which Voice Is the Agent on attribution when the channels cannot answer it, and Twenty Minutes, Not Two Quarters on keeping telephony out of the critical path.
Curious what is in the 95% you never hear?
Book a demo and we will walk you through the platform — how the reviews work, what the reports contain, and how the evidence trail is built.
Book a Demo