Case Study 12Configuration · Multi-campaign floors

One Rubric Does Not Fit Every Campaign: The Configuration Gap

A collections floor, an inbound support desk and an outbound sales team are three different jobs. Scoring them on one rubric produces numbers that are technically consistent and operationally meaningless. This case study is about per-team configuration, and the keyword algebra that stops it becoming a loophole.

Additive
risk words, never replacing
Logged
every approval that hid a severe flag
No
re-scoring of already-audited calls
1
config read by prompt, linter and rollup
Two overlapping keyword sets: defaults plus a workspace's additions, with an explicit allow list subtracted and the subtraction logged.

Key findings

  1. Collections, support and sales are three different jobs. One rubric across them produces numbers that are consistent and meaningless.
  2. Per-team configuration multiplies the surface area of reporting, benchmarks and support — which is why vendors defer it.
  3. Risk vocabulary is additive: a workspace can add words but cannot delete defaults by misconfiguring. No workspace can lose detection by accident.
  4. When an approved word mutes something the model rated severe, that suppression is logged — because a client could otherwise quietly hide their own bad behaviour.

Three jobs, one scorecard

A collections floor, an inbound support desk and an outbound sales team are three different jobs with three different definitions of a good call.

Score them against one rubric and you get numbers that are technically consistent and operationally meaningless. The collections team is penalised for not closing a sale. The sales team is penalised for not verifying identity against a mandate they do not operate under. The support desk is measured on a talk ratio calibrated for outbound.

Everyone can see the numbers are wrong. Nobody can say precisely how wrong, because the rubric is applied identically and therefore looks fair. And the first time a team lead successfully argues that the scorecard does not describe their campaign, the scorecard stops being used for that campaign — and then, by precedent, for the others.

A consistent rubric applied to inconsistent work is not fairness. It is a category error with a spreadsheet attached.

Why vendors defer it

Because per-team configuration multiplies the surface area of everything downstream.

The moment two teams can be scored differently, every report needs to know which configuration produced it, every benchmark needs a scope, every cross-team comparison needs a caveat, and every support conversation starts with “which workspace?”. Trend lines break when a configuration changes mid-quarter unless somebody thought about that in advance.

It is a lot of work that produces no new capability — only correctness. So it gets deferred, and the usual compromise is a single global configuration, or per-tenant settings that require a support ticket to change and therefore never change.

What we changed: the workspace is the unit of analysis

A workspace maps to a team, a campaign or a department, and it carries the configuration that decides what its numbers mean.

Setting Effect
PII masking On/off per workspace; redacts names, account numbers and ID numbers from transcripts
Primary language English or other, including Hindi and code-switched Hinglish
Filename format The pattern used to parse metadata out of dialler filenames
Flagged / approved keywords Per-workspace brand-risk vocabulary
Call-stage weights Opening / handling / closing — must sum to 100
Volume target Calls per agent per day
Duration benchmark Average handling time
Talk ratio The healthy band, commonly cited at 40–60%
Cross-talk benchmark Interruption threshold
Processing limits Minimum duration, maximum duration, percentage processed

The stated purpose is one sentence, and it is the right test for whether a configuration surface is worth building:

Workspaces exist so that every number downstream means what you think it means.

The keyword algebra, and why it is adversarial

Per-workspace risk vocabulary is where this could quietly become a loophole, so it is worth showing the design rather than describing it.

effective_risk = (default ∪ workspace.extra_risk) − workspace.allow
approved       = workspace.allow

Five constraints, each of which exists for a stated reason:

Additive, never replacing. A workspace can add risk words. It cannot delete the defaults by misconfiguring, by importing a shorter list, or by clearing a field. No workspace can lose detection by accident.

The allow list is the only removal mechanism — which makes every removal an explicit, named, auditable act rather than a side effect.

A suppression log records when an approved word muted something the model rated severe. This is the part that matters. Without it, a client could approve a term and quietly stop hearing about their own worst calls, and nobody — including us — would be able to tell the difference between a legitimate false-positive fix and a cover-up. With it, the approval is still allowed and the consequence is still visible.

No re-scoring of already-audited calls. A configuration change applies going forward. Historical reports stay explainable, which means a report from March still says what it said in March, and a trend line is not quietly rewritten by a settings change in July.

One shared configuration is read by the analysis prompt, the linter and the rollup — so the three cannot drift apart and start disagreeing about what counts as risky.

And two operational details: the configuration has a safe fallback — if the fetch fails, use the default and log loudly; if the database is down, stop rather than proceed on guesses — and client-supplied words are treated as a prompt-injection boundary, sanitised and capped, because they are text from a customer being injected into an AI prompt.

What we do not claim

We do not claim the configuration surface is simple. It is genuinely the richest part of the product, and a floor with eight campaigns has real setup work to do. We would rather have that conversation than ship one global rubric.

We do not claim per-workspace settings make cross-team comparison automatic. Two workspaces with different weights produce scores that are internally comparable and cross-comparable only with care. That is a property of the underlying reality, not a limitation we can engineer away.

Where to start

Take one week of scores from your busiest campaign and one from your quietest. Ask whether the same rubric produced both, and whether it should have.

If the answer is “the same rubric, and no, it should not have”, you have found the reason your QA numbers are argued about rather than acted on.

Book a demo and we will walk you through the workspace configuration and the suppression log on a real risk-word list.


Related: Two Supervisors, Two Scores on the rubric this configures, and Where Your Recordings Live on the residency and masking settings alongside it.

Curious what is in the 95% you never hear?

Book a demo and we will walk you through the platform — how the reviews work, what the reports contain, and how the evidence trail is built.

Book a Demo