Case Study 03Language · Hindi, English & Hinglish
The Language Tax: Why English-First Tooling Quietly Fails on Hinglish
Indian agents switch between Hindi, English and Hinglish inside a single sentence. Confronted with a code-switched call, English-first conversation intelligence does not fail loudly — it returns a plausible score built from the fraction it understood. This case study is about what it takes to judge a call in the language it was spoken.
- Native
- the language judgement happens in
- 3
- languages inside one sentence
- After
- where translation sits, not before
- 600+
- risk behaviours, written for this market
Key findings
- Code-switching is not an edge case on an Indian floor. It is the default register, and it happens mid-clause.
- English-first tooling fails silently: it scores the fraction it parsed and reports a confident number for the whole call.
- Translation is not the problem. Translating before judgement is — that is where sarcasm, politeness and threat get flattened.
- Acoustic emotion is a signal, not a validated measurement in Indian languages. We say so rather than let it be assumed.
One sentence, three languages
“Sir, aapka EMI due hai, I’m calling from the recovery department.”
That is not a curiosity. On an Indian collections or support floor it is the default register — agents switch between Hindi, English and Hinglish constantly, unconsciously, and in the middle of a clause. They are not choosing a language for the call. They are using all of them at once, the way the customer does.
Almost every conversation-intelligence platform on the market was built on English phrase matching. Confronted with a call like that one, it does something worse than failing: it fails quietly.
It parses what it recognises, scores that, and returns a number. The dashboard fills with confident-looking metrics. Nothing anywhere indicates that half the conversation was never analysed, because from the tool’s point of view nothing went wrong — it processed everything it could see.
Why this is architectural, not a tuning problem
The instinct is to treat this as a coverage-of-vocabulary issue: add Hindi terms to the dictionary and the problem shrinks. It does not, for two reasons.
A keyword dictionary is written in a language. Proximity queries over English phrases cannot score a Hindi utterance badly — they cannot score it at all. There is no partial credit. The clause simply is not in the index, so it contributes nothing to any score, and its absence is indistinguishable from silence.
And the dictionary is usually the vendor’s most valuable asset. Years of tuning, thousands of curated phrases, the thing the product is actually sold on. Replacing it is the most expensive decision the vendor can make, so it does not get made. Instead the market gets a translation shim bolted to the front, which we will come back to, because it is the wrong shape.
There is a third problem that nobody mentions in the demo. Off-the-shelf risk taxonomies in this category are overwhelmingly US healthcare and insurance in origin — Medicare Advantage disclaimers, ACA exchange language, TCPA and DNC rules, BBB complaint handling. Those are real, well-built taxonomies. They are also close to useless on a Mumbai recovery floor, where the risk vocabulary is unauthorised payment channels, credit-score coercion, community shaming and property threats. A tool can be simultaneously excellent and entirely inapplicable.
The shim that looks like a solution
The common workaround is to translate the call to English, then run the existing English analysis over the translation.
It demos well. It is also where the meaning goes.
Sarcasm does not survive machine translation. Neither does politeness register — and in Hindi, politeness register is doing enormous work: the difference between a firm reminder and an implied threat often lives entirely in verb form and honorific, both of which flatten to the same neutral English sentence. Idiom becomes literal. An indirect threat, which is the only kind a trained collections agent makes, becomes a polite request for payment.
So the analysis scores a conversation that nobody actually had. And because the translated text is fluent and plausible, there is no signal that anything was lost.
What we changed: judge native, display English
The decision underneath our analysis layer is a single ordering constraint, and everything else follows from it.
Analysis is judged on the native transcript. Never on a machine translation. Nuance and severity live in the language the words were spoken in, so that is where judgement happens.
Translation still exists — we are not pretending otherwise, and a claim that there is no translation anywhere in the system would be false. The point is where it sits:
We judge the call in the language it was spoken. Translation happens after the analysis, for your report — never before it, where the nuance would be lost.
The translated layer is stored additively, alongside the native transcript rather than replacing it. The native text remains the source of truth; nothing is overwritten. Your report reads in English, because your compliance officer reads in English. The judgement that produced it read Hindi.
That is a more interesting claim than “no translation layer,” and it is one we can defend line by line.
Two things that had to be rebuilt underneath it
Working out who the agent is. Speaker attribution used to lean on English-only opening patterns, which is hopeless when the call opens with namaste, mera naam, bol raha, or की तरफ से. Opening-cue scoring now runs over transliterated and Devanagari cues in the first forty-five seconds, plus collections vocabulary — loan, emi, बकाया, रिकवरी. Get this wrong and every downstream score is attached to the wrong person, which is its own case study.
Risk vocabulary written for this market. Not translated into it — written for it. Unauthorised payment channels (UPI, Paytm, PhonePe, GPay), credit and identity coercion (CIBIL, Aadhaar card, PAN card), community shaming (village council, village head, neighbours, the watchman), property threats (seal your house, sell your gold). These are the phrases that carry regulatory risk on an Indian floor, and none of them appear in a taxonomy built in Hartford.
Transcription runs on Whisper large-v3-turbo with the language auto-detected, in transcribe mode rather than translate mode — so the native transcript is preserved rather than being silently Englished at the door.
One thing we will not claim
Our acoustic emotion model is fine-tuned on English-language data. Its quality on real Indian-language telephony has not been validated against a labelled sample.
So we describe emotion as a signal that contributes to a picture, and we do not describe it as a validated measurement for Hindi. If a vendor tells you their emotion detection is accurate in Indian languages, ask what it was validated against. It is a fair question and it has a specific answer.
We would rather be the vendor that draws that line in public than the one that gets asked about it in month four.
Where to start
The test is quick and you can run it on tooling you already have. Take three genuinely code-switched calls — not English calls with a Hindi greeting, but real mid-clause switching. Run them through whatever is scoring your floor today. Then read the transcript against the audio and mark what got dropped.
Most teams are surprised by how much of the risky material sits in the half that was never scored. That is not a coincidence: agents under pressure switch to their first language, and pressure is exactly where the compliance risk is.
Book a demo and we will show you how the platform handles Hinglish — the native transcript beside the English report.
Related: What You Cannot Prove on the evidence trail, and The 5% Illusion on the coverage gap underneath both.
Curious what is in the 95% you never hear?
Book a demo and we will walk you through the platform — how the reviews work, what the reports contain, and how the evidence trail is built.
Book a Demo