AI handles the language-translation surface of multilingual customer service well: it reads and replies in many languages, often with quality close to a human translator on common languages. It handles the deeper layer (cultural register, idiom, escalation norms in each language, low-resource languages) badly. The right pattern is to scope AI's multilingual role to where it genuinely works.
A customer messages support in Mandarin. The AI replies in Mandarin, fluently. The customer asks a follow-up that involves a regional turn of phrase. The AI either misreads it or replies in a register that feels stiff for the locale. Multilingual fluency at the surface masked a real gap underneath.
What people in the field are saying
Service Matters covers multi-language contact-centre infrastructure in "Demystifying orchestration: the key...": routing customers to language-appropriate handling is part of the orchestration problem, and AI changes which routing decisions stay rules and which become judgement calls.
What does multilingual AI do well?
Translation of routine content into common languages (Spanish, French, Mandarin, Portuguese, German). Answering FAQ-style questions in those languages from the same knowledge base. Maintaining a customer's preferred language across a session. Voice AI for routine, well-defined calls in the supported languages.
Where does it fail?
Low-resource languages: dialects, regional minorities, languages with limited training data. The AI reverts to a more common language or produces text that is grammatically correct but tone-deaf. Code-switching: the customer mixes two languages mid-conversation (common in many regions); the AI handles one part well and drops the other. Cultural register: politeness norms, escalation language, and how feedback is given vary by language; the AI rarely calibrates.
What about non-English speakers calling an English-trained AI?
This is the most damaging failure mode. The AI's speech recognition is weaker on accented English, the answer quality is lower, and the customer is statistically more likely to need escalation. A contact centre that does not measure performance per accent or per language will not notice; the average looks fine because the dominant accent is fine.
What is the practical pattern?
Scope the AI's multilingual role explicitly. For routine contacts in well-resourced languages, let the AI handle the whole conversation. For complex or sensitive contacts in any language, escalate. For under-resourced languages, route to a human in the customer's language if at all possible, rather than to an AI that will sound fluent and miss the substance. Measure performance per language and per accent; do not trust the average.
How would I start doing this?
Look at your contact mix by language. Pick the top one or two languages and treat them as the AI's launch scope. Hold the rest with the existing human team while you measure. Add language coverage as the data justifies, not as the vendor demo promised.
Related: the field note on the accent gap in voice AI, how well voice AI works, and where AI belongs in the channel mix.