The 30-second answer
If 100 customers start with the AI and 40 of them end up on a human's screen, your escalation rate is 40%. Same dataset as containment, viewed from the other side.
Why it matters more than containment on its own
A low containment rate looks bad on a dashboard. A high escalation rate looks bad on a dashboard. But neither tells you the right story without two extra pieces:
- When the AI escalates. Early (the AI says "this looks complex, let me get a person" before frustrating the customer) is good. Late (after the customer has typed the same thing three times and is now angry) is bad. Same escalation rate, completely different customer experience.
- What the human receives at the handoff. A clean handover means the human sees the full transcript, knows what the AI tried, and starts where the AI left off. A bad handover dumps the customer on a new agent with no context — and the customer has to repeat everything.
An AI agent with a 45% escalation rate that escalates early with clean context is providing real value. An AI agent with a 25% escalation rate that escalates late with no context is destroying it. The dashboard number doesn't distinguish those two.
How vendors tune the escalation threshold
Every AI-agent vendor (Intercom Fin, Decagon, Sierra, Ada) has a tunable threshold for when the bot gives up and asks for a human. The trade-off is direct:
- Lower the threshold (escalate faster) → higher escalation rate, fewer frustrated customers, more human cost.
- Raise the threshold (try harder before escalating) → lower escalation rate, better headline containment, more customers stuck in dead-end loops.
Most vendors default to a setting that optimises for their containment number in case studies. That's not necessarily the setting that optimises for your customers. Worth tuning early in a deployment.
What to measure (escalation quality, not just rate)
Four things, in order of how often they get measured (least to most):
- Escalation timing. What turn of the conversation did the escalation happen at? Most vendors report this; few teams look.
- Handoff completeness. Does the human see the AI's full transcript and the customer's identity? Surprisingly often: no.
- Repeat-context rate. When the customer reaches a human, do they have to re-explain the problem? If yes, the handoff is broken.
- CSAT post-escalation. A successful escalation should produce higher CSAT than an AI-only resolution on the same kind of issue (the human gets to demonstrate real care). If it doesn't, the friction of the handoff is eating the goodwill.
The handoff cost question
Every escalation costs human time. A vendor will quote you a cost-per-contact figure to estimate AI savings; the honest version subtracts the cost of escalations from that figure. If your AI escalates 40% of contacts and the average escalated contact takes 50% longer to resolve than a non-AI baseline (because the human now has to undo whatever the AI tried), the savings calculation gets thinner fast.
This is where recontact piles on top: a contact that got escalated, handled, and then recontacted within 72 hours is the most expensive contact in your dataset. The ones that matter for ROI aren't the contained ones — they're the escalated-then-recontacted ones, because those are the ones where the AI added cost without producing value.