Good AI customer service feels like the customer got what they came for, fast, without having to think about the AI. They were not made to explain themselves three times. They were not pushed to a human when the AI could have helped, or pushed to a bot when they needed a human. They left with their problem solved and a sense that they were taken seriously. Most of these are invisible when present and obvious when absent.
A customer needs to change their shipping address an hour before a parcel ships. They message support. Two minutes later, the address is changed and confirmed. They forget the interaction within a day. That is good AI customer service. The same customer, in a worse system, calls back later because the change did not actually happen, and tells three friends about it. That is bad AI customer service. The two are distinguished by whether the customer's problem was solved and whether the customer felt respected while it was happening.
What people in the field are saying
Blake Morgan's "Experience is everything: the CX strategy..." argues that customer experience is everything because the rest of the business reads from the same scoreboard. AI customer service that is genuinely good lifts that score; AI that is fluent but unhelpful does not.
What are the markers of good?
Five, mostly invisible. The customer got the right answer or the right action. The interaction was short. The AI handed off cleanly when it needed to, with full context. The customer felt heard, especially if they were upset. The customer did not have to come back about the same thing later.
What are the markers of bad?
The customer had to repeat themselves to escalate. The AI's answer sounded right and was wrong. The action was confirmed and did not happen. The customer felt processed rather than helped. The customer came back about the same issue within a week.
What does this mean for measurement?
CSAT measures the in-the-moment feeling, which AI is good at producing whether it deserves to or not. The honest metrics are downstream: did the customer come back, did they renew, did they recommend the company afterward. Measuring those is harder; measuring CSAT alone is the trap.
What does this mean for design?
Design the AI for the customer's goal, not for the system's metric. If the metric is "resolve in chat," the AI will resolve in chat even when escalation was the right move. If the goal is "solve the customer's problem," the AI will escalate when it should and resolve when it can. The right metric follows from the right goal, not the other way round.
What is the customer's actual test?
Whether they would use the service again, willingly, the next time they have the same kind of problem. Run that test on a cohort that went through your AI. If most of them came back to a human channel next time, the AI was not good, no matter what CSAT said. If they came back to the AI, it was.
Related: why customers distrust AI bots, the CX belief gap, and the right balance between AI, self-service, and human support.