Wrap-up codes (the categories an agent assigns to a closed contact: "billing question," "shipping delay," "subscription cancellation") are now generated by AI from the conversation rather than typed by the agent at the end of the call. This is faster and more consistent, but it also exposes that the rubric itself was often inconsistent across agents. The AI does not reveal new problems; it makes existing problems visible.

A contact centre's reporting shows that "general inquiry" is the most common wrap-up code. Before AI, that meant agents picked it when they did not have time to find the right code. With AI auto-coding, the same category is now applied based on actual conversation content, and the share drops sharply because most conversations have a more specific topic. The reporting was always wrong; AI revealed which way.

What people in the field are saying

CX Decoded's "The AI metrics mirage..." includes the wrap-up-code question in its broader critique: many of the metrics a contact centre runs on were honest in spirit and dishonest in practice, and AI exposes the dishonesty before fixing it.

What does AI do well here?

Reads the conversation, extracts the topic, applies the code from the available rubric. Consistent across agents and across shifts. Faster (no after-call work for the agent on this dimension). Multi-label where appropriate (a contact can have a primary topic and a secondary one).

What does it do badly?

Forces a contact into a code when none of the available codes fit, instead of flagging that the rubric needs a new code. Codes by surface topic ("billing") when the underlying issue is something else ("the bill is right but the customer doesn't understand it"). Misses subtlety that an experienced agent would have caught.

What gets exposed?

Three things. Codes that were over-used because agents could not find the right one. Codes that overlap and should be merged. Topics the rubric does not have a code for but should. The cleanup work that always needed doing finally has the data to support it.

What is the implication for reporting?

Year-over-year comparisons on wrap-up codes break in the year you introduce AI auto-coding. The historic data was wrong in a different direction than the new data. Comparisons must be restated, or use a parallel window where both methods ran. Pretending the comparison is valid produces wrong conclusions.

What is the practical first step?

Run AI auto-coding in shadow mode for a quarter. Compare the AI's codes to the agents' codes per case. The diffs are where the rubric was unclear. Clean the rubric (merge, split, add missing codes) before switching to AI as the source of truth.

Related: what AI changes about contact-centre QA, why AI containment numbers are misleading, and what AI does to FCR.