Simulated Customer Conversations Could Keep AI Support Skills Improving
Companies that deploy AI customer-support agents often equip them with “Skills,” portable modules that package domain knowledge and handling procedures. Most Skills are written by hand or generated in a single model pass, and they rarely improve from the failed conversations they cause. A paper from Tencent Cloud Andon and Zhejiang University argues that the main obstacle to sustained improvement is the feedback used to guide revisions. The authors call their system SkillEvo.
Earlier self-improvement methods typically test a Skill with a single question-and-answer exchange. The authors contend that this feedback runs dry. Once the first revision fixes the gaps visible in the opening message, the signal for further changes fades, and problems that surface only across several turns stay hidden. SkillEvo instead uses a simulated user, built from real human-handled support tickets, that holds a multi-turn conversation with the agent. The simulated user works from a list of key and minor intents, a set of known facts, and a shifting emotional state. An intent state machine ensures each intent is raised and addressed before the conversation ends. A separate attribution step then sorts failures into knowledge gaps, capability limits, or evaluation noise, and only knowledge gaps guide revisions.
The authors also add a governance layer, which they say is needed because a single score can reject a bad revision without revealing its cause. The layer compares each revised Skill against both the original production version and the previous round. It rejects candidates that delete stable facts and flags knowledge bloat, broken cross-references, and vague restatements of specific values. One mechanism the paper describes shows the logic: if the simulated user never raises a key intent, the failure is attributed to the simulation rather than the Skill, and that sample is excluded from the agent’s score.
Across six cloud-service categories, 9 production Skills, and 98 reference files, the task success rate (TSR), measured on a held-out evaluation set, rose from 30.0 percent for the original Skills to 81.8 percent after four rounds. That is a gain of 51.8 percentage points. Compared with self-reflection, in which a model edits its own Skill without external feedback, the final SkillEvo score was 23.0 points higher, reaching 58.8 percent. Compared with single-turn QA-driven evolution, it was 15.4 points higher, reaching 66.4 percent. Removing governance lowered TSR to 78.6 percent. The authors report that cumulative knowledge bloat reached 2.8 percent with governance and 16.2 percent without it. The regression rate, the share of previously passing tickets that later failed, fell from 28.2 to 21.1 percent across rounds.
The study has clear limits. The ticket data cannot be released because of privacy and commercial-confidentiality constraints, so outside researchers cannot test the findings directly. The work comes from one company’s support operation. The simulator’s realism was judged by two experts on 200 dialogues, who found 95.3 percent agreement. The agent’s accuracy on intents already raised was 71.1 percent, a figure the authors describe as a measure of knowledge gaps rather than whole-conversation resolution. The tables report single values without confidence intervals. The authors also state that no revision reaches production without human confirmation.
The work suggests that the feedback loop, not only the editing step, may limit how far automated Skill maintenance can go. Whether the gains hold across other domains, longer evolution sequences, or different simulated users remains untested.
Chat about this paper
To chat about this paper, you'll need a free Gemini API key from Google AI Studio.
Your API key will be stored securely in your browser's local storage.