Teaching AI Agents to Write Their Own Skills Helps, but Gains Are Uneven
AI agents increasingly rely on “skills,” structured documents that give them step-by-step instructions, workflows, and domain knowledge for specialized tasks. Prebuilt skills cannot anticipate every new job, so researchers are asking whether agents can generate their own skills from experience. A new study from Carnegie Mellon University and Amazon AGI introduces SkillLearnBench, a benchmark for testing these methods, and finds that they help on average but unevenly.
The benchmark includes 20 verified tasks across 15 sub-domains, drawn from a community taxonomy of real-world skill use. Each task is designed so that an agent without skills usually fails, but succeeds when given human-written reference skills. Each task also has multiple instances with varied inputs, which tests whether a generated skill transfers beyond the case it was created from. The researchers evaluate skills at three levels: the quality of the skill text, the agent’s execution trajectory, and the final task outcome.
The study compares four methods for generating skills from a single seed instance. One-Shot writes a skill in one pass. Self Feedback has the agent attempt the task, review its own trajectory, and revise. Teacher Feedback lets an expert model, which can see the human-authored skills, answer the agent’s questions after failed attempts without revealing the answer. Skill Creator follows a multi-stage pipeline of analysis, edge-case investigation, writing, and validation. Each method ran on six models from the Claude and Gemini families, and a fixed solver, Claude Sonnet 4.6, then used the resulting skills.
The clearest illustration of the mechanism involves Self Feedback on Productivity Tools tasks with Claude Sonnet 4.6, tracked over four rounds. Coverage of the key steps stayed flat, alignment with the reference trajectory steadily declined, and accuracy rose after the first revision before falling sharply. The authors attribute this drift to self-revision reshuffling content without new information. Teacher Feedback, by contrast, substantially raised coverage in its first round, and accuracy grew through round four.
Averaged across the six models, accuracy was 10.17% with no skill and 74.50% with human-authored skills. The four methods scored 30.44% (One-Shot), 31.08% (Self Feedback), 27.47% (Teacher Feedback), and 27.33% (Skill Creator). Each beat the no-skill baseline by roughly 17 to 21 percentage points, but Self Feedback’s edge over One-Shot was only 0.64 percentage points. The best single configuration, One-Shot with Claude Sonnet 4.6, reached 38.83%, which the authors say covers about 45% of the gap between no-skill and human-authored performance. Results also varied by domain: all four methods scored below the no-skill baseline of 16.67% on Utilities tasks. Stronger models did not reliably produce better skills.
The study has clear limits. It contains 20 tasks and 100 instances, uses a single solving agent, and relies on an LLM judge (GPT-5-mini) for the skill-quality and trajectory scores. Skill Creator was the most frequently used method, at 84.47% usage, yet it trailed on accuracy, which the authors read as evidence that adoption alone is not enough. Iterative methods also cost more computation, though Self Feedback used the fewest tokens on average.
The authors conclude that skill generation will need to ground skills in core task logic and ensure agents reliably follow them. The open-sourced benchmark gives researchers a shared way to measure progress, though its small size means the findings describe these tasks rather than agent learning in general.
Chat about this paper
To chat about this paper, you'll need a free Gemini API key from Google AI Studio.
Your API key will be stored securely in your browser's local storage.