New Framework Turns Plain-Language Task Descriptions Into Optimized LLM Skills
Researchers at Vanderbilt University, the University of Southern California, and Adobe have proposed Prompt2Skill, a multi-agent framework that builds task-specific “skills” for large language models from a natural-language description alone. The problem it targets is practical: skills are external text artifacts placed in a model’s context at inference time to supply procedures and domain knowledge, but expert-authored skills are expensive, and existing automated skill optimizers typically require a curated, labeled training set that users may not have. Prompt2Skill instead asks whether a bare request such as “I want a model that reads a passage and answers questions about it” can be enough.
The framework splits the work between two agents. A Data Agent turns the prompt into a task specification, including input–output contract and answer format, then retrieves candidate datasets from sources such as Hugging Face and Wikipedia. It adapts retrieved records into task examples or synthesizes new ones when needed, validating them before use. A Skill Agent starts from a seed skill and refines it in a closed loop: the target model runs on a reflection set, a reflector model proposes edits based on failures, and each candidate edit is accepted only if it meets a paired statistical screen on a fresh acceptance batch. A frozen validation set is consulted once for final selection. The result is a SKILL.md artifact supplied at inference; the target model’s weights are not updated.
The paper gives a concrete example of why model-specific optimization matters. An off-the-shelf skill scored 0.048 on Llama-3.2-1B-Instruct’s SQuAD exact-match benchmark, down from 0.207 for direct prompting. Prompt2Skill, by contrast, raised that same score to 0.492.
Across four benchmarks—SearchQA, SQuAD, AIME, and SpreadsheetBench—and five target models—Qwen3-8B, Qwen3-32B, Llama-3.2-1B-Instruct, Claude Haiku 4.5, and GPT-5.5—the paper reports the highest average relative improvement for every target model. The per-model average relative gains were 29.1% for Qwen3-8B, 26.1% for Qwen3-32B, 22.0% for Claude Haiku 4.5, and 24.3% for GPT-5.5. Llama-3.2-1B-Instruct, which could not complete any SpreadsheetBench task, showed a 93.5% average relative gain over its three reported text benchmarks. Across 19 model–benchmark pairs, the authors calculate average relative improvements of 36.1% over direct prompting and 83.0% over off-the-shelf skills. These are averages of relative changes, not pooled accuracy or percentage-point gains. In a matched-budget comparison with SkillOpt using Qwen3-8B as solver and GPT-5.5 as editor, Prompt2Skill improved SearchQA exact match by 2.93 percentage points and SQuAD exact match/F1 by 1.20/0.61 points.
The authors also note limitations. Optimization still runs on proxy data and a proxy metric, which may not match the user’s true task distribution or evaluation criterion; the paper describes distribution and evaluator mismatch as separate sources of error. The acceptance rule uses a per-candidate screening threshold rather than a conventional 5% significance test and does not adjust for multiple comparisons across candidates or rounds. Llama’s missing SpreadsheetBench result also limits that model’s aggregate. The broader promise is that users could specify specialized tasks in plain language and obtain reusable inference-time skills without fine-tuning, but the evidence comes from four benchmarks and automatically acquired proxy data, so it supports a promising direction rather than general deployment.
Chat about this paper
To chat about this paper, you'll need a free Gemini API key from Google AI Studio.
Your API key will be stored securely in your browser's local storage.