2026-10-02
AI for Software Development
AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines
Relevance: AgSpec directly targets coding agent pipelines, improving generation throughput through retrieval-based speculative decoding. It manages session, workspace, and global corpora while adapting draft lengths online from verification feedback. This is infrastructure for AI-assisted software development, relevant to code completion and generation efficiency. The method is evaluated on repository-level multi-agent coding benchmarks, showing up to 4.37x faster generation. It addresses the practical deployment of AI coding assistants and agents, fitting the AI for Software Development topic.
💡 Summary 📄 Full paper
Harness Engineering
Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions
Relevance: Prompt2Skill builds skills as external artifacts that LLMs consume at inference time, derived from natural language task descriptions alone. It refines these skills in a closed loop of reflective editing, without requiring a curated training set. Skills are key harness components, and this work addresses creating and iterating editable harness artifacts. The framework targets specialized domains and emerging tasks, optimizing skills for the specific consuming model. It is relevant to harness engineering by automating the construction and improvement of prompt/skill scaffolding around a foundation model.
💡 Summary 📄 Full paper
Human-in-the-Loop Evaluation
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
Relevance: ScholarCatalyst builds a benchmark by having 184 lead authors label which candidates did or could have advanced their completed projects, each with a detailed rationale. This is an expert annotation protocol for evaluating retrieval systems, where domain experts provide judgments on model outputs. The retrieval task uses author-provided judgments to measure agentic search and embedding retrieval. By keeping researchers meaningfully involved in assessing relevance and inspiration, the work contributes directly to human-in-the-loop evaluation methodology for scientific agents and retrieval models.
💡 Summary 📄 Full paper
Human-AI Collaboration
No paper recommendations for this topic.
Simulated Users
OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
Relevance: OpenTumorBoard provides a benchmark with specialist roles and a BOARD SIMULATION setting where an LLM generates an entire multidisciplinary discussion and reaches consensus. This uses LLM-based persona agents to simulate specialist discussants in a clinical decision-making setting. It evaluates model alignment with recorded board conclusions and expert-reviewed cases. While it simulates experts rather than end users, it is relevant to simulated people and persona agents for evaluating interactive, multi-role AI systems. It also enables model adaptation from real-world discussion trajectories.
💡 Summary 📄 Full paper