2026-09-25
AI for Software Development
Schrรถdingerโs Code Repository: Have LLMs Learned SWE-bench or Memorized It?
Relevance: This paper directly examines repository-level coding agents and the SWE-bench benchmark, asking whether strong performance reflects robust software engineering reasoning or memorization of canonical repository cues. It proposes dynamically instantiated repository representations that preserve executable behavior while changing naming, layout, and implementation patterns. This is central to AI for software development because it evaluates coding agents under realistic generalization conditions and highlights exploration and localization costs. It also has HCI implications for how developers interpret coding-agent evaluations, inspect failures, and design tools that support reliable repository reasoning.
๐ก Summary ๐ Full paper
Harness Engineering
Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents
Relevance: This paper explicitly centers model-harness co-evolution for real-world mobile planner agents. It orchestrates memory, skills, and tools at runtime, then feeds structured action feedback and preserved failure traces back into coordinated model and harness adaptation. It also uses a human-gated data flywheel. This is core harness engineering: designing, composing, and iterating the operational scaffold around a foundation model. It is relevant to HCI because it addresses how teams inspect execution evidence, steer agent improvements, and maintain reliable agent behavior across data production, training, and deployment loops.
๐ก Summary ๐ Full paper
World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
Relevance: World Action Agent presents a multi-agent harness through which vision-language models pilot robots using basic tools. The visual action workspace, editable action rehearsal, Imagination Agent, Skill Agent, and evidence-based review make the harness itself a central contribution. It supports human teaching and converts interaction traces into training data for smaller models. This is highly relevant to harness engineering because it addresses composing tools, previewing and revising actions, acquiring skills, and improving the agent scaffold around a foundation model. It also has HCI implications for inspecting and correcting agent behavior.
๐ก Summary ๐ Full paper
Human-in-the-Loop Evaluation
StudentBench: AI and human tutoring yield equivalent GRE learning gains
Relevance: StudentBench is a human-in-the-loop evaluation platform and study that measures learning gains across 2,383 participants receiving AI tutoring, human tutoring, or no tutoring. It also collects 2,028 pairwise rubric evaluations from expert human tutors comparing LLM-generated lesson plans and practice problems. The paper directly advances evaluation methodology involving humans by combining large-scale participant outcomes with expert annotation and preference judgments. It shows how human evaluation can separate AI tutors across pedagogical dimensions and compare AI tutoring against expert human tutoring in a controlled, high-stakes learning setting.
๐ก Summary ๐ Full paper
HappyWorld-Bench
Relevance: HappyWorld-Bench is a comprehensive benchmark for world models that evaluates generative construction, consistency, responsiveness, and modification under interaction. It builds HappyWorld-Arena to organize human A/B comparisons and derive model-level Elo ratings across video, spatial, and embodied world models. These human judgments complement automated behavioral metrics. The core contribution is evaluation methodology for interactive AI systems, with humans assessing reliability as agents explore and modify generated worlds. It is highly relevant to human-in-the-loop evaluation because it makes human preference and comparative judgment central to benchmarking world models.
๐ก Summary ๐ Full paper
Human-AI Collaboration
PUBG Ally: A Conversational Embodied Agent as an AI Teammate
Relevance: PUBG Ally is an embodied AI teammate that reasons, acts, and speaks alongside human players in PUBG: BATTLEGROUNDS. It combines agentic tool use with real-time game control, maintaining context, deciding what to say, and issuing high-level actions that steer movement and combat. The paper studies teammate quality through player feedback, preference comparisons, and live-service surveys. This is directly about human-AI collaboration: shared control, role allocation, communication, trust, and failure recovery in a demanding real-time team task. It also examines how players perceive the AI as a teammate or companion rather than only a tool.
๐ก Summary ๐ Full paper
Knowledge Pull Requests for Continual Document Authoring
Relevance: Knowledge Pull Requests introduce a framework for continual document authoring that integrates new knowledge through claim extraction, routing, conflict flagging, and a ChangeLog separating knowledge changes from text diffs. This supports human-AI co-authoring by making each AI-driven revision interpretable and reviewable. It is relevant to human-AI collaboration through mixed-initiative writing, transparency, and preserving human agency in document revision workflows. The framework aims to help people understand what knowledge changed and how the text changed, enabling collaborative authoring where AI proposes updates but humans can inspect, accept, or revise them.
๐ก Summary ๐ Full paper
Simulated Users
No paper recommendations for this topic.