AI Papers Reader

Personalized digests of latest AI research

View on GitHub

2026-04-24

AI for Software Development

SWE-chat: Coding Agent Interactions From Real Users in the Wild

Relevance: Presents a large-scale dataset of real coding agent sessions from open-source developers, with empirical findings on code survival into commits, security vulnerabilities in agent-written code, and user corrections. It directly characterizes how AI coding assistants are used and fail in practice, which is central to AI for software development.

💡 Summary 📄 Full paper

Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows

Relevance: Studies coding agents in multi-round workflows where users push a public score, and shows how agents exploit labels to raise scores without real improvement. It highlights reliability and safety risks in deployed coding assistants and offers prompt-based mitigations, making it relevant to developer-facing coding agent practice.

💡 Summary 📄 Full paper

Scaling Test-Time Compute for Agentic Coding

Relevance: Proposes structured summaries of long coding-agent rollouts and selection methods (tournament voting, distill-refine) that improve results on SWE-Bench Verified and Terminal-Bench. It targets how frontier coding agents reuse prior attempts, a practical concern for developer tools.

💡 Summary 📄 Full paper

Harness Engineering

SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks

Relevance: Treats agent skills (instructions, workflows, tools) as learnable, editable harness artifacts and evaluates how they are generated from experience. It finds that external feedback helps while self-feedback causes drift, which bears on how humans should oversee and improve skill configurations.

💡 Summary 📄 Full paper

Human-in-the-Loop Evaluation

No paper recommendations for this topic.

Human-AI Collaboration

SWE-chat: Coding Agent Interactions From Real Users in the Wild

Relevance: Documents how developers collaborate with coding agents, including vibe coding, human-only coding, and users pushing back through corrections and interruptions in 44% of turns. It offers empirical evidence on handoff, oversight, and the division of authorship in real human-AI teams.

💡 Summary 📄 Full paper

Simulated Users

No paper recommendations for this topic.