AI Papers Reader

Personalized digests of latest AI research

View on GitHub

2026-07-24

AI for Software Development

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

Relevance: Presents a multi-domain benchmark for coding agents spanning repository-level engineering, front-end development, office workflows, and security. Tasks are reverse-engineered from real commits and pull requests, and the suite runs on two agent harnesses (CodeBuddy Code and Claude Code). It directly evaluates AI-assisted software development and is useful for understanding how coding agents perform on realistic developer work.

💡 Summary 📄 Full paper

ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders

Relevance: Evaluates coding agents that turn fuzzy product requirements into working software through clarification, planning, and debugging. It grounds ambiguity in real open-source repositories and uses multi-dimensional diagnostics, including functional correctness and design quality. It targets the vibe-coding workflow that is central to modern AI-assisted software development.

💡 Summary 📄 Full paper

NVIDIA-labs OO Agents: Native Python Object-Oriented Agents

Relevance: Proposes an agent framework in which an agent is a Python object, with methods as actions, fields as state, and docstrings as prompts. Because agent code is testable, traceable, and refactorable like ordinary software, it fits developer workflows. It is evaluated on SWE-bench Verified, a software-engineering benchmark.

💡 Summary 📄 Full paper

Harness Engineering

AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

Relevance: Directly addresses observability and debugging of the agent harness. It provides a Detect-Attribute-Recover-Rerun loop, trajectory-based root-cause diagnosis, a Python library, CLI, web console, and a shareable Error Hub. This matches the emphasis on inspecting why an agent acted as it did and on making failures actionable for operators.

💡 Summary 📄 Full paper

DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines

Relevance: Builds a platform where an agent constructs editable pipeline artifacts, with DataFlow-Skills as procedural guidance and a WebUI that syncs conversational authoring with a visual DAG editor. The persistent, editable, inspectable workflow is a clear example of a harness that humans can compose, steer, and version.

💡 Summary 📄 Full paper

NVIDIA-labs OO Agents: Native Python Object-Oriented Agents

Relevance: Makes agent behavior testable, traceable, and refactorable through a programming model in which prompts, state, and actions live in one object. It also exposes model-callable harness APIs for context and events, which supports editable and inspectable harness components.

💡 Summary 📄 Full paper

Human-in-the-Loop Evaluation

EduPanel: A Three-Agent LLM Judge for Teaching Videos – Reliability, Complementarity, and Human Trust Calibration

Relevance: Centers evaluation methodology on expert studies that compare an LLM judge’s reliability with human experts. Expert feedback improves scoring accuracy, and experts can detect unreliable outputs, which reflects a hybrid pipeline in which humans calibrate and audit an automated judge.

💡 Summary 📄 Full paper

Human-AI Collaboration

No paper recommendations for this topic.

Simulated Users

ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders

Relevance: Uses an automated User Agent to simulate a user who gradually reveals hidden constraints from a fuzzy requirement. The paper explicitly designs the simulator to be reproducible and to avoid inventing new requirements or leaking implementation details, so it addresses simulator fidelity and reliability in interactive evaluation.

💡 Summary 📄 Full paper

EduPanel: A Three-Agent LLM Judge for Teaching Videos – Reliability, Complementarity, and Human Trust Calibration

Relevance: Conditions its evaluation on learner personas and includes learner-persona analyses. This is a persona-based stand-in for different learners, and the paper also examines how well such simulated assessments match human experts.

💡 Summary 📄 Full paper