AI Papers Reader

Personalized digests of latest AI research

View on GitHub

2025-10-24

AI for Software Development

When β€œCorrect” Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?

Relevance: Directly examines code agents (SWE-agent, OpenHands, ChatGPT, Claude) that automatically fix bugs in real repositories. It shows that patches can pass all tests yet contain vulnerabilities, exposing a gap in current evaluation of AI-generated code. This is a central concern for developer-facing AI tools that propose fixes and code changes.

πŸ’‘ Summary πŸ“„ Full paper

From Charts to Code: A Hierarchical Benchmark for Multimodal Models

Relevance: Benchmarks LMMs on generating plotting code from charts and tables, including reproduction, editing, and generation from long tables. It targets a practical code-generation workflow in which users give instructions and models produce executable code, and it evaluates both code correctness and rendered-output fidelity.

πŸ’‘ Summary πŸ“„ Full paper

Harness Engineering

No paper recommendations for this topic.

Human-in-the-Loop Evaluation

ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge

Relevance: Uses over 7,000 response-criterion pairs judged by human experts with PhD and MBA backgrounds as the ground truth. It then builds LLM judges calibrated to those human judgments, which is a hybrid pipeline in which humans anchor and audit automated evaluation in a high-stakes professional domain.

πŸ’‘ Summary πŸ“„ Full paper

Human-AI Collaboration

No paper recommendations for this topic.

Simulated Users

No paper recommendations for this topic.