2025-10-24
AI for Software Development
When βCorrectβ Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?
Relevance: Directly examines code agents (SWE-agent, OpenHands, ChatGPT, Claude) that automatically fix bugs in real repositories. It shows that patches can pass all tests yet contain vulnerabilities, exposing a gap in current evaluation of AI-generated code. This is a central concern for developer-facing AI tools that propose fixes and code changes.
π‘ Summary π Full paper
From Charts to Code: A Hierarchical Benchmark for Multimodal Models
Relevance: Benchmarks LMMs on generating plotting code from charts and tables, including reproduction, editing, and generation from long tables. It targets a practical code-generation workflow in which users give instructions and models produce executable code, and it evaluates both code correctness and rendered-output fidelity.
π‘ Summary π Full paper
Harness Engineering
No paper recommendations for this topic.
Human-in-the-Loop Evaluation
ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
Relevance: Uses over 7,000 response-criterion pairs judged by human experts with PhD and MBA backgrounds as the ground truth. It then builds LLM judges calibrated to those human judgments, which is a hybrid pipeline in which humans anchor and audit automated evaluation in a high-stakes professional domain.
π‘ Summary π Full paper
Human-AI Collaboration
No paper recommendations for this topic.
Simulated Users
No paper recommendations for this topic.