AI Papers Reader

Personalized digests of latest AI research

View on GitHub

Coding Agents Struggle With Scientific Software, New Benchmark Finds

Scientific research increasingly runs on code. Simulations, instrument-data pipelines, and screening tools for chemical or biological candidates all depend on software, so a faulty patch can corrupt more than a programโ€™s output. It can also undermine the evidence behind a conclusion. Most evaluations of AI coding agents report a single aggregate success rate, which says little about why agents fail on scientific code. Researchers at Shanghai Innovation Institute and Fudan University have introduced SWE-bench Science, a benchmark designed to examine those failures.

The benchmark contains 119 tasks drawn from 98 GitHub repositories across 20 scientific domains, with chemistry the largest at 24 tasks. The tasks fall into three paradigms. Issue-driven tasks ask an agent to repair a known defect (52 tasks). Expert-exploratory tasks require the agent to identify an unknown root cause from an observed scientific anomaly (49 tasks). Engineering-integration tasks require connecting several modules into a complete workflow (18 tasks). Agents work in a snapshot of the repository with a problem statement and a public test they can run while debugging. Their submitted patch is then checked against private tests they cannot see. The headline metric, Pass@1, counts a task as solved only if every private test passes.

The paper illustrates a common failure with a task about a crystal modeled two ways, using a primitive cell with sampled k-points and a larger supercell, whose energies should agree per primitive cell. In one example, a patch passed the single public check but failed seven of ten private scientific cases, including tests of FFT mesh parity and k-point ordering. It therefore scored zero on Pass@1. The authors describe this as an incomplete repair that matched the visible symptom without restoring the underlying physics.

The best-performing configuration, Claude Code running Claude-Opus-5 at maximum reasoning depth, reached 47.90% Pass@1, while its public score was 96.64%. Other results were DeepSeek-V4-Pro at 42.02%, GPT-5.6-sol at 40.34%, Kimi-K3 at 35.29%, and Qwen3.5-397B at 14.29%, the lowest of the eight configurations tested. The gap between public and private scores is the central point of the paper: visible tests alone overstate correctness.

The authors sorted unsuccessful attempts into four mechanisms: deficits in scientific knowledge, misguided exploration or surface-level repair, incomplete system integration, and failure to generalize beyond observed cases. The distribution varied by model. Claude-Opus-5 had only two surface-repair errors, while DeepSeek-V4-flash had 48 integration errors and only six generalization errors.

To test the role of scientific guidance, the team removed explanations such as scientific rationales and expected properties from 91 tasks, keeping the code, symptoms, and tests. For GPT-5.6-sol, Pass@1 fell from 36.26% to 31.87%, a drop of 4.39 percentage points. For DeepSeek-V4-flash, it rose from 16.48% to 23.08%, a gain of 6.60 points. The authors suggest that well-grounded guidance can constrain a repair in useful ways, while poorly aligned guidance can encourage anchoring on a supplied explanation. They also note that the weaker model appeared to benefit more, though that interpretation rests on only two models.

These findings carry real caveats. The ablation covers two models, and the authors report no statistical significance and no established causal effect. Each scientific domain contains few tasks, which limits cross-domain comparisons. The authors also acknowledge that their analysis of how agents use domain knowledge remains preliminary.

Still, the benchmark gives researchers a structured way to separate visible-test success from complete correctness. Its results suggest that giving agents more scientific information does not reliably produce better repairs, and that agents must connect such information to executable evidence. SWE-bench Science offers a testbed for that question rather than a settled answer to it.