AI Papers Reader

Personalized digests of latest AI research

View on GitHub

Coding Agents Gain From Smarter Use of Their Own Past Attempts

Large language models often improve when they are allowed to spend more computation at inference time, such as by generating several answers and picking the best one. That approach works well when each answer is short enough to compare directly. Autonomous coding agents, which read files, edit code, run commands, and react to errors, do not fit this pattern. Each attempt, which the authors call a rollout, can span dozens of steps and produce a long trajectory that is difficult to compare or reuse. A team from Meta Superintelligence Labs and several universities argues that the central obstacle is representation: how prior experience is condensed into a form that later computation can use.

The researchers convert each rollout into a compact structured summary that keeps its hypotheses, progress, and failure modes while discarding repetitive terminal output and dead-end exploration. These summaries support two forms of scaling. For parallel scaling, they introduce Recursive Tournament Voting (RTV). Summaries are compared in small groups of two, with a language model acting as judge casting multiple votes per comparison, and winners advance until a single rollout remains. Selection does not use test results. For sequential scaling, they adapt Parallel-Distill-Refine (PDR), in which new rollouts are conditioned on summaries from earlier attempts. Their combined pipeline runs 16 rollouts, uses RTV to choose the four strongest summaries, runs 16 fresh rollouts conditioned on those four, and applies RTV once more to select a final answer.

The paper gives a concrete illustration of how much the quality of context matters. On SWE-Bench Verified, a benchmark built from real GitHub issues, Claude-4.5-Opus’s second-round rollouts averaged 0.1% success when all four context summaries came from failed attempts. When all four came from successful attempts, the average rose to 99.2%.

On SWE-Bench Verified, using the mini-SWE-agent scaffold, Claude-4.5-Opus’s average pass@1 (the share of tasks solved on a single attempt, averaged across rollouts) rose from 70.94% to 77.60%, a gain of 6.66 percentage points. On Terminal-Bench v2.0, using the Terminus 1 scaffold, the same model rose from 46.95% to 59.09%, a gain of 12.14 percentage points, or about 26% in relative terms. Gemini-3.1-Pro improved by 4.35 points on SWE-Bench Verified and 12.28 points on Terminal-Bench. The authors report consistent gains across five frontier models. In a 100-task ablation on SWE-Bench Verified, Gemini-3.1-Pro’s second-round average reached 73.75% when refined from a single prior summary, 76.94% when refined from four randomly chosen summaries, and 79.25% when refined from four summaries selected by RTV. The authors attribute the ordering to the quality of the context.

The approach has costs and open questions. Each task involves two sets of 16 rollouts plus several rounds of judging, and the paper does not report dollar or wall-clock costs. Judge accuracy also varied. Gemini-3.1-Pro was the least accurate judge, and the authors note that API failures during its final experiments may have contributed, which left its final-stage gain at just 0.44 points. The authors also caution that their judge-accuracy comparisons across models are not controlled. Terminal-Bench results cover 88 of its 89 tasks.

Still, the authors say that the second-round rollouts needed roughly half as many steps, and that their results suggest that test-time scaling for long-horizon agents depends on how past experience is represented, selected, and reused. The findings come from the authors’ own experiments, and independent replication would strengthen them.