Mara Chain Turns Failed AI Edits Into Stepping Stones
Improving a deployed AI system increasingly means rewriting the prompts, scripts, and configuration files that guide it, rather than retraining the underlying model. A common method is a propose-evaluate-select loop: a language model suggests a modified version, an evaluator scores it, and only versions that clear an acceptance bar are kept. A new preprint from researchers at Ant Group and affiliated institutions argues that this filter discards too much. Rejected candidates, the authors write, often hold partial fixes or useful diagnoses, and without them later proposals keep rediscovering the same failures.
The authors call their approach Mara Chain. Instead of discarding a rejected candidate, the method keeps its rollout traces (records of each test run), its remaining failures, a structured analysis, and its history of changes. It then generates a sequence of descendants, each informed by the accumulated evidence. The chain runs to a fixed depth, five by default, and its best node is accepted only if it clears the same bar as the original candidate. If no node does, the chain is dropped. To keep the candidate pool from growing without limit, the method also removes candidates that other candidates dominate across validation tasks and retains only the three with the highest mean validation scores.
The paper’s clearest example involves a TerminalBench 2.1 task, sanitize-git-repo, which requires finding a hidden token in a repository. Without the chain, an early proposal contained a partial fix but failed the acceptance bar, so it was discarded. Later proposals began from different ideas, and the task remained failed. With the chain, the descendants moved from prose guidance to a script, then corrected a line-filtering step that hid the token, then matched tokens by length, and finally printed only the matched token. That lineage passed the task. The trace was run on GLM-5.
On AppWorld, which tests agent skills, Mara Chain reached a validation score of 0.8 after 3,280 rollouts, compared with 9,514 for GEPA, a reduction of 65.5 percent. ACE and SkillOpt-Lite never reached that score. On the Challenge test split, Mara Chain scored 76.7 on task-grade completion, against 74.1 for ACE and 69.3 for GEPA. On TerminalBench 2.1, an 89-task benchmark, the optimized harness reached a 71.9 percent pass rate, 20.2 percentage points above AHE and 22.5 points above Meta-Harness. On the MuSiQue retrieval benchmark, test nDCG@10 rose from 0.301 to 0.405 over a hand-written pipeline. Removing the chain lowered the TerminalBench pass rate from 71.9 to 57.3 percent in a single-slot ablation.
The authors acknowledge several limits. Each failed proposal can trigger up to five extra rollouts, so the method spends additional compute on rejected ideas. The authors also note that refinement works on one batch at a time and may overfit that batch’s failure modes, and that the diagnosis logic is hand-designed and fixed throughout the search. Gains were also uneven: on AppWorld’s easier Normal split, removing the chain left task-grade completion unchanged. Across models, average improvements over an empty skill were 40.3 points on GLM-5, 32.7 on DeepSeek-V4-Pro, and 23.4 on Qwen3.5-397B-A17B. The authors attribute the plateau near 0.87 on AppWorld partly to model capability and bounded depth.
The work suggests that the logs of failed attempts can be a useful resource in automated system improvement rather than waste. Its results, however, come from three benchmarks, one proposer agent, and comparisons that depend on each baseline’s configuration. They show that retaining failed candidates helped in these settings, not that the approach will generalize to all artifact-optimization problems.
Chat about this paper
To chat about this paper, you'll need a free Gemini API key from Google AI Studio.
Your API key will be stored securely in your browser's local storage.