AI Papers Reader

Personalized digests of latest AI research

View on GitHub

Toolkit Aims to Turn LLM Agent Failures Into Tested Repairs

Large language model (LLM) agents, which plan, call external tools, and hand work to other agents, are being used for longer and more complex tasks. When they fail, the visible error often appears many steps after the decision that caused it, and current tracing tools mainly show what happened rather than why. A team led by researchers at the University of Illinois Urbana-Champaign has released AgentDebugX, an open-source toolkit that links failure detection, root-cause attribution, suggested repair, and rerunning into one loop.

The central idea is that a diagnosis matters only if it improves the next run. AgentDebugX converts logs from frameworks such as LangGraph and CrewAI into a common trajectory format, so any diagnostic method can analyze the same execution. Detection flags failures using rule checks and an LLM judge. Attribution then traces each symptom backward to the agent and step most likely responsible. Recovery proposes a fix, and the rerun is saved beside the original so the two can be compared. Proposed fixes are advisory and require human or policy approval before use.

At the center of the system is DeepDebug, a multi-turn diagnostic agent. It reads the full trace, follows handoffs backward in multi-agent runs or bisects the step range in single-agent runs, cross-examines any conflicting candidate steps, and produces a report naming the responsible agent and step, with supporting evidence and one suggested fix. The authors illustrate the difficulty with a common pattern: an agent returns a wrong final answer because it dropped a planning constraint much earlier, and the error surfaces only after several reasonable-looking actions.

On the Who&When benchmark of 184 failed traces, using the qwen3.5-9b model, DeepDebug identified both the responsible agent and the exact mistake step in 28.8% of cases. The strongest single-pass baseline, All-at-Once, reached 21.7%. Responsible-agent accuracy rose from 47.8% to 56.0%. The advantage was model-dependent: with qwen3.6-27b, strict accuracy was 38.0% against 36.4%, and on hosted models the authors found that a single global reading performed as well or better. The gains were concentrated in traces longer than 40 events, but that subgroup contained only 26 traces, which the authors treat as descriptive evidence.

The recovery experiment used GAIA, a benchmark of general assistant tasks. A vanilla qwen3.5-9b research agent solved 55.8% of validation tasks and failed 73. After one rerun guided by DeepDebug’s diagnosis, 13 of those 73 failures were repaired. Three decoupled self-correction baselines, which received the failure context and a generic summary but not DeepDebug’s localization, repaired between 4 and 6. Overall accuracy rose from 55.8% to 63.6%, a gain of 7.8 percentage points.

The costs and caveats are substantial. On a 25-trace sample, DeepDebug used about 12.8K tokens per trace compared with 8.1K for a single whole-trace pass, roughly 1.6 times as much. The GAIA comparison tests the full recipe rather than attribution alone, and the authors note that the benefit of the shared error corpus, which can serve as debugging memory, has not yet been evaluated. Redaction of sensitive data relies partly on pattern matching, which the authors acknowledge cannot guarantee removal, and they warn that diagnostic labels can be wrong.

The work suggests that better fault attribution can translate into measurable gains in repair, though the evidence comes from two benchmarks and a small set of configurations. Because the toolkit is open source and framework-independent, it may help teams compare diagnostic methods on shared evidence. Whether it generalizes beyond these settings remains an open question.