AI Papers Reader

Personalized digests of latest AI research

View on GitHub

Repo0 Lets AI Agents Revise a Software Architecture as They Build It

Large language model agents can now write useful code snippets and resolve bugs in existing projects, but building a complete software project from a plain-language description remains much harder. Researchers from Shanghai Jiao Tong University, Chongqing University, and collaborators argue that a central obstacle is architecture. Most existing systems draft a repository’s structure once, at the outset, and then treat that plan as fixed. The authors contend that a project’s modularity, meaning how cleanly its components divide responsibilities, often becomes apparent only after code is written, so an early blueprint can lock in weak boundaries.

Their system, called Repo0, maintains an evolving architectural record. The record, which the authors call a Dual-Directed-Acyclic-Graph (Dual-DAG), has three parts: a requirement-level graph linking functional requirements that should be considered together, a component-level graph of implementation modules and their dependencies, and an alignment table that ties each requirement to the components that realize it. Starting from a requirements document, a language model extracts high-level requirements, breaks them into sub-requirements, and proposes initial components. The system then repeatedly applies structural actions, including splitting a component with loosely related responsibilities, merging components whose responsibilities overlap heavily, and revising component descriptions. Two measures trigger these actions. Cohesion measures how densely a component’s requirements are functionally linked, and coupling is the overlap between two components’ requirement sets. The loop stops when a full round yields no eligible split or merge. Only then does code generation begin, using test-driven development.

The paper offers a concrete example from a generated statistics library. A component called Experimental Namespace Manager was overloaded, so the system split it into two components, a Namespace & Registry Manager and an Access Control & Compatibility Manager, each mapped to its own file. The authors present this as an illustration of the mechanism rather than a measured outcome.

The authors evaluated Repo0 on RepoCraft, a benchmark of six Python repositories with paraphrased names. In the main text they report three of them, using GPT-5 mini and DeepSeek V3.2. Against RPG, a graph-based planning baseline, Repo0 posted higher Functionality Coverage, the share of reference functional categories matched, and higher Pass Rate, the share of tasks passing adapted ground-truth tests. With GPT-5 mini on django, coverage rose from 60.42% to 80.50%, a gain of 20.08 percentage points, and pass rate rose from 47.33% to 74.36%, a gain of 27.03 percentage points. With DeepSeek V3.2 on statsmodels, pass rate rose from 39.29% to 69.03%, a 29.74-point gain. Percentage-point differences like these are not relative percentages; the relative improvements are larger on lower-baseline measures.

In an ablation study on requests, removing the structural-evolution loop caused the largest drop among the components tested, cutting coverage by 5.68 points and pass rate by 8.47 points. The authors also compared the metric-guided loop with versions in which a language model chose structural actions for one to five fixed rounds. On statsmodels with GPT-5 mini, the metric-guided version scored higher, which the authors attribute to explicit convergence criteria preventing over-fragmentation.

Cost is a tradeoff. Under DeepSeek V3.2, Repo0’s total costs were $21.28 for requests, $41.12 for statsmodels, and $100.06 for django, with evaluation dominating the django total. The authors report lower end-to-end cost than RPG in most settings but note exceptions.

Several limitations remain. The study covers one language and one benchmark, and the authors’ thresholds were tuned on two held-out repositories and checked against golden architectures by hand. Because the language model still rewrites boundaries and descriptions, results depend on its reasoning. The work suggests that treating architecture as a revisable state can improve generated repositories, but it does not establish how the approach performs on other languages or on projects with human-level complexity.