Study of Real Coding Sessions Finds Agents Write Much Code, but Much of It Gets Thrown Away
AI coding agents, which autonomously read files, run terminal commands, and edit code in response to a developer’s requests, are being adopted rapidly. Yet researchers have had little systematic evidence on how developers use them, how often their output is kept, or where they go wrong. A new study from Stanford University, which introduces a dataset called SWE-chat, offers one of the first looks at those questions using sessions recorded from real users.
The dataset, described by Joachim Baumann and colleagues, contains nearly 18,000 sessions from more than 700 public GitHub repositories, with over 229,000 user prompts and 2 million agent tool calls. Developers opted in by installing Entire.io, an open-source tool that logs agent transcripts and links them to git commits, attributing each line of code to either the human or the agent. The authors note that their sample reflects early adopters who chose to log their work, so it may not represent all coding agent users. About 69% of the data comes from Claude Code.
The researchers sorted sessions into three modes: human-only, where the human wrote all committed code (25.0% of sessions); collaborative, where both contributed (33.9%); and “vibe coding,” a term for sessions in which the agent wrote more than 99% of committed code (41.1%). The authors report that the vibe-coding share more than doubled over their eight-month observation window, from under 20% of sessions in February to over 50% in recent months.
The study’s central efficiency finding is that only 59.4% of agent-produced code survived into user commits. Vibe-coded sessions kept a larger share, at 65.0%, compared with 55.9% for collaborative sessions. But vibe coding was more expensive per committed line: a median of $2.71 per 100 lines, versus $1.33 for human-only and $0.92 for collaborative sessions. The researchers also ran the static-analysis tool Semgrep on commits. Vibe-coded commits introduced 0.47 security findings per 1,000 lines, about 3.8 times the human-only rate of 0.12 and 3.2 times the collaborative rate of 0.15.
User behavior also stood out. Users pushed back after roughly 46% of agent turns, a figure that held across coding modes, and interrupted agents in a smaller share of turns. Agents asked for clarification in only about 3% of turns. Session success, rated by an automated judge on a 0–100 scale, averaged 73.2.
One low-rated session illustrates a failure mode the authors describe. In it, the agent made 42 tool calls searching for code it never found, and the session was scored poorly. The paper does not give the model behind this session.
The study has important caveats. Its cost and security comparisons are correlational, and the authors suggest the higher survival rate in vibe coding may reflect lower user scrutiny rather than better output. The security measure counts static-analysis findings, not confirmed exploits. Session labels came from large language model judges validated against human annotations with moderate-to-high agreement, and turn-level oversight figures come only from Claude Code.
Still, the authors argue the findings point to gaps that controlled benchmarks may miss, including the value of realistic evaluations, better interaction design, and user simulators trained on real sessions. Because the dataset is designed to keep growing, the researchers say it may also track how these patterns change over time.
Chat about this paper
To chat about this paper, you'll need a free Gemini API key from Google AI Studio.
Your API key will be stored securely in your browser's local storage.