AI Papers Reader

Personalized digests of latest AI research

View on GitHub

Speeding Up Coding Agents by Reusing Their Own Trail

Coding agents rarely solve repository-level tasks in one pass. They work through multi-turn sessions, with specialized agents editing files, running tests, and revising failed attempts. Each extra turn can improve the answer but lengthens the wait. Speculative decoding can cut that wait without changing output: a drafter proposes tokens, and the target model verifies them in parallel, committing accepted ones. Retrieval-based methods draft by copying continuations from existing text. But according to a new preprint from KAIST, existing retrieval drafters overlook much of what coding agents reuse and choose draft lengths poorly.

AGSPEC, the authors’ framework, leaves the retrieval engine unchanged and supplies policies around it. It builds three corpora by source: session (prompts, tool outputs, generated tokens), workspace (repository artifacts opened during the session), and a fixed global reference corpus. Crucially, workspace files are indexed both as they are and in the form the agent emits them. Under unified-diff editing, for example, a coder copying a line into a patch adds a space or minus prefix absent from the original file. Without those emission-form copies, retrieved file text may not match the generated patch.

AGSPEC also controls draft length. It profiles each agent’s typical accepted length offline to set caps, then adapts a scale online using verification feedback. Global-corpus matches must beat live session/workspace matches by a five-token margin.

On SWE-bench Verified and TeamBench, the authors tested AGSPEC with Devstral-24B, Gemma3-27B, and Qwen3.6-27B against autoregressive decoding, five retrieval-based drafters, and EAGLE-3. AGSPEC achieved the highest or second-highest throughput in all evaluated settings, reaching 2.27–4.37× autoregressive throughput at batch size 1 and 1.08–4.76× at batch size 16. Its throughput averaged 18.0% higher than the fastest prior method. With SAM-Decoding as the retrieval engine, AGSPEC raised Gemma3-27B’s SWE-bench speedup over autoregressive decoding from 2.72× to 3.61× at batch size 16.

On draft length, a fixed 32-token cap was 14% faster than 8 at batch size 1 but 30% slower at batch size 16, because rejected drafts consume shared verification compute. AGSPEC was fastest at both sizes, 25% and 16% above the 8-token cap. In one coder call, enabling the online scale cut rejected tokens by 81.1% while adding only 11 decoding steps. AGSPEC also improved throughput on LiveCodeBench, which lacks a repository, and Terminal-Bench, a single-agent setting.

The approach is not free. The paper reports a global corpus using 2.5–4.3 GB of host memory with 9–15 seconds of load time, and retrieval adding 1.1–5.5% of step time at batch size 16. Session and workspace corpora may hold sensitive code, prompts, or tool outputs; deployments should isolate and discard them per session. The authors also found no meaningful benefit from letting workspace data grow across sessions, leaving many sessions on the same repository for future work.

The results suggest that speculative decoding for coding agents should track the agent pipeline—what text sessions produce and how agents emit it—rather than treating each request in isolation. They do not establish that AGSPEC helps every coding workload; gains varied by model, engine, agent, and turn.