AI Papers Reader

Personalized digests of latest AI research

View on GitHub

Benchmark Tests Whether Coding Agents Can Build Software From Vague Requests

Coding agents are moving beyond completing isolated functions toward building whole software projects, often from a loose description of what a user wants. A new benchmark from researchers at Fudan University, Meituan, and Singapore Management University, among others, aims to measure how well these systems handle that gap. The authors argue that most existing tests give agents fully specified tasks, which does not match “vibe coding” workflows, in which requirements emerge through conversation.

The benchmark, ICAE-Bench, contains 480 tasks across 12 programming languages, along with a 50-task subset called ICAE-Bench-Lite for faster experiments. Each task begins with a fuzzy product requirements document (PRD), a description that omits some details. The coding agent works in a container that includes the runtime environment but not the original source code or tests. It may ask a simulated User Agent up to 16 questions. That agent answers only from benchmark-authored records of hidden requirements, so it cannot invent new ones or leak implementation details. Tasks are drawn from real open-source repositories, and finished projects are scored with black-box tests that check outputs against given inputs, regardless of language.

The paper’s file-parser example shows the mechanism. The initial request asks for a parser without listing every supported format or edge case. Representative examples can be recovered by asking the right questions, but hidden tests also probe malformed and boundary inputs that the agent never sees. An agent can therefore reproduce visible behavior while failing the hidden cases, a pattern the authors describe as typical.

On the full benchmark, the best overall pass rate was 38.2%, achieved by Claude-Opus-4.8 under the Claude Code framework; GPT-5.5 scored 37.2%. Pass rates on public examples were consistently higher than on enhanced, harder cases. Giving models the complete requirements document, an upper-bound reference, produced the best results for four of six models on the Lite subset, and interaction recovered only part of that gap. Switching from Claude Code to OpenHands lowered overall scores for every model, by 5.5 to 21.8 percentage points. GPT-5.5 fell from 53.3% to 31.5%, while Claude-Opus-4.8 fell from 48.2% to 42.7%. In a GLM-5.1 experiment, placing ready-to-run public test files in the workspace raised the overall pass rate from 37.4% to 61.8%, though constraint coverage slightly declined. The authors attribute the gain to making information executable rather than adding requirements.

The findings carry several caveats. Recovering more constraints did not reliably improve correctness. A richer execution environment lowered GLM-5.1’s pass rate from 37.4% to 28.4% while its repositories grew to 674.1% of the reference code’s lines of code. Model rankings on the Lite subset correlated only moderately with the full benchmark (Spearman’s ρ = 0.71), so the authors caution that Lite supports ablations but not definitive rankings. Critic-model scores for semantic similarity, API similarity, and design quality agreed only moderately with human raters, with Pearson correlations between 0.37 and 0.47. Results also depend on the User Agent’s underlying model.

The work’s significance lies in reframing evaluation around clarification, requirement retention, and verification rather than static completion. The authors conclude that the main bottleneck is not asking questions but carrying clarified requirements into coherent implementations. Because the results come from one benchmark and a fixed set of models and frameworks, they suggest a difficulty pattern rather than a settled ranking of coding agents.