Tencent Releases Open Benchmark Testing Coding Agents on Office, Web and Security Work
Coding agents are increasingly asked to do more than edit software. They may be asked to build a web page, reconcile a spreadsheet, or analyze a security artifact. Existing tests have not kept pace. SWE-bench and its verified variant draw heavily on public GitHub issues, so a model’s score may partly reflect exposure to those issues, and the tasks are mostly single-issue bug fixes. A team from Tencent and affiliated groups has released Tencent WorkBuddy Bench, a 260-task suite designed to address both problems.
The suite has four subsets: Code (80 repository-level software tasks), Web (70 front-end tasks), Office (50 file and data workflows), and Security (60 red-team and blue-team tasks). Each task is packaged in a uniform directory format and runs in an isolated container without internet access. The authors say each task was reverse-engineered from a real commit, pull request, or business scenario and rewritten as a short, informal request. The rewritten prompt omits the root cause and the reference fix, so it cannot be recovered by searching for the original issue. The authors are explicit that this does not make the suite contamination-free. Models may have seen the original public code, and any public release is exposed to later training. They plan to manage that exposure through dataset versioning.
Grading differs by subset. Code uses hidden tests, which are withheld from the agent while it works but published with the dataset. Web uses rubric items checked by rule scripts, by language or vision models, and by an agent that interacts with the running page. Office combines rule checks with an LLM judge that scores binary rubric items. Security uses a deterministic scoring script. Because the instruments differ, the authors state that scores should not be compared across subsets, and they report no suite-wide average.
The paper’s examples show what the tasks demand. One Code request, voiced by a product manager, asks the agent to compare conversion and revenue for a checkout experiment and to exclude purchases that happen long afterward. The request does not name the relevant file or the attribution window. The authors report that early runs often failed in two ways: agents looped on editing test files until they timed out, or they edited the wrong files in a large codebase. They attribute these failures to navigation and grounding rather than to code synthesis.
The paper’s leaderboard, shown in Figure 1, reports percentage scores for seven models under the benchmark’s harnesses. Claude Opus 4.8 scored 74.4% on Code and 82.4% on Office, the highest in each column, and 68.1% on Web, just ahead of HY-3 at 67.7% and GLM-5.2 at 67.4%. GLM-5.2 led Security at 76.3%. DeepSeek-V4-Pro scored 58.9% on Code, a gap of 15.5 percentage points from Claude Opus 4.8 on that subset. Kimi K2.7 has no Security score. The figure also includes an overall column, but the text says no suite-wide average is reported, and the paper does not explain how that column was computed, so this story does not rely on it. The report also does not describe repeated runs or uncertainty estimates, so differences of a few points should be read cautiously.
The work’s main contribution is a public, auditable benchmark that spans more kinds of professional work than bug-fixing suites typically do, along with a construction method meant to make prompts harder to find online. Its limits are also clear. The leaderboard is a single snapshot created by the benchmark’s authors, its subsets cannot be compared directly, and the authors’ claim of contamination resistance is limited to prompt construction and ongoing versioning.
Chat about this paper
To chat about this paper, you'll need a free Gemini API key from Google AI Studio.
Your API key will be stored securely in your browser's local storage.