Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

Hugging Face Daily Papers Papers

Summary

This paper introduces Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents designed to resist data contamination by reverse-engineering tasks from real commits and business scenarios, covering Code, Web, Office, and Security domains.

We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running distribution-informed coding-agent tasks across four work domains - Code, Web, Office, and Security. Rather than adapting public issue text, every task is reverse-engineered from a real commit, pull request, or business scenario and rewritten as a short, colloquial, role-played request, so that a task's prompt is not recoverable by web-searching the underlying issue, pull request, or commit thread. Because the dataset is released openly - task directories, environment images, evaluation harness, tests, and reference solutions - contamination resistance rests on this construction together with dataset versioning rather than on secrecy. The four subsets - repository-level engineering, front-end development, office and business workflows, and red-/blue-team security - probe complementary facets of real work, each with its own verification style. All are packaged in a uniform task-directory format and run, under a uniform and reproducible protocol, on two agent harnesses (CodeBuddy Code and Claude Code); the full open release makes the benchmark reproducible end to end and directly auditable, since any third party can re-run each task and inspect its content. Because each subset uses a different scoring instrument, scores are not comparable across subsets and the suite reports no suite-wide average. We report a cross-model leaderboard across several model families.
Original Article
View Cached Full Text

Cached at: 07/24/26, 05:06 AM

Paper page - Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

Source: https://huggingface.co/papers/2607.20911 Authors:

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

WeintroduceTencentWorkBuddyBench,amulti-domainevaluationsuiteforcodingagents;thisreportdocumentsitsconstructionmethodology,scoringprotocol,andacross-modelleaderboard.Atitscoreisaunifiedevaluationframeworkforconstructingandrunningdistribution-informedcoding-agenttasksacrossfourworkdomains-Code,Web,Office,andSecurity.Ratherthanadaptingpublicissuetext,everytaskisreverse-engineeredfromarealcommit,pullrequest,orbusinessscenarioandrewrittenasashort,colloquial,role-playedrequest,sothatatask’spromptisnotrecoverablebyweb-searchingtheunderlyingissue,pullrequest,orcommitthread.Becausethedatasetisreleasedopenly-taskdirectories,environmentimages,evaluationharness,tests,andreferencesolutions-contaminationresistancerestsonthisconstructiontogetherwithdatasetversioningratherthanonsecrecy.Thefoursubsets-repository-levelengineering,front-enddevelopment,officeandbusinessworkflows,andred-/blue-teamsecurity-probecomplementaryfacetsofrealwork,eachwithitsownverificationstyle.Allarepackagedinauniformtask-directoryformatandrun,underauniformandreproducibleprotocol,ontwoagentharnesses(CodeBuddyCodeandClaudeCode);thefullopenreleasemakesthebenchmarkreproducibleendtoendanddirectlyauditable,sinceanythirdpartycanre-runeachtaskandinspectitscontent.Becauseeachsubsetusesadifferentscoringinstrument,scoresarenotcomparableacrosssubsetsandthesuitereportsnosuite-wideaverage.Wereportacross-modelleaderboardacrossseveralmodelfamilies.

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2607\.20911

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2607.20911 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2607.20911 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2607.20911 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?

Hugging Face Daily Papers

CODA-BENCH is a new benchmark for evaluating code agents on data-intensive tasks, bridging the gap between code-centric and data-centric evaluations. It includes over 1,000 tasks from 31 communities, with realistic data scale and noise, revealing that even top agents achieve only 61.1% success rate.

SWE-Bench Pro V2 (9 minute read)

TLDR AI

SWE-Bench Pro V2 is an updated benchmark for evaluating AI agents in software engineering, featuring 642 tasks across 11 repositories with improved evaluation protocols and contamination controls.

ProgramBench (5 minute read)

TLDR AI

ProgramBench is a new benchmark that evaluates AI agents' ability to reconstruct complete software projects from compiled binaries and documentation without access to source code or decompilation tools.