Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
Summary
This paper introduces Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents designed to resist data contamination by reverse-engineering tasks from real commits and business scenarios, covering Code, Web, Office, and Security domains.
View Cached Full Text
Cached at: 07/24/26, 05:06 AM
Paper page - Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
Source: https://huggingface.co/papers/2607.20911 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
WeintroduceTencentWorkBuddyBench,amulti-domainevaluationsuiteforcodingagents;thisreportdocumentsitsconstructionmethodology,scoringprotocol,andacross-modelleaderboard.Atitscoreisaunifiedevaluationframeworkforconstructingandrunningdistribution-informedcoding-agenttasksacrossfourworkdomains-Code,Web,Office,andSecurity.Ratherthanadaptingpublicissuetext,everytaskisreverse-engineeredfromarealcommit,pullrequest,orbusinessscenarioandrewrittenasashort,colloquial,role-playedrequest,sothatatask’spromptisnotrecoverablebyweb-searchingtheunderlyingissue,pullrequest,orcommitthread.Becausethedatasetisreleasedopenly-taskdirectories,environmentimages,evaluationharness,tests,andreferencesolutions-contaminationresistancerestsonthisconstructiontogetherwithdatasetversioningratherthanonsecrecy.Thefoursubsets-repository-levelengineering,front-enddevelopment,officeandbusinessworkflows,andred-/blue-teamsecurity-probecomplementaryfacetsofrealwork,eachwithitsownverificationstyle.Allarepackagedinauniformtask-directoryformatandrun,underauniformandreproducibleprotocol,ontwoagentharnesses(CodeBuddyCodeandClaudeCode);thefullopenreleasemakesthebenchmarkreproducibleendtoendanddirectlyauditable,sinceanythirdpartycanre-runeachtaskandinspectitscontent.Becauseeachsubsetusesadifferentscoringinstrument,scoresarenotcomparableacrosssubsetsandthesuitereportsnosuite-wideaverage.Wereportacross-modelleaderboardacrossseveralmodelfamilies.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2607\.20911
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.20911 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.20911 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.20911 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?
CODA-BENCH is a new benchmark for evaluating code agents on data-intensive tasks, bridging the gap between code-centric and data-centric evaluations. It includes over 1,000 tasks from 31 communities, with realistic data scale and noise, revealing that even top agents achieve only 61.1% success rate.
SWE-Bench Pro V2 (9 minute read)
SWE-Bench Pro V2 is an updated benchmark for evaluating AI agents in software engineering, featuring 642 tasks across 11 repositories with improved evaluation protocols and contamination controls.
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
SWE-bench Science introduces a repository-level benchmark for evaluating coding agents on scientific software repair tasks, revealing failure mechanisms and mixed effects of scientific guidance.
SWE Context Bench just proved something I think a lot of coding agent users already feel
A new benchmark paper 'SWE Context Bench' tests whether coding agents can reuse knowledge across tasks, highlighting a gap in existing benchmarks that only evaluate isolated problem-solving. The author discusses solutions like external memory and mentions tools such as langmem, mem0, supermemory, and Greplica.
ProgramBench (5 minute read)
ProgramBench is a new benchmark that evaluates AI agents' ability to reconstruct complete software projects from compiled binaries and documentation without access to source code or decompilation tools.