TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

Hugging Face Daily Papers Papers

Summary

TraceDance is an automated system that builds targeted benchmarks from real-world agent deployment traces to evaluate undesirable behaviors in LLMs, achieving high construction rates and revealing significant weaknesses in current frontier models.

An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable behaviors. For efficient construction, Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash large language model (LLM), while the Anchor Synthesis Loop generates and revises specifications for custom behaviors. The benchmarks use decision-point continuation to evaluate an LLM's next turn at a recorded decision point with a behavior-specific rubric, without a reference answer or environment replay. Experiments in coding and general tool use draw on 252,557 sessions and produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests. Both human annotators confirm the requested behavior in 84% of sampled instances, and the automated grader's agreement with human pass/fail judgments is comparable to that between the annotators. Nine frontier LLMs achieve a mean pass rate of only 26.7%, showing that they still struggle to respond appropriately at the evaluated decision points. Analysis across behavior-specific benchmarks further reveals weaknesses in how current LLMs behave as agents. By turning deployment problems into targeted benchmarks, TraceDance could serve as a key component of the recursive self-improvement (RSI) loop.
Original Article
View Cached Full Text

Cached at: 09/29/26, 04:11 AM

Paper page - TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

Source: https://huggingface.co/papers/2609.33295 Published on Sep 27

#1 Paper of the day Authors:

,

,

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

Anagentcancompleteataskwhileexhibitingundesirablebehaviorduringexecution.Developersneedtestsforthespecificbehaviorsencounteredindeployment,beyondfixedbenchmarksuites.WepresentTraceDance,anagentsystemthatconstructstargetedbenchmarksfromdeploymenttracesforuser-specifiedundesirablebehaviors.Forefficientconstruction,Anchor-and-Confirmcombinesprogrammableretrievalwithcandidate-levelconfirmationbyaFlashlargelanguagemodel(LLM),whiletheAnchorSynthesisLoopgeneratesandrevisesspecificationsforcustombehaviors.Thebenchmarksusedecision-pointcontinuationtoevaluateanLLM’snextturnatarecordeddecisionpointwithabehavior-specificrubric,withoutareferenceanswerorenvironmentreplay.Experimentsincodingandgeneraltoolusedrawon252,557sessionsandproduce107benchmarkswith4,125instances,fulfilling95.3%ofbuild-targetrequests.Bothhumanannotatorsconfirmtherequestedbehaviorin84%ofsampledinstances,andtheautomatedgrader’sagreementwithhumanpass/failjudgmentsiscomparabletothatbetweentheannotators.NinefrontierLLMsachieveameanpassrateofonly26.7%,showingthattheystillstruggletorespondappropriatelyattheevaluateddecisionpoints.Analysisacrossbehavior-specificbenchmarksfurtherrevealsweaknessesinhowcurrentLLMsbehaveasagents.Byturningdeploymentproblemsintotargetedbenchmarks,TraceDancecouldserveasakeycomponentoftherecursiveself-improvement(RSI)loop.

View arXiv pageView PDFProject pageGitHubAdd to collection

Get this paper in your agent:

hf papers read 2609\.33295

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.33295 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.33295 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.33295 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents

arXiv cs.AI

BenchTrace is a benchmark for evaluating the self-evolution abilities of LLM agents, focusing on reflection and controlled evolution through a dataset of 1,821 annotated episodes and two evaluation tasks: Reflection Evaluation and Evolution Evaluation. Experiments with Qwen3-32B and GPT-4.1 show both models struggle, with a main bottleneck in diagnosis and issues in generalization and forgetting.

TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation

arXiv cs.CL

TRACE Bench is a task-driven agentic checklist evaluation framework for roleplay, decomposing role profiles into checklists, using a user agent for natural conversation, and tracing scores back to checklist items and dialogue evidence. It achieves 99.91% coverage, outperforming the MiniMax Role-play Benchmark's 73.74%, and supports closed-loop benchmark evolution across 26 models.