TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces
Summary
TraceDance is an automated system that builds targeted benchmarks from real-world agent deployment traces to evaluate undesirable behaviors in LLMs, achieving high construction rates and revealing significant weaknesses in current frontier models.
View Cached Full Text
Cached at: 09/29/26, 04:11 AM
Paper page - TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces
Source: https://huggingface.co/papers/2609.33295 Published on Sep 27
#1 Paper of the day Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Anagentcancompleteataskwhileexhibitingundesirablebehaviorduringexecution.Developersneedtestsforthespecificbehaviorsencounteredindeployment,beyondfixedbenchmarksuites.WepresentTraceDance,anagentsystemthatconstructstargetedbenchmarksfromdeploymenttracesforuser-specifiedundesirablebehaviors.Forefficientconstruction,Anchor-and-Confirmcombinesprogrammableretrievalwithcandidate-levelconfirmationbyaFlashlargelanguagemodel(LLM),whiletheAnchorSynthesisLoopgeneratesandrevisesspecificationsforcustombehaviors.Thebenchmarksusedecision-pointcontinuationtoevaluateanLLM’snextturnatarecordeddecisionpointwithabehavior-specificrubric,withoutareferenceanswerorenvironmentreplay.Experimentsincodingandgeneraltoolusedrawon252,557sessionsandproduce107benchmarkswith4,125instances,fulfilling95.3%ofbuild-targetrequests.Bothhumanannotatorsconfirmtherequestedbehaviorin84%ofsampledinstances,andtheautomatedgrader’sagreementwithhumanpass/failjudgmentsiscomparabletothatbetweentheannotators.NinefrontierLLMsachieveameanpassrateofonly26.7%,showingthattheystillstruggletorespondappropriatelyattheevaluateddecisionpoints.Analysisacrossbehavior-specificbenchmarksfurtherrevealsweaknessesinhowcurrentLLMsbehaveasagents.Byturningdeploymentproblemsintotargetedbenchmarks,TraceDancecouldserveasakeycomponentoftherecursiveself-improvement(RSI)loop.
View arXiv pageView PDFProject pageGitHubAdd to collection
Get this paper in your agent:
hf papers read 2609\.33295
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.33295 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.33295 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.33295 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories
TraceGraph is a graph-based framework that constructs shared decision landscapes from multi-model agent trajectories, enabling diagnosis of failure regions and improvement via trap-aware recovery pipelines.
TRACE: Trajectory Reasoning through Adaptive Cross-Step Evidence Aggregation for LLM Agents
TRACE is a monitoring framework for long-horizon LLM agent trajectories that uses a Triage-Inspect-Judge loop to connect evidence across temporally distant actions, achieving high recall and F1 on evasive sabotage detection tasks.
SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse
Presents SkillTrace, a multi-trace provenance auditing framework for LLM-agent skill reuse that extracts expression, implementation, and operational traces, achieving strong accuracy on a benchmark and enabling large-scale wild audits.
BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents
BenchTrace is a benchmark for evaluating the self-evolution abilities of LLM agents, focusing on reflection and controlled evolution through a dataset of 1,821 annotated episodes and two evaluation tasks: Reflection Evaluation and Evolution Evaluation. Experiments with Qwen3-32B and GPT-4.1 show both models struggle, with a main bottleneck in diagnosis and issues in generalization and forgetting.
TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation
TRACE Bench is a task-driven agentic checklist evaluation framework for roleplay, decomposing role profiles into checklists, using a user agent for natural conversation, and tracing scores back to checklist items and dialogue evidence. It achieves 99.91% coverage, outperforming the MiniMax Role-play Benchmark's 73.74%, and supports closed-loop benchmark evolution across 26 models.