TimeEvo: 故障驱动的时间序列智能体自进化
摘要
TimeEvo 是一种故障驱动的自进化方法,用于时间序列智能体,它能够诊断能力差距、综合工具,并在任务和骨干网络上提高准确性。
查看缓存全文
缓存时间: 2026/09/24 09:19
# TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent
Source: [https://arxiv.org/html/2609.27277](https://arxiv.org/html/2609.27277)
Jie YangYan ZhengAffiliation:Visa ResearchJiarui SunAffiliation:Visa ResearchXiran FanAffiliation:Visa ResearchJunpeng WangAffiliation:Visa ResearchLiang WangAffiliation:Visa ResearchZelin XuAffiliation:Visa ResearchAffiliation:University of FloridaQinghua LiuAffiliation:Visa ResearchAffiliation:The Ohio State UniversityZhengyu FangAffiliation:Visa ResearchAffiliation:Case Western Reserve UniversityYiwei CaiAffiliation:Visa ResearchPhilip S\. YuAffiliation:University of Illinois at Chicago
###### Abstract
Time series agents answer analytical questions by calling external tools, and which tools they carry is decided by people before the agent runs\. However, we identify two failures in this setup\.Human–Agent Tool Misalignment: a library of 21 expert\-curated tools helps on some tasks and hurts on others, dropping anomaly accuracy under every backbone we test\.Silent Harm: one round of generic self\-revision changes 147 answers and breaks 56 of them, while the final score moves by less than a point\. Both follow from the same gap: whether a tool helps is decided question by question at runtime, while tools are supplied in advance and judged by a single average\. To address this, we propose TimeEvo, which clusters an agent’s diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence\-only tools that fill them, and admits the candidate library only through a paired admission gate\. Experiments on ten time series QA tasks and three backbones show that TimeEvo, starting from an empty library, improves accuracy on every task and every backbone, and that a library grown on a cheap model still gains when it is installed into stronger ones\. Code is available at[https://github\.com/Muyiiiii/TimeEvo](https://github.com/Muyiiiii/TimeEvo)\.
$\\dagger$$\\dagger$footnotetext:Corresponding author\.## 1Introduction
Time series analysis\([Hu et al\.,](https://arxiv.org/html/2609.27277#bib.bib1);[Beqari et al\., 2024](https://arxiv.org/html/2609.27277#bib.bib36)\)is a fundamental problem with broad impact in real\-world applications such as finance\([Yang et al\., 2025b](https://arxiv.org/html/2609.27277#bib.bib2);[Cheng et al\., 2025](https://arxiv.org/html/2609.27277#bib.bib41)\), traffic\([Hu et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib3)\), and healthcare\([Yang et al\., 2025a](https://arxiv.org/html/2609.27277#bib.bib4)\), and spans a wide range of tasks, including forecasting\([Yang et al\., 2026a](https://arxiv.org/html/2609.27277#bib.bib5);[Wang et al\., 2025](https://arxiv.org/html/2609.27277#bib.bib47)\), imputation\([Yang et al\., 2026b](https://arxiv.org/html/2609.27277#bib.bib6)\), and time series understanding\([Xie et al\., 2024](https://arxiv.org/html/2609.27277#bib.bib7)\)\. To address this diverse task landscape, model\-based approaches follow two broad designs\. Dedicated models\([Liu et al\., 2024](https://arxiv.org/html/2609.27277#bib.bib8);[Zeng et al\., 2023](https://arxiv.org/html/2609.27277#bib.bib9)\)provide efficient and competitive performance on well\-defined objectives, whereas multimodal language models\([Jin et al\., 2024](https://arxiv.org/html/2609.27277#bib.bib16);[Xie et al\., 2024](https://arxiv.org/html/2609.27277#bib.bib7)\)treat time series as an additional modality for natural\-language interaction and open\-ended questions\([Kong et al\., 2025](https://arxiv.org/html/2609.27277#bib.bib14)\)\. Despite their different interfaces, both acquire analytical capabilities through task\-oriented training and encode them in model parameters\. Therefore, supporting a new task or capability typically requires corresponding data construction, training, or tuning\([Dong et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib39);[Kong et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib46)\)\.
Large language model \(LLM\) agents\([Yao et al\., 2022](https://arxiv.org/html/2609.27277#bib.bib17);[Zhu et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib37)\)offer an alternative: rather than internalizing every time series operation in model parameters\([Dong et al\., 2025](https://arxiv.org/html/2609.27277#bib.bib40)\), they interpret user requests, plan analyses, and invoke external analytical tools\([Wu et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib10)\)\. Across tasks ranging from forecasting\([Weng et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib11)\)to open\-ended time series understanding\([Kong et al\., 2025](https://arxiv.org/html/2609.27277#bib.bib14);[Yu et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib38)\), they synthesize tool outputs and intermediate evidence into coherent natural\-language responses\([Ye et al\., 2024](https://arxiv.org/html/2609.27277#bib.bib19);[Zhao et al\., 2025](https://arxiv.org/html/2609.27277#bib.bib12)\)\. This design implements many domain\-specific analytical operations as reusable external modules that can be inspected, recombined, and modified without retraining the underlying model\([Wu et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib10)\)\. Building on this modular design, recent agent frameworks further refine how agents use these external modules by distilling execution experience into reusable guidance\([Liu et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib15);[Ye et al\., 2024](https://arxiv.org/html/2609.27277#bib.bib19)\)\.
However, task\-specific development persists: human effort does not disappear but shifts from model training\([Fan et al\., 2026a](https://arxiv.org/html/2609.27277#bib.bib45);[Fan et al\., 2026b](https://arxiv.org/html/2609.27277#bib.bib44);[Fang et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib35)\)to configuring prompts, workflows, and tools for tasks anticipated before deployment\([Hao et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib43)\)\. It remains unclear whether supplying an agent with more tools makes it better at the tasks those tools were written for, and whether the agent’s own revisions can be trusted to improve it\([Miao et al\., 2025](https://arxiv.org/html/2609.27277#bib.bib42)\)\. To investigate, we analyze the errors and repair attempts of a frozen\-LLM time series agent \(TSAgent\) across time series QA tasks as shown in[Fig\.1](https://arxiv.org/html/2609.27277#S1.F1)\. Our analysis reveals two key phenomena:
Figure 1:Two failure phenomena in TSAgent\.We compare the same frozen agent with and without 21 expert\-curated tools\(a\), and before and after one generic self\-refinement pass\(b\), pairing every question across ten QA tasks and three backbone LLMs\. Improved and degraded tasks are shadedblueandorangein \(a\), and fixed and broken answers are shown ingreenandredin \(b\)\.- ❶Human–Agent Tool Misalignment:*what humans prefer to supply is not what the agent benefits from\.*Equipping the same frozen agent with a library of 21 expert\-curated analysis tools of TimeART\([Wu et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib10)\)makes it better on some tasks and worse on others \([Fig\.1](https://arxiv.org/html/2609.27277#S1.F1)a\)\. The largest gains and the largest losses both recur in the same direction under all three backbone LLMs\. The losses are not a coverage problem: the library includes a dedicated anomaly\-detection tool, yet the anomaly task drops by 5\.8 to 8\.8 points under every backbone\. T3 shares the same task family with the improving T2 and T4\([Weng et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib11)\), yet it is degraded under all three backbones\. The misalignment is structural: the supply is fixed per task before deployment, but whether a tool helps is decided question by question at runtime\.
- ❷Silent Harm:*self\-revision breaks answers that the final score never shows\.*A single pass of the Self\-Refine\([Madaan et al\., 2023](https://arxiv.org/html/2609.27277#bib.bib20)\)changes 147 answers across the ten tasks: 49 are fixed, 56 previously correct answers break, and the remaining 42 swap one wrong answer for another\. Yet the final score moves by less than one point \([Fig\.1](https://arxiv.org/html/2609.27277#S1.F1)b\)\. The pattern is not specific to one model: on two further backbones the pass touches 97 and 163 answers, breaks 45 and 56 of them, and again never moves the final score by a full point\. On T4 alone, the pass on luna changes 14 answers, fixes 0, and breaks 11\. The silence is arithmetic: inside a single average, the harm silently erases the fixes, and the final score never moves\.
Together, these findings show that the tools agents are given do not match what they need at runtime, while unexamined self\-revision silently erases much of what it fixes\. This leads to a central question:*Can a time series agent supply its own missing tools from diagnosed failures while measurably protecting what it already answers correctly from silent harm?*
To answer this question, we proposeTimeEvo, a self\-evolving agent for time series analysis\. The key idea is to let the agent discover the tools it actually needs from the failures it makes, and to trust no self\-made update until it proves that it helps more than it harms\. Specifically, TimeEvo clusters the agent’s errors into failure buckets, plans a measurement contract for each, and synthesizes evidence\-only tools to fulfill them\. The whole candidate library must then pass two levels of validation, and a failing library is pruned and retested once\. Ultimately, the agent evolves its own task\-specific tool library, in which every tool has earned its place without harming silently\.
Our main contributions are summarized as follows:
- •We identify two phenomena in time series agents,Human–Agent Tool MisalignmentandSilent Harm: what humans prefer to supply is not what the agent benefits from, and self\-revision silently erases much of what it fixes\.
- •We proposeTimeEvo, a self\-evolving agent that discovers the tools it needs from the failures it makes, and admits every self\-made update only after two levels of validation\.
- •Extensive experiments on ten time series QA benchmarks show that TimeEvo improves its base agent on every task, with pooled gains from\+1\.5\+1\.5to\+14\.7\+14\.7points\. Neither expert\-curated tools nor generic self\-refinement reproduces this\.
## 2Related Work
### 2\.1Time Series Analysis
Time series analysis has long been dominated by specialized models built for individual tasks such as forecasting\([Yang et al\., 2026a](https://arxiv.org/html/2609.27277#bib.bib5)\), classification\([Dempster et al\., 2020](https://arxiv.org/html/2609.27277#bib.bib30)\), and anomaly detection\([Liu and Paparrizos, 2024](https://arxiv.org/html/2609.27277#bib.bib31)\)\. Pretrained foundation models \(e\.g\., Chronos\([Ansari et al\., 2025](https://arxiv.org/html/2609.27277#bib.bib21)\), TimesFM\([Das et al\., 2023](https://arxiv.org/html/2609.27277#bib.bib22)\), and Moirai\([Liu et al\., 2025a](https://arxiv.org/html/2609.27277#bib.bib23)\)\) extend this line by transferring forecasting capability across domains\. Furthermore, the emergence of LLMs has enabled more flexible and powerful approaches to time series analysis\. Multimodal language models such as ChatTS\([Xie et al\., 2024](https://arxiv.org/html/2609.27277#bib.bib7)\)and TimeLLM\([Jin et al\., 2024](https://arxiv.org/html/2609.27277#bib.bib16)\)take time series as an additional input modality and answer open\-ended questions end to end\. Yet facing a new task, these models still depend on collecting task\-specific data and training or tuning the model\. In contrast, agent\-based systems \(e\.g\., TS\-Reasoner\([Ye et al\., 2024](https://arxiv.org/html/2609.27277#bib.bib19)\), TS\-Agent\([Liu et al\., 2025b](https://arxiv.org/html/2609.27277#bib.bib24)\), TimeCopilot\([Garza and Rosillo, 2025](https://arxiv.org/html/2609.27277#bib.bib25)\)\) offer a training\-free alternative that moves analytical capabilities outside the model: the agent plans over external tools and synthesizes their outputs into natural\-language answers\. Training\-free, however, does not mean effort\-free: human effort merely shifts from model training to designing prompts, workflows, and tools before deployment\. To reduce this reliance on human design, we propose TimeEvo, which treats diagnosed failures as the supervision signal for evolving the capability library itself\.
### 2\.2Self\-Evolving Agents
Self\-evolution has been extensively explored for general LLM agents\. One line, exemplified by Reflexion\([Shinn et al\., 2023](https://arxiv.org/html/2609.27277#bib.bib18)\)and EvolveR\([Wu et al\., 2025](https://arxiv.org/html/2609.27277#bib.bib26)\), distills experience from past trajectories into reusable lessons or principles\. Another line maintains growing skill libraries such as Voyager\([Wang et al\., 2023](https://arxiv.org/html/2609.27277#bib.bib27)\), and recent work like SkillOpt\([Yang et al\., 2026c](https://arxiv.org/html/2609.27277#bib.bib28)\)and CoEvoSkills\([Zhang et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib29)\)further tests learned skills before deployment with held\-out scores or per\-task checks\. Within time series, TimeClaw distills usage experience over a fixed tool library\([Liu et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib15)\)\. However, these frameworks admit an update through the final score or a per\-task check at best, and some skip validation entirely\. The admission never hinges on how many previously correct answers the update breaks\. Inside a single final score, such breaks cancel against the fixes, so an update can cause real harm while the score barely moves, exactly the Silent Harm identified in our analysis\. This motivates TimeEvo, which evolves the tool library from diagnosed failures and admits every update only after two levels of validation that weigh what it fixes against what it breaks\.
## 3Method
Figure 2:Overview of TimeEvo\.Each round turns the agent’s own failures into new tools: errors are clustered into failure buckets, each bucket is planned into a measurement contract, and contracts are synthesized into evidence\-only candidate tools\. The candidate library is admitted only after two levels of validation, a per\-tool prefilter and a whole\-library paired gate\. A failing library is pruned and retested once\. Admitted tools join the agent’s library, while rejected rounds leave the agent untouched\. At inference, the frozen agent answers first, and admitted tools review only in\-scope questions, keeping or replacing the answer\.### 3\.1Problem Definition
Each time series question is a triplex=\(q,𝐒,𝒪\)x=\(q,\\mathbf\{S\},\\mathcal\{O\}\): a natural\-language questionqqthat may carry event context, an observed input𝐒\\mathbf\{S\}collecting one or two primary series together with any named covariate or candidate series the question refers to, and a finite option set𝒪\\mathcal\{O\}\. An agent is a frozen LLM equipped with a libraryℒ\\mathcal\{L\}of executable analysis tools, and it answers by combining tool evidence with its own reasoning to select one option from𝒪\\mathcal\{O\}\. We write𝒜ℒ\\mathcal\{A\}\_\{\\mathcal\{L\}\}for the agent carryingℒ\\mathcal\{L\}, and the base agent starts from an empty libraryℒ0\\mathcal\{L\}\_\{0\}\(a pre\-installed forecasting root is kept only as an ablation\)\. The gold answer is determined by the observed input alone, so whether the agent answers a question correctly can be checked automatically\. A final score alone cannot judge a library updateℒ→ℒ′\\mathcal\{L\}\\rightarrow\\mathcal\{L\}^\{\\prime\}, which is the Silent Harm of[Fig\.1](https://arxiv.org/html/2609.27277#S1.F1)b\. We therefore track what an update fixes and what it breaks on a question setDD:
helped\(D\)\\displaystyle\\mathrm\{helped\}\(D\)=\{x∈D:v\(x\)=0,v′\(x\)=1\},\\displaystyle=\\\{x\\in D:v\(x\)=0,\\ v^\{\\prime\}\(x\)=1\\\},\(1\)harmed\(D\)\\displaystyle\\mathrm\{harmed\}\(D\)=\{x∈D:v\(x\)=1,v′\(x\)=0\},\\displaystyle=\\\{x\\in D:v\(x\)=1,\\ v^\{\\prime\}\(x\)=0\\\},wherevvandv′v^\{\\prime\}denote correctness underℒ\\mathcal\{L\}andℒ′\\mathcal\{L\}^\{\\prime\}on the same questions\. Given train, validation, and test splitsDtrain,Dval,DtestD\_\{\\mathrm\{train\}\},D\_\{\\mathrm\{val\}\},D\_\{\\mathrm\{test\}\}, the goal is to evolveℒ\\mathcal\{L\}from the agent’s failures onDtrainD\_\{\\mathrm\{train\}\}so that helped outnumbers harmed on unseen questions, whileDtestD\_\{\\mathrm\{test\}\}is never used for any decision\.
### 3\.2Overview and Frozen Residual Inference
TimeEvo runs in rounds, and each round turns the agent’s current failures into new tools \([Fig\.2](https://arxiv.org/html/2609.27277#S3.F2)\)\. A round has four stages: diagnose the failures on the training split into measurement contracts, synthesize evidence\-only tools for the contracts, screen each tool with a cheap prefilter, and submit the surviving library as a whole to a paired admission gate\. Afterrraccepted rounds the deployed library is a chainℒr=\(ℒ0,𝒯1,…,𝒯r\)\\mathcal\{L\}\_\{r\}=\(\\mathcal\{L\}\_\{0\},\\mathcal\{T\}\_\{1\},\\ldots,\\mathcal\{T\}\_\{r\}\), where𝒯k\\mathcal\{T\}\_\{k\}is the tool set admitted in roundkk, and a question is threaded through the stages:
𝒜ℒk\(x\)=\{ρk\(x,Ek\(x\),𝒜ℒk−1\(x\)\)ifx∈scope\(𝒯k\),𝒜ℒk−1\(x\)otherwise,\\mathcal\{A\}\_\{\\mathcal\{L\}\_\{k\}\}\(x\)=\\begin\{cases\}\\rho\_\{k\}\\big\(x,\\ E\_\{k\}\(x\),\\ \\mathcal\{A\}\_\{\\mathcal\{L\}\_\{k\-1\}\}\(x\)\\big\)&\\text\{if \}x\\in\\mathrm\{scope\}\(\\mathcal\{T\}\_\{k\}\),\\\\\[2\.0pt\] \\mathcal\{A\}\_\{\\mathcal\{L\}\_\{k\-1\}\}\(x\)&\\text\{otherwise\},\\end\{cases\}\(2\)wherescope\(𝒯k\)\\mathrm\{scope\}\(\\mathcal\{T\}\_\{k\}\)matches the stage’s declared scope fields \(task types, evidence types, single\- or dual\-series input, and an anti\-scope\),Ek\(x\)E\_\{k\}\(x\)is the numeric evidence the stage’s tools compute onxx, andρk\\rho\_\{k\}is the same agent reviewing the incoming answer with that evidence\. At answer time, as the*Frozen Residual Inference*strip of[Fig\.2](https://arxiv.org/html/2609.27277#S3.F2)shows, the agent consults a structured scope catalog and inspects at most three shortlisted tools per question\. The review is conservative by construction: the incoming answer is the default, and a replacement is honored only when a stage tool actually ran on the question\. If a tool errors, is never invoked, or the review does not explicitly decide to replace, the question falls back to the incoming answer\. Out\-of\-scope questions are therefore never degraded\. In\-scope harm remains possible, and the admission gate measures it alongside repair and screens candidate libraries for positive net improvement on validation data\.
### 3\.3Failure\-Driven Analysis and Planning
As shown in the*Failure\-driven Analysis*panel of[Fig\.2](https://arxiv.org/html/2609.27277#S3.F2), roundrrfirst replays the current library on the training split, collects the questions it answers wrong, and partitions them into failure buckets:
ℱr=\{x∈Dtrain:vr−1\(x\)=0\},ℱr=ℬ1∪⋯∪ℬJ,\\mathcal\{F\}\_\{r\}=\\\{x\\in D\_\{\\mathrm\{train\}\}:v\_\{r\-1\}\(x\)=0\\\},\\qquad\\mathcal\{F\}\_\{r\}=\\mathcal\{B\}\_\{1\}\\cup\\cdots\\cup\\mathcal\{B\}\_\{J\},\(3\)wherevr−1v\_\{r\-1\}denotes correctness under the current libraryℒr−1\\mathcal\{L\}\_\{r\-1\}and the buckets are disjoint\. The agent builds this partition in two steps: it first defines a small set of failure categories from a sample of the errors, then assigns every error to exactly one category\. Errors on which a tool fired and errors that no tool reached are clustered separately, because the former indicate a defective or misused tool while the latter indicate a coverage gap\. The definition step never sees task metadata, so buckets group errors by the missing capability rather than by question type\. Oversized buckets are split once, and buckets with fewer than two errors are set aside\.
Each bucketℬj\\mathcal\{B\}\_\{j\}is then diagnosed by a local planner \(the*Planner*panel in[Fig\.2](https://arxiv.org/html/2609.27277#S3.F2)\), which must return a*measurement contract*cjc\_\{j\}: the diagnosed root cause, the required measurement, the inputs it needs, a success criterion, and a scope with an anti\-scope of applicability\. The contract also declares one library action, typed by the diagnosed failure mode:createa new tool when evidence is missing,refinean existing tool with usage guidance when its evidence is ignored or misread,rescopea tool that fires on the wrong questions, andretirea tool that should leave the library\. Because buckets are planned in parallel, a global planner coordinates the proposals afterwards: duplicated measurements are merged and overlapping scopes are narrowed, while a merge of incompatible contracts is rejected mechanically and falls back to the independent per\-bucket plans\. Planning is thus committed to measurements rather than answers: what a bucket receives is a contract for the evidence its failures lack, not a rule for how to answer them\.
### 3\.4Candidate Synthesis and Prefilter
As shown in the*Candidate Synthesis*panel of[Fig\.2](https://arxiv.org/html/2609.27277#S3.F2), each coordinated contractcjc\_\{j\}receives up to three sequential synthesis attempts:
tj\(i\)=Synth\(cj,fj\(i−1\)\),tj\(i\):\(𝐒,q\)↦𝐞∈ℝdj,i≤3,t\_\{j\}^\{\(i\)\}=\\mathrm\{Synth\}\\big\(c\_\{j\},\\ f\_\{j\}^\{\(i\-1\)\}\\big\),\\qquad t\_\{j\}^\{\(i\)\}:\(\\mathbf\{S\},q\)\\mapsto\\mathbf\{e\}\\in\\mathbb\{R\}^\{d\_\{j\}\},\\qquad i\\leq 3,\(4\)wherefj\(i−1\)f\_\{j\}^\{\(i\-1\)\}is the structured feedback of the previous failed attempt \(fj\(0\)=∅f\_\{j\}^\{\(0\)\}=\\emptyset\), and the delivered tooltjt\_\{j\}is the first passing attempt, one deterministic Python function\. A tool is evidence\-only: it never outputs a verdict, an option string, or a dataset\-specific constant\. It may read a covariate only when the question mentions it, and when its input is unavailable it abstains rather than fabricates a value\. A sandbox that permits only a fixed allowlist of operations enforces determinism and isolation, and a candidate that fails any of these checks is discarded\. A tool is not limited to plain arithmetic on the series: it may call the raw Chronos\-2 model\([Ansari et al\., 2025](https://arxiv.org/html/2609.27277#bib.bib21)\)as a forecasting primitive, and when its contract requests a decision boundary it may carry a*train\-only calibration*, a shallow decision tree fitted on training rows alone and compiled into its source:
gj=Fit\(\{\(tj\(x\),y\(x\)\):x∈Dtrain\}\),tj←tj⊕gjiffBalAcc\(gj\)≥0\.55,g\_\{j\}=\\mathrm\{Fit\}\\big\(\\\{\(\\,t\_\{j\}\(x\),\\ y\(x\)\\,\):x\\in D\_\{\\mathrm\{train\}\}\\\}\\big\),\\qquad t\_\{j\}\\leftarrow t\_\{j\}\\oplus g\_\{j\}\\ \\ \\text\{iff\}\\ \\ \\mathrm\{BalAcc\}\(g\_\{j\}\)\\geq 0\.55,\(5\)wherey\(x\)y\(x\)is the gold answer,⊕\\oplusdenotes compiling the fitted tree into the tool’s source code, andBalAcc\\mathrm\{BalAcc\}is balanced accuracy on an internal train\-side split with at least 16 samples per class\. Thresholds are therefore measured on training rows rather than invented by the agent\.
The first level of validation is a per\-tool prefilter, run on at most 32 held\-out errorsℋj⊂ℬj\\mathcal\{H\}\_\{j\}\\subset\\mathcal\{B\}\_\{j\}of the tool’s own bucket\. A candidate survives only if it delivers at least one*attributable*repair: a question fixed with the tool actually attached, actually executed, and visible in the execution record\. It must also break few previously correct answers and rarely fail to execute\. Two consecutive attempts without an attributable repair send the bucket back for re\-diagnosis, which may change the artifact kind, for example from a new tool to usage guidance\. The prefilter only discards clearly useless candidates: every statistical decision is reserved for the admission gate, the only decision whose mistakes would reach deployment\.
### 3\.5Paired Admission
The second level of validation \(the*Paired Admission*panel in[Fig\.2](https://arxiv.org/html/2609.27277#S3.F2)\) decides admission for the round’s surviving tools𝒯r\\mathcal\{T\}\_\{r\}as a whole rather than tool by tool, because tools that look harmless in isolation can interfere once deployed together\. The current library and the candidate library answer the full validation split in one paired run, and only questions that both sides complete are compared\. Applying[Eq\.1](https://arxiv.org/html/2609.27277#S3.E1)toDvalD\_\{\\mathrm\{val\}\}givesh=\|helped\(Dval\)\|h=\|\\mathrm\{helped\}\(D\_\{\\mathrm\{val\}\}\)\|andm=\|harmed\(Dval\)\|m=\|\\mathrm\{harmed\}\(D\_\{\\mathrm\{val\}\}\)\|\. Over thennpaired questions this yields the gate statistic:
δ=h−mn,δlb=δ−Z⋅max\(h\+m,1\)n,Z=1\.96,\\delta=\\frac\{h\-m\}\{n\},\\qquad\\delta\_\{\\mathrm\{lb\}\}=\\delta\-Z\\cdot\\frac\{\\sqrt\{\\max\(h\+m,\\,1\)\}\}\{n\},\\qquad Z=1\.96,\(6\)and the library passes strictly whenh\>0h\>0andδlb\>0\\delta\_\{\\mathrm\{lb\}\}\>0\. A clean 3\-versus\-0 round on a small split still fails, because its lower bound stays negative\. A single\-round run has no later rounds in which to accumulate evidence, so a supplementary rule admits a library that satisfies:
δ\>0,h−m≥max\(3,⌈0\.02n⌉\),mn≤0\.10,\\delta\>0,\\qquad h\-m\\ \\geq\\ \\max\\big\(3,\\ \\lceil 0\.02\\,n\\rceil\\big\),\\qquad\\frac\{m\}\{n\}\\ \\leq\\ 0\.10,\(7\)that is, a positive paired delta, a net repair of at least three questions or two percent of the split, and an absolute harm rate within ten percent\.
When the gate rejects the library, TimeEvo attempts one repair before giving up\. The validation records show which tool fixed or broke which question, so every updateuureceives its own countshuh\_\{u\}andmum\_\{u\}, and an update is removed only when it is clearly harmful:
𝒫=\{u∈𝒯r:mu\>0∧\(hu=0∨mu−hu≥3\)\}\.\\mathcal\{P\}=\\big\\\{\\,u\\in\\mathcal\{T\}\_\{r\}\\ :\\ m\_\{u\}\>0\\ \\wedge\\ \\big\(h\_\{u\}=0\\ \\vee\\ m\_\{u\}\-h\_\{u\}\\geq 3\\big\)\\,\\big\\\}\.\(8\)A tool is left alone whenever the evidence against it is thin: one that changed no answers may simply have met no matching questions, and one whose counts are small and mixed cannot be told apart from noise\. The pruned library is retested in exactly one more full paired run and commits only if it strictly improves the lower bound\. There is no second prune, because repeated retesting would slowly fit the library to the validation split\. Whether a library passes at once or only after the prune, it then reconciles its scope fields against how its tools behaved, and because that edit changes the library, it passes the same paired gate one final time before joining the chain\. A rejected round changes nothing\.
## 4Experiments
Table 1:Main comparison on the ten tasks across three backbones\. Each cell shows test accuracy with the within\-run paired delta \(pp\) in parentheses, whereboldmarks the best method per column within a block andunderlinethe second best\. The last block installs the library that GPT\-5\.6\-luna evolved into three stronger models, and itsshadedrows carry that library and are not ranked\.### 4\.1Experimental Settings
Tasks\.The formal suite covers ten time series QA tasks from six public sources: TemporalBench T1–T4\([Weng et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib11)\), TimeSeriesExam\([Cai et al\., 2024](https://arxiv.org/html/2609.27277#bib.bib13)\), Merrill\([Merrill et al\., 2024](https://arxiv.org/html/2609.27277#bib.bib32)\), TimeMQA Anomaly and Classification\([Kong et al\., 2025](https://arxiv.org/html/2609.27277#bib.bib14)\), MMTS Match\([Yin et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib33)\), and TSAQA Data\-Transformation\([Jing et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib34)\)\. Split details are in Appendix[A](https://arxiv.org/html/2609.27277#A1)\.
Protocol\.Every run uses the same frozen protocol: empty root library, one evolution round, train\-only calibration, and a report\-only test split\. We evaluate on three backbones, GPT\-5\.6\-luna, GPT\-5\.6\-terra, and GPT\-5\.4\-mini, with all four LM roles switched together\. Every number we report is a within\-run paired delta \([Eq\.1](https://arxiv.org/html/2609.27277#S3.E1)\)\.
Baselines\.We compare against five non\-evolving baselines: a fixed Chronos\-2 forecasting tool\([Ansari et al\., 2025](https://arxiv.org/html/2609.27277#bib.bib21)\), the 21 expert\-curated tools of TimeART\([Wu et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib10)\), Self\-Refine\([Madaan et al\., 2023](https://arxiv.org/html/2609.27277#bib.bib20)\), five\-sample self\-consistency voting, and few\-shot in\-context learning with four training examples\. Each tests one alternative explanation of the gains: that a strong pre\-installed tool, a hand\-engineered library, one more round of reflection, the same budget spent on sampling, or direct contact with the training data would suffice\. All are measured with the same paired protocol on the same splits\.
### 4\.2Overall Comparison
Figure 3:Human tools, and the transfer of an evolved library\.\(a\): Test delta on TimeMQA Anomaly for the 21 expert\-curated tools of TimeART and for the evolved library\.\(b\): Ten\-task mean accuracy of three stronger models, carrying the library GPT\-5\.6\-luna evolved\.\(c\): Per\-task gain of those models against the gain luna itself obtained with the same library\. Diagonal marks equal gain\.We compare TimeEvo against the five baselines on all ten tasks, under each of the three backbones\. Results are in[Table1](https://arxiv.org/html/2609.27277#S4.T1), and[Fig\.3](https://arxiv.org/html/2609.27277#S4.F3)a takes a closer look at TimeMQA Anomaly, where the human library fails\. Based on these results, we summarize our observations \(Obs\.\) as follows:
Obs\. ❶: TimeEvo is the only method that improves every task on every backbone\.All thirty of its task results are positive\. Its mean gain is the largest in every block:\+8\.79\+8\.79on luna,\+8\.73\+8\.73on mini, and\+7\.21\+7\.21on terra, the strongest backbone of the three\. Each backbone evolves its own library from its own failures, so the gains come from the loop rather than from one lucky set of tools\. These results answer the first half of the central question of[Section1](https://arxiv.org/html/2609.27277#S1): an agent can supply its own missing tools from its own diagnosed failures\.
Obs\. ❷: Human\-supplied tools and generic self\-improvement pay for their peaks with losses elsewhere\.The fixed Chronos\-2 tool is strongly positive only on the two forecasting tiers it was built for \(T2/T4, up to\+22\.22\+22\.22\) and flat or negative elsewhere\. As shown in[Fig\.3](https://arxiv.org/html/2609.27277#S4.F3)a, the TimeART library includes a dedicated anomaly tool yet loses−6\.25/−8\.75/−5\.79\-6\.25/\-8\.75/\-5\.79on Anomaly under every backbone, and one Self\-Refine pass breaks nearly as many answers as it fixes \([Fig\.1](https://arxiv.org/html/2609.27277#S1.F1)b\)\. Voting with five samples spends five times the budget for ten\-task means of\+0\.79/−0\.21/\+1\.64\+0\.79/\-0\.21/\+1\.64, and few\-shot ICL, which reads the same training split our method learns from, is strongly negative on Classification under all three backbones\. None of these alternatives is free of human design effort\.
### 4\.3Transfer Analysis
We test whether a library evolved on one model still helps a different one\. For each task we install the library that GPT\-5\.6\-luna evolved into Claude\-opus\-5, Claude\-sonnet\-5, and GPT\-5\.6\-sol, and no test data is used to pick it\. From the last block of[Table1](https://arxiv.org/html/2609.27277#S4.T1)and from[Fig\.3](https://arxiv.org/html/2609.27277#S4.F3)b\-c, we observe:
Obs\. ❸: A stronger model is not a substitute for an evolved library\.As shown in[Fig\.3](https://arxiv.org/html/2609.27277#S4.F3)b, without tools the three expensive models score no better than the cheapest one, with everything between 54\.6 and 57\.1\. Installing luna’s library lifts each of them on the ten\-task mean by\+10\.50\+10\.50,\+10\.65\+10\.65, and\+8\.27\+8\.27, and only one of the thirty task results is negative\. As a result, we can conclude that what they were missing was not a better model but the right tools, and that what the library carries is task knowledge rather than model\-specific habits\.
Obs\. ❹: The library gains more on models stronger than the one that grew it\.As shown in[Fig\.3](https://arxiv.org/html/2609.27277#S4.F3)c, 21 of the 30 task results sit above the diagonal, so the same library usually gains more on the stronger model than it did on luna\. On Anomaly the two Claude models gain\+26\.25\+26\.25and\+26\.50\+26\.50, against\+14\.69\+14\.69on luna itself\. The tools compute the same evidence in every case, and what differs is how deeply a model can exploit it\. Therefore, growing the tools need not happen on an expensive model at all\.
Figure 4:Ablations on TimeMQA Anomaly and TSAQA\-DT under two backbones\.\(a\): Test accuracy with one component removed, each bar rising from that run’s own no\-tool base \(blue\)\. A red cross marks accuracy the library lost, and a hatched bar marks a gate\-rejected round\.\(b\): Answers fixed and broken by the update, under the protected review and under unconditional overwrite of the same library\.\(c\): Sub\-task accuracy with and without the evolved tool, on two in\-scope sub\-tasks and one out\-of\-scope \(shaded\)\. Full numbers in Appendix[D](https://arxiv.org/html/2609.27277#A4)\.
### 4\.4Ablation and General Analysis
We first remove one component at a time, then open the deployed libraries and check how the gate behaves\. We ablate on two datasets, Anomaly and TSAQA\-DT, each under mini and terra, and the four cells were chosen before the runs\. Against the full method we compare six changes, named as in[Fig\.4](https://arxiv.org/html/2609.27277#S4.F4)\. \(1\)w/o cali\.drops the train\-only decision trees, which only the Anomaly libraries carry\. \(2\)w/o gateforce\-accepts the candidate library\. \(3\)w/o scoperemoves the structured scopes\. \(4\)w/ chronosand \(5\)w/ 21 toolsreplace the empty root by a pre\-installed forecasting tool and by the 21 expert\-curated tools of TimeART\. \(6\)w/o residual chainoverwrites the base answers unconditionally instead of reviewing them\. From[Figs\.4](https://arxiv.org/html/2609.27277#S4.F4)and[5](https://arxiv.org/html/2609.27277#S4.F5), we observe:
Obs\. ❺: Every component earns its place, and none can be removed for free\.The full method reaches the highest test accuracy in all four cells\. Calibration and the gate cost the most: without calibration both Anomaly runs deploy nothing, and without the gate the mini run deploys a library that ends1\.251\.25below its own base, where the full method ends16\.2516\.25above\. Removing structured scopes costs1\.001\.00to9\.259\.25points\. In[Fig\.4](https://arxiv.org/html/2609.27277#S4.F4)b, overwriting fixes 16 to 25 more answers per cell than reviewing does, and breaks far more: 59 versus 19, 32 versus 22, 26 versus 3, and 28 versus 10\. The chain adds no repairs itself, but keeps a good library from breaking answers that were already right\.
Obs\. ❻: A pre\-installed root is not a shortcut\.As shown in[Fig\.4](https://arxiv.org/html/2609.27277#S4.F4)a, both alternative roots end below the empty root in all four cells: the Chronos wrapper by3\.043\.04to18\.5018\.50points and the 21 expert\-curated tools of TimeART by8\.258\.25to21\.5021\.50\. The human library is the weaker start of the two, and evolution from it is rejected by the gate on both mini tasks, and on TSAQA\-DT under terra it ends0\.500\.50below its own base\. Tools written before deployment do not become the right tools when an agent evolves on top of them, which is theHuman–Agent Tool Misalignmentof Phenomenon ❶ appearing inside our own method\.
Obs\. ❼: The admitted tools are targeted and used exactly where they claim\.As illustrated in[Fig\.4](https://arxiv.org/html/2609.27277#S4.F4)c, on TimeMQA Classification the deployed tool computes exactly the band powers that the question’s criteria name, and the out\-of\-scope activity sub\-task stays at exactly 45\.0\. On TSAQA\-DT, the libraries evolved under all three backbones independently rediscover the same Fourier and wavelet family, so the measurement contract, not a lucky sample, determines what gets built\. 25 of the 29 deployed libraries are pure tools, and the other four admitted a textual patch, because the gate arbitrates by measured repair and harm rather than by artifact type\.
Figure 5:What the agent builds, and how the gate behaves\.\(a\): One question from the report\-only test trace: the tool the agent built, the numbers it returned, and the answer before and after\.\(b\): Gate calibration across all formal runs, validation lower bound against deployed test delta, with the report\-only diagnostic for rejected rounds\.\(c\): Admitted gain per round on a three\-round run \(terra, Anomaly\)\. Two further cases in Appendix[B](https://arxiv.org/html/2609.27277#A2)\.Obs\. ❽: The repairs are single missing measurements, not better reasoning\.As shown in[Fig\.5](https://arxiv.org/html/2609.27277#S4.F5)a, the base agent fails this question without computing anything, reading a global summary of a series whose overall spread looks ordinary\. The evolved tool supplies exactly one measurement and the answer flips\. The raw level gap between the two adjacent windows is0\.0250\.025, but standardized it is10σ10\\,\\sigma, and the variance ratio of0\.0001450\.000145says the second window is almost flat\. The two cases in Appendix[B](https://arxiv.org/html/2609.27277#A2)fail the same way: the base agent is not bad at arithmetic but simply never does it\.
Obs\. ❾: The gate errs visibly, in both directions, and the loop stops itself\.As shown in[Fig\.5](https://arxiv.org/html/2609.27277#S4.F5)b, every strict\-region acceptance \(δlb\>0\\delta\_\{\\mathrm\{lb\}\}\>0\) lands test\-positive\. The borderline band contains one false positive \(a T1 library at−1\.27\-1\.27, the only accepted library that lands test\-negative\), and the rejected region contains one conservative false negative \(a Merrill library whose report\-only diagnostic scored\+5\.00\+5\.00\)\. A rejected round rolls back cleanly and counts as zero\. As illustrated in[Fig\.5](https://arxiv.org/html/2609.27277#S4.F5)c, on a three\-round run, round 1 admits new tools \(validation gain\+18\.50\+18\.50\), round 2 admits only usage guidance \(\+2\.75\+2\.75\), and round 3 is rejected outright\. Every tool in the final library carries a round\-1 identifier, so the tools stabilize in one round and later rounds only teach how to use them\.
## 5Conclusion
This paper studies two failures of tool\-based time series agents: the tools humans supply are not the ones the agent needs at runtime, and an agent’s own revisions break answers that the final score hides\. To address them, we introduce TimeEvo, which turns diagnosed failures into evidence\-only tools and admits a candidate library only when a paired comparison shows it fixes more than it breaks, rolling the round back otherwise\. Experiments on ten time series QA tasks and three backbones show that TimeEvo, starting from an empty library, improves accuracy on every task and every backbone, and that a library grown on a cheap model still gains when it is installed into stronger ones\. An agent’s own failures are therefore a better specification for tools than human anticipation\.
## References
- Ansariet al\.\(2025\)A\. F\. Ansari, O\. Shchur, J\. Küken, A\. Auer, B\. Han, P\. Mercado, S\. S\. Rangapuram, H\. Shen, L\. Stella, X\. Zhang,et al\.Chronos\-2: from univariate to universal forecasting\.arXiv preprint arXiv:2510\.15821\.Cited by:[§2\.1](https://arxiv.org/html/2609.27277#S2.SS1.p1.1),[§3\.4](https://arxiv.org/html/2609.27277#S3.SS4.p1.2),[§4\.1](https://arxiv.org/html/2609.27277#S4.SS1.p3.1)\.
- Beqariet al\.\(2024\)J\. Beqari, J\. Powell, J\. Hurd, A\. L\. Potter, M\. McCarthy, D\. Srinivasan, D\. Wang, J\. Cranor, L\. Zhang, K\. Webster,et al\.A pilot study using machine learning algorithms and wearable technology for the early detection of postoperative complications after cardiothoracic surgery\.Annals of surgery281\(3\),pp\. 514\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p1.1)\.
- Caiet al\.\(2024\)Y\. Cai, A\. Choudhry, M\. Goswami, and A\. DubrawskiTimeseriesexam: a time series understanding exam\.arXiv preprint arXiv:2410\.14752\.Cited by:[Table 2](https://arxiv.org/html/2609.27277#A1.T2.2.6.2),[Appendix A](https://arxiv.org/html/2609.27277#A1.p3.1.1),[§4\.1](https://arxiv.org/html/2609.27277#S4.SS1.p1.1)\.
- Chenget al\.\(2025\)Z\. Cheng, J\. Yang, Y\. Song, D\. Cheng, G\. Yang, and B\. WangNeighbor\-enhanced graph pre\-training and prompt learning framework for fraud detection\.InProceedings of the 34th ACM International Conference on Information and Knowledge Management,pp\. 5617–5625\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p1.1)\.
- Daset al\.\(2023\)A\. Das, W\. Kong, R\. Sen, and Y\. ZhouA decoder\-only foundation model for time\-series forecasting\.arXiv preprint arXiv:2310\.10688\.Cited by:[§2\.1](https://arxiv.org/html/2609.27277#S2.SS1.p1.1)\.
- Dempsteret al\.\(2020\)A\. Dempster, F\. Petitjean, and G\. I\. WebbROCKET: exceptionally fast and accurate time series classification using random convolutional kernels: a\. dempster et al\.\.Data Mining and Knowledge Discovery34\(5\),pp\. 1454–1495\.Cited by:[§2\.1](https://arxiv.org/html/2609.27277#S2.SS1.p1.1)\.
- Donget al\.\(2026\)H\. Dong, Q\. Feng, K\. Jiang, H\. Ye, X\. Zhang, and G\. SongAgent\-valuebench: a comprehensive benchmark for evaluating agent values\.External Links:2605\.10365,[Link](https://arxiv.org/abs/2605.10365)Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p1.1)\.
- Donget al\.\(2025\)H\. Dong, W\. Zhu, G\. Song, and L\. WangAuroRA: breaking low\-rank bottleneck of lora with nonlinear mapping\.InAdvances in Neural Information Processing Systems,D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(Eds\.\),Vol\.38, Main Conference,pp\. 36929–36961\.External Links:[Document](https://dx.doi.org/10.52202/085713-1241),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/file/34aae5b7991a4fb4fe3081ca4f94549c-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p2.1)\.
- Fanet al\.\(2026a\)S\. Fan, N\. Elhendawy, J\. Sun, K\. Fang, K\. Zhang, Y\. Wang, and L\. ChengMOSAIC: module discovery via sparse additive identifiable causal learning for scientific time series\.arXiv preprint arXiv:2605\.05524\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p3.1)\.
- Fanet al\.\(2026b\)S\. Fan, K\. Zhang, and L\. ChengTrace: trajectory recovery for continuous mechanism evolution in causal representation learning\.arXiv preprint arXiv:2601\.21135\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p3.1)\.
- Fanget al\.\(2026\)Z\. Fang, D\. Wang, A\. L\. Potter, L\. Zhang, Y\. Zhang, B\. Rettner, A\. Zhu, M\. McCarthy, Q\. Guo, D\. Srinivasan,et al\.A novel wearables\-based sleep quality index to quantify postoperative sleep quality\.The Journal of Thoracic and Cardiovascular Surgery\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p3.1)\.
- Garza and Rosillo \(2025\)A\. Garza and R\. RosilloTimeCopilot\.arXiv preprint arXiv:2509\.00616\.Cited by:[§2\.1](https://arxiv.org/html/2609.27277#S2.SS1.p1.1)\.
- Haoet al\.\(2026\)H\. Hao, D\. Min, Z\. Zhang, Y\. Zhang, M\. Xu, Y\. Ge, and L\. ChengPOISE: position\-aware undetectable skill injection on llm agents\.arXiv preprint arXiv:2606\.07943\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p3.1)\.
- \[14\]Y\. Hu, J\. Yang, X\. Dai, W\. Cai, K\. Ding, Y\. Li, Q\. Liu, E\. Ma, Z\. Qu, Y\. Wang,et al\.The landscape of agentic time series systems: architectures, reliability, and frontiers\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p1.1)\.
- Huet al\.\(2026\)Y\. Hu, J\. Yang, T\. Zhou, P\. Liu, Y\. Tang, R\. Jin, and L\. SunBridging past and future: distribution\-aware alignment for time series forecasting\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 126148–126172\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p1.1)\.
- Jinet al\.\(2024\)M\. Jin, S\. Wang, L\. Ma, Z\. Chu, J\. Zhang, X\. Shi, P\. Chen, Y\. Liang, Y\. Li, S\. Pan,et al\.Time\-llm: time series forecasting by reprogramming large language models\.InInternational conference on learning representations,Vol\.2024,pp\. 23857–23880\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.27277#S2.SS1.p1.1)\.
- Jinget al\.\(2026\)B\. Jing, S\. Chen, L\. Zheng, B\. Liu, Z\. Li, J\. Zou, T\. Wei, Z\. Liu, Z\. Zeng, R\. Qiu,et al\.Tsaqa: time series analysis question and answering benchmark\.InProceedings of the Fifth Workshop on Generation, Evaluation and Metrics \(GEM\),pp\. 944–979\.Cited by:[Table 2](https://arxiv.org/html/2609.27277#A1.T2.2.11.2),[Appendix A](https://arxiv.org/html/2609.27277#A1.p7.1.1),[§4\.1](https://arxiv.org/html/2609.27277#S4.SS1.p1.1)\.
- Konget al\.\(2025\)Y\. Kong, Y\. Yang, Y\. Hwang, W\. Du, S\. Zohren, Z\. Wang, M\. Jin, and Q\. WenTime\-mqa: time series multi\-task question answering with context enhancement\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 29736–29753\.Cited by:[Table 2](https://arxiv.org/html/2609.27277#A1.T2.2.8.2),[Table 2](https://arxiv.org/html/2609.27277#A1.T2.2.9.2),[Appendix A](https://arxiv.org/html/2609.27277#A1.p5.1.1),[§1](https://arxiv.org/html/2609.27277#S1.p1.1),[§1](https://arxiv.org/html/2609.27277#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.27277#S4.SS1.p1.1)\.
- Konget al\.\(2026\)Y\. Kong, Q\. Yao, Y\. Nie, Y\. Li, Y\. Shao, S\. Zohren, A\. Vettoruzzo, J\. Vanschoren, M\. Jin, and Q\. WenTimeSage\-mt: a multi\-turn benchmark for evaluating agentic time series reasoning\.arXiv preprint arXiv:2606\.01498\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p1.1)\.
- Liuet al\.\(2025a\)C\. Liu, T\. Aksu, J\. Liu, X\. Liu, H\. Yan, Q\. Pham, S\. Savarese, D\. Sahoo, C\. Xiong, and J\. LiMoirai 2\.0: when less is more for time series forecasting\.arXiv preprint arXiv:2511\.11698\.Cited by:[§2\.1](https://arxiv.org/html/2609.27277#S2.SS1.p1.1)\.
- Liuet al\.\(2026\)H\. Liu, D\. Li, R\. Jiang, J\. Deng, W\. Ye, and Y\. SekimotoTimeClaw: a time\-series ai agent with exploratory execution learning\.arXiv preprint arXiv:2605\.10038\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.27277#S2.SS2.p1.1)\.
- Liuet al\.\(2025b\)P\. Liu, E\. Fons, S\. Vyetrenko, D\. Borrajo, V\. K\. Potluru, and M\. VelosoTs\-agent: a time series reasoning agent with iterative statistical insight gathering\.InFirst Workshop on Foundations of Reasoning in Language Models,Cited by:[§2\.1](https://arxiv.org/html/2609.27277#S2.SS1.p1.1)\.
- Liu and Paparrizos \(2024\)Q\. Liu and J\. PaparrizosThe elephant in the room: towards a reliable time\-series anomaly detection benchmark\.Advances in Neural Information Processing Systems37,pp\. 108231–108261\.Cited by:[§2\.1](https://arxiv.org/html/2609.27277#S2.SS1.p1.1)\.
- Liuet al\.\(2024\)Y\. Liu, T\. Hu, H\. Zhang, H\. Wu, S\. Wang, L\. Ma, and M\. LongItransformer: inverted transformers are effective for time series forecasting\.InInternational conference on learning representations,Vol\.2024,pp\. 11116–11140\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p1.1)\.
- Madaanet al\.\(2023\)A\. Madaan, N\. Tandon, P\. Gupta, S\. Hallinan, L\. Gao, S\. Wiegreffe, U\. Alon, N\. Dziri, S\. Prabhumoye, Y\. Yang,et al\.Self\-refine: iterative refinement with self\-feedback\.Advances in neural information processing systems36,pp\. 46534–46594\.Cited by:[item ❷](https://arxiv.org/html/2609.27277#S1.I1.ix2.p1.1),[§4\.1](https://arxiv.org/html/2609.27277#S4.SS1.p3.1)\.
- Merrillet al\.\(2024\)M\. A\. Merrill, M\. Tan, V\. Gupta, T\. Hartvigsen, and T\. AlthoffLanguage models still struggle to zero\-shot reason about time series\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 3512–3533\.Cited by:[Table 2](https://arxiv.org/html/2609.27277#A1.T2.2.7.2),[Appendix A](https://arxiv.org/html/2609.27277#A1.p4.1.1),[§4\.1](https://arxiv.org/html/2609.27277#S4.SS1.p1.1)\.
- Miaoet al\.\(2025\)C\. Miao, H\. P\. Zou, Y\. Li, Y\. Chen, Y\. Wang, F\. Wang, Y\. Li, W\. Yang, B\. He, X\. Zhang,et al\.Recode\-h: a benchmark for research code development with interactive human feedback\.arXiv preprint arXiv:2510\.06186\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p3.1)\.
- Shinnet al\.\(2023\)N\. Shinn, F\. Cassano, A\. Gopinath, K\. Narasimhan, and S\. YaoReflexion: language agents with verbal reinforcement learning\.Advances in neural information processing systems36,pp\. 8634–8652\.Cited by:[§2\.2](https://arxiv.org/html/2609.27277#S2.SS2.p1.1)\.
- Wanget al\.\(2023\)G\. Wang, Y\. Xie, Y\. Jiang, A\. Mandlekar, C\. Xiao, Y\. Zhu, L\. Fan, and A\. AnandkumarVoyager: an open\-ended embodied agent with large language models\.arXiv preprint arXiv:2305\.16291\.Cited by:[§2\.2](https://arxiv.org/html/2609.27277#S2.SS2.p1.1)\.
- Wanget al\.\(2025\)H\. Wang, L\. Pan, Y\. Shen, Z\. Chen, D\. Yang, Y\. Yang, S\. Zhang, X\. Liu, H\. Li, and D\. TaoFredf: learning to forecast in the frequency domain\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 6893–6922\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p1.1)\.
- Wenget al\.\(2026\)M\. Weng, D\. Cao, W\. Yang, Y\. Sharma, and Y\. LiuTemporalbench: a benchmark for evaluating llm\-based agents on contextual and event\-informed time series tasks\.InProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\. 2,pp\. 9997–10008\.Cited by:[Table 2](https://arxiv.org/html/2609.27277#A1.T2.2.2.2),[Table 2](https://arxiv.org/html/2609.27277#A1.T2.2.3.2),[Table 2](https://arxiv.org/html/2609.27277#A1.T2.2.4.2),[Table 2](https://arxiv.org/html/2609.27277#A1.T2.2.5.2),[Appendix A](https://arxiv.org/html/2609.27277#A1.p2.1.1),[item ❶](https://arxiv.org/html/2609.27277#S1.I1.ix1.p1.1),[§1](https://arxiv.org/html/2609.27277#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.27277#S4.SS1.p1.1)\.
- Wuet al\.\(2025\)R\. Wu, X\. Wang, J\. Mei, P\. Cai, D\. Fu, C\. Yang, L\. Wen, X\. Yang, Y\. Shen, Y\. Wang,et al\.Evolver: self\-evolving llm agents through an experience\-driven lifecycle\.arXiv preprint arXiv:2510\.16079\.Cited by:[§2\.2](https://arxiv.org/html/2609.27277#S2.SS2.p1.1)\.
- Wuet al\.\(2026\)X\. Wu, J\. Lu, Z\. Li, X\. Qiu, J\. Hu, C\. Guo, C\. S\. Jensen, and B\. YangTimeart: towards agentic time series reasoning via tool\-augmentation\.arXiv preprint arXiv:2601\.13653\.Cited by:[item ❶](https://arxiv.org/html/2609.27277#S1.I1.ix1.p1.1),[§1](https://arxiv.org/html/2609.27277#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.27277#S4.SS1.p3.1)\.
- Xieet al\.\(2024\)Z\. Xie, Z\. Li, X\. He, L\. Xu, X\. Wen, T\. Zhang, J\. Chen, R\. Shi, and D\. PeiChatts: aligning time series with llms via synthetic data for enhanced understanding and reasoning\.arXiv preprint arXiv:2412\.03104\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.27277#S2.SS1.p1.1)\.
- Yanget al\.\(2026a\)J\. Yang, Y\. Hu, Y\. Li, K\. Zhang, K\. Ding, and P\. S\. YuFrom observations to states: latent time series forecasting\.arXiv preprint arXiv:2602\.00297\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.27277#S2.SS1.p1.1)\.
- Yanget al\.\(2025a\)J\. Yang, Y\. Hu, K\. Zhang, L\. Niu, P\. S\. Yu, and K\. DingRevisiting multivariate time series forecasting with missing values\.arXiv preprint arXiv:2509\.23494\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p1.1)\.
- Yanget al\.\(2026b\)J\. Yang, K\. Zhang, G\. Zhang, P\. S\. Yu, and K\. DingGlocal information bottleneck for time series imputation\.Advances in Neural Information Processing Systems38,pp\. 104452–104484\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p1.1)\.
- Yanget al\.\(2025b\)J\. Yang, R\. Zhang, Z\. Cheng, D\. Cheng, G\. Yang, and B\. WangGrad: guided relation diffusion generation for graph augmentation in graph fraud detection\.InProceedings of the ACM on Web Conference 2025,pp\. 5308–5319\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p1.1)\.
- Yanget al\.\(2026c\)Y\. Yang, Z\. Gong, W\. Huang, Q\. Yang, Z\. Zhou, Z\. Huang, Y\. Li, X\. Gao, Q\. Dai, B\. Liu,et al\.Skillopt: executive strategy for self\-evolving agent skills\.arXiv preprint arXiv:2605\.23904\.Cited by:[§2\.2](https://arxiv.org/html/2609.27277#S2.SS2.p1.1)\.
- Yaoet al\.\(2022\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. CaoReact: synergizing reasoning and acting in language models\.arXiv preprint arXiv:2210\.03629\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p2.1)\.
- Yeet al\.\(2024\)W\. Ye, W\. Yang, D\. Cao, Y\. Zhang, L\. Tang, J\. Cai, and Y\. LiuTs\-reasoner: domain\-oriented time series inference agents for reasoning and automated analysis\.arXiv preprint arXiv:2410\.04047\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.27277#S2.SS1.p1.1)\.
- Yinet al\.\(2026\)Y\. Yin, Z\. Xiao, M\. Li, Y\. Liu, S\. Nan, Y\. He, R\. Wang, Z\. Zhang, Q\. Liao, and Y\. GuMmts\-bench: a comprehensive benchmark for time series understanding and reasoning\.arXiv preprint arXiv:2602\.08588\.Cited by:[Table 2](https://arxiv.org/html/2609.27277#A1.T2.2.10.2),[Appendix A](https://arxiv.org/html/2609.27277#A1.p6.1.1),[§4\.1](https://arxiv.org/html/2609.27277#S4.SS1.p1.1)\.
- Yuet al\.\(2026\)F\. Yu, T\. Feng, D\. Min, L\. Cheng, G\. Liu, and T\. ZhouTSRouter: dynamic modality\-model selection for time series reasoning\.arXiv preprint arXiv:2607\.08940\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p2.1)\.
- Zenget al\.\(2023\)A\. Zeng, M\. Chen, L\. Zhang, and Q\. XuAre transformers effective for time series forecasting?\.InProceedings of the AAAI conference on artificial intelligence,Vol\.37,pp\. 11121–11128\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p1.1)\.
- Zhanget al\.\(2026\)H\. Zhang, S\. Fan, H\. P\. Zou, Y\. Chen, Z\. Wang, J\. Zhou, C\. Li, W\. Huang, Y\. Yao, K\. Zheng,et al\.Coevoskills: self\-evolving agent skills via co\-evolutionary verification\.arXiv preprint arXiv:2604\.01687\.Cited by:[§2\.2](https://arxiv.org/html/2609.27277#S2.SS2.p1.1)\.
- Zhaoet al\.\(2025\)H\. Zhao, X\. Zhang, J\. Wei, Y\. Xu, Y\. He, S\. Sun, and C\. YouTimeseriesscientist: a general\-purpose ai agent for time series analysis\.arXiv preprint arXiv:2510\.01538\.Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p2.1)\.
- Zhuet al\.\(2026\)S\. Zhu, S\. Fan, X\. Wang, W\. Wu, K\. Zhou, and B\. HuangRSIAgent: autonomous exploration for recursive self\-improvement in new environments\.External Links:2609\.15364,[Link](https://arxiv.org/abs/2609.15364)Cited by:[§1](https://arxiv.org/html/2609.27277#S1.p2.1)\.
## Appendix ADataset Details
Table 2:The ten\-task suite: data sources and frozen train/validation/test sizes\. Merrill is downsampled from its official 5779/1000/1036 split, and every other task uses its full frozen pool\. Splits are grouped by source series, so questions from the same series never cross splits\.The ten tasks come from six public sources, and every source is adapted to one row contract: a question, a closed option set whose gold answer appears verbatim among the options, one or two primary series, and optional named covariate or candidate series\. This appendix describes each source and the split design, and[Table2](https://arxiv.org/html/2609.27277#A1.T2)lists the frozen sizes\.
TemporalBench \(T1–T4\)\([Weng et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib11)\)\.The four tiers share one source pool and are split by source sample, so all questions expanded from the same sample stay in the same split\. T1 asks property questions about the observed history \(trend, volatility, seasonality, anomaly\), with label spaces read from the source prompts rather than invented\. T2 and T4 are forecasting\-style multiple\-choice questions, distinguished only by whether event context is present, and T3 packs several sub\-questions per sample with clinical covariates such as temperature and respiratory rate\. All series are univariate with a median length of 300 points, and only the observed history is ever exposed, never another tier’s future targets\.
TimeSeriesExam\([Cai et al\., 2024](https://arxiv.org/html/2609.27277#bib.bib13)\)\.An exam of time series knowledge and the only task with two\-series questions \(99 single and 50 dual in the test split\)\. Its series are the longest in the suite \(median 1024 points\)\. The split groups all questions that share the exact same input series and stratifies by question type, difficulty, option count, and even the position of the correct option, using no model outputs\.
Merrill\([Merrill et al\., 2024](https://arxiv.org/html/2609.27277#bib.bib32)\)\.Scenario matching: every question shows one series and four candidate natural\-language scenario descriptions, under one shared question text\. We keep the official train/validation/test files and downsample them to 240/120/240 for cost\. Series lengths span 12 to 1460 points, the widest range in the suite\.
Time\-MQA \(Anomaly and Classification\)\([Kong et al\., 2025](https://arxiv.org/html/2609.27277#bib.bib14)\)\.The source embeds each series inside the question text\. We extract the longest numeric list into a machine\-readable series so tools can compute on it, and turn the closed\-set free\-text answer into explicit options\. From 37,000 source rows, one representative row is kept per unique series, and one group whose duplicate series carried contradictory labels is excluded entirely\. Anomaly is a two\-way normal\-versus\-anomaly task with the shortest series in the suite \(8 to 64 points, median 16\)\. Classification mixes a six\-way activity\-recognition half with a two\-way freeze\-of\-gait half, over accelerometer snippets of 9 to 30 points\.
MMTS Match\([Yin et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib33)\)\.Each question shows a target series and four candidate series and asks which candidate is most similar\. The four candidates are carried as named covariate series, and the task is kept as a four\-way choice rather than rewritten into easier pairwise comparisons\. The 400 source rows contain only 112 unique target series, so whole target groups are assigned to one split\.
TSAQA \(Data\-Transformation\)\([Jing et al\., 2026](https://arxiv.org/html/2609.27277#bib.bib34)\)\.Closed\-set questions about transformed series \(multiple\-choice and true\-or\-false halves\), with options lifted from the source text and one series per question\. Row identifiers encode the official split, so our frozen manifest never crosses the official train/validation/test boundaries\.
Leakage control and heterogeneity\.Across the suite, splits are grouped at the series level \(all questions sharing a series move together\), duplicate series are collapsed to one representative, contradictory groups are dropped, official split boundaries are preserved, and future windows are never exposed to the agent or its tools\. Series lengths span two orders of magnitude \(8 to 1460 points\) and option counts range from two to six, so absolute accuracies are not comparable across tasks\. All comparisons in this paper are therefore within\-task paired deltas\.
## Appendix BAdditional Case Studies
Figure 6:Two further repairs\.Left: the agent picked the Fourier spectrum that looked right, and the synthesized tool computed the real one, which matches the correct option to three decimals\.Right: the agent abstained because it never compared the two windows, and the tool measured the dispersion drop directly\.Both cases below are single questions from the report\-only test traces of accepted rounds, quoted with the tool’s own returned values, and both are drawn in[Fig\.6](https://arxiv.org/html/2609.27277#A2.F6)\. For each case we give the question, the answer before and after the round, the failure the planner diagnosed, the measurement contract it wrote, and the numbers the synthesized tool returned on that question\.
Comparing options by eye\.On TSAQA\-DT \(tsaqa\_test\_30686, terra\), the question asks which of the listed option vectors is the Fourier transform of the given series\. The base agent answers \(A\), whose leading coefficients\[0,23\.09,39\.60,13\.86,6\.90\]\[0,23\.09,39\.60,13\.86,6\.90\]are the largest and therefore look the most like a spectrum, and the gold answer is \(B\),\[0,18\.55,18\.96,12\.00,11\.07\]\[0,18\.55,18\.96,12\.00,11\.07\]\. The planner diagnoses the bucket as a candidate\-comparison error, namely that the solver never produces aligned quantitative comparisons between the requested transform and the candidates, and writes a contract for candidate\-wise normalized residuals after computing the requested representation and aligning each candidate by valid length and coefficient index\. The synthesized tooltransform\_reference\_pathsreturns a series length of 81 and a magnitude path of\[0\.0005,18\.5500,18\.9578,12\.0046,11\.0670,7\.5275,…\]\[0\.0005,18\.5500,18\.9578,12\.0046,11\.0670,7\.5275,\\ldots\], which matches option \(B\) to three decimals at every index, and the evolved agent answers \(B\)\.
Abstaining without computing\.On TemporalBench T2 \(tsb\_000032\_T2\_volatility\_change, mini\), the question gives a history window and asks whether volatility over the forecast horizon rises, falls, or cannot be determined\. The base agent answers \(D\) Uncertain, which is the fallback option rather than a wrong measurement, and the gold answer is \(B\) decreased\. The planner diagnoses that the agent never computed the cross\-window spread comparison the question needs, and writes a contract for a deterministic comparison of forecast\-horizon dispersion against historical dispersion\. The synthesized toolforecast\_horizon\_dispersion\_gapreturns a history MAD of10\.000010\.0000against a forecast MAD of4\.85444\.8544, a delta of−5\.1456\-5\.1456, a relative delta of−0\.5146\-0\.5146, and 48 points in each window, and the evolved agent answers \(B\)\.
Neither repair required better reasoning about time series\. In both cases the agent had already read the question correctly and simply had no number to read off, and one deterministic measurement was enough to settle the answer\.
## Appendix CFrozen Protocol and Round Procedure
This appendix records the complete frozen configuration and the step\-by\-step round procedure for reproducibility\.
### C\.1Round Procedure
1. 0\.Root library\.The root is configurable and immutable\.emptyinstalls no tool and is the frozen protocol used for every reported result, whilechronosinstalls a Chronos\-2 wrapper and is retained only as an ablation\.
2. 1\.Base pass\.The current chain answers the full training split under a semantic\-fingerprint cache\. This yields the base accuracy of the round and the error set that everything downstream is built from\.
3. 2\.Failure clustering\.Errors are partitioned by whether a tool fired on them, and each partition is clustered independently by a triage LM in two stages\. A clustering failure invalidates the round as measurement\-invalid, which is a rollback rather than a quality verdict\. - •Partition\.Thetool\_firedside collects questions where a tool ran and the answer was still wrong, which points at a defective or misused tool, and theno\_toolside collects questions no tool reached, which points at a coverage gap\. An empty root puts every error on the second side\. - •Define\.The stage samples at most 150 errors and returns at most eight categories, each carrying a short label and a criteria string, with at most two retries\. - •Assign\.Every error is sent in batches of 32 under a strict JSON contract that must cover each id in the batch, with at most two retries per batch and at most three passes over what remains\. Small batches are deliberate, since one long response that fails validation would otherwise void a large batch\. - •Resplit and drop\.A bucket larger than 64 errors is split once and not recursively\. Errors that stay unassigned, and errors in buckets with fewer than two members, are dropped\.
4. 3\.Two\-level planning\.A local planner diagnoses each bucket independently, and a global planner then coordinates the proposals before anything is synthesized\. - •Local planner\.It sees the bucket’s error dossier, at most 12 support examples, and at most 48 background examples, and returns one proposal card\. - •Descriptive fields\.The action \(create,refine,rescope, orretire\), the artifact kind, the target tool, the failure type, the root cause, and the focus\. These carry context only\. - •Contract fields\.Required measurement, required signature, success criterion, scope, anti\-scope, matcher, and an optional calibration request\. These are checked repeatedly downstream\. - •Global planner\.It merges proposals that ask for the same measurement and narrows scopes that overlap\. Merging incompatible contracts is rejected mechanically rather than discouraged in the prompt, and after two failed attempts the round falls back to the validated local plans as singleton groups and continues\.
5. 4\.Synthesis\.Each work group gets at mostKupdate=3K\_\{\\mathrm\{update\}\}\{=\}3sequential attempts\. - •What the synthesizer sees\.The current library, rendered separately, together with a digest carrying the bucket dossier and its support examples, the planner’s diagnosis, the contract, a compact timeline of previous failed attempts, and the numerically strongest previous attempt kept apart so that a weak later retry cannot displace it\. - •Hard constraints\.The prompt fixes a single\-function output with nowhileloops and a zero\-argument signature\. - •Checks before execution\.A single\-function check, a smoke call built from the signature, and a context\-selection check against the contract\. - •Calibration\.When a contract asks for a decision boundary, the tool may carry a shallow decision tree of kindbinary\_decision\_tree\_v1, fitted on training rows alone with validation and test never read and answers stored only as hashes\. The tree needs at least 16 samples per class and an internal 80/20 selection split grouped by input fingerprint, and it is accepted only at an internal balanced accuracy of at leastβ=0\.55\\beta=0\.55\. An accepted tree is AST\-compiled into the tool source, and a calibrated tool may not call the forecasting primitive\. - •Retry policy\.A contract\-violating artifact receives one targeted repair, an undeliverable input signature is adapted between retries, and two consecutive zero\-repair outcomes trigger a replan that may change the artifact kind, for example from a new tool to usage guidance\.
6. 5\.Candidate prefilter\.Each candidate is screened cheaply before any candidate reaches the round\-level gate\. - •Sampling\.The first 32 held\-out errors of the candidate’s own bucket, taken in a deterministic order that interleaves the members of a coordinated work group so that a bounded screen does not sample only the largest source bucket\. - •Attribution\.A newly created tool is first checked for signature deliverability on those questions, while lifecycle updates are attributed through action\-specific traces instead, because narrowing or rescoping can help precisely by preventing delivery\. - •Criterion\.At least one attributable repair, meaning a question fixed with the tool bound, executed, and visible in the trace, within coarse caps on harm and execution errors\. - •Stopping\.A repair rate of at least0\.400\.40stops the search early, and otherwise the best positive candidate after three attempts still reaches the gate\.
7. 6\.Whole\-library gate\.The current library and the candidate library each answer the full validation split, and only questions both arms completed are compared, so helped and harmed are counted on identical questions\. - •Statistic\.Withnnpaired questions the gate computesδ=\(h−m\)/n\\delta=\(h\-m\)/nandδlb=δ−1\.96max\(h\+m,1\)/n\\delta\_\{\\mathrm\{lb\}\}=\\delta\-1\.96\\sqrt\{\\max\(h\+m,1\)\}/n\. - •Strict path\.A library passes whenh\>0h\>0andδlb\>0\\delta\_\{\\mathrm\{lb\}\}\>0\. - •Single\-round path\.A single\-round run has no later round in which to accumulate evidence, so a supplementary rule admits a library whenδ\>0\\delta\>0, the net repair is at leastmax\(3,⌈0\.02n⌉\)\\max\(3,\\lceil 0\.02n\\rceil\), and the absolute harm rate is at most0\.100\.10\. - •Invalidity\.An execution error rate above2%2\\%renders the measurement invalid, which is separated from a quality failure\.
8. 7\.One\-shot attribution pruning\.On a failed gate with clearly harmful updates, all of them and their dependents are deleted in one transaction, the pruned library is re\-verified in exactly one more full paired run, and it is deployed only on strict improvement together with a passing gate, else the round rolls back\. There is no second prune\.
9. 8\.Chain write and report\-only test\.An accepted library first reconciles its declared scope fields against how its tools actually behaved, and because that edit changes the library it passes the same paired gate one final time\. The round then appends exactly one repair stage to the chain\. The test split is answered once per round for reporting only, and on a rollback the best failed candidate may receive one diagnostic test run that participates in no selection\.
### C\.2Frozen Configuration
[Table3](https://arxiv.org/html/2609.27277#A3.T3)lists every knob that was fixed before the first run and never touched again, so that a reader can reproduce a round without reading the code\. Three groups appear in the table\. The first fixes the models and the shape of a round: all four agent roles run on the same backbone, synthesis runs at a higher reasoning effort because it writes code, the root library is empty, and a run is a single round with at most three synthesis attempts per work group\. The second fixes the two levels of validation: the prefilter samples 32 held\-out errors and demands one attributable repair, while the gate compares the whole library on the full validation split atZ=1\.96Z=1\.96and admits either through the strict lower bound or through the single\-round rule with a net floor of three and a harm ceiling of0\.100\.10\. The third fixes how the agent answers: it consults a structured scope catalog before inspecting at most three tools, tools return numbers rather than verdicts, and no hand\-written prompt is added for any task\. Two entries deserve emphasis, because they are what makes a rejected round harmless:AUTO\_PRUNEis off so that only one bounded prune can run, andTEST\_REJECTED\_CANDIDATEis on so that a rolled\-back library still receives one diagnostic test run that participates in no selection\.
Table 3:The frozen protocol\. Every reported run uses these settings without exception\.
## Appendix DFull Ablation Results
### D\.1Component Ablations
[Table4](https://arxiv.org/html/2609.27277#A4.T4)is the numeric table behind[Fig\.4](https://arxiv.org/html/2609.27277#S4.F4)a, with one row per configuration on each of the four cells\. The two*Val*columns are what the gate actually saw before it made its decision: the paired delta on the validation split, its lower bound, and the helped and harmed counts behind them\. Reading them next to*Final*shows how well the gate’s own evidence predicted the deployed outcome, which is the calibration plotted in[Fig\.5](https://arxiv.org/html/2609.27277#S4.F5)b\. The*Rejected diagnostic*column is filled only for rounds the gate rejected\. Such a round deploys nothing, so its final accuracy equals its base, and the diagnostic reports what the rejected library would have scored had it been deployed, measured once and used in no decision\.*Final*carries the gain the round itself produced,*Gap to full*compares it against the full method on the same cell, and*Tools*gives the size of the deployed library, with*rejected*marking a round that shipped none\.
Table 4:Component ablations on TimeMQA Anomaly \(main\) and TSAQA\-DT \(reproduction\), each under two backbones\. Val columns show what the gate saw \(paired delta, lower bound, helped/harmed\)\. Base and Final are that run’s own test accuracy before and after the round, with the gain the round itself produced in parentheses, and the last column gives the gap to the full method\. A rejected library is not deployed, so its final accuracy equals its base and the diagnostic column reports what the rejected library would have scored\. Calibration is marked N/A on TSAQA\-DT, whose libraries carry no calibration trees\.
### D\.2Residual Chain against Overwrite
The residual chain is not a removed component but a second way of applying the same library, so[Table5](https://arxiv.org/html/2609.27277#A4.T5)reports it separately\. Both arms deploy the same tools on the same questions, and the only difference is whether an admitted tool reviews the base answer or replaces it outright\. The two arms are what[Fig\.4](https://arxiv.org/html/2609.27277#S4.F4)b draws\.
Table 5:The residual chain against unconditional overwrite\. Both arms deploy the same library on the same questions, and the only difference is whether the library reviews the base answer or replaces it outright\. Fixed counts previously wrong answers the update repaired, and harmed counts previously correct answers it broke\.
### D\.3Evolution Funnel across Tasks and Backbones
[Table6](https://arxiv.org/html/2609.27277#A4.T6)summarizes attrition across the one\-round process, and[Table7](https://arxiv.org/html/2609.27277#A4.T7)opens the full trace of one complete GPT\-5\.6\-luna run per task, one row per task\. Every field is read from that run’s sole round, with no cross\-run imputation and no missing fields\.
Across the 30 runs, 4,898 of 4,910 errors were assigned to a bucket: only 12 remained unassigned and none were discarded for an undersized bucket\. Coordination reduced 197 local proposals to 194 contracts, while nine of 203 failure buckets, spread across seven runs, produced no candidate\. The prefilter retained 227 of 483 synthesis attempts \(47\.0%\), after which 113 artifacts were deployed: 108 tools, five prompt patches, and no usage\-only artifacts\. The recorded final outcomes were 22 passes, seven passes after one\-shot pruning \(17 artifacts removed in total\), and one rollback\. Thus most cost occurs at executable synthesis, the prefilter, and the library gate rather than during clustering or coordination\.
Funnel width reflects error diversity rather than the eventual payoff\. For example, the T3 run has 367 errors, 11 buckets, 31 synthesis attempts, and six deployed artifacts, yet a test gain of\+7\.43\+7\.43pp, while the MMTS Match run has only 28 errors, three buckets, six attempts, and two deployed artifacts, yet gains\+18\.75\+18\.75pp\. More failures therefore create more search work, but not necessarily more useful headroom\.
Table 6:Aggregate one\-round evolution funnel over the 30 task–backbone runs\. Totals count events across runs, and medians and ranges are computed per run\. Yield compares each row with the stage above it\.Table 7:One\-round evolution funnel of one complete GPT\-5\.6\-luna run per task, recorded end to end for process analysis\.*Errors*is the base pass’s error count on the training split and*Tagged*the errors assigned to a bucket\.*Local*and*Contracts*are proposals before and after coordination,*Synth\.*counts every synthesis attempt,*Prefilter*the attempts that passed it, and*Deploy T/P*the deployed tools and prompt patches\.*Final gate*records a direct pass, a rollback, or the number of artifacts removed by the one\-shot prune before a passing re\-verification\. Each test delta is that run’s own paired delta rather than the number reported in[Table1](https://arxiv.org/html/2609.27277#S4.T1)\.相似文章
EvoTest:面向自我改进智能体系统的进化式测试时学习
EvoTest 引入了 J-TTL,一个衡量智能体测试时学习能力的基准,并提出了一个进化框架,其中 Actor 智能体玩游戏,而 Evolver 智能体在不进行微调的情况下迭代改进系统的提示、记忆和超参数。该方法在基于复杂文本的游戏中表现出优于基于反思和记忆的基线方法的性能。
EVOTS: 用于时间序列预测的进化Transformer搜索
提出了一种进化神经架构搜索框架(EvoTS),用于发现任务自适应的类Transformer模型,用于多变量时间序列预测。该方法使用模块化基因组表示,并在ETT基准数据集上取得了竞争性的性能。
EvoDS:具备技能学习与上下文管理的自演化自主数据科学智能体
EvoDS 是一款自演化自主数据科学智能体,通过强化学习驱动的技能获取与自适应上下文压缩进行改进,在基准测试上超越开源智能体 28.9%。
EvoMaster:构建可进化大规模自主科学智能体的基础框架
# 论文页面 - EvoMaster:构建可进化大规模自主科学智能体的基础框架 来源:[https://huggingface.co/papers/2604.17406](https://huggingface.co/papers/2604.17406) 作者:,,,,,,,,,,,,,,,,,,,,, ## 摘要 EvoMaster 是一个可扩展、自我进化的智能体框架,专为大规模科学发现设计,支持在实验周期中迭代优化假设并持续积累知识。大语言模型与智能体的融合正在催生“智能体科学”新时代。
FlowEvo:通过工作流与可执行技能的协同演化实现自演化智能体
FlowEvo是一个免训练框架,使得大语言模型智能体能够在推理时协同演化可复用技能和工作流,在ALFWorld、HumanEval和GSM8K等基准测试中实现了最先进的准确性和效率。