When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents
Summary
This paper frames recovery in LLM agent harnesses as a causal decision problem, separating rescue from harm outcomes, and introduces the Causal Intervention Router (CIR), which raises Qwen3-14B success on long-horizon ALFWorld tasks from 70.33% to 73.33% by selectively triggering environment refresh.
View Cached Full Text
Cached at: 10/02/26, 09:46 AM
# When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents
Source: [https://arxiv.org/html/2610.00372](https://arxiv.org/html/2610.00372)
Shuyao XiaoAffiliation:School of Artificial Intelligence, Beijing Normal UniversityAffiliation:Ke HoldingsEmail:[xiaoshuyao@mail\.bnu\.edu\.cn](mailto:)Xuan ChenAffiliation:Ke HoldingsKe ChaoAffiliation:School of Artificial Intelligence, Beijing Normal UniversityMing CuiAffiliation:Ke HoldingsFeifei QianAffiliation:School of Artificial Intelligence, Beijing Normal UniversityChaoyang MeiAffiliation:Ke HoldingsFanlin MengAffiliation:Ke HoldingsZiming YuAffiliation:School of Artificial Intelligence, Beijing Normal UniversityJunxi YinAffiliation:Ke Holdings
###### Abstract
Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors\. Yet recovery is usually judged only by average task success\. This hides an important tension\. The same operation can rescue a failing trajectory or disrupt one that would otherwise succeed\. We frame recovery as a causal decision problem\. Starting from the same execution state, we compare what happens with and without recovery, separate rescue from harm, and study how the value of recovery changes over time\. We then introduce the Causal Intervention Router \(CIR\), a lightweight policy that uses information available before recovery to decide when intervention is worthwhile\. On long\-horizon ALFWorld tasks with Qwen3\-14B, CIR raises success from 70\.33% to 73\.33%, a gain of3\.003\.00percentage points\. It leaves all evaluated trajectories with correct observations untouched\. Additional controls show that the benefit of recovery cannot be explained solely by the new observation returned by the environment\. These results provide a practical way to evaluate recovery and apply it selectively\.
## 1Introduction
Large language model agents do not act alone\. An external harness passes observations between the model and its environment, maintains context, and coordinates long sequences of actions\([Packer et al\., 2023](https://arxiv.org/html/2610.00372#bib.bib1);[Wu et al\., 2024a](https://arxiv.org/html/2610.00372#bib.bib2);[Yang et al\., 2024](https://arxiv.org/html/2610.00372#bib.bib3);[Yao et al\., 2026](https://arxiv.org/html/2610.00372#bib.bib27)\)\. If an observation is outdated or missing, the agent may act on the wrong information, and the error can spread through later steps\([Yao et al\., 2023](https://arxiv.org/html/2610.00372#bib.bib5);[Shah, 2026](https://arxiv.org/html/2610.00372#bib.bib6);[Ma et al\., 2026](https://arxiv.org/html/2610.00372#bib.bib29)\)\. A harness can respond by querying the environment again, which is called*refresh*\.
Refresh is not always helpful\. It may turn a failure into a success, but it may also derail a trajectory that would otherwise finish the task\. We call these outcomes*rescue*and*harm*\. Existing methods improve recovery through reflection, external feedback, or search\([Shinn et al\., 2023](https://arxiv.org/html/2610.00372#bib.bib9);[Madaan et al\., 2023](https://arxiv.org/html/2610.00372#bib.bib10);[Gou et al\., 2024](https://arxiv.org/html/2610.00372#bib.bib11);[Zhou et al\., 2024](https://arxiv.org/html/2610.00372#bib.bib12)\)\. Evaluations, however, usually report an overall success rate or identify where failures occur\([Ma et al\., 2024](https://arxiv.org/html/2610.00372#bib.bib14);[Cemri et al\., 2025](https://arxiv.org/html/2610.00372#bib.bib4);[Zhang et al\., 2025](https://arxiv.org/html/2610.00372#bib.bib15)\)\. These summaries do not show which trajectories recovery changes\. The same average gain can hide very different numbers of rescues and harms\.
This creates two challenges\. First, an ordinary run shows only one outcome\. The agent either recovers or continues without recovery\([Holland, 1986](https://arxiv.org/html/2610.00372#bib.bib7)\)\. Without the alternative outcome, we cannot tell whether recovery caused the final result\. Second, timing matters\. After an error, later actions and feedback change the situation\. An intervention that helps now may be ineffective or harmful later\. A useful evaluation must compare both choices from the same state and repeat that comparison at different times\.
Together, these challenges expose a gap between evaluating a recovery operation and deciding when to use it\. A positive average effect does not imply that recovery is safe to apply broadly\. The same average can result from a few reliable rescues or from many rescues offset by many harms\. A modest average can also hide a large benefit in a recognizable subset of states\. The central question is whether refresh works for a given trajectory and when its expected benefit outweighs its risk\.
Figure[1](https://arxiv.org/html/2610.00372#S1.F1)shows our approach\. We first introduce controlled observation errors\. From the same execution point, we create two continuations, one with refresh and one without\. This paired comparison reveals whether refresh rescues the task, harms it, or leaves the outcome unchanged\. We also compare refresh variants to identify what drives the effect\. We then use these paired outcomes to train the Causal Intervention Router \(CIR\)\. CIR decides whether to refresh using information available at decision time\. It also limits harm to trajectories with correct observations\.
The paired design serves two purposes\. As an evaluation tool, it measures the effect of recovery while holding the earlier trajectory fixed\. As a source of supervision, it reveals the states in which recovery changes the final outcome\. The same paired outcomes used for evaluation also train CIR\. At test time, it uses only information available in a live run\.
Experiments on long\-horizon ALFWorld tasks show that the value of refresh varies across errors, intervention times, and trajectories\. They also show that a newly returned observation is not the sole source of the benefit\. CIR turns this evidence into a selective policy\. For Qwen3\-14B, it improves success from 70\.33% to 73\.33% and does not refresh any evaluated clean trajectory\. Our main contributions are as follows\.
1. 1\.We introduce a paired framework for evaluating harness recovery\. By continuing from the same pre\-recovery state with and without an intervention, the framework attributes changes in the final outcome to recovery and separates rescue from harm\. Comparisons across times and recovery variants show when refresh helps and which parts of the operation matter\.
2. 2\.We introduce CIR, which uses the paired evidence to decide when recovery is worth the risk\. It requires no retraining of the underlying agent and uses only information available before intervention\. On Qwen3\-14B, CIR improves overall success by3\.003\.00percentage points\. The gain reaches9\.339\.33points when the agent receives a more severely outdated observation\. All evaluated clean trajectories remain untouched\.
Figure 1:Overview of our framework\. \(A\) We introduce controlled observation errors and try recovery at different times\. \(B\) From the same execution point, paired runs with and without refresh reveal rescue, harm, and unchanged outcomes\. \(C\) CIR uses information available before intervention to decide whether the expected benefit of refresh outweighs its risk\.
## 2Related Work
#### Agent Harnesses and Recovery Mechanisms\.
Agent harnesses organize interactions among models, tools, and environments, so their design can directly affect task performance\([Yao et al\., 2026](https://arxiv.org/html/2610.00372#bib.bib27)\)\. AutoGen\([Wu et al\., 2024a](https://arxiv.org/html/2610.00372#bib.bib2)\)provides composable conversation patterns, SWE\-agent\([Yang et al\., 2024](https://arxiv.org/html/2610.00372#bib.bib3)\)studies agent and computer interfaces, and StateFlow\([Wu et al\., 2024b](https://arxiv.org/html/2610.00372#bib.bib8)\)uses explicit states to manage long tasks\. Recovery methods revise execution through reflection, feedback, or search\([Li et al\., 2026](https://arxiv.org/html/2610.00372#bib.bib30)\)\. Reflexion\([Shinn et al\., 2023](https://arxiv.org/html/2610.00372#bib.bib9)\)learns from task feedback, Self\-Refine\([Madaan et al\., 2023](https://arxiv.org/html/2610.00372#bib.bib10)\)revises outputs iteratively, CRITIC\([Gou et al\., 2024](https://arxiv.org/html/2610.00372#bib.bib11)\)verifies outputs with external tools, and LATS\([Zhou et al\., 2024](https://arxiv.org/html/2610.00372#bib.bib12)\)combines feedback with tree search\. These methods improve how agents recover, but their aggregate evaluations do not identify the trajectories changed by recovery or the successful trajectories that recovery disrupts\. This gap matters because self\-correction without reliable feedback can degrade an initially correct response\([Huang et al\., 2024](https://arxiv.org/html/2610.00372#bib.bib13)\)\. We instead measure rescue and harm from the same starting state and use those outcomes to learn when recovery should be applied\.
#### Agent Failure Diagnosis and Causal Intervention\.
Process\-level evaluation studies where and why agent failures occur\([Barke et al\., 2026](https://arxiv.org/html/2610.00372#bib.bib28)\)\. AgentBoard\([Ma et al\., 2024](https://arxiv.org/html/2610.00372#bib.bib14)\)tracks fine\-grained progress, MAST\([Cemri et al\., 2025](https://arxiv.org/html/2610.00372#bib.bib4)\)provides a failure taxonomy for multi\-agent systems, Who&When\([Zhang et al\., 2025](https://arxiv.org/html/2610.00372#bib.bib15)\)identifies responsible agents and steps, and CatchBench\([Zhao et al\., 2026](https://arxiv.org/html/2610.00372#bib.bib16)\)audits failures from configurations, live prefixes, and completed traces\. Counterfactual approaches move from describing failures toward tracing or repairing them\. In memory\-augmented dialogue,[Xiao et al\. \(2026\)](https://arxiv.org/html/2610.00372#bib.bib31)use four counterfactual trajectories to separate error propagation through memory updates and subsequent questions\. Causal Agent Replay\([Shah, 2026](https://arxiv.org/html/2610.00372#bib.bib6)\)estimates the contribution of individual steps, CausalFlow\([Bonagiri et al\., 2026](https://arxiv.org/html/2610.00372#bib.bib17)\)generates targeted repairs, HarnessFix\([Chen et al\., 2026](https://arxiv.org/html/2610.00372#bib.bib18)\)diagnoses and patches harness flaws, and DoVer\([Ma et al\., 2025](https://arxiv.org/html/2610.00372#bib.bib19)\)tests debugging hypotheses through intervention\. These methods localize causes or repair observed failures\. They do not directly answer our question\. From the same current state, will a given recovery operation improve or damage the final outcome? When should it be applied? We answer these questions with paired runs and a policy that limits observed harm on clean trajectories\.
This distinction separates our work from both failure localization and repair generation\. We keep the recovery operation fixed and study its causal value across states and times\. This makes rescue and harm directly comparable and allows the same evidence used for evaluation to train a deployment\-time decision rule\.
## 3Paired Counterfactual Evaluation of Recovery
An aggregate success rate cannot tell us whether recovery changes a particular trajectory\. We therefore compare recovery with continued execution from the same task state\. Controlled changes to the observation make the error explicit, while paired runs isolate the recovery decision and reveal both rescue and harm \(Figure[1](https://arxiv.org/html/2610.00372#S1.F1)\(A–B\)\)\.
### 3\.1Perturbation and Recovery Setting
The unit of evaluation is a task instance, denoted byii\. At a predefined point in the run, the harness keeps the environment unchanged and modifies only the observation shown to the agent\. The observation conditioneeis clean, stale, or missing\. A clean observation preserves the current information\. A stale observation replaces it with information from an earlier state, and a missing observation removes its content\. In this way, we change the information available to the agent without changing the underlying task\.
The recovery delayddis the number of environment actions between the modified observation and the recovery operationmm\. A refresh queries the environment again and gives the agent the current observation\. Replanning instead asks the agent to reconsider its next actions using information already in context\. The triple\(e,d,m\)\(e,d,m\)therefore specifies what information the agent receives, when recovery occurs, and which recovery operation is used\. Section[5\.1](https://arxiv.org/html/2610.00372#S5.SS1)gives the concrete settings\.
### 3\.2Counterfactual Comparison
For each taskiiand observation conditionee, execution continues until delaydd\. We then branch the run from exactly the same state\. One branch continues without recovery\. The other applies operationmm\. The branches share the environment state and the full execution history before this decision\.
Because the two branches use the same state, history, model, and decoding settings, their difference is attributable to the recovery decision for the states studied here\.
Using potential\-outcome notation\([Rubin, 1974](https://arxiv.org/html/2610.00372#bib.bib20)\), letYi0\(e\)Y\_\{i\}^\{0\}\(e\)andYim\(e,d\)Y\_\{i\}^\{m\}\(e,d\)denote the binary task outcomes without recovery and with operationmmat delaydd, respectively\. Under deterministic replay, the no\-recovery continuation is shared across delay comparisons, so we suppressddinYi0\(e\)Y\_\{i\}^\{0\}\(e\)\. The instance\-level effect of recovery is
τim\(e,d\)=Yim\(e,d\)−Yi0\(e\)\.\\tau\_\{i\}^\{m\}\(e,d\)=Y\_\{i\}^\{m\}\(e,d\)\-Y\_\{i\}^\{0\}\(e\)\.\(1\)
### 3\.3Rescue and Harm
We distinguish two directions of recovery effects\. A*rescue*occurs when recovery turns a failure into a success, whereas*harm*occurs when recovery turns a success into a failure\.
Let𝟏\[⋅\]\\mathbf\{1\}\[\\cdot\]denote the indicator function, which equals 1 when the enclosed condition holds and 0 otherwise\. The rescue indicatorℛim\(e,d\)\\mathcal\{R\}\_\{i\}^\{m\}\(e,d\)and harm indicatorℋim\(e,d\)\\mathcal\{H\}\_\{i\}^\{m\}\(e,d\)for instanceiiare defined as
ℛim\(e,d\)=\[Yi0\(e\)=0,Yim\(e,d\)=1\],ℋim\(e,d\)=\[Yi0\(e\)=1,Yim\(e,d\)=0\]\.\\mathcal\{R\}\_\{i\}^\{m\}\(e,d\)=\\mathbf\{1\}\\\!\\left\[Y\_\{i\}^\{0\}\(e\)=0,\\;Y\_\{i\}^\{m\}\(e,d\)=1\\right\],\\mathcal\{H\}\_\{i\}^\{m\}\(e,d\)=\\mathbf\{1\}\\\!\\left\[Y\_\{i\}^\{0\}\(e\)=1,\\;Y\_\{i\}^\{m\}\(e,d\)=0\\right\]\.\(2\)If both continuations succeed, recovery preserves success\. If both fail, recovery leaves the failure unchanged\. Together with rescue and harm, these cases form the four possible outcomes of a paired comparison\. The instance\-level recovery effect can therefore be written as
τim\(e,d\)=ℛim\(e,d\)−ℋim\(e,d\)\.\\tau\_\{i\}^\{m\}\(e,d\)=\\mathcal\{R\}\_\{i\}^\{m\}\(e,d\)\-\\mathcal\{H\}\_\{i\}^\{m\}\(e,d\)\.\(3\)The average recovery effect equals the rescue rate minus the harm rate\. The same net effect can arise from rare changes or from frequent changes in both directions\. The average alone cannot distinguish these cases\. The paired outcomes also provide supervision for learning when recovery is more likely to rescue than harm a trajectory\.
This view changes what counts as a successful recovery method\. A method should not be judged only by how often it succeeds after intervention, because many of those trajectories may already have succeeded without intervention\. Instead, evaluation should ask whether the method changes the outcome, in which direction, and at what point in the trajectory\. The paired design answers all three questions with the same set of outcomes\.
## 4From Evaluation to Selective Recovery
Paired evaluation tells us whether recovery helped only after both branches have finished\. In a live run, the harness must decide before either outcome is known\. CIR estimates the value of recovery from the current state\. It refreshes only when the expected benefit justifies the risk\. A refresh obtains a new observation from the environment and adds it to the context\. CIR refreshes at most once per trajectory \(Figure[1](https://arxiv.org/html/2610.00372#S1.F1)\(C\)\)\.
CIR makes a simple distinction\. Detecting that something looks wrong is different from predicting that a particular response will help\. An anomaly detector addresses the first question\. A recovery policy must also answer the second\. Some erroneous trajectories are not rescued by refresh, and some apparently normal trajectories can be harmed by it\. CIR combines error detection with separate predictions of task success under refresh and continued execution\.
### 4\.1Decision Objective
LetXi,dX\_\{i,d\}denote the information available before recovery at delaydd\. The decisionπ\(Xi,d\)∈\{0,1\}\\pi\(X\_\{i,d\}\)\\in\\\{0,1\\\}indicates whether to refresh or continue for now\. LetYiπ\(e\)Y\_\{i\}^\{\\pi\}\(e\)denote the final binary outcome under this sequential policy\. If no refresh is triggered,Yiπ\(e\)=Yi0\(e\)Y\_\{i\}^\{\\pi\}\(e\)=Y\_\{i\}^\{0\}\(e\)\. Otherwise, it equals the outcome after refresh at the policy\-selected delay\.
Letϵ∈\[0,1\]\\epsilon\\in\[0,1\]denote the allowed harm rate on clean trajectories, and let𝔼\[⋅\]\\mathbb\{E\}\[\\cdot\]denote expectation over task instances under the specified observation condition\. We denote an optimal policy byπ⋆\\pi^\{\\star\}and formulate selective recovery as
π⋆\\displaystyle\\pi^\{\\star\}=argmaxπ𝔼\[Yiπ\(e\)−Yi0\(e\)∣e≠clean\],\\displaystyle=\\arg\\max\_\{\\pi\}\\mathbb\{E\}\\\!\\left\[Y\_\{i\}^\{\\pi\}\(e\)\-Y\_\{i\}^\{0\}\(e\)\\mid e\\neq\\mathrm\{clean\}\\right\],\(4\)subject to\\displaystyle\\text\{subject to\}𝔼\[Yi0\(e\)\(1−Yiπ\(e\)\)∣e=clean\]≤ϵ\.\\displaystyle\\mathbb\{E\}\\\!\\left\[Y\_\{i\}^\{0\}\(e\)\\bigl\(1\-Y\_\{i\}^\{\\pi\}\(e\)\\bigr\)\\mid e=\\mathrm\{clean\}\\right\]\\leq\\epsilon\.The objective rewards improvement when the observation is wrong, relative to never recovering\. Maximizing gain alone could encourage frequent refreshes that disrupt correct executions\. The second line limits this risk\. Its value is the fraction of clean trajectories that succeed without recovery but fail under the policy\. At runtime, neither possible outcome is known\. CIR must therefore estimate the value of refresh from information available before the decision\.
### 4\.2Estimating Recovery Value
Detecting a suspicious observation is not enough\. An error may exist even when refresh would not improve the final outcome\. CIR separates two questions\. Did the agent receive a faulty observation? How likely is the task to succeed with or without refresh? It answers them using execution behavior, consistency between actions and observations, task progress, and contextual information\.
Letxxdenote a value ofXi,dX\_\{i,d\}, and letPr\(⋅∣⋅\)\\Pr\(\\cdot\\mid\\cdot\)denote conditional probability\. We useperr\(x\)p\_\{\\mathrm\{err\}\}\(x\)for the probability that the trajectory has received an erroneous observation,p0\(x\)p\_\{0\}\(x\)for the probability of success without subsequent recovery, andp1\(x\)p\_\{1\}\(x\)for the probability of success after immediate refresh\. Denoting refresh byref\\mathrm\{ref\}and its final outcome at delayddbyYiref\(e,d\)Y\_\{i\}^\{\\mathrm\{ref\}\}\(e,d\), we define
perr\(x\)\\displaystyle p\_\{\\mathrm\{err\}\}\(x\)=Pr\(e≠clean∣Xi,d=x\),\\displaystyle=\\Pr\\\!\\left\(e\\neq\\mathrm\{clean\}\\mid X\_\{i,d\}=x\\right\),\(5\)p0\(x\)\\displaystyle p\_\{0\}\(x\)=Pr\(Yi0\(e\)=1∣Xi,d=x\),\\displaystyle=\\Pr\\\!\\left\(Y\_\{i\}^\{0\}\(e\)=1\\mid X\_\{i,d\}=x\\right\),p1\(x\)\\displaystyle p\_\{1\}\(x\)=Pr\(Yiref\(e,d\)=1∣Xi,d=x\)\.\\displaystyle=\\Pr\\\!\\left\(Y\_\{i\}^\{\\mathrm\{ref\}\}\(e,d\)=1\\mid X\_\{i,d\}=x\\right\)\.CIR estimates these quantities with three models supervised by the observation condition, the no\-recovery outcome, and the refresh outcome in the paired data from Section[3](https://arxiv.org/html/2610.00372#S3)\. The two outcome models follow a T\-learner\-style design and estimate the two decision outcomes separately\([Künzel et al\., 2019](https://arxiv.org/html/2610.00372#bib.bib22)\)\. This separation prevents high anomaly probability from being treated as evidence that refresh is beneficial\. For brevity, we use the same notation for the probabilities and their estimates\.
To account for potential rescue and harm, we construct a rescue scoresR\(x\)s\_\{\\mathrm\{R\}\}\(x\)and a harm scoresH\(x\)s\_\{\\mathrm\{H\}\}\(x\)\.
sR\(x\)=\(1−p0\(x\)\)p1\(x\),sH\(x\)=p0\(x\)\(1−p1\(x\)\)\.s\_\{\\mathrm\{R\}\}\(x\)=\\bigl\(1\-p\_\{0\}\(x\)\\bigr\)p\_\{1\}\(x\),\\qquad s\_\{\\mathrm\{H\}\}\(x\)=p\_\{0\}\(x\)\\bigl\(1\-p\_\{1\}\(x\)\\bigr\)\.\(6\)The rescue score is high when continued execution is likely to fail and refresh is likely to succeed\. The harm score is high when continued execution is likely to succeed and refresh is likely to fail\. These scores help identify states in which refresh is more promising\. The marginal success probabilities do not determine the probability that the same task is rescued or harmed\([Tian and Pearl, 2000](https://arxiv.org/html/2610.00372#bib.bib23)\)\. We therefore do not interpret these scores as rescue or harm probabilities\.
Letλ\>0\\lambda\>0denote the harm penalty\. Letuλ\(x\)u\_\{\\lambda\}\(x\)denote the utility obtained by subtracting the weighted harm score from the rescue score\.
uλ\(x\)=sR\(x\)−λsH\(x\)\.u\_\{\\lambda\}\(x\)=s\_\{\\mathrm\{R\}\}\(x\)\-\\lambda s\_\{\\mathrm\{H\}\}\(x\)\.\(7\)Whenλ=1\\lambda=1, this expression reduces top1\(x\)−p0\(x\)p\_\{1\}\(x\)\-p\_\{0\}\(x\), the difference in predicted success probabilities between refresh and no recovery\. Increasingλ\\lambdalowers the utility of states with high harm scores, making refresh less likely to be triggered in those states\. Thus,λ\\lambdacontrols the weight assigned to harm in each decision, whereasϵ\\epsilonlimits the fraction of all clean trajectories that the policy turns from success into failure\.
### 4\.3Dynamic Recovery Decisions
The value of recovery can change as the agent takes more actions and receives more feedback\. CIR therefore reconsiders refresh at several candidate times instead of committing to one fixed delay\. At each time, it compares an immediate refresh with continuing without later recovery\. If it chooses to continue, it can reconsider after the state has changed\.
Letpfail\(x\)=1−p0\(x\)p\_\{\\mathrm\{fail\}\}\(x\)=1\-p\_\{0\}\(x\)denote the predicted probability of failure without subsequent recovery\. Letα\\alpha,β\\beta, andγ\\gammadenote the thresholds for the error probability, failure probability, and utility score\. Letdmind\_\{\\min\}denote the earliest allowed refresh delay\. The indicator function𝟏\[⋅\]\\mathbf\{1\}\[\\cdot\]was defined in Section[3](https://arxiv.org/html/2610.00372#S3)\. The decision rule before any refresh has occurred is
π\(Xi,d\)=\[d≥dmin,perr\(Xi,d\)≥α,pfail\(Xi,d\)≥β,uλ\(Xi,d\)≥γ\]\.\\pi\(X\_\{i,d\}\)=\\mathbf\{1\}\\\!\\left\[\\begin\{aligned\} &d\\geq d\_\{\\min\},\\\\ &p\_\{\\mathrm\{err\}\}\(X\_\{i,d\}\)\\geq\\alpha,\\\\ &p\_\{\\mathrm\{fail\}\}\(X\_\{i,d\}\)\\geq\\beta,\\\\ &u\_\{\\lambda\}\(X\_\{i,d\}\)\\geq\\gamma\\end\{aligned\}\\right\]\.\(8\)CIR checks candidate times in chronological order and refreshes at the first time when all conditions are satisfied\. No further recovery is triggered afterward\. If no candidate time satisfies the conditions, the task proceeds without recovery\.
We select the decision thresholds on development data\. Among the settings whose observed harm rate on clean trajectories does not exceedϵ\\epsilon, we choose the one with the largest net gain under erroneous observations\. We then apply that fixed policy to new tasks\. This procedure is related to empirical welfare maximization\([Kitagawa and Tetenov, 2018](https://arxiv.org/html/2610.00372#bib.bib21)\), with paired recovery outcomes defining the benefit and the clean\-trajectory results defining the risk limit\.
## 5Experiments
Our experiments answer three questions\. When does refresh help or harm? Which parts of refresh produce the benefit? Can CIR use this evidence to recover selectively on new tasks? We first describe the common setup and then answer each question\.
### 5\.1Experimental Setup
#### Task, agent, and perturbations\.
We use ALFWorld\([Shridhar et al\., 2021](https://arxiv.org/html/2610.00372#bib.bib24)\), where an agent completes long household tasks through text actions and observations\. Our agent is Qwen3\-14B\([Yang et al\., 2025](https://arxiv.org/html/2610.00372#bib.bib25)\), used with a fixed ReAct\-style harness\([Yao et al\., 2023](https://arxiv.org/html/2610.00372#bib.bib5)\)and greedy decoding\. Each run is limited to 49 model turns and 64 environment actions\. We generated the agent trajectories on eight NVIDIA H200 GPUs\. We first run each task without an injected error or recovery\. We call this the*factual trajectory*\. The observation is modified after 25% of the environment actions in that trajectory\.
The primary perturbation is a two\-step stale observation, which replaces the current observation with the observation from two environment actions earlier\. Additional analyses use a one\-step stale observation from one environment action earlier and a missing observation whose content is replaced by a fixed unavailable\-observation message\. A refresh executes an additionallookaction and appends the returned current observation to the transcript\. We test recovery delaysd∈\{0,1,2,4\}d\\in\\\{0,1,2,4\\\}\.
#### Paired runs and data partitions\.
Each comparison starts two runs from the same environment state and model input, with identical model and decoding settings\. The runs differ only in whether the assigned recovery operation is applied\. A task is*prefix\-feasible*if it reaches the chosen intervention point and contains enough earlier history to construct the required stale observation\. We also verify that the two runs have identical histories before they branch\.
Each cohort serves a distinct role\. The primary and independent cohorts test whether recovery effects replicate\. The mechanism cohort isolates the components of refresh\. The evaluation cohort tests CIR\. Some analyses focus on tasks whose factual trajectory succeeds\. To check that the conclusions do not depend on this filter, we also report complete cohorts that retain every prefix\-feasible task\. This gives 106 tasks in the primary cohort and 100 separately executed tasks in the independent cohort\. These cohorts cover the available prefix\-feasible tasks, not the full ALFWorld distribution\. The labels*factual success*and*factual failure*describe the unmodified run\. Rescue and harm always compare recovery with no recovery under the same observation condition\.
The primary factual\-success analysis retains 93 of the 106valid\_unseentasks, while the independent cohort contributes 89 factual\-success tasks\. Mechanism controls use a separate 74\-task set, of which 67 pass the paired\-analysis protocol checks detailed in Appendix[A](https://arxiv.org/html/2610.00372#A1)\.
CIR is fitted on the 93 primary tasks and evaluated on a separate set of 76valid\_seentasks\. The namesvalid\_seenandvalid\_unseenrefer to ALFWorld splits\. No task appears in both fitting and evaluation\. Of the evaluation tasks, 75 are prefix\-feasible\. Each is tested under four observation conditions, giving 300 evaluation episodes\. We apply the fitted models and decision thresholds without further tuning\. All tasks remain in the denominator at every delay\. If a task has already ended, we carry its final outcome forward and do not apply recovery\.
#### Estimation and metrics\.
Our main quantity is the paired difference in task success between recovery and no recovery, measured in percentage points \(pp\)\. At delaydd, we also compare this effect between two\-step stale and clean observations\. We report the paired outcome counts from Section[3\.3](https://arxiv.org/html/2610.00372#S3.SS3)\. Confidence intervals use a task\-level cluster bootstrap\. This keeps all conditions from the same task together in each resample\([Efron, 1979](https://arxiv.org/html/2610.00372#bib.bib26)\)\. For learned policies, we additionally report how often the policy intervenes and how often it harms a clean trajectory\.
#### CIR fitting and comparison policies\.
CIR fits threeℓ2\\ell\_\{2\}\-regularized logistic models withC=0\.1C=0\.1to estimate the observation\-error probability and the two outcome probabilities in Eq\. equation[5](https://arxiv.org/html/2610.00372#S4.E5)\. Inputs contain only information available before a candidate intervention, including action and reasoning repetition, no\-op behavior, observation novelty, consistency between actions and observations, task progress, and context\-budget statistics\.
To keep model fitting separate from threshold selection, we divide the 93 training tasks into five groups\. In turn, we fit the models on four groups and predict the held\-out group\. This gives every training task an out\-of\-fold \(OOF\) prediction from models that did not see that task\. We choose the decision thresholds from these predictions and then refit the models on all training tasks\. All model parameters and thresholds are fixed before test evaluation\. We useλ=2\\lambda=2,ϵ=0\.02\\epsilon=0\.02,dmin=1d\_\{\\min\}=1, and thresholdsα=0\.876\\alpha=0\.876,β=0\.052\\beta=0\.052, andγ=0\\gamma=0\. During evaluation, CIR checksd∈\{1,2,4\}d\\in\\\{1,2,4\\\}in order and refreshes at most once\.
We compare CIR with never refreshing and with a sequential anomaly baseline\. The anomaly baseline uses the same error detector at each candidate time but does not predict whether refresh will help\. Its threshold is selected from the same OOF predictions, objective, and clean\-harm limit used for CIR\. We also include a version whose training\-time intervention rate is matched to CIR\. Neither baseline is adjusted on the test set\.
### 5\.2Recovery Depends on State and Timing
Figure 2:Paired recovery effects in the 93\-task factual\-success cohort across delays \(a\) and in the complete prefix\-feasible primary \(n=106n=106\) and independent \(n=100n=100\) cohorts \(b\)\. “Stale” denotes the two\-step stale condition\. Error bars are task\-level cluster\-bootstrap 95% confidence intervals\.We first ask whether the same refresh operation has different effects in different states and at different times\. We begin with the 93 primary tasks whose factual trajectory succeeds\. Under clean observations, continued execution reproduces a known success, so any failure after refresh is directly interpretable as harm\. Figure[2](https://arxiv.org/html/2610.00372#S5.F2)\(a\) shows this cohort\. An immediate refresh lowers success by6\.456\.45percentage points under clean observations \(95% CI\[−11\.83,−2\.15\]\[\-11\.83,\-2\.15\]\)\. It raises success by7\.537\.53points under two\-step stale observations \(95% CI\[−3\.23,18\.28\]\[\-3\.23,18\.28\]\)\. The difference between these effects is13\.9813\.98points \(95% CI\[2\.15,24\.73\]\[2\.15,24\.73\]\)\. The stale\-minus\-clean difference remains positive at later delays\. The estimates are18\.2818\.28,15\.0515\.05, and13\.9813\.98points at delays 1, 2, and 4\.
In this cohort, continuing without recovery under a clean observation reproduces the successful factual trajectory, soYi0\(clean\)=1Y\_\{i\}^\{0\}\(\\mathrm\{clean\}\)=1by construction\. Refresh can preserve or harm that success, but it cannot create a rescue\. We therefore repeat the analysis with factual failures included\.
The independent 89\-task factual\-success cohort shows the same ordering at every tested delay\. The effect under stale observations is more positive than the effect under clean observations, with differences ranging from8\.998\.99to14\.6114\.61points\. At delay 2, refresh changes success by\+10\.11\+10\.11points under stale observations and−4\.49\-4\.49points under clean observations\.
We next include every prefix\-feasible task in both cohorts\. This includes 13 factual failures in the primary cohort and 11 in the independent cohort\. Figure[2](https://arxiv.org/html/2610.00372#S5.F2)\(b\) summarizes the results\. In the 106\-task primary cohort, refresh changes success by−3\.77\-3\.77points for clean observations \(95% CI\[−9\.43,0\.94\]\[\-9\.43,0\.94\]\) and\+6\.60\+6\.60points for stale observations \(95% CI\[−2\.83,16\.04\]\[\-2\.83,16\.04\]\)\. The difference is10\.3810\.38points \(95% CI\[0\.00,20\.75\]\[0\.00,20\.75\]\)\. In the 100\-task independent cohort, the corresponding effects are0\.000\.00and\+4\.00\+4\.00points\. Their difference is4\.004\.00points \(95% CI\[−5\.00,14\.00\]\[\-5\.00,14\.00\]\)\. The estimated effect is therefore more favorable under stale observations in both complete cohorts, including tasks with failed factual trajectories\.
The paired counts show what these averages hide\. In the complete primary cohort, refresh produces 2 rescues and 6 harms under clean observations, compared with 17 rescues and 10 harms under stale observations\. At delay 1 in the factual\-success cohort, the positive average under stale observations combines 15 rescues with 4 harms\. Under clean observations, refresh turns 6 successes into failures\. Appendix[B](https://arxiv.org/html/2610.00372#A2)gives the full breakdown\.
Taken together, these results show that recovery value is not a fixed property of the refresh operation\. It depends on the information available to the agent, the state reached after the error, and the time at which recovery is attempted\. This explains why a single unconditional rule can hide useful interventions and unnecessary ones behind the same average success rate\.
### 5\.3Why Does Refresh Help?
Figure 3:Mechanism controls under clean observations at delay 0 \(a\) and two\-step stale observations at delay 4 \(b\)\. Points are effects relative to no recovery\. Error bars are task\-level cluster\-bootstrap 95% confidence intervals\.We next ask what part of refresh produces the gain\. A full refresh queries the environment and shows the returned observation to the agent\. A second variant, called content\-ablated refresh, makes the same query but hides the returned content\. A third variant asks the agent to replan without querying the environment\. Under clean observations at delay 0, these operations change success by−8\.96\-8\.96,−5\.97\-5\.97, and−14\.93\-14\.93percentage points, respectively \(Figure[3](https://arxiv.org/html/2610.00372#S5.F3)\)\. Under two\-step stale observations at delay 4, the changes are\+13\.43\+13\.43points \(95% CI\[4\.48,22\.39\]\[4\.48,22\.39\]\),\+10\.45\+10\.45points \(95% CI\[1\.49,20\.90\]\[1\.49,20\.90\]\), and\+1\.49\+1\.49points \(95% CI\[−8\.96,11\.98\]\[\-8\.96,11\.98\]\)\. Full refresh rescues ten tasks\. Content\-ablated refresh also rescues eight of them\. The returned observation is therefore not required for those shared rescues\. Appendix[A](https://arxiv.org/html/2610.00372#A1)gives the task\-level comparison\. This result matters because refresh is often treated as if it were equivalent to supplying better observation content\. Our controls show that refresh has two distinct components\. It queries the environment and then shows the returned content to the agent\. The returned content is unnecessary for most shared rescues\. Evaluating only the full operation would miss this distinction\.
### 5\.4CIR Learns When to Recover
We evaluate CIR on all 75 held\-out prefix\-feasible tasks, whether or not their factual trajectory succeeds\. Across 300 episodes, never refreshing succeeds on70\.33%70\.33\\%, while CIR reaches73\.33%73\.33\\%\. The gain is3\.003\.00percentage points \(95% CI\[0\.67,5\.67\]\[0\.67,5\.67\]\), with 11 rescues and 2 harms\.
Figure 4:Selective recovery on 75 held\-out prefix\-feasible tasks\. Each task is evaluated under four observation conditions, giving 300 episodes\. Points show effects relative to never refreshing\. Error bars are task\-level cluster\-bootstrap 95% confidence intervals\. CIR intervenes in 16\.3% of episodes\.Table 1:Test performance of CIR by observation condition on all 75 held\-out prefix\-feasible tasks\.The effect varies by observation condition \(Figure[4](https://arxiv.org/html/2610.00372#S5.F4)and Table[1](https://arxiv.org/html/2610.00372#S5.T1)\)\. CIR never intervenes when the observation is clean, so success is unchanged\. Its largest gain occurs when the observation is two steps out of date\. Success rises from60\.00%60\.00\\%to69\.33%69\.33\\%, an improvement of9\.339\.33points \(95% CI\[1\.33,17\.33\]\[1\.33,17\.33\]\) from 9 rescues and 2 harms\. The gains for one\-step stale and missing observations are each1\.331\.33points\. CIR intervenes in 31 of the 75 two\-step stale episodes and in only 1 of the 75 missing\-observation episodes\. It concentrates recovery where the observed benefit is largest\.
CIR intervenes in16\.3%16\.3\\%of episodes\. The sequential anomaly baseline intervenes in12\.7%12\.7\\%and improves success by1\.671\.67points \(95% CI\[0\.33,3\.33\]\[0\.33,3\.33\]\), with 6 rescues and 1 harm\. The rate\-matched anomaly baseline intervenes in15\.7%15\.7\\%and improves success by2\.002\.00points \(95% CI\[0\.67,3\.67\]\[0\.67,3\.67\]\), with 6 rescues and no harms\. CIR has the highest point estimate\. Its margins over the two anomaly baselines are1\.331\.33points \(95% CI\[−0\.33,3\.33\]\[\-0\.33,3\.33\]\) and1\.001\.00point \(95% CI\[−0\.67,2\.67\]\[\-0\.67,2\.67\]\), respectively\. The3\.003\.00\-point gain over never refreshing is supported by a task\-level sign\-flip test \(p=0\.034p=0\.034\)\. The rate\-matched comparison helps separate selective recovery from simply intervening more often\. CIR and the rate\-matched anomaly policy refresh a similar fraction of episodes\. CIR produces 11 rescues rather than 6 and attains the highest observed success rate\. This pattern is consistent with CIR’s design\. The policy considers the predicted outcome of the intervention, not only whether the current observation appears unusual\. CIR does not trigger on any of the 75 clean episodes, preserving the never\-refresh success rate throughout the evaluated clean cohort\.
#### Discussion\.
Recovery should be evaluated as a decision, not only as an operation\. Paired continuations reveal which trajectories recovery rescues or harms and how its value changes with state and timing\. They provide more information than post\-recovery success alone\. They also show why anomaly detection is insufficient\. An unusual observation does not imply that refresh will help\. CIR predicts outcomes under refresh and continued execution before acting\. A harness can therefore estimate recovery value first and intervene only when the expected value is positive\.
## 6Conclusion
Average success alone hides how recovery changes individual trajectories\. Our paired evaluation shows that the value of refresh depends on the observation, the timing, and the trajectory\. It also separates the tasks that recovery rescues from those it harms\. Controlled refresh variants show that the newly returned observation is not the only source of benefit\. CIR turns this evidence into a practical decision\. Without retraining the underlying agent, it improves success by3\.003\.00percentage points across 300 episodes from 75 held\-out tasks\. Its largest gain occurs when observations are more severely outdated, and it attains the highest observed success rate among the policies we compare\. These results show how paired causal evidence can support more selective recovery in the harness of LLM agents\.
### AI use statement
In this work, we used generative AI tools only for language polishing, including improving grammar, clarity, and fluency of the manuscript\. We have not used generative AI tools for generating research ideas, developing the methodology, designing experiments, analyzing results, or drawing scientific conclusions; other disclosure categories are not applicable to this work\. We reviewed all AI\-assisted edits and verified that they did not alter the technical content or scientific claims\. We take full responsibility for the final content of this work, including all text, claims, and artifacts\.
### Ethics Statement
All experiments are conducted in controlled benchmark environments and do not involve human subjects or private user data\. We do not identify additional ethical risks beyond those commonly associated with LLM agent research\.
## References
- Barkeet al\.\(2026\)S\. Barke, A\. Goyal, A\. Khare, A\. Singh, S\. Nath, and C\. BansalAgentRx: diagnosing ai agent failures from execution trajectories\.External Links:2602\.02475,[Link](https://arxiv.org/abs/2602.02475)Cited by:[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px2.p1.1)\.
- Bonagiriet al\.\(2026\)A\. Bonagiri, D\. Borkar, G\. J\. Anderias, S\. Rafatirad, and H\. HomayounCausalFlow: causal attribution and counterfactual repair for LLM agent failures\.Note:arXiv preprint arXiv:2605\.25338External Links:2605\.25338,[Document](https://dx.doi.org/10.48550/arXiv.2605.25338),[Link](https://arxiv.org/abs/2605.25338)Cited by:[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px2.p1.1)\.
- Cemriet al\.\(2025\)M\. Cemri, M\. Z\. Pan, S\. Yang, L\. A\. Agrawal, B\. Chopra, R\. Tiwari, K\. Keutzer, A\. Parameswaran, D\. Klein, K\. Ramchandran, M\. A\. Zaharia, J\. E\. Gonzalez, and I\. StoicaWhy do multi\-agent LLM systems fail?\.InAdvances in Neural Information Processing Systems,Vol\.38\.External Links:[Document](https://dx.doi.org/10.52202/085713-4082),[Link](https://papers.nips.cc/paper_files/paper/2025/hash/b1041e52d3be19f0a9bc491657488e4a-Abstract-Datasets_and_Benchmarks_Track.html)Cited by:[§1](https://arxiv.org/html/2610.00372#S1.p2.1),[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px2.p1.1)\.
- Chenet al\.\(2026\)M\. Chen, J\. Wang, Z\. Liu, Y\. Wang, H\. Zheng, and Q\. WangFrom failed trajectories to reliable LLM agents: diagnosing and repairing harness flaws\.Note:arXiv preprint arXiv:2606\.06324External Links:2606\.06324,[Document](https://dx.doi.org/10.48550/arXiv.2606.06324),[Link](https://arxiv.org/abs/2606.06324)Cited by:[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px2.p1.1)\.
- Efron \(1979\)B\. EfronBootstrap methods: another look at the jackknife\.The Annals of Statistics7\(1\),pp\. 1–26\.External Links:[Document](https://dx.doi.org/10.1214/aos/1176344552),[Link](https://doi.org/10.1214/aos/1176344552)Cited by:[§5\.1](https://arxiv.org/html/2610.00372#S5.SS1.SSS0.Px3.p1.1)\.
- Gouet al\.\(2024\)Z\. Gou, Z\. Shao, Y\. Gong, Y\. Shen, Y\. Yang, N\. Duan, and W\. ChenCRITIC: large language models can self\-correct with tool\-interactive critiquing\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2305.11738)Cited by:[§1](https://arxiv.org/html/2610.00372#S1.p2.1),[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px1.p1.1)\.
- Holland \(1986\)P\. W\. HollandStatistics and causal inference\.Journal of the American Statistical Association81\(396\),pp\. 945–960\.External Links:[Document](https://dx.doi.org/10.1080/01621459.1986.10478354),[Link](https://www.jstor.org/stable/2289064)Cited by:[§1](https://arxiv.org/html/2610.00372#S1.p3.1)\.
- Huanget al\.\(2024\)J\. Huang, X\. Chen, S\. Mishra, H\. S\. Zheng, A\. W\. Yu, X\. Song, and D\. ZhouLarge language models cannot self\-correct reasoning yet\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2310.01798)Cited by:[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px1.p1.1)\.
- Kitagawa and Tetenov \(2018\)T\. Kitagawa and A\. TetenovWho should be treated? empirical welfare maximization methods for treatment choice\.Econometrica86\(2\),pp\. 591–616\.External Links:[Document](https://dx.doi.org/10.3982/ECTA13288),[Link](https://onlinelibrary.wiley.com/doi/10.3982/ECTA13288)Cited by:[§4\.3](https://arxiv.org/html/2610.00372#S4.SS3.p3.1)\.
- Künzelet al\.\(2019\)S\. R\. Künzel, J\. S\. Sekhon, P\. J\. Bickel, and B\. YuMetalearners for estimating heterogeneous treatment effects using machine learning\.Proceedings of the National Academy of Sciences116\(10\),pp\. 4156–4165\.External Links:[Document](https://dx.doi.org/10.1073/pnas.1804597116),[Link](https://www.pnas.org/doi/10.1073/pnas.1804597116)Cited by:[§4\.2](https://arxiv.org/html/2610.00372#S4.SS2.p2.3)\.
- Liet al\.\(2026\)S\. Li, Y\. Wang, P\. Li, Z\. Wei, and Y\. TangReSeek: a self\-correcting framework for search agents with instructive rewards\.InProceedings of the 43rd International Conference on Machine Learning,Cited by:[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px1.p1.1)\.
- Maet al\.\(2024\)C\. Ma, J\. Zhang, Z\. Zhu, C\. Yang, Y\. Yang, Y\. Jin, Z\. Lan, L\. Kong, and J\. HeAgentBoard: an analytical evaluation board of multi\-turn LLM agents\.InAdvances in Neural Information Processing Systems,Vol\.37\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2024/hash/877b40688e330a0e2a3fc24084208dfa-Abstract-Datasets_and_Benchmarks_Track.html)Cited by:[§1](https://arxiv.org/html/2610.00372#S1.p2.1),[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px2.p1.1)\.
- Maet al\.\(2025\)M\. Ma, J\. Zhang, F\. Yang, Y\. Kang, Q\. Lin, S\. Rajmohan, and D\. ZhangDoVer: intervention\-driven auto debugging for LLM multi\-agent systems\.Note:arXiv preprint arXiv:2512\.06749External Links:2512\.06749,[Document](https://dx.doi.org/10.48550/arXiv.2512.06749),[Link](https://arxiv.org/abs/2512.06749)Cited by:[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px2.p1.1)\.
- Maet al\.\(2026\)Z\. Ma, H\. Huang, S\. Zou, Y\. Wang, S\. Yang, Y\. Hu, F\. Wei, and X\. ChuLongHorizon\-harness: advancing long\-horizon agents for real\-world tasks\.External Links:2608\.01964,[Link](https://arxiv.org/abs/2608.01964)Cited by:[§1](https://arxiv.org/html/2610.00372#S1.p1.1)\.
- Madaanet al\.\(2023\)A\. Madaan, N\. Tandon, P\. Gupta, S\. Hallinan, L\. Gao, S\. Wiegreffe, U\. Alon, N\. Dziri, S\. Prabhumoye, Y\. Yang, S\. Gupta, B\. P\. Majumder, K\. Hermann, S\. Welleck, A\. Yazdanbakhsh, and P\. ClarkSelf\-Refine: iterative refinement with self\-feedback\.InAdvances in Neural Information Processing Systems,Vol\.36\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/91edff07232fb1b55a505a9e9f6c0ff3-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2610.00372#S1.p2.1),[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px1.p1.1)\.
- Packeret al\.\(2023\)C\. Packer, S\. Wooders, K\. Lin, V\. Fang, S\. G\. Patil, I\. Stoica, and J\. E\. GonzalezMemGPT: towards LLMs as operating systems\.Note:arXiv preprint arXiv:2310\.08560External Links:2310\.08560,[Link](https://arxiv.org/abs/2310.08560)Cited by:[§1](https://arxiv.org/html/2610.00372#S1.p1.1)\.
- Rubin \(1974\)D\. B\. RubinEstimating causal effects of treatments in randomized and nonrandomized studies\.Journal of Educational Psychology66\(5\),pp\. 688–701\.External Links:[Document](https://dx.doi.org/10.1037/h0037350),[Link](https://doi.org/10.1037/h0037350)Cited by:[§3\.2](https://arxiv.org/html/2610.00372#S3.SS2.p3.1)\.
- Shah \(2026\)J\. ShahCausal Agent Replay: counterfactual attribution for LLM\-agent failures\.Note:arXiv preprint arXiv:2606\.08275External Links:2606\.08275,[Document](https://dx.doi.org/10.48550/arXiv.2606.08275),[Link](https://arxiv.org/abs/2606.08275)Cited by:[§1](https://arxiv.org/html/2610.00372#S1.p1.1),[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px2.p1.1)\.
- Shinnet al\.\(2023\)N\. Shinn, F\. Cassano, A\. Gopinath, K\. Narasimhan, and S\. YaoReflexion: language agents with verbal reinforcement learning\.InAdvances in Neural Information Processing Systems,Vol\.36\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/1b44b878bb782e6954cd888628510e90-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2610.00372#S1.p2.1),[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px1.p1.1)\.
- Shridharet al\.\(2021\)M\. Shridhar, X\. Yuan, M\. Côté, Y\. Bisk, A\. Trischler, and M\. HausknechtALFWorld: aligning text and embodied environments for interactive learning\.InInternational Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2010.03768)Cited by:[§5\.1](https://arxiv.org/html/2610.00372#S5.SS1.SSS0.Px1.p1.1)\.
- Tian and Pearl \(2000\)J\. Tian and J\. PearlProbabilities of causation: bounds and identification\.Annals of Mathematics and Artificial Intelligence28,pp\. 287–313\.External Links:[Document](https://dx.doi.org/10.1023/A%3A1018912507879),[Link](https://link.springer.com/article/10.1023/A:1018912507879)Cited by:[§4\.2](https://arxiv.org/html/2610.00372#S4.SS2.p3.2)\.
- Wuet al\.\(2024a\)Q\. Wu, G\. Bansal, J\. Zhang, Y\. Wu, B\. Li, E\. Zhu, L\. Jiang, X\. Zhang, S\. Zhang, J\. Liu, A\. H\. Awadallah, R\. W\. White, D\. Burger, and C\. WangAutoGen: enabling next\-gen LLM applications via multi\-agent conversations\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=BAakY1hNKS)Cited by:[§1](https://arxiv.org/html/2610.00372#S1.p1.1),[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px1.p1.1)\.
- Wuet al\.\(2024b\)Y\. Wu, T\. Yue, S\. Zhang, C\. Wang, and Q\. WuStateFlow: enhancing LLM task\-solving through state\-driven workflows\.InFirst Conference on Language Modeling,External Links:[Link](https://arxiv.org/abs/2403.11322)Cited by:[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px1.p1.1)\.
- Xiaoet al\.\(2026\)S\. Xiao, S\. Wang, X\. Chen, K\. Chao, M\. Cui, F\. Qian, F\. Meng, C\. Mei, C\. Jiang, Q\. Ouyang, and J\. YiWhen errors become memories: causal pathway tracing in multi\-turn memory\-augmented llms\.External Links:2608\.30198,[Link](https://arxiv.org/abs/2608.30198)Cited by:[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px2.p1.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. QiuQwen3 technical report\.Note:arXiv preprint arXiv:2505\.09388External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§5\.1](https://arxiv.org/html/2610.00372#S5.SS1.SSS0.Px1.p1.1)\.
- Yanget al\.\(2024\)J\. Yang, C\. E\. Jimenez, A\. Wettig, K\. Lieret, S\. Yao, K\. Narasimhan, and O\. PressSWE\-agent: agent\-computer interfaces enable automated software engineering\.InAdvances in Neural Information Processing Systems,Vol\.37\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2024/hash/5a7c947568c1b1328ccc5230172e1e7c-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2610.00372#S1.p1.1),[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px1.p1.1)\.
- Yaoet al\.\(2023\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. CaoReAct: synergizing reasoning and acting in language models\.InInternational Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2210.03629)Cited by:[§1](https://arxiv.org/html/2610.00372#S1.p1.1),[§5\.1](https://arxiv.org/html/2610.00372#S5.SS1.SSS0.Px1.p1.1)\.
- Yaoet al\.\(2026\)Y\. Yao, X\. Tan, C\. Liu, Y\. Li, Z\. Wang, W\. Yu, Z\. Tan, Y\. Tian, G\. Zhao, L\. Sun, X\. Zhang, and T\. YangHarness\-bench: measuring harness effects across models in realistic agent workflows\.External Links:2605\.27922,[Link](https://arxiv.org/abs/2605.27922)Cited by:[§1](https://arxiv.org/html/2610.00372#S1.p1.1),[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2025\)S\. Zhang, M\. Yin, J\. Zhang, J\. Liu, Z\. Han, J\. Zhang, B\. Li, C\. Wang, H\. Wang, Y\. Chen, and Q\. WuWhich agent causes task failures and when? on automated failure attribution of LLM multi\-agent systems\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 76583–76599\.External Links:[Link](https://proceedings.mlr.press/v267/zhang25cq.html)Cited by:[§1](https://arxiv.org/html/2610.00372#S1.p2.1),[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px2.p1.1)\.
- Zhaoet al\.\(2026\)Y\. Zhao, M\. Li, R\. Li, P\. Z\. Wang, S\. Jiang, L\. Pang, X\. Xiao, and X\. HuCatchBench: when can an agent failure be caught?\.Note:arXiv preprint arXiv:2608\.22808External Links:2608\.22808,[Document](https://dx.doi.org/10.48550/arXiv.2608.22808),[Link](https://arxiv.org/abs/2608.22808)Cited by:[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px2.p1.1)\.
- Zhouet al\.\(2024\)A\. Zhou, K\. Yan, M\. Shlapentokh\-Rothman, H\. Wang, and Y\. WangLanguage agent tree search unifies reasoning, acting, and planning in language models\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 62138–62160\.External Links:[Link](https://proceedings.mlr.press/v235/zhou24r.html)Cited by:[§1](https://arxiv.org/html/2610.00372#S1.p2.1),[§2](https://arxiv.org/html/2610.00372#S2.SS0.SSS0.Px1.p1.1)\.
## Appendix ATask\-Level Overlap between Recovery Operations
We analyze the joint outcomes of full refresh and content\-ablated refresh under two\-step stale observations at delayd=4d=4\. Both outcomes are classified relative to the same no\-recovery outcome\. The mechanism\-control records contain 74 tasks, and 67 satisfy the protocol checks\. Two of these tasks terminate successfully before the scheduled recovery, so neither refresh operation is executed\. They remain in the 67\-task aggregate analysis in Section[5\.3](https://arxiv.org/html/2610.00372#S5.SS3), but they are excluded from the joint analysis of executed interventions\. Figure[5](https://arxiv.org/html/2610.00372#A1.F5)shows the joint task\-level outcomes\. Table[2](https://arxiv.org/html/2610.00372#A1.T2)reports the paired effects\. The remaining 65 tasks share the same task identity and exact token prefix before recovery in both branches\. No environment\-replay mismatches were recorded\.
Figure 5:Joint recovery outcomes under two\-step stale observations at delayd=4d=4\. Rows denote full refresh and columns denote content\-ablated refresh, each classified relative to the same no\-recovery outcome\. The analysis includes 65 tasks with matching execution prefixes and excludes two tasks that terminate successfully before recovery\. Eight rescues are shared, two occur only under full refresh, and one occurs only under content\-ablated refresh\.The outcome categories agree on 59 of 65 tasks\. These include eight shared rescues, 38 stable successes, and 13 stable failures\. The six disagreements include two full\-only rescues, one content\-ablated\-only rescue, two harms avoided by full refresh, and one harm introduced by full refresh relative to content\-ablated refresh\. The high overall agreement therefore includes 51 tasks whose outcomes neither operation changes\. The eight shared rescues show that the returned content is not necessary for these observed benefits\.
Table 2:Paired success effects on the 65 tasks where both recovery operations can be executed\. Intervals use 10,000 task\-level cluster\-bootstrap resamples\. All values are in percentage points\.The paired net gains are nine tasks for full refresh versus no recovery, seven for content\-ablated refresh versus no recovery, and two for full refresh versus content\-ablated refresh\. Removing the two already\-terminated tasks changes the denominator from 67 to 65 but does not change these gains\. This explains the difference from the aggregate effects in Section[5\.3](https://arxiv.org/html/2610.00372#S5.SS3)\. Both refresh variants produce positive gains relative to no recovery\. Full refresh has a3\.083\.08\-pp point estimate over content\-ablated refresh\.
## Appendix BAdditional Recovery Results
Figure[6](https://arxiv.org/html/2610.00372#A2.F6)gives the complete paired\-outcome decomposition for the 93\-task factual\-success cohort at recovery delayd=1d=1\. Under stale observations, refresh produces 15 rescues and 4 harms, alongside 59 stable successes and 15 stable failures\. Under clean observations, 87 trajectories remain successful and 6 are harmed\. This decomposition illustrates how the positive stale\-state average arises from trajectory\-level changes that aggregate success rates alone conceal\.
Figure 6:Paired outcomes at recovery delayd=1d=1in the 93\-task factual\-success cohort\. Each bar partitions trajectories into stable success, rescue, harm, and stable failure\.Similar Articles
EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents
Introduces EvoHarness-RL, a framework that learns runtime harness policies for long-horizon LLM agents, enabling them to construct and update external state (belief, progress, experience) during task execution. Using Qwen3-8B on ALFWorld, it achieves 96.9% success and reveals harness annealing and evolution dynamics.
ReFlect: An Effective Harness System for Complex Long-Horizon LLM Reasoning
This paper introduces ReFlect, a training-free harness system that wraps LLMs with deterministic error detection and recovery logic to improve performance on complex, long-horizon reasoning tasks.
best of the best agentic harnesses do this…
The author shares insights on building effective agent harnesses: the best ones minimize LLM reliance for trivial tasks and reserve LLMs for complex reasoning, distinguishing genuine harnesses from simple wrappers.
Stop Comparing LLM Agents Without Disclosing the Harness
This position paper argues that in long-horizon LLM agent tasks, the execution harness often determines performance more than the model itself, and current benchmarks misattribute harness-level gains to model improvements. It proposes a harness-aware evaluation framework with disclosure standards and variance decomposition protocols.
EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses
EvoUndo introduces a framework for evaluating and ensuring recoverability in self-modifying LLM agents, showing that reliable recovery requires co-designing verification, state grounding, and recovery language expressivity.