Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories

arXiv cs.AI Papers

Summary

This paper identifies post-retrieval reuse as a bottleneck for long-horizon agent memory and proposes query-conditioned reuse (QCR), a simple target-bound note format, showing improved success and token efficiency across WebArena, WorkArena, and AppWorld.

arXiv:2608.12847v1 Announce Type: new Abstract: Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. We identify this post-retrieval reuse step as a distinct bottleneck for long-horizon trajectory memory and formulate an evaluation framework that holds candidate retrieval, target state, model, decoding, and tool budget fixed while varying the support delivered to the agent. We instantiate the framework with query-conditioned reuse (QCR), a deliberately simple target-bound note that records a reusable procedure, bindings to recover, applicability conditions, and verification requirements. QCR serves to test the reuse hypothesis rather than to claim a universally preferred memory format. Across 2,391 target instances in WebArena, WorkArena, and AppWorld, QCR reaches 62.3% average Success, 10.7 points above Full Trajectory, while using 48.9% fewer online tokens. Summary reranking selects a reusable memory for 94.8% of targets, placing end-task Success within 1.8 points of an oracle reusable selector. Analyses by trajectory length and source--target binding shift show that direct trajectory injection loses much of its utility as traces grow longer or source-specific values change, whereas target-bound support preserves a larger share of the measured gain. The resulting framework separates retrieval quality from the problem of turning retrieved experience into safe, useful support for a new task.
Original Article
View Cached Full Text

Cached at: 08/14/26, 09:28 AM

# Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories
Source: [https://arxiv.org/html/2608.12847](https://arxiv.org/html/2608.12847)
Heng WangLingling ZhangMuye HuangXinyu ZhangJiashuai LiuHang YanRongman Xu

###### Abstract

Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed\. We identify this post\-retrieval reuse step as a distinct bottleneck for long\-horizon trajectory memory and formulate an evaluation framework that holds candidate retrieval, target state, model, decoding, and tool budget fixed while varying the support delivered to the agent\. We instantiate the framework with query\-conditioned reuse \(QCR\), a deliberately simple target\-bound note with a workflow invariant, bindings to re\-obtain, applicability conditions, and a verification guardrail\. QCR serves to test the reuse hypothesis rather than to claim a universally preferred memory format\. Across 2,391 target instances in WebArena, WorkArena, and AppWorld, QCR reaches 62\.3% average Success, 10\.7 points above Full Trajectory, while using 48\.9% fewer online tokens\. Summary reranking selects a reusable memory for 94\.8% of targets, placing end\-task Success within 1\.8 points of an oracle reusable selector\. Analyses by trajectory length and source–target binding shift show that direct trajectory injection loses much of its utility as traces grow longer or source\-specific values change, whereas target\-bound support preserves a larger share of the measured gain\. The resulting framework separates retrieval quality from the problem of turning retrieved experience into safe, useful support for a new task\.

## 1Introduction

Figure 1:A design hypothesis: the bottleneck shifts with memory item length\.For short facts or episodes, retrieving the relevant item often supplies what the target needs\. Long trajectories can save more exploration, but their target\-side value depends on reuse, rebinding, and execution\.QCRaddresses this post\-retrieval step\.Agent memory has progressed from carrying a short interaction history to maintaining external stores, structured records, retrieval indices, and learned memory operations\. These systems aim to let an agent bring past experience into a later decision rather than rediscovering the same information or procedure from scratch\. Recent methods make histories easier to retain, organize, and retrieve\([19](https://arxiv.org/html/2608.12847#bib.bib20);[37](https://arxiv.org/html/2608.12847#bib.bib37);[8](https://arxiv.org/html/2608.12847#bib.bib12);[2](https://arxiv.org/html/2608.12847#bib.bib10);[30](https://arxiv.org/html/2608.12847#bib.bib31);[11](https://arxiv.org/html/2608.12847#bib.bib7);[33](https://arxiv.org/html/2608.12847#bib.bib8);[14](https://arxiv.org/html/2608.12847#bib.bib9)\), while long\-context benchmarks test whether systems recover evidence over extended interactions\([1](https://arxiv.org/html/2608.12847#bib.bib1);[17](https://arxiv.org/html/2608.12847#bib.bib18);[27](https://arxiv.org/html/2608.12847#bib.bib28);[23](https://arxiv.org/html/2608.12847#bib.bib25);[10](https://arxiv.org/html/2608.12847#bib.bib14)\)\. This progress makes historical experience accessible\. It does not yet establish that the experience helps an agent solve its current query\. This distinction matters because many memory evaluations naturally end at access: can a system retain an item, rank it for a query, or answer a question about an earlier interaction? LongBench, LoCoMo, and LongMemEval make these access questions measurable across long contexts and conversations\([1](https://arxiv.org/html/2608.12847#bib.bib1);[17](https://arxiv.org/html/2608.12847#bib.bib18);[27](https://arxiv.org/html/2608.12847#bib.bib28)\)\. They are necessary tests, but they leave open a second question for an acting agent: once an experience has entered the context, does it improve the new task that the agent must now complete?

For short, self\-contained memory items, retrieval and reuse are often nearly the same operation\. If a query needs a fact, a local instruction, or a compact episode, returning the right item usually supplies the evidence needed for the next answer or action\. Retrieval quality is therefore a useful proxy for memory utility in these settings\. Long\-horizon task experience has a different structure\. A successful trajectory may encode a valuable tool workflow, decision rule, and verification sequence, so it can save far more exploration than a short fact\. Yet it also carries source\-specific users, objects, paths, dates, observations, and failed branches\. Finding such a trajectory does not tell the agent which part transfers, which binding has expired, or which checks must be repeated before it acts\. This is the regime targeted by interactive agent environments: browser, API, and database tasks couple many actions to a changing state and evaluate their final consequences\([38](https://arxiv.org/html/2608.12847#bib.bib38);[6](https://arxiv.org/html/2608.12847#bib.bib5);[24](https://arxiv.org/html/2608.12847#bib.bib24);[31](https://arxiv.org/html/2608.12847#bib.bib33)\)\. A literal trace can place the agent in the right subsystem while still applying an obsolete argument to the wrong current object\.

Figure[1](https://arxiv.org/html/2608.12847#S1.F1)expresses the resulting bottleneck shift\. Moving from short facts through episodes to long trajectories, a memory item can carry more of the work that an agent would otherwise repeat\. The main difficulty also moves rightward along the memory pipeline\. Once a relevant long trajectory has been found, the target agent must extract the procedure that still applies, recover current bindings, reject source details that no longer hold, and verify the new final state\. Retrieval remains necessary, but it is no longer enough\. The difficulty is also not merely a context\-window problem\. Giving an actor a long raw trace can bury a current objective beneath old observations and incidental branches, a failure consistent with evidence that models do not always use the relevant portion of a long context reliably\([15](https://arxiv.org/html/2608.12847#bib.bib16)\)\. Experience\-memory methods accordingly extract reflections, skills, workflows, or reusable knowledge from prior runs\([22](https://arxiv.org/html/2608.12847#bib.bib23);[35](https://arxiv.org/html/2608.12847#bib.bib35);[26](https://arxiv.org/html/2608.12847#bib.bib27)\); what remains unclear is how to assess the value of that extracted support for a later query\.

This observation changes the basic evaluation question\. Memory should be judged by whether past experience improves the solution of the current query: higher verified completion, less unnecessary exploration, or lower online cost without discarding necessary checks\. A high retrieval score alone cannot answer this question for long trajectories\. Consider an earlier multi\-app workflow for search, verification, and artifact creation\. A later request may preserve that workflow while changing the person, file, date, or environment state\. Starting from scratch wastes the earlier run; replaying it can copy an obsolete recipient, path, or state assumption\. The useful support is a target\-bound account of the procedure, the bindings to recover, and the checks that remain necessary\. We call this operation query\-conditioned reuse\.

This criterion separates memory utility from both retrieval quality and source fidelity\. A support object may faithfully preserve a source trajectory yet harm the target by carrying stale bindings forward; conversely, an aggressively short summary may save tokens while omitting the precondition that prevents an invalid action\. The relevant comparison therefore holds the available history fixed and asks which use of that history gives the target agent the best verified outcome for its own environment\.

We introduce an end\-to\-end setting that evaluates this post\-retrieval operation directly\. A unified frozen bank contains verified historical trajectories\. For each target, the retriever returns candidate experiences and a shared ranker selects one record before any memory condition runs\. Full Trajectory, Generic Summary, andQCRtherefore receive the same selected experience but use it differently\. Target Success, Milestone completion, API calls, and online tokens then measure whether that experience actually helps the current query\. Our contribution is a problem formulation and evaluation protocol for this boundary, together withQCR, a minimal target\-conditioned support transformation\. The analyses test how selected\-memory length and source–target binding shift change the utility of direct reuse\.

## 2Related Work

### 2\.1Long\-Context and Retrieval\-Oriented Memory

Memory systems study storage, indexing, updating, and context delivery, including MemGPT, MemoryBank, HippoRAG, Mem0, and A\-MEM\([19](https://arxiv.org/html/2608.12847#bib.bib20);[37](https://arxiv.org/html/2608.12847#bib.bib37);[8](https://arxiv.org/html/2608.12847#bib.bib12);[2](https://arxiv.org/html/2608.12847#bib.bib10);[30](https://arxiv.org/html/2608.12847#bib.bib31)\)\. LongMemEval, LongBench, LoCoMo, MemBench, and MemoryAgentBench evaluate retention and access over long or incremental interactions\([27](https://arxiv.org/html/2608.12847#bib.bib28);[1](https://arxiv.org/html/2608.12847#bib.bib1);[17](https://arxiv.org/html/2608.12847#bib.bib18);[23](https://arxiv.org/html/2608.12847#bib.bib25);[10](https://arxiv.org/html/2608.12847#bib.bib14)\); RAG supplies the standard retrieve\-then\-condition pattern\([13](https://arxiv.org/html/2608.12847#bib.bib11)\), and surveys organize memory operations and representations\([34](https://arxiv.org/html/2608.12847#bib.bib34);[7](https://arxiv.org/html/2608.12847#bib.bib2)\)\. Transformer\-XL and Memorizing Transformers make much longer access possible\([3](https://arxiv.org/html/2608.12847#bib.bib4);[28](https://arxiv.org/html/2608.12847#bib.bib29)\), but more context does not guarantee that a model uses the relevant portion\([15](https://arxiv.org/html/2608.12847#bib.bib16)\)\. We therefore ask a downstream question: after a trajectory has been selected, what support lets an agent use it on a new task?

### 2\.2Trajectory and Procedural Experience

Prior experience can appear as reflection, a skill, a script, or a retrieved trajectory: Reflexion, Generative Agents, Voyager, ReAct, ExpeL, Synapse, and Agent Workflow Memory instantiate these choices\([22](https://arxiv.org/html/2608.12847#bib.bib23);[20](https://arxiv.org/html/2608.12847#bib.bib21);[25](https://arxiv.org/html/2608.12847#bib.bib26);[32](https://arxiv.org/html/2608.12847#bib.bib32);[35](https://arxiv.org/html/2608.12847#bib.bib35);[36](https://arxiv.org/html/2608.12847#bib.bib36);[26](https://arxiv.org/html/2608.12847#bib.bib27)\)\. SAM uses state\-adaptive cues, Agentic Memory learns memory operations, and OCR\-Memory trades representation for faithful long\-history access\([11](https://arxiv.org/html/2608.12847#bib.bib7);[33](https://arxiv.org/html/2608.12847#bib.bib8);[14](https://arxiv.org/html/2608.12847#bib.bib9)\)\. We instead fix the candidate set and selected trajectory, then measure whether its representation changes a later target’s outcome or online cost\.

### 2\.3Long\-Horizon Agent Evaluation

AgentBench, AppWorld, WebArena, WorkArena, andτ\\tau\-bench provide interactive and verifiable settings for multi\-step agency\([16](https://arxiv.org/html/2608.12847#bib.bib17);[24](https://arxiv.org/html/2608.12847#bib.bib24);[38](https://arxiv.org/html/2608.12847#bib.bib38);[6](https://arxiv.org/html/2608.12847#bib.bib5);[31](https://arxiv.org/html/2608.12847#bib.bib33)\)\. Mind2Web, AndroidWorld, WebVoyager, GAIA, OSWorld, and SWE\-bench broaden this coverage across web, mobile, multimodal, and software tasks\([5](https://arxiv.org/html/2608.12847#bib.bib3);[21](https://arxiv.org/html/2608.12847#bib.bib22);[9](https://arxiv.org/html/2608.12847#bib.bib13);[18](https://arxiv.org/html/2608.12847#bib.bib19);[29](https://arxiv.org/html/2608.12847#bib.bib30);[12](https://arxiv.org/html/2608.12847#bib.bib15)\)\. These benchmarks usually score isolated execution\. Our unified bank instead lets a target search same\- and cross\-environment history while retaining its native verifier, so the outcome measures whether prior experience reduces target work without becoming an identical replay\.

## 3Task Setting: Query\-Conditioned Trajectory Reuse

Figure[2](https://arxiv.org/html/2608.12847#S3.F2)defines the evaluation unit used in this paper\. The unit is a*source–target pair*: a verified historical trajectory and a later task that preserves a reusable procedure while changing the values needed to carry it out\. The source is not a demonstration to replay\. It is a record of one successful interaction that may contain a procedure, failed branches, and state checks\. The target is a new task with its own initial state, tool feedback, and verifier\. Memory helps only when the agent extracts the part of the source that still applies and re\-obtains the values that no longer do\.

![Refer to caption](https://arxiv.org/html/2608.12847v1/Fig2.png)Figure 2:Evaluation pipeline for query\-conditioned trajectory reuse\.Offline, verified historical rollouts populate one unified memory bank\. A binding\-aware rewrite creates a target query with a related workflow but new target\-specific values\. Online, a fixed retriever returns histories from the frozen bank, after which the agent executes only against the target environment and its own verifier\. The comparison measures target Success, Milestone completion, API calls, and non\-overlapping online tokens\.### 3\.1Problem Definition

For a target tasktt, letqtq\_\{t\}denote its natural\-language query and letot,0o\_\{t,0\}be the initial observation\. Before the target starts, the agent has access to a frozen bankℬ\\mathcal\{B\}of verified historical trajectories\. A fixed retrieverRRreturns the same top\-kkrecords for every compared method,

Zt=R⁡\(qt,ot,0,ℬ\)\.Z\_\{t\}=R\(q\_\{t\},o\_\{t,0\},\\mathcal\{B\}\)\.GivenZtZ\_\{t\}, the target query, and the initial observation, a reuse mechanismρ\\rhowrites a support objectrt=ρ⁡\(Zt,qt,ot,0\)r\_\{t\}=\\rho\(Z\_\{t\},q\_\{t\},o\_\{t,0\}\)\. The acting agent then produces a new target trajectoryτ^t\\hat\{\\tau\}\_\{t\}from\(qt,ot,0,rt\)\(q\_\{t\},o\_\{t,0\},r\_\{t\}\)and receives the target’s own verifier outcome\. This setup separates two questions that are often conflated: whether the bank retrieves a potentially relevant history, and whether the agent can turn that history into an action plan for the current task\.

The target evaluation rewards verified completion and penalizes online work\. We therefore report success or verifier score together with API calls and token cost\. A longer or more literal memory representation is not preferred by definition; it must reduce the work of solving the target\. All methods start from the same cachedZtZ\_\{t\}and the same ranker\-selected record, so a difference cannot arise from retrieval or source selection\. They differ only in the representation and use of that selected experience\.

### 3\.2Offline Memory Construction

The bank stores episodes rather than author\-written skills or task summaries\. We construct one unified frozen bank of 623 verified historical trajectories from successful source\-task executions across WebArena, WorkArena, and AppWorld\. An agent solves each source task through its native interface, and a rollout entersℬ\\mathcal\{B\}only after that environment’s checker accepts its final state\. We jointly index all trajectories in one mixed memory pool rather than partitioning records by task family or environment\. The supplementary material reports the bank’s source\-benchmark composition and complete manifest\.

Each retained record contains the source instruction, the ordered sequence of observations and actions, tool calls with their arguments and returned observations, terminal artifacts when present, and the verifier result\. We retain the environment, task identifier, and rollout configuration for audits, but exclude these provenance fields from retrieval representations and target prompts\. The bank can therefore contain detours and failed attempts that occurred before a successful final state, as it would in a deployed agent system\.

We construct the bank from task families of intermediate difficulty\. A no\-memory baseline is run repeatedly on candidate tasks, and we retain families that yield both verified successes and verified failures\. Trivial tasks leave little room for prior experience to matter, whereas tasks with no successful rollouts provide no verified trajectory to store\. Successful source runs from the retained families form the frozen bank used at evaluation time\.

### 3\.3Binding\-Aware Target Construction

Starting from each verified source trajectoryτs\\tau\_\{s\}, we create up to four target variants by retaining the source’s intended workflow while replacing one or more target\-specific bindings at different divergence levels\. A binding is a value that an agent must ground in the current task or environment, such as an entity, a user, a record identifier, a file or location, a date, a parameter, or the relevant current state\. The rewrite may preserve a procedure such as “inspect, validate, modify, and verify,” but it never licenses the agent to copy the source values into the target\. Some trajectories cannot support every rewrite pattern because of task\-specific constraints\. The final benchmark therefore contains 2,391 valid target task instances, or 3\.84 target variants per historical trajectory on average\. The supplementary material reports target counts by benchmark and selected\-memory\-length group\.

The target starts from its own state and is checked by its own executable verifier or pre\-specified rubric\. Literal replay ofτs\\tau\_\{s\}can therefore fail even when its workflow remains useful; the agent must discover the current bindings through target\-side observations and tool calls\. We retain the source–target relation only as sealed audit metadata for relevance annotation and error analysis\. It does not appear in bank records, retrieval indices, prompts, or target\-agent inputs\.

### 3\.4Frozen Retrieval and Evaluation Boundary

After bank construction, all verified source episodes are pooled into one snapshot and frozen\. The retriever indexes only visible source instructions and trajectory descriptors, not environment labels, task identities, sealed source–target relations, verifier diagnostics, or target\-conditioned summaries\. Givenqtq\_\{t\}andot,0o\_\{t,0\}, a fixed embedding retriever produces a top\-55candidate setZtZ\_\{t\}\. A lightweight ranking stage then selects one trajectory fromZtZ\_\{t\}; we cache both the candidate set and the selected record before target execution and reuse them for Full Trajectory, Generic Summary, andQCRconditions\. The retriever may return a transferable workflow, a superficial match, or nothing useful\. No condition receives an oracle source trajectory\.

All compared conditions share the target query and state, frozen bank, cachedZtZ\_\{t\}, ranker\-selected trajectory, acting model, decoding configuration, and tool budget\.QCRmay read only the selected trajectory,qtq\_\{t\}, andot,0o\_\{t,0\}while writingrtr\_\{t\}; it cannot call the environment privately or replace the cached selection\. In partially observable settings, the acting agent performs any required target\-side discovery and pays the resulting cost\.

For each target run, letIbaseI\_\{\\mathrm\{base\}\}be the non\-memory acting prompt,ImemI\_\{\\mathrm\{mem\}\}the retrieved context shown to the acting agent,IsynI\_\{\\mathrm\{syn\}\}andOsynO\_\{\\mathrm\{syn\}\}the support\-synthesis tokens, andOactO\_\{\\mathrm\{act\}\}the acting\-agent output\. We report

Conline=Ibase\+Imem\+Isyn\+Osyn\+Oact,C\_\{\\mathrm\{online\}\}=I\_\{\\mathrm\{base\}\}\+I\_\{\\mathrm\{mem\}\}\+I\_\{\\mathrm\{syn\}\}\+O\_\{\\mathrm\{syn\}\}\+O\_\{\\mathrm\{act\}\},alongside API calls\. The API\-reported acting input isIbase\+ImemI\_\{\\mathrm\{base\}\}\+I\_\{\\mathrm\{mem\}\}, so we report it as a breakdown rather than add it twice\. Source\-rollout cost belongs to offline bank construction and is logged separately\.

Because a target can have more than one useful predecessor, we audit retrieval without assuming a single oracle source\. Blind annotators label each returned record as irrelevant, surface\-related but procedurally unusable, workflow\-relevant, or highly actionable\. We report*usable\-memory@1*,*usable\-memory@k*, and the rank of the first workflow\-relevant record\. These diagnostics tell us whether retrieval exposes history that could help; the end\-to\-end metrics determine whether the agent actually converts that opportunity into a successful, efficient target run\.

## 4A Minimal Query\-Conditioned Reuse Framework

### 4\.1Design Principle

The framework is intentionally simple\. It does not replace the memory store, retriever, or acting agent\. Instead, it inserts one operation after retrieval: produce compact reuse support explicitly conditioned on the target query and current state\. This makes the framework a diagnostic intervention\. If it helps after the same retrieved records have been fixed, the improvement supports the claim that the missing operation is reuse rather than storage alone\.

### 4\.2Reusable Support

GivenZtZ\_\{t\},qtq\_\{t\}, andot,0o\_\{t,0\},QCRproduces a short support object with four fields: \(i\) a workflow invariant, \(ii\) bindings that must be re\-obtained in the target, \(iii\) applicability conditions \(including when to decline reuse\), and \(iv\) a verification guardrail\. These fields are a minimal implementation choice, not a claim that one universal memory schema is optimal\. The support must be substantially shorter than the retrieved set and must not reveal target answers or hidden evaluator information\.

The workflow field keeps only the action pattern that the target still needs: for example, inspect the current object, verify the relevant condition, carry out a modification, and validate the result\. The re\-obtain field blocks a common misuse of trajectory memory\. A source trace can mention a user name, file path, account, object identifier, or prior artifact that helped solve the source task but says nothing about the target value\. The reuse note names that dependency without supplying the old value as an answer\. Applicability and verification fields retain the reason an earlier agent paused, changed branch, or validated\. They make non\-reuse a valid outcome when target state violates a source precondition, rather than encouraging an agent to replay history\.

This representation deliberately stays small\. A learned graph, a hierarchy of summaries, or a new persistent store may outperform it later\. They would not, by themselves, establish whether the useful operation lies before or after retrieval\. The minimal design keeps the intervention legible: given the same retrieved records, the target agent receives either the raw history or a target\-conditioned account of how to use it\.

### 4\.3Target Execution

The acting agent receives either no memory, the selected raw trajectory, its generic source\-only summary, or theQCRsupport object\. A lightweight ranker uses compact descriptors of the cached top\-55records and the target query to select one record before any target condition runs\. Generic summaries are produced offline from each historical record alone, before the target query arrives; their length budget matches theQCRsupport budget\. Thus the memory conditions share the sameZtZ\_\{t\}, selected trajectory, target query, initial target observation, tool access, and target\-side generation budget\. They differ only in how the same selected experience is represented and used\. In the evaluation, the acting model, decoding settings, tool budget, and target rollouts are held fixed within each target across conditions\. We report the cost of producingQCRsupport separately and in the online reuse total\. This avoids treating a reduction in target\-agent output alone as a free gain\.

We use DeepSeek\-V4\-Pro\([4](https://arxiv.org/html/2608.12847#bib.bib6)\)for both historical and target runs\. Prompts state that historical information is advisory rather than an instruction to replay actions literally\. A raw\-trajectory baseline can inspect and exploit its history, whereasQCRreceives no extra tool privileges, target state, or verifier hints\. The comparison changes only the use of the same selected historical experience: raw delivery, source\-only compression, or target\-bound support construction\.

## 5Experiments

### 5\.1Protocol and Metrics

We compare four conditions under the frozen\-retrieval protocol in Section 3:No Memory,Generic Summary,Full Trajectory, andQCR\. For every target, an embedding retriever returns the same top\-55historical trajectories for the three memory conditions\. A lightweight ranker selects one trajectory from that shared set before any condition runs\. Full Trajectory supplies the selected record directly, Generic Summary supplies its source\-only summary, andQCRwrites its query\-conditioned support object from the same selected record\. Thus, the acting agent never receives five long trajectories at once, and Table[1](https://arxiv.org/html/2608.12847#S5.T1)isolates how the selected experience is represented and used\. The selection diagnostic below evaluates the ranker separately\.

We report verified success, milestone completion, API calls, and non\-overlapping online tokens\. Online tokens include the acting prompt, selected memory, support\-synthesis input and output, and acting output\. To measure reuse rather than task difficulty alone, the stratified analyses report

U=Smemory−Sno​memory\.U=S\_\{\\mathrm\{memory\}\}\-S\_\{\\mathrm\{no\\ memory\}\}\.Candidate relevance and final\-memory relevance were judged against the sealed source–target relation and a reusable\-workflow annotation\. A paired trajectory is the historical trajectory used to construct the target; a reusable trajectory may be another history if it supplies a valid procedure for the target\.

### 5\.2Benchmarks, Model, and Controls

We evaluate end\-to\-end target execution in WebArena\([38](https://arxiv.org/html/2608.12847#bib.bib38)\), WorkArena\([6](https://arxiv.org/html/2608.12847#bib.bib5)\), and AppWorld\([24](https://arxiv.org/html/2608.12847#bib.bib24)\)\. Each environment supplies its own initial state, tool interface, and success or milestone checker, while retrieval searches the same unified mixed memory bank across environment boundaries\. We use DeepSeek\-V4\-Pro\([4](https://arxiv.org/html/2608.12847#bib.bib6)\)for source rollouts, summary ranking, support synthesis, and target execution\. Within a target, every condition shares the target suite, initial state, cached retrieval result, ranker\-selected trajectory, model, decoding configuration, tool budget, and verifier\. The comparison therefore changes only the representation and use of the selected experience after retrieval\.

### 5\.3Memory Representations and Run Records

No Memoryreceives no historical record\.Generic Summaryreceives a source\-only summary prepared before the target query arrives, so it cannot encode target answers or current environment observations\.Full Trajectoryreceives the ranker\-selected historical record with its ordered observations, actions, tool arguments, and returned outputs\. A common ranker sees compact descriptors of the top\-55candidates and selects one source trajectory for every memory condition\.QCRthen writes four short fields from that selected trajectory: the workflow invariant, bindings to re\-obtain, applicability conditions, and a verification guardrail\. The acting agent receives that note rather than the raw trajectory or an oracle source–target mapping\.

For each run, we log the cached candidate identifiers, selected memory, verifier outcome, milestone score, API calls, and non\-overlapping token components\. Online\-token accounting includes the base acting prompt, delivered memory, support\-synthesis input and output, and acting output; it excludes the offline cost of collecting verifier\-approved source trajectories\. These records make the reported efficiency comparison traceable to the same target run rather than to different retrieval outcomes\.

The supplement supplies the prompt template, configuration tables, annotation definitions, and audit\-ledger schema needed to interpret the reported tables and diagnostics\.

### 5\.4End\-to\-End Performance

Table 1:End\-to\-end performance across 2,391 target instances\.Success and Milestone are percentages; API Calls and Online Tokens are means\.Table[1](https://arxiv.org/html/2608.12847#S5.T1)shows that historical experience helps, but the procedure used after candidate retrieval matters\. Generic Summary gains 9\.5 success points over No Memory, while Full Trajectory gains a further 3\.7 points\.QCRreaches 62\.3% success, 10\.7 points above Full Trajectory, with the fewest API calls among the memory conditions\. Its 9\.4k online tokens are about half of the 18\.4k required by direct trajectory injection\. The result therefore does not come from sending the actor more historical context\. The same ranking holds in WebArena, WorkArena, and AppWorld:QCRis best on both Success and Milestone in all six environment\-specific comparisons\. Relative to Full Trajectory, its Success margin is 10\.9 points in WebArena, 10\.8 in WorkArena, and 10\.4 in AppWorld\. The corresponding gains over No Memory are 23\.2, 23\.8, and 24\.7 points\. This consistency matters because the three environments differ in interaction modality and state observability; the effect is not carried by one easier benchmark\.

### 5\.5Retrieval and Single\-Memory Selection

Figure[3](https://arxiv.org/html/2608.12847#S5.F3)separates candidate coverage from the decision about what to inject\. The embedding retriever places the paired trajectory in the top five for 95\.6% of targets and at least one reusable trajectory for 97\.8%\. Its top\-one paired accuracy, however, is only 78\.9%\. Ranking candidate summaries against the target raises final paired\-memory accuracy to 91\.7% and final reusable\-memory accuracy to 94\.8%; only 5\.2% of selected memories are irrelevant\. The mean reciprocal rank of the paired trajectory is 0\.87\. Relative to direct top\-11retrieval, reranking gains 12\.8 points in paired accuracy and 12\.4 points in reusable\-memory accuracy\.

The selection ablation gives the same picture at the task level\. Directly using the retriever’s first item lowers success to 56\.1%, and selecting a random top\-five item lowers it to 44\.8%\. The ranking prompt reaches 62\.3%, only 1\.8 points below an oracle that selects a reusable candidate\. Candidate selection still leaves headroom, but it is not the main source of end\-task failure in this setting\. Figure[3](https://arxiv.org/html/2608.12847#S5.F3)visualizes both parts of this result: broad top\-55coverage enables reranking, and the reranked choice closes most of the gap to oracle end\-task success\. The 6\.2\-point improvement over retriever top\-11shows that choosing the memory, rather than increasing the number of injected trajectories, accounts for the gain\.

Figure 3:Why summary reranking matters\. Left: the top\-55candidate set has high paired and reusable coverage, while direct top\-11retrieval is less reliable\. Reranking compact candidate summaries restores final\-memory quality without presenting five full trajectories to the acting agent\. Right: the resulting end\-task success nearly matches oracle reusable selection\.
### 5\.6Sensitivity to Selected\-Memory Length

We partition target instances by the effective\-action length of the ranker\-selected memory trajectory: Short has 5–10 actions, Medium 11–20, Long 21–35, and Very Long more than 35\. The same selected memory defines a length group for every compared condition, including No Memory\. No\-Memory Success falls from 55\.2% in the Short group to 18\.9% in the Very Long group\. The groups therefore differ substantially in task difficulty, so Table[2](https://arxiv.org/html/2608.12847#S5.T2)reports within\-group utility rather than raw success\.

Direct injection degrades steeply: Full Trajectory falls from \+18\.4 points for short histories to \+2\.9 for very long ones\. Generic Summary retains a little more of its initially smaller gain, but its utility never reaches that ofQCR\. Query\-conditioned support also becomes less useful as histories lengthen, yet it retains \+13\.2 points for very long trajectories and 60\.3% of its short\-trajectory utility\. Full Trajectory retains only 15\.8% of its short\-trajectory utility; Generic Summary retains 32\.4%\. Because the length groups also differ in no\-memory difficulty, this is an association under the registered construction rather than a causal estimate of length alone\.

Table 2:Memory utility by selected\-memory trajectory length\.Entries are percentage\-point gains over No Memory within each length group\.
### 5\.7Binding Shift Is Associated with the Reuse Gap

We next vary the number and type of target\-specific bindings rewritten from the source task\. A small rewrite changes one local binding; a medium rewrite changes two or three bindings or one central constraint; a large rewrite changes at least four bindings, or both the target entity and initial environment state\. The no\-rewrite condition preserves the original intent and tests same\-intent recovery\.

Table[3](https://arxiv.org/html/2608.12847#S5.T3)identifies the failure mode suggested by the main result\. When no binding changes, Full Trajectory has high utility \(\+26\.9\)\. Under a large rewrite, its utility shrinks to \+2\.2, and the generic summary reaches only \+5\.3\.QCRdeclines as the target moves further from the source, but preserves \+20\.1 points under the largest shift\. The gain comes with fewer stale\-binding errors: at large shift, direct trajectories produce stale bindings on 46\.9% of targets, compared with 10\.9% forQCR, while correct rebinding rises from 31\.7% to 77\.8%\. We count a stale\-binding error when an action, output, or tool argument repeats a source\-specific value that conflicts with the target query or target\-side observation\. Under large shift, Full Trajectory retains 8\.2% of its no\-shift utility \(2\.2/26\.9\), whereasQCRretains 67\.9% \(20\.1/29\.6\)\. The method does not make binding shift disappear; it reduces the rate at which stale source values displace current\-task evidence\.

Table 3:Memory utility under binding shift\.Entries are percentage\-point gains over No Memory within each rewrite level\.
### 5\.8Interpretation

The four results form a consistent account\. Historical trajectories offer useful procedural information, since both memory baselines beat No Memory\. Candidate retrieval and single\-memory selection are accurate enough that a substantial part of the remaining loss occurs after a relevant trajectory reaches the actor\. The length and rewrite analyses associate weaker direct reuse with long source traces and larger target differences\.QCRkeeps the workflow while requiring the actor to recover current bindings, a mechanism consistent with its higher success at lower online cost\.

Selection and reuse are separate stages: a ranker decides whether usable history reaches the actor, while the delivered representation determines whether it can be applied without copying source\-side values\.

## 6Discussion and Limitations

The experiments isolate the value of a verified prior trajectory after it has entered the memory pipeline\. Candidate selection is already strong: reranking selects reusable history for 94\.8% of targets and trails oracle reusable selection by 1\.8 success points\. The results are consistent with a remaining post\-selection cost when long histories and changed bindings expose the actor to source details that no longer apply\. This distinction matters for memory\-system design\. A store may preserve a complete record for evidence and provenance, while the acting prompt should contain a compact, target\-bound account of the reusable procedure\.

The study has a narrow boundary\. It evaluates successful source trajectories, a single selected memory, and controlled source–target binding shifts; it does not measure naturally recurring task histories, partial failures, multi\-memory composition, or open\-ended memory acquisition\. Those settings may change both the available procedures and the state that an agent must recover\. The reported token savings also do not mean that every task should use fewer tokens: safety\-sensitive targets can require additional checks\. We measure verified completion, but not irreversible side effects or policy violations caused by a reused trajectory\. We therefore treat cost as one outcome beside verified completion, rather than as a goal on its own\.

The comparison holds the retrieved candidate set, acting model, decoding, tool budget, and target\-state access fixed\. It attributes differences to the representation and use of the same selected trajectory; it does not by itself isolate every field of the support schema\. Future work can replace the embedding retriever or learn a reuse policy, but it should retain this accounting boundary and test whether the resulting help reduces target work without importing stale source bindings\.

## 7Conclusion

We study agent memory by asking how past experience helps a current query\. Long completed trajectories help, but raw delivery and a generic source\-only summary leave useful target\-specific work unresolved\. Given the same selected historical trajectory,QCRraises success to 62\.3% and reduces online tokens by 48\.9% relative to Full Trajectory\. The selection analysis shows that a reusable memory is available for 94\.8% of targets after ranking\. As histories become longer or target bindings move farther from the source, raw\-trajectory utility falls, while target\-bound support retains a larger measured gain\.

Memory systems should therefore preserve rich records in storage while giving the actor a compact, target\-bound account of the procedure, the bindings it must recover, and the checks that still apply\. The setting leaves room for better retrievers and learned reuse policies, but it makes their test clear: they must improve target success without hiding the cost of the help\.

This framing also changes how memory baselines should be interpreted\. A raw trajectory is not simply a stronger version of a short summary because it contains more tokens; it is an intervention that exposes an actor to both useful procedure and obsolete state\. Conversely, a short description is not automatically useful merely because it is cheap\. The relevant question is whether the information sent after retrieval lets the target agent take fewer unnecessary actions while still checking the values that changed\. By fixing candidate retrieval, target state, model, decoding, and tool budget, the present comparison evaluates that question after candidate retrieval has ended\. Future systems can use different stores, retrievers, or learned support writers, but should report the same distinction between what was retrieved, what was selected, what was delivered to the actor, and what the actor verified in the target environment\.

## References

- Baiet al\.\(2024\)Y\. Bai, X\. Lv, J\. Zhang, H\. Lyu, J\. Tang, Z\. Huang, Z\. Du, X\. Liu, A\. Zeng, L\. Hou, Y\. Dong, J\. Tang, and J\. LiLongBench: a bilingual, multitask benchmark for long context understanding\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,pp\. 3119–3137\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.172)Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- Chhikaraet al\.\(2025\)P\. Chhikara, D\. Khant, S\. Aryan, T\. Singh, and D\. YadavMem0: building production\-ready AI agents with scalable long\-term memory\.External Links:2504\.19413Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- Daiet al\.\(2019\)Z\. Dai, Z\. Yang, Y\. Yang, J\. Carbonell, Q\. V\. Le, and R\. SalakhutdinovTransformer\-xl: attentive language models beyond a fixed\-length context\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,pp\. 2978–2988\.External Links:[Document](https://dx.doi.org/10.18653/v1/P19-1285)Cited by:[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- DeepSeek \(2026\)DeepSeekDeepSeek api documentation: deepseek\-v4\-pro\.Note:https://api\-docs\.deepseek\.com/updates/Accessed July 29, 2026Cited by:[§4\.3](https://arxiv.org/html/2608.12847#S4.SS3.p2.1),[§5\.2](https://arxiv.org/html/2608.12847#S5.SS2.p1.1)\.
- Denget al\.\(2023\)X\. Deng, Y\. Gu, B\. Zheng, S\. Chen, S\. Stevens, B\. Wang, H\. Sun, and Y\. SuMind2Web: towards a generalist agent for the web\.External Links:2306\.06070Cited by:[§2\.3](https://arxiv.org/html/2608.12847#S2.SS3.p1.1)\.
- Drouinet al\.\(2024\)A\. Drouin, M\. Gasse, M\. Caccia, I\. H\. Laradji, M\. Del Verme, T\. Marty, L\. Boisvert, M\. Thakkar, Q\. Cappart, D\. Vazquez, N\. Chapados, and A\. LacosteWorkArena: how capable are web agents at solving common knowledge work tasks?\.External Links:2403\.07718Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.12847#S2.SS3.p1.1),[§5\.2](https://arxiv.org/html/2608.12847#S5.SS2.p1.1)\.
- Duet al\.\(2025\)Y\. Du, W\. Huang, D\. Zheng, Z\. Wang, S\. Montella, M\. Lapata, K\. Wong, and J\. Z\. PanRethinking memory in AI: taxonomy, operations, topics, and future directions\.External Links:2505\.00675Cited by:[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- Gutiérrezet al\.\(2024\)B\. J\. Gutiérrez, Y\. Shu, Y\. Gu, M\. Yasunaga, and Y\. SuHippoRAG: neurobiologically inspired long\-term memory for large language models\.InAdvances in Neural Information Processing Systems,External Links:[Document](https://dx.doi.org/10.52202/079017-1902)Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- Heet al\.\(2024\)H\. He, W\. Yao, K\. Ma, W\. Yu, Y\. Dai, H\. Zhang, Z\. Lan, and D\. YuWebVoyager: building an end\-to\-end web agent with large multimodal models\.External Links:2401\.13919Cited by:[§2\.3](https://arxiv.org/html/2608.12847#S2.SS3.p1.1)\.
- Huet al\.\(2025\)Y\. Hu, Y\. Wang, and J\. McAuleyEvaluating memory in LLM agents via incremental multi\-turn interactions\.External Links:2507\.05257Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- Huet al\.\(2026\)Y\. Hu, H\. Qian, S\. Wang, J\. Liu, Z\. Zhao, J\. Tan, Z\. Liu, and Z\. DouSAM: state\-adaptive memory for long\-horizon reasoning agent\.External Links:2605\.24468Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p1.1),[§2\.2](https://arxiv.org/html/2608.12847#S2.SS2.p1.1)\.
- Jimenezet al\.\(2024\)C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. R\. NarasimhanSWE\-bench: can language models resolve real\-world GitHub issues?\.InInternational Conference on Learning Representations,Cited by:[§2\.3](https://arxiv.org/html/2608.12847#S2.SS3.p1.1)\.
- Lewiset al\.\(2020\)P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Kuttler, M\. Lewis, W\. Yih, T\. Rocktäschel, S\. Riedel, and D\. KielaRetrieval\-augmented generation for knowledge\-intensive NLP tasks\.InAdvances in Neural Information Processing Systems,Vol\.33\.Cited by:[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- Liet al\.\(2026\)J\. Li, Y\. Zhang, X\. Yang, J\. Qu, J\. Xu, S\. Yang, J\. Ding, and E\. C\. NgaiOCR\-memory: optical context retrieval for long\-horizon agent memory\.External Links:2604\.26622Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p1.1),[§2\.2](https://arxiv.org/html/2608.12847#S2.SS2.p1.1)\.
- Liuet al\.\(2024a\)N\. F\. Liu, K\. Lin, J\. Hewitt, A\. Paranjape, M\. Bevilacqua, F\. Petroni, and P\. LiangLost in the middle: how language models use long contexts\.Transactions of the Association for Computational Linguistics12,pp\. 157–173\.External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00638)Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p3.1),[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- Liuet al\.\(2024b\)X\. Liu, H\. Yu, H\. Zhang, Y\. Xu, X\. Lei, H\. Lai, Y\. Gu, H\. Ding, K\. Men, K\. Yang, S\. Zhang, X\. Deng, A\. Zeng, Z\. Du, C\. Zhang, S\. Shen, T\. Zhang, Y\. Su, H\. Sun, M\. Huang, Y\. Dong, and J\. TangAgentBench: evaluating LLMs as agents\.InInternational Conference on Learning Representations,Cited by:[§2\.3](https://arxiv.org/html/2608.12847#S2.SS3.p1.1)\.
- Maharanaet al\.\(2024\)A\. Maharana, D\. Lee, S\. Tulyakov, M\. Bansal, F\. Barbieri, and Y\. FangEvaluating very long\-term conversational memory of LLM agents\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,pp\. 13851–13870\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.747)Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- Mialonet al\.\(2024\)G\. Mialon, C\. Fourrier, C\. Swift, T\. Wolf, Y\. LeCun, and T\. ScialomGAIA: a benchmark for general AI assistants\.InInternational Conference on Learning Representations,Cited by:[§2\.3](https://arxiv.org/html/2608.12847#S2.SS3.p1.1)\.
- Packeret al\.\(2023\)C\. Packer, S\. Wooders, K\. Lin, V\. Fang, S\. G\. Patil, I\. Stoica, and J\. E\. GonzalezMemGPT: towards LLMs as operating systems\.External Links:2310\.08560Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- Parket al\.\(2023\)J\. S\. Park, J\. C\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. BernsteinGenerative agents: interactive simulacra of human behavior\.InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology,pp\. 1–22\.External Links:[Document](https://dx.doi.org/10.1145/3586183.3606763)Cited by:[§2\.2](https://arxiv.org/html/2608.12847#S2.SS2.p1.1)\.
- Rawleset al\.\(2025\)C\. Rawles, S\. Clinckemaillie, Y\. Chang, J\. Waltz, G\. Lau, M\. Fair, A\. Li, W\. E\. Bishop, W\. Li, F\. Campbell\-Ajala, D\. K\. Toyama, R\. J\. Berry, D\. Tyamagundlu, T\. P\. Lillicrap, and O\. RivaAndroidWorld: a dynamic benchmarking environment for autonomous agents\.InInternational Conference on Learning Representations,Cited by:[§2\.3](https://arxiv.org/html/2608.12847#S2.SS3.p1.1)\.
- Shinnet al\.\(2023\)N\. Shinn, F\. Cassano, A\. Gopinath, K\. Narasimhan, and S\. YaoReflexion: language agents with verbal reinforcement learning\.InAdvances in Neural Information Processing Systems,Vol\.36\.Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p3.1),[§2\.2](https://arxiv.org/html/2608.12847#S2.SS2.p1.1)\.
- Tanet al\.\(2025\)H\. Tan, Z\. Zhang, C\. Ma, X\. Chen, Q\. Dai, and Z\. DongMemBench: towards more comprehensive evaluation on the memory of LLM\-based agents\.InFindings of the Association for Computational Linguistics: ACL 2025,External Links:2506\.21605Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- Trivediet al\.\(2024\)H\. Trivedi, T\. Khot, M\. Hartmann, R\. Manku, V\. Dong, E\. Li, S\. Gupta, A\. Sabharwal, and N\. BalasubramanianAppWorld: a controllable world of apps and people for benchmarking interactive coding agents\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,pp\. 16022–16076\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.850)Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.12847#S2.SS3.p1.1),[§5\.2](https://arxiv.org/html/2608.12847#S5.SS2.p1.1)\.
- Wanget al\.\(2023\)G\. Wang, Y\. Xie, Y\. Jiang, A\. Mandlekar, C\. Xiao, Y\. Zhu, L\. Fan, and A\. AnandkumarVoyager: an open\-ended embodied agent with large language models\.External Links:2305\.16291Cited by:[§2\.2](https://arxiv.org/html/2608.12847#S2.SS2.p1.1)\.
- Wanget al\.\(2024\)Z\. Z\. Wang, J\. Mao, D\. Fried, and G\. NeubigAgent workflow memory\.External Links:2409\.07429Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p3.1),[§2\.2](https://arxiv.org/html/2608.12847#S2.SS2.p1.1)\.
- Wuet al\.\(2025\)D\. Wu, H\. Wang, W\. Yu, Y\. Zhang, K\. Chang, and D\. YuLongMemEval: benchmarking chat assistants on long\-term interactive memory\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- Wuet al\.\(2022\)Y\. Wu, M\. N\. Rabe, D\. Hutchins, and C\. SzegedyMemorizing transformers\.InInternational Conference on Learning Representations,Cited by:[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- Xieet al\.\(2024\)T\. Xie, D\. Zhang, J\. Chen, X\. Li, S\. Zhao, R\. Cao, T\. J\. Hua, Z\. Cheng, D\. Shin, F\. Lei, Y\. Liu, Y\. Xu, S\. Zhou, S\. Savarese, C\. Xiong, V\. Zhong, and T\. YuOSWorld: benchmarking multimodal agents for open\-ended tasks in real computer environments\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 52040–52094\.Cited by:[§2\.3](https://arxiv.org/html/2608.12847#S2.SS3.p1.1)\.
- Xuet al\.\(2025\)W\. Xu, Z\. Liang, K\. Mei, H\. Gao, J\. Tan, and Y\. ZhangA\-mem: agentic memory for LLM agents\.External Links:2502\.12110Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- Yaoet al\.\(2025\)S\. Yao, N\. Shinn, P\. Razavi, K\. Narasimhan, and Sierraτ\\tau\-bench: a benchmark for tool\-agent\-user interaction in real\-world domains\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.12847#S2.SS3.p1.1)\.
- Yaoet al\.\(2023\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. CaoReAct: synergizing reasoning and acting in language models\.InInternational Conference on Learning Representations,Cited by:[§2\.2](https://arxiv.org/html/2608.12847#S2.SS2.p1.1)\.
- Yuet al\.\(2026\)Y\. Yu, L\. Yao, Y\. Xie, Q\. Tan, J\. Feng, Y\. Li, and L\. WuAgentic memory: learning unified long\-term and short\-term memory management for large language model agents\.External Links:2601\.01885Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p1.1),[§2\.2](https://arxiv.org/html/2608.12847#S2.SS2.p1.1)\.
- Zhanget al\.\(2024\)Z\. Zhang, X\. Bo, C\. Ma, R\. Li, X\. Chen, Q\. Dai, J\. Zhu, Z\. Dong, and J\. WenA survey on the memory mechanism of large language model based agents\.External Links:2404\.13501Cited by:[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- Zhaoet al\.\(2024\)A\. Zhao, D\. Huang, Q\. Xu, M\. Lin, Y\. Liu, and G\. HuangExpeL: LLM agents are experiential learners\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.38,pp\. 19632–19642\.Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p3.1),[§2\.2](https://arxiv.org/html/2608.12847#S2.SS2.p1.1)\.
- Zhenget al\.\(2024\)L\. Zheng, R\. Wang, X\. Wang, and B\. AnSynapse: trajectory\-as\-exemplar prompting with memory for computer control\.InInternational Conference on Learning Representations,Cited by:[§2\.2](https://arxiv.org/html/2608.12847#S2.SS2.p1.1)\.
- Zhonget al\.\(2024\)W\. Zhong, L\. Guo, Q\. Gao, H\. Ye, and Y\. WangMemoryBank: enhancing large language models with long\-term memory\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.38,pp\. 19724–19731\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v38i17.29946)Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.12847#S2.SS1.p1.1)\.
- Zhouet al\.\(2024\)S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried, U\. Alon, and G\. NeubigWebArena: a realistic web environment for building autonomous agents\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2608.12847#S1.p2.1),[§2\.3](https://arxiv.org/html/2608.12847#S2.SS3.p1.1),[§5\.2](https://arxiv.org/html/2608.12847#S5.SS2.p1.1)\.

## Appendix ATask Construction and Retrieval Protocol

The study uses a unified bank of verified historical trajectories\. For each target, retrieval returns a shared top\-55candidate set, and the ranker selects one trajectory before any memory condition runs\. The target state, model, decoding settings, tool budget, random seed, and verifier remain fixed across conditions\. Candidates may come from any of the three environments, but environment labels are excluded from the retrieval representation\.

Table A1:Memory\-bank and evaluation settings\.
Table A2:Source\-trajectory and target\-task inventory after exclusions\. Target counts enumerate unique target instances, not seed\-expanded runs\. Each reported result averages the three seed\-matched runs\. “Paired target” denotes the target constructed from the listed source trajectory; it does not identify the record ultimately selected by the ranker\.
### A\.1Binding\-Shift Construction

The no\-rewrite condition preserves the source intent\. Small, medium, and large rewrites change one local binding, two or three bindings or one central constraint, and at least four bindings or both the target entity and initial environment state, respectively\. The rewrite procedure preserves the task family while requiring the acting agent to ground values in the target\.

Table A3:Distribution of source–target binding rewrites\.
Table A4:Retrieval and reranking configuration\. The candidate bank is pooled across environments, while environment labels are excluded from the retrieval input\.

## Appendix BModel Configuration

All conditions use DeepSeek\-V4\-Pro for generation and ranking\. The source rollout and target actor sample at temperature0\.20\.2; all support\-writing and selection steps use deterministic decoding\.

Table A5:Model and decoding settings\.

## Appendix CAnnotation and Rebinding Analysis

A stale\-binding error occurs when an action, output, or tool argument repeats a source\-specific value that conflicts with the target query or a target\-side observation\. Correct rebinding requires the actor to recover the target value before using it\. Annotators inspect the source record, target task, target\-side observations, and executed action trace rather than inferring a label from final success alone\.

Table A6:Rebinding analysis for large binding rewrites\.
Table A7:Annotation statistics for source–target construction and rebinding audits\. Candidate relevance is labelled at the returned\-record level using the rubric in Table[A9](https://arxiv.org/html/2608.12847#A5.T9); its count is not included in the 2,391 source–target audit units\.

## Appendix DQCR Support Construction

QCR writes a compact target\-bound note from the selected trajectory, target query, and initial target observation\. It does not access the environment or the verifier while preparing the note\. The actor must ground every required target value through the query, current observation, or a later tool call\.

Table A8:QCR support schema\.
### D\.1Support\-Writer Prompt Template

The following prompt template specifies the information supplied to the QCR writer in this study\.

> You receive one historical trajectory, the current target query, and the initial target observation\. Write a short support note with four labeled fields: \(1\) workflow invariant, \(2\) bindings to re\-obtain, \(3\) applicability conditions, and \(4\) verification guardrail\. Treat historical identifiers, paths, users, dates, tool outputs, and environment state as source\-side evidence, not target answers\. Do not infer hidden target state, call tools, or copy a historical binding into the target\. If the source procedure does not apply, state the condition that blocks reuse\.

The actor receives the note as advisory information\. It may follow the workflow only after checking the target\-side preconditions, and it must recover each listed binding before submitting an action that depends on it\.

## Appendix ERetrieval Labels and Selection Details

Blind relevance annotation distinguishes a lexical match from a trajectory that supplies an executable procedure\. The source–target pairing remains sealed metadata and does not appear in the retrieved record, ranker input, or acting prompt\.

Table A9:Relevance labels used for retrieval and selection audits\.
Table A10:Retrieval and single\-memory selection diagnostics\. Coverage and accuracy entries are proportions of unique target instances\. End\-task Success entries average within environment and seed, then take the unweighted mean across WebArena, WorkArena, and AppWorld\.

## Appendix FStratified Utility Results

The following tables give the values underlying the length and binding\-shift analyses in the main paper\. Each entry reports the percentage\-point change in Success relative to No Memory within the same stratum\.

Table A11:Memory utility by selected\-memory trajectory length\.
Table A12:Memory utility under binding shift\.

## Appendix GPer\-Run Audit\-Ledger Schema

The evaluation ledger schema specifies one row for each target\-condition execution\. It keeps the retrieval boundary auditable: the same candidate set and selected record must appear across the compared memory conditions for a given target\. The schema also separates a failed execution from an invalid reuse decision\.

Table A13:Fields retained for every target\-condition run\.
Source\-rollout cost remains outside the online total because it belongs to offline bank construction\. Online accounting includes the non\-memory acting input, delivered memory, QCR\-writer input and output, and actor output\. The acting API input contains the first two terms only, so it is recorded as a breakdown rather than counted again\.

### G\.1Information Boundary

The retriever indexes visible source instructions and trajectory descriptors\. It does not index environment labels, task identities, sealed source–target relations, verifier diagnostics, or target\-conditioned summaries\. QCR may read only the selected source record, target query, and initial target observation; it cannot query the environment privately, replace the selected record, or see hidden evaluator information\. The acting agent pays for any target\-side discovery through its own tool calls\.

#### Case\-study reporting\.

Qualitative examples should show the source procedure and the changed target binding side by side\. A useful pair contains one Full Trajectory failure that copies a stale value and one QCR success that recovers the value from current evidence\. Private identifiers, credentials, and hidden evaluator material must be redacted before release\.

Table A14:Required layout for each released source–target case study\.

Similar Articles