Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents
Summary
VerMem is a framework for unified memory management in LLM agents, using local and global verifiers with reinforcement learning to jointly control long-term and short-term memory. It outperforms strong baselines across five benchmarks with improved efficiency-performance trade-offs.
View Cached Full Text
Cached at: 08/05/26, 07:39 AM
# Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents
Source: [https://arxiv.org/html/2608.03137](https://arxiv.org/html/2608.03137)
Xiaolong Sun1Qichao Wang2Hangyu Li3Liang Chen1,∗ 1Sun Yat\-Sen University, Guangzhou, China 2Nanyang Technological University, Singapore 3Tencent, Shenzhen, China sunxlong@mail2\.sysu\.edu\.cnqichao001@e\.ntu\.edu\.sg masonhyli@tencent\.comchenliang6@mail\.sysu\.edu\.cn ∗Corresponding author
###### Abstract
Large language model \(LLM\) agents must retain reusable information, control a bounded active context, and recover earlier evidence during long\-horizon interaction\. Existing methods commonly optimize long\-term memory \(LTM\) and short\-term memory \(STM\) separately, while unified policies are often trained primarily with trajectory\-level feedback, which provides weak credit for individual memory decisions\. We present Verifiable Memory \(VerMem\), a framework that represents LTM, active context, and episodic history as distinct states and controls them with one memory operation policy\. Seven atomic operations let the policy add, revise, or soft\-delete LTM entries; retrieve LTM into the active context; filter or summarize the active context; and restore selected episodic fragments\. VerMem is initialized by supervised fine\-tuning and trained with a three\-stage reinforcement\-learning curriculum\. The local verifier scores executable memory transitions, and a global verifier assesses evidence coherence and terminal\-memory consistency after task completion\. These scores are combined with programmatically computed task, evidence\-recall, efficiency, and constraint signals through hierarchical credit assignment\. The verifiers are used only during training\. Across five benchmarks and two LLM backbones, VerMem achieves the best result on the vast majority of reported metrics and consistently outperforms strong memory baselines\. Under controlled online\-token budgets on three interactive benchmarks, it also achieves the strongest efficiency–performance frontier among the compared methods\. Code is available at[https://github\.com/Sun\-SYSU\-24/VerMem](https://github.com/Sun-SYSU-24/VerMem)\.
## 1Introduction
In long\-horizon tasks involving multi\-step interaction, tool use, and complex reasoning, the performance of large language model \(LLM\) agents depends not only on single\-step generation, but also on retaining and using task\-relevant information across decisions\. Prior work usesagent memoryto describe the information that an agent can attend to and use at a given time\(Yu et al\.,[2026](https://arxiv.org/html/2608.03137#bib.bib28)\)\. Such information includes stored past executions and state representations formed from interaction history\(Xiong et al\.,[2026](https://arxiv.org/html/2608.03137#bib.bib23); Goodyear et al\.,[2025](https://arxiv.org/html/2608.03137#bib.bib4)\)\. Agent memory is commonly divided into long\-term memory and short\-term memory\(Yu et al\.,[2026](https://arxiv.org/html/2608.03137#bib.bib28)\)\. Long\-term memory \(LTM\) persistently stores user\- or task\-specific knowledge for reuse in later turns, stages, or tasks\(Jiang et al\.,[2024](https://arxiv.org/html/2608.03137#bib.bib6); Zhong et al\.,[2024](https://arxiv.org/html/2608.03137#bib.bib29)\)\. Short\-term memory \(STM\) contains information in the current input context and supports ongoing reasoning, action selection, and context control\(Wu et al\.,[2025b](https://arxiv.org/html/2608.03137#bib.bib22); Gao et al\.,[2025](https://arxiv.org/html/2608.03137#bib.bib3)\)\. LTM preserves information across time, while STM keeps the current reasoning state usable\. Their coordination is therefore central to long\-horizon agent reasoning\.
Existing systems often optimize only one side of this process\. STM\-oriented methods organize or compress the current reasoning state\(Qian et al\.,[2026](https://arxiv.org/html/2608.03137#bib.bib14); Li et al\.,[2025](https://arxiv.org/html/2608.03137#bib.bib8)\), whereas LTM\-oriented methods build external stores for persistent recall\(Zhong et al\.,[2024](https://arxiv.org/html/2608.03137#bib.bib29); Xu et al\.,[2025](https://arxiv.org/html/2608.03137#bib.bib24)\)\. Stored information, however, affects a decision only after it enters the active context at the right time, and reusing task\-misaligned experience can propagate errors\(Xiong et al\.,[2026](https://arxiv.org/html/2608.03137#bib.bib23)\)\. The AgeMem policy introduced byYu et al\. \([2026](https://arxiv.org/html/2608.03137#bib.bib28)\)places LTM and STM operations in one policy, but its STM actions focus on retrieval, filtering, and summarization\. Long\-horizon tasks may also require an earlier observation, tool output, or intermediate conclusion to return from episodic history after it has left the active context\. This operation differs from retrieving persistent knowledge and from compressing recent context\.
Learning such behavior introduces a second difficulty\.Yan et al\. \([2026](https://arxiv.org/html/2608.03137#bib.bib25)\),Zhou et al\. \([2025](https://arxiv.org/html/2608.03137#bib.bib30)\), andYu et al\. \([2026](https://arxiv.org/html/2608.03137#bib.bib28)\)optimize memory decisions with downstream outcomes or trajectory\-level feedback\. These signals compare complete rollouts, but do not identify which atomic transition was useful or harmful\. Process supervision and transition\-level credit improve intermediate feedback\(Setlur et al\.,[2025](https://arxiv.org/html/2608.03137#bib.bib17); Luo et al\.,[2025](https://arxiv.org/html/2608.03137#bib.bib10)\), whileMiroMind Team \([2026](https://arxiv.org/html/2608.03137#bib.bib12)\)combines local and global verification during inference\. Operation\-level and trajectory\-level verification have not been jointly used to train a unified memory operation policy\.
These limitations leave three challenges for a unified and trainable memory operation policy\.\(A\) Heterogeneous memory coordination\.LTM determines what should be stored, revised, or removed\. STM determines what should enter, remain in, or leave the active context\. The two memory types operate over different states and time scales\. Coordinating them within a shared decision process while preserving their functional boundaries remains difficult\.\(B\) Historical\-context recovery\.The current prompt is only a budgeted view of the full task history\. Early observations, tool outputs, and intermediate conclusions may become relevant again after the subgoal changes\. Recovering the relevant fragments without reintroducing large amounts of irrelevant history is therefore nontrivial\.\(C\) Multi\-granularity credit assignment\.The utility of a memory operation is often delayed and depends on later operations\. The correct write may become useful only after a later retrieval, while an incorrect filter may cause failure several steps later\. Training must assign operation\-specific credit while preserving alignment with the complete task objective\.
To address these challenges, we proposeVerifiable Memory \(VerMem\), shown in Figure[1](https://arxiv.org/html/2608.03137#S3.F1)\. VerMem preserves LTM, active context, and episodic history as distinct states, while one memory operation policy coordinates LTM maintenance, STM control, and historical\-context recovery through seven atomic tools\. Training uses a supervised fine\-tuning \(SFT\) warmup followed by a three\-stage reinforcement learning curriculum for LTM, STM, and their joint use\. The local verifier evaluates each realized memory transition, and the global verifier evaluates the completed trajectory\. Their separately normalized advantages are combined at every memory decision\. Both verifiers are removed during inference\. Across five long\-horizon benchmarks and two backbones, VerMem consistently improves task performance and achieves a stronger efficiency–performance frontier\.
The main contributions of this work are summarized as follows:
- ∙\\bulletWe proposeVerMem, a unified agent memory management framework\.VerMem coordinates LTM maintenance and STM control through a single memory operation policy while preserving distinct states for LTM, active context, and episodic history\. Its atomic action space further supports episode\-level context selection\.
- ∙\\bulletWe introduce local–global verifier\-guided hierarchical credit assignment\.VerMemconstructs operation\-level local advantages and trajectory\-level global advantages from stateful multi\-step rollouts\. The two signals are combined at each memory decision, allowing operation quality and final task utility to jointly guide policy updates\.
- ∙\\bulletWe conduct systematic evaluations across diverse complex tasks\.We compareVerMemwith strong memory baselines on five benchmarks and evaluate its efficiency–performance trade\-off under controlled online\-token budgets\. Ablations further examine unified LTM/STM management, local and global credit signals, and the multi\-component reward design\.
## 2Related Work
### 2\.1Long\-Term Memory \(LTM\)
Long\-term memory research studies how agents preserve, organize, and reuse historical information beyond the active context\.LangChain Team \([2025](https://arxiv.org/html/2608.03137#bib.bib7)\)describes LangMem as a set of modular primitives for recording, searching, and consolidating persistent knowledge\.Zhong et al\. \([2024](https://arxiv.org/html/2608.03137#bib.bib29)\)introduce MemoryBank for long\-term recall and user profiling through continual memory updates\. The Mem0 framework ofChhikara et al\. \([2025](https://arxiv.org/html/2608.03137#bib.bib2)\)extracts, consolidates, and retrieves salient information from extended conversations, while the A\-Mem framework ofXu et al\. \([2025](https://arxiv.org/html/2608.03137#bib.bib24)\)organizes memory through dynamic indexing, linking, and evolution\. These systems improve persistent\-state construction, but their primary concern is how information is stored and organized\. Stored knowledge affects current decisions only after it enters the active context at the appropriate stage\. Incorrect or task\-misaligned experiences may also propagate errors when reused\(Xiong et al\.,[2026](https://arxiv.org/html/2608.03137#bib.bib23)\)\. This motivates joint control of persistent\-memory maintenance and active\-context use rather than isolated optimization of LTM\.
### 2\.2Short\-Term Memory \(STM\)
Short\-term memory research focuses on maintaining the active context required by ongoing reasoning under a bounded context budget\.Qian et al\. \([2026](https://arxiv.org/html/2608.03137#bib.bib14)\)organize reasoning traces and transient tool outputs into a compact executive memory\.Li et al\. \([2025](https://arxiv.org/html/2608.03137#bib.bib8)\)develop structured schemata and activate query\-relevant information for long\-document understanding\.Wu et al\. \([2025a](https://arxiv.org/html/2608.03137#bib.bib21)\)periodically compress growing interaction histories into concise reasoning states and further adapt agents to reason over these summaries with ReSum\-GRPO\. These methods improve context usability, but they remain centered on the current input, document, or ongoing trajectory\. In long\-horizon tasks, useful evidence may remain in the episodic history after it leaves the active context\. Retrieval, filtering, and summarization alone do not fully specify which earlier episode should return to the current reasoning state\.VerMemaddresses this gap by selecting task\-relevant fragments from episodic history and materializing them into the active context as an explicit STM decision\.
### 2\.3Learning Memory Policies with Verification
Reinforcement learning turns memory management into a learnable decision process\.Yan et al\. \([2026](https://arxiv.org/html/2608.03137#bib.bib25)\)learn structured updates,Huo et al\. \([2026](https://arxiv.org/html/2608.03137#bib.bib5)\)expose atomic memory operations,Zhou et al\. \([2025](https://arxiv.org/html/2608.03137#bib.bib30)\)jointly learn consolidation and reasoning, andYu et al\. \([2026](https://arxiv.org/html/2608.03137#bib.bib28)\)train a unified LTM/STM policy\. These methods primarily optimize memory behavior with answer\-level or trajectory\-level objectives\. Process supervision evaluates intermediate reasoning progress\(Setlur et al\.,[2025](https://arxiv.org/html/2608.03137#bib.bib17)\), whileLuo et al\. \([2025](https://arxiv.org/html/2608.03137#bib.bib10)\)andPeng et al\. \([2026](https://arxiv.org/html/2608.03137#bib.bib13)\)assign credit to agent transitions or hierarchical decisions\.MiroMind Team \([2026](https://arxiv.org/html/2608.03137#bib.bib12)\)applies local and global verification during inference\. VerMem instead uses operation\-level and trajectory\-level verification signals to train the memory operation policy\.
## 3Method
We propose Verifiable Memory \(VerMem\), a unified framework that maintains persistent LTM, a budgeted active context, and episodic history as distinct states\. As shown in Figure[1](https://arxiv.org/html/2608.03137#S3.F1), one memory operation policy coordinates LTM maintenance, STM control, and historical\-context recovery\. The policy is initialized with SFT and optimized through a three\-stage reinforcement learning curriculum\. The local verifier supplies operation\-level semantic scores, while the composite global branch supplies trajectory\-level credit from programmatic and global\-verifier components\. Both verifiers are removed during inference\.
Figure 1:Overview of VerMem\. Left: the memory operation policy coordinates long\-term and short\-term memory tools to maintain long\-term memory and the active context\. Historical\-context recovery restores relevant episodic history for the LLM, whose actions and answers are recorded for subsequent steps\. Right: a three\-stage curriculum generatesKKstateful rollouts\. The local verifier evaluates executable memory transitions, while the global verifier assesses evidence coherence and terminal\-memory consistency after task completion\. These scores contribute to separately normalized local and global advantages; constraint costs remain separate\. The verifier branch is used only during training\.### 3\.1Problem Formulation
Unified memory operation policy formulation\.The fixed task specification is denoted byqq\. At memory\-decision steptt, the long\-term memoryMtM\_\{t\}, active contextCtC\_\{t\}, and same\-task episodic historyHtH\_\{t\}define the conceptual state and its policy\-visible serialization:
st\\displaystyle s\_\{t\}=\(q,Mt,Ct,Ht\),\\displaystyle=\(q,M\_\{t\},C\_\{t\},H\_\{t\}\),\(1\)xt\\displaystyle x\_\{t\}=g\(st\)\.\\displaystyle=g\(s\_\{t\}\)\.Here,MtM\_\{t\}stores persistent information for reuse in later steps or stages,CtC\_\{t\}is the budgeted context visible to the task solver, andHtH\_\{t\}preserves the ordered record of observations, task actions, tool outputs, intermediate conclusions, and memory decisions from the same task\. SeparatingCtC\_\{t\}fromHtH\_\{t\}allows information to leave the active context without being removed from that record\. The functionggretains the task specification and current task state, then serializes budgeted views ofMtM\_\{t\},CtC\_\{t\}, and same\-taskHtH\_\{t\}within the 8,192\-token policy\-state limit used by the protocol\. It does not concatenateHtH\_\{t\}without bound\. Equation \([1](https://arxiv.org/html/2608.03137#S3.E1)\) therefore distinguishes the conceptual state from the bounded input presented to the policy\.
Letvtv\_\{t\}denote an operation type andξt\\xi\_\{t\}its structured arguments\. The complete generated memory command isut=\(vt,ξt\)u\_\{t\}=\(v\_\{t\},\\xi\_\{t\}\), wherevt∈𝒱∪\{∅\}v\_\{t\}\\in\\mathcal\{V\}\\cup\\\{\\varnothing\\\}, and𝒱\\mathcal\{V\}denotes the seven operation types formalized in Equation \([6](https://arxiv.org/html/2608.03137#S3.E6)\)\. The null command has empty arguments\. The complete command and memory transition are
ut\\displaystyle u\_\{t\}=\(vt,ξt\),\\displaystyle=\(v\_\{t\},\\xi\_\{t\}\),\(2\)ut\\displaystyle u\_\{t\}∼πθ\(⋅∣xt\),\\displaystyle\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x\_\{t\}\),s~t\\displaystyle\\widetilde\{s\}\_\{t\}=𝒯m\(st,ut\)\.\\displaystyle=\\mathcal\{T\}\_\{m\}\(s\_\{t\},u\_\{t\}\)\.For a valid non\-null command,𝒯m\\mathcal\{T\}\_\{m\}commits the specified memory transition\. Forvt=∅v\_\{t\}=\\varnothing,MtM\_\{t\}andCtC\_\{t\}are unchanged,and the memory\-state transition is an identity transition, sos~t=st\\tilde\{s\}\_\{t\}=s\_\{t\}\. The identity decision is retained in the rollout record but does not create a memory\-state modification\. The null command remains a trainable decision in the rollout record\. Rejected non\-null commands leaveMtM\_\{t\}andCtC\_\{t\}unchanged and record the failed decision and its reason inHtH\_\{t\}\. They also incur a constraint cost\. Thus, Equation \([2](https://arxiv.org/html/2608.03137#S3.E2)\) definess~t\\widetilde\{s\}\_\{t\}on every branch\.
After the memory transition, the task solver acts on the resulting active context, the environment returns an observation, and a separate interaction transition constructs the next memory\-decision state:
bt\\displaystyle b\_\{t\}∼πϕtask\(⋅∣q,C~t\),\\displaystyle\\sim\\pi\_\{\\phi\}^\{\\mathrm\{task\}\}\(\\cdot\\mid q,\\widetilde\{C\}\_\{t\}\),zt\+1\\displaystyle z\_\{t\+1\}∼𝒫env\(⋅∣bt\),\\displaystyle\\sim\\mathcal\{P\}\_\{\\mathrm\{env\}\}\(\\cdot\\mid b\_\{t\}\),\(3\)st\+1\\displaystyle s\_\{t\+1\}=𝒯e\(s~t,bt,zt\+1\)\.\\displaystyle=\\mathcal\{T\}\_\{e\}\(\\widetilde\{s\}\_\{t\},b\_\{t\},z\_\{t\+1\}\)\.The transition𝒯e\\mathcal\{T\}\_\{e\}appends the task action and resulting observation toHtH\_\{t\}and constructs the next active context\. The memory transition in Equation \([2](https://arxiv.org/html/2608.03137#S3.E2)\) and the task/environment transition in Equation \([3](https://arxiv.org/html/2608.03137#S3.E3)\) have distinct roles\. Repeating them produces a task trajectoryτ\\tau\. Unified management shares the state and task objective but does not merge LTM, active context, and episodic history into one storage structure\.
Progressive training and hierarchical credit\.The policy follows
θSFT→θA→θB→θC=θ⋆\.\\theta\_\{\\mathrm\{SFT\}\}\\rightarrow\\theta\_\{A\}\\rightarrow\\theta\_\{B\}\\rightarrow\\theta\_\{C\}=\\theta^\{\\star\}\.\(4\)Phase A trains LTM maintenance, Phase B trains STM control under distractors, and Phase C trains their coordination in complete tasks\. Equation \([4](https://arxiv.org/html/2608.03137#S3.E4)\) shows the parameter curriculum\. For decisionttin trajectorykk, the local verifier produces a raw operation score for an executable command, while the global branch combines post\-trajectory verifier scores with programmatic task, evidence\-recall, and efficiency signals\. Normalization converts the two branches intoAt,klocalA\_\{t,k\}^\{\\mathrm\{local\}\}andAkglobalA\_\{k\}^\{\\mathrm\{global\}\}\. The executor separately records hard violations asct,kc\_\{t,k\}\. Their credit is
At,khier=λlocalAt,klocal\+λglobalAkglobal−λconstraintct,k\.A\_\{t,k\}^\{\\mathrm\{hier\}\}=\\lambda\_\{\\mathrm\{local\}\}A\_\{t,k\}^\{\\mathrm\{local\}\}\+\\lambda\_\{\\mathrm\{global\}\}A\_\{k\}^\{\\mathrm\{global\}\}\-\\lambda\_\{\\mathrm\{constraint\}\}c\_\{t,k\}\.\(5\)We setλlocal=λglobal=λconstraint=1\\lambda\_\{\\mathrm\{local\}\}=\\lambda\_\{\\mathrm\{global\}\}=\\lambda\_\{\\mathrm\{constraint\}\}=1\. The local term distinguishes executable commands within one trajectory, while the global term preserves alignment with the completed task\. Equation \([5](https://arxiv.org/html/2608.03137#S3.E5)\) keeps the local, global, and constraint channels distinct, including for rejected commands\.
### 3\.2Atomic Memory Tools
The operation\-type sets are
𝒱LTM\\displaystyle\\mathcal\{V\}\_\{\\mathrm\{LTM\}\}=\{Add,Update,\\displaystyle=\\\{\\textsc\{Add\},\\textsc\{Update\},\(6\)Delete\},\\displaystyle\\hskip 11\.99998pt\\textsc\{Delete\}\\\},𝒱STM\\displaystyle\\mathcal\{V\}\_\{\\mathrm\{STM\}\}=\{Retrieve,Filter,\\displaystyle=\\\{\\textsc\{Retrieve\},\\textsc\{Filter\},SelectEpisode,Summarize\},\\displaystyle\\hskip 11\.99998pt\\textsc\{SelectEpisode\},\\textsc\{Summarize\}\\\},𝒱\\displaystyle\\mathcal\{V\}=𝒱LTM∪𝒱STM\.\\displaystyle=\\mathcal\{V\}\_\{\\mathrm\{LTM\}\}\\cup\\mathcal\{V\}\_\{\\mathrm\{STM\}\}\.VerMem exposes these seven atomic tools, summarized in Table[1](https://arxiv.org/html/2608.03137#S3.T1)and formalized in Equation \([6](https://arxiv.org/html/2608.03137#S3.E6)\), together with the null action∅\\varnothing\. The LTM tools add, revise, or soft\-delete persistent information\. The STM tools retrieve LTM content, filter or summarize the active context, and restore relevant same\-task fragments fromHtH\_\{t\}through SelectEpisode\. The unified action space lets the policy choose between persistent\-memory maintenance and active\-context control from the current state\. Detailed schemas, preconditions, and transition checks are given in Supplementary Material, Section A\.
ToolTargetFunctionAddLTMAdd toMtM\_\{t\}UpdateLTMCreate revised versions inMtM\_\{t\}DeleteLTMSoft\-delete entries inMtM\_\{t\}RetrieveSTMRetrieveMt→CtM\_\{t\}\\rightarrow C\_\{t\}FilterSTMFilterCtC\_\{t\}SelectEpisodeSTMRestoreHt→CtH\_\{t\}\\rightarrow C\_\{t\}SummarizeSTMSummarizeCtC\_\{t\}Table 1:Atomic memory tools in VerMem\.
### 3\.3Training Pipeline
Curriculum and SFT warmup\.VerMem uses an SFT warmup followed by three reinforcement learning phases\. The supervised set𝒟SFT=\{\(xi,vi⋆,ξi⋆\)\}i=1N\\mathcal\{D\}\_\{\\mathrm\{SFT\}\}=\\\{\(x\_\{i\},v\_\{i\}^\{\\star\},\\xi\_\{i\}^\{\\star\}\)\\\}\_\{i=1\}^\{N\}contains validated targets for the seven tools and∅\\varnothing, whereξi⋆\\xi\_\{i\}^\{\\star\}denotes structured arguments andui⋆=\(vi⋆,ξi⋆\)u\_\{i\}^\{\\star\}=\(v\_\{i\}^\{\\star\},\\xi\_\{i\}^\{\\star\}\)\. The objective is
ℒSFT=−𝔼\(xi,vi⋆,ξi⋆\)∼𝒟SFT\[logπθ\(vi⋆,ξi⋆∣xi\)\]\.\\mathcal\{L\}\_\{\\mathrm\{SFT\}\}=\-\\mathbb\{E\}\_\{\(x\_\{i\},v\_\{i\}^\{\\star\},\\xi\_\{i\}^\{\\star\}\)\\sim\\mathcal\{D\}\_\{\\mathrm\{SFT\}\}\}\\left\[\\log\\pi\_\{\\theta\}\(v\_\{i\}^\{\\star\},\\xi\_\{i\}^\{\\star\}\\mid x\_\{i\}\)\\right\]\.\(7\)Equation \([7](https://arxiv.org/html/2608.03137#S3.E7)\) trains both operation selection and structured argument generation\. Boundary templates and candidates from a frozen DeepSeek\-V3\.2\(Liu et al\.,[2025](https://arxiv.org/html/2608.03137#bib.bib9)\)teacher are retained only when they satisfy the tool schema and transition constraints\. The warmup teaches tool selection, argument generation, and valid structured outputs, but does not determine the delayed task utility of a memory decision\. Data construction, filtering, and coverage are detailed in Supplementary Material, Section D\.
Phase\-wise reinforcement learning\.Phase A optimizes𝒱LTM∪\{∅\}\\mathcal\{V\}\_\{\\mathrm\{LTM\}\}\\cup\\\{\\varnothing\\\}on construction, revision, deletion, and no\-change decisions\. Phase B optimizes𝒱STM∪\{∅\}\\mathcal\{V\}\_\{\\mathrm\{STM\}\}\\cup\\\{\\varnothing\\\}on retrieval, filtering, summarization, and historical\-context recovery under distractors\. Phase C enables𝒱∪\{∅\}\\mathcal\{V\}\\cup\\\{\\varnothing\\\}in complete multi\-step tasks, where LTM and STM operations may alternate\. Only memory operation policy parameters are transferred across phases;MtM\_\{t\},CtC\_\{t\}, andHtH\_\{t\}are initialized independently for every training instance\.
Stateful trajectory collection\.At policy updatennof phasep∈\{A,B,C\}p\\in\\\{A,B,C\\\}, the current policy samplesKKtrajectories from the same phase\-specific initial state:
τk,p\(q,n\)\\displaystyle\\tau\_\{k,p\}^\{\(q,n\)\}∼ℙp,n\(τ∣s0,p\(q\)\),k=1,…,K,\\displaystyle\\sim\\mathbb\{P\}\_\{p,n\}\\left\(\\tau\\mid s\_\{0,p\}^\{\(q\)\}\\right\),\\quad k=1,\\ldots,K,\(8\)𝒢q,p\(n\)\\displaystyle\\mathcal\{G\}\_\{q,p\}^\{\(n\)\}=\{τ1,p\(q,n\),…,τK,p\(q,n\)\}\.\\displaystyle=\\\{\\tau\_\{1,p\}^\{\(q,n\)\},\\ldots,\\tau\_\{K,p\}^\{\(q,n\)\}\\\}\.Each rollout stores both the post\-memory states~t\\widetilde\{s\}\_\{t\}and the post\-interaction statest\+1s\_\{t\+1\}, so that memory\-transition credit is not conflated with the subsequent task action\. The candidate group in Equation \([8](https://arxiv.org/html/2608.03137#S3.E8)\) supports subsequent local and global credit assignment\. Phase\-specific states, termination rules, and token masks are provided in Supplementary Material, Sections D and F\.
### 3\.4Local and Global Verifier\-Guided Hierarchical Credit Assignment
We optimize the memory operation policy with local and global credit signals\. For each taskqq, theKKtrajectories sampled from the same initial state form the candidate group𝒢q=\{τ1\(q\),…,τK\(q\)\}\\mathcal\{G\}\_\{q\}=\\\{\\tau\_\{1\}^\{\(q\)\},\\ldots,\\tau\_\{K\}^\{\(q\)\}\\\}\. Letℬop\\mathcal\{B\}\_\{\\mathrm\{op\}\}contain every memory decision collected in the current update, including null and rejected commands, and letℬexec⊆ℬop\\mathcal\{B\}\_\{\\mathrm\{exec\}\}\\subseteq\\mathcal\{B\}\_\{\\mathrm\{op\}\}contain valid non\-null commands and valid null decisions\. Structural, safety, and budget violations are recorded separately asct,kc\_\{t,k\}\. For\(q,k,t\)∈ℬexec\(q,k,t\)\\in\\mathcal\{B\}\_\{\\mathrm\{exec\}\}, the local verifier producesrt,klocalr\_\{t,k\}^\{\\mathrm\{local\}\}, which evaluates task relevance, evidence grounding, local progress, and information fidelity\. The null decision is treated as a valid identity command and is scored with its own rubric\. Rejected commands are not sent to the local verifier; their local advantage is set to zero\. After task termination, the global trajectory branch produces the composite scorerkglobalr\_\{k\}^\{\\mathrm\{global\}\}\. Its verifier and programmatic components are defined below\.
The composite global score is normalized within the candidate group for the same task\. Local scores are normalized by operation type\. The decision subset and two advantages are
ℬv\\displaystyle\\mathcal\{B\}\_\{v\}=\{\(q,k,t\)∈ℬexec:vt,k=v\},\\displaystyle=\\\{\(q,k,t\)\\in\\mathcal\{B\}\_\{\\mathrm\{exec\}\}:v\_\{t,k\}=v\\\},\(9\)Akglobal\\displaystyle A\_\{k\}^\{\\mathrm\{global\}\}=rkglobal−μqglobalσqglobal\+ϵ,\\displaystyle=\\frac\{r\_\{k\}^\{\\mathrm\{global\}\}\-\\mu\_\{q\}^\{\\mathrm\{global\}\}\}\{\\sigma\_\{q\}^\{\\mathrm\{global\}\}\+\\epsilon\},At,klocal\\displaystyle A\_\{t,k\}^\{\\mathrm\{local\}\}=rt,klocal−μvt,klocalσvt,klocal\+ϵ\.\\displaystyle=\\frac\{r\_\{t,k\}^\{\\mathrm\{local\}\}\-\\mu\_\{v\_\{t,k\}\}^\{\\mathrm\{local\}\}\}\{\\sigma\_\{v\_\{t,k\}\}^\{\\mathrm\{local\}\}\+\\epsilon\}\.The global statistics are computed from theKKtrajectories in𝒢q\\mathcal\{G\}\_\{q\}, while local statistics are computed from operations of the same type in the current update batch\. This operation\-wise normalization reduces scale differences across tool\-specific verifier rubrics\. When an update contains too few instances of an operation type, phase\-specific running statistics are used\. The corresponding thresholds and update rules are provided in Supplementary Material, Section F\. The local expression in Equation \([9](https://arxiv.org/html/2608.03137#S3.E9)\) is defined only for decisions inℬexec\\mathcal\{B\}\_\{\\mathrm\{exec\}\}\. For rejected commands, it is not evaluated andAt,klocal=0A\_\{t,k\}^\{\\mathrm\{local\}\}=0\.
At each memory decision, Equation \([5](https://arxiv.org/html/2608.03137#S3.E5)\) combines the two advantages with the constraint cost\. The local advantage is assigned only to the current atomic operation, while the global advantage provides task\-level supervision to all operations in the trajectory\. Different memory decisions within one trajectory can therefore receive different learning signals\. Poorly rated operations in a successful trajectory therefore need not receive the trajectory’s full positive credit, while useful operations in an unsuccessful trajectory may retain a positive local contribution\. The final sign still depends on the weighted sum of the local, global, and constraint terms\. Rejected commands are not evaluated by the local semantic verifier, soAt,klocal=0A\_\{t,k\}^\{\\mathrm\{local\}\}=0\. They still receive the trajectory\-levelAkglobalA\_\{k\}^\{\\mathrm\{global\}\}, whilect,kc\_\{t,k\}provides their direct operation\-level penalty\.
Letyt,k,1:Lt,ky\_\{t,k,1:L\_\{t,k\}\}be the generated tokens that serialize the complete commandut,ku\_\{t,k\}, including its structured arguments, and letht,k,j=\(xt,k,yt,k,<j\)h\_\{t,k,j\}=\(x\_\{t,k\},y\_\{t,k,<j\}\)be the token prefix\. The token\-level policy ratio and KL term are
ρt,k,j\(θ\)\\displaystyle\\rho\_\{t,k,j\}\(\\theta\)=πθ\(yt,k,j∣ht,k,j\)πθold\(yt,k,j∣ht,k,j\),\\displaystyle=\\frac\{\\pi\_\{\\theta\}\(y\_\{t,k,j\}\\mid h\_\{t,k,j\}\)\}\{\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(y\_\{t,k,j\}\\mid h\_\{t,k,j\}\)\},\(10\)dt,k,jKL\\displaystyle d\_\{t,k,j\}^\{\\mathrm\{KL\}\}=DKL\(πθ\(⋅∣ht,k,j\)∥πref\(⋅∣ht,k,j\)\)\.\\displaystyle=D\_\{\\mathrm\{KL\}\}\\\!\\left\(\\pi\_\{\\theta\}\(\\cdot\\mid h\_\{t,k,j\}\)\\,\\\|\\,\\pi\_\{\\mathrm\{ref\}\}\(\\cdot\\mid h\_\{t,k,j\}\)\\right\)\.Here,θold\\theta\_\{\\mathrm\{old\}\}denotes the policy frozen before the current update, andπref\\pi\_\{\\mathrm\{ref\}\}is the phase\-specific reference policy\. The coefficientϵclip\\epsilon\_\{\\mathrm\{clip\}\}controls PPO clipping, whileβKL\\beta\_\{\\mathrm\{KL\}\}weights the KL regularization term\. All command tokens shareAt,khierA\_\{t,k\}^\{\\mathrm\{hier\}\}, but the clipped ratio and KL term in Equation \([10](https://arxiv.org/html/2608.03137#S3.E10)\) are evaluated token by token\. The operation\-normalized objective is
ℓt,k,j\(θ\)=\\displaystyle\\ell\_\{t,k,j\}\(\\theta\)=\{\}min\(ρt,k,j\(θ\)At,khier,\\displaystyle\\min\\big\(\\rho\_\{t,k,j\}\(\\theta\)A\_\{t,k\}^\{\\mathrm\{hier\}\},\(11\)clip\(ρt,k,j\(θ\),1−ϵclip,1\+ϵclip\)At,khier\)\\displaystyle\\operatorname\{clip\}\(\\rho\_\{t,k,j\}\(\\theta\),1\-\\epsilon\_\{\\mathrm\{clip\}\},1\+\\epsilon\_\{\\mathrm\{clip\}\}\)A\_\{t,k\}^\{\\mathrm\{hier\}\}\\big\)−βKLdt,k,jKL,\\displaystyle\-\\beta\_\{\\mathrm\{KL\}\}d\_\{t,k,j\}^\{\\mathrm\{KL\}\},𝒥GRPO\(θ\)=\\displaystyle\\mathcal\{J\}\_\{\\mathrm\{GRPO\}\}\(\\theta\)=\{\}1\|ℬop\|∑\(q,k,t\)∈ℬop1Lt,k∑j=1Lt,kℓt,k,j\(θ\)\.\\displaystyle\\frac\{1\}\{\|\\mathcal\{B\}\_\{\\mathrm\{op\}\}\|\}\\sum\_\{\(q,k,t\)\\in\\mathcal\{B\}\_\{\\mathrm\{op\}\}\}\\frac\{1\}\{L\_\{t,k\}\}\\sum\_\{j=1\}^\{L\_\{t,k\}\}\\ell\_\{t,k,j\}\(\\theta\)\.Equation \([11](https://arxiv.org/html/2608.03137#S3.E11)\) uses the clipped surrogate from PPO and candidate\-group normalization from GRPO\(Schulman et al\.,[2017](https://arxiv.org/html/2608.03137#bib.bib16); Shao et al\.,[2024](https://arxiv.org/html/2608.03137#bib.bib18)\)\. The memory operation policy loss is applied only to tokens that serialize the atomic memory command\. Task\-solver outputs and environment observations are excluded from this loss\. Task specifications, state\-serialization tokens, and padding tokens are also excluded\. Operations at different time steps retain different advantages, and the factor1/Lt,k1/L\_\{t,k\}prevents commands with longer arguments from dominating the update\. Loss masks and stabilization settings are provided in Supplementary Material, Sections B and F\. The local verifier, global verifier, and reward computation are used only during training\.
### 3\.5Reward Function Design
The training signal has three explicitly separated channels: semantic quality for executable memory commands, global utility for the completed trajectory, and hard execution constraints\. LetTkT\_\{k\}be the number of memory\-decision opportunities inτk\\tau\_\{k\}, including null and rejected commands, and letℰk\\mathcal\{E\}\_\{k\}index its executable commands, including valid null decisions\. Null decisions count toward the decision horizon but not as memory\-tool calls\. Keeping the three channels separate prevents a hard rejection from also receiving an undefined or duplicated semantic penalty\.
Local operation reward\.The local component is
Rlocal\(τk\)=1Tk∑t∈ℰkrt,klocal\.R^\{\\mathrm\{local\}\}\(\\tau\_\{k\}\)=\\frac\{1\}\{T\_\{k\}\}\\sum\_\{t\\in\\mathcal\{E\}\_\{k\}\}r\_\{t,k\}^\{\\mathrm\{local\}\}\.\(12\)Equation \([12](https://arxiv.org/html/2608.03137#S3.E12)\) normalizes by all memory\-decision opportunities while summing only scores produced for executable decisions\. It is therefore zero whenℰk\\mathcal\{E\}\_\{k\}is empty and is used for analysis and monitoring, not as a substitute for the operation\-wise advantages\. The score evaluates only the current atomic memory operation and its realized state transition\. It measures task relevance, evidence grounding, local progress, and information fidelity\. For LTM operations, it evaluates whether information has persistent value, whether an update agrees with new evidence, and whether a deletion is justified\. For STM operations, it evaluates whether retrieval or episode\-level context selection supplies missing information and whether filtering or summarization preserves useful evidence while reducing context load\. Structural validity is excluded from this semantic score and is handled by the auxiliary channel\.
Composite global trajectory reward\.The composite global branch is computed after task termination:
rkglobal=\\displaystyle r\_\{k\}^\{\\mathrm\{global\}\}=\{\}14\(rktask\+rkevid\+rkstate\+rkeff\)\.\\displaystyle\\tfrac\{1\}\{4\}\\big\(r\_\{k\}^\{\\mathrm\{task\}\}\+r\_\{k\}^\{\\mathrm\{evid\}\}\+r\_\{k\}^\{\\mathrm\{state\}\}\+r\_\{k\}^\{\\mathrm\{eff\}\}\\big\)\.\(13\)All four terms lie in\[0,1\]\[0,1\]\. The task scorerktaskr\_\{k\}^\{\\mathrm\{task\}\}and annotated supporting\-fact recallrksupr\_\{k\}^\{\\mathrm\{sup\}\}are computed programmatically\. The global verifier returns only an evidence\-coherence scorevkcohv\_\{k\}^\{\\mathrm\{coh\}\}and a terminal\-memory\-consistency scorevkstatev\_\{k\}^\{\\mathrm\{state\}\}\. The evidence and state components are
rkevid\\displaystyle r\_\{k\}^\{\\mathrm\{evid\}\}=12\(rksup\+vkcoh\),\\displaystyle=\\tfrac\{1\}\{2\}\\big\(r\_\{k\}^\{\\mathrm\{sup\}\}\+v\_\{k\}^\{\\mathrm\{coh\}\}\\big\),\(14\)rkstate\\displaystyle r\_\{k\}^\{\\mathrm\{state\}\}=vkstate\.\\displaystyle=v\_\{k\}^\{\\mathrm\{state\}\}\.Equations \([13](https://arxiv.org/html/2608.03137#S3.E13)\) and \([14](https://arxiv.org/html/2608.03137#S3.E14)\) make the boundary explicit: the global verifier supplies onlyvkcohv\_\{k\}^\{\\mathrm\{coh\}\}andvkstatev\_\{k\}^\{\\mathrm\{state\}\}, while the composite global branch also includes programmatic task correctness, supporting\-fact recall, and efficiency\. Terminal\-memory consistency requires current active LTM and context to agree with observed evidence and treatsHtH\_\{t\}as an ordered provenance record: superseded or failed events may remain in history if their status is explicit\. Since policy training uses only HotpotQA,rktaskr\_\{k\}^\{\\mathrm\{task\}\}is derived from answer correctness\. Environment success, planning progress, and goal\-state satisfaction are used as evaluation metrics on the downstream agent benchmarks rather than as policy\-training rewards\.
LetC¯konline∈\[0,1\]\\overline\{C\}\_\{k\}^\{\\mathrm\{online\}\}\\in\[0,1\]denote the programmatically computed online cost from tokens, task steps, and executed non\-null memory\-tool calls\. These components are separately min–max normalized within theKKcandidates for the same task and then averaged\. The efficiency term is
rkeff=𝕀\[rktask≥δ\]\(1−C¯konline\)\.r\_\{k\}^\{\\mathrm\{eff\}\}=\\mathbb\{I\}\\left\[r\_\{k\}^\{\\mathrm\{task\}\}\\geq\\delta\\right\]\\left\(1\-\\overline\{C\}\_\{k\}^\{\\mathrm\{online\}\}\\right\)\.\(15\)We useδ=0\.5\\delta=0\.5\. This gate prevents low\-cost failures from receiving an efficiency reward in Equation \([15](https://arxiv.org/html/2608.03137#S3.E15)\)\.
Auxiliary constraints\.The auxiliary cost is
Paux\(τk\)=1Tk∑t=1Tkct,k\.P\_\{\\mathrm\{aux\}\}\(\\tau\_\{k\}\)=\\frac\{1\}\{T\_\{k\}\}\\sum\_\{t=1\}^\{T\_\{k\}\}c\_\{t,k\}\.\(16\)Equation \([16](https://arxiv.org/html/2608.03137#S3.E16)\) averages direct constraint costs over the memory\-decision horizon\. The step cost records discrete or hard violations, including invalid tools or arguments, failed state transitions, unsupported writes or updates, unjustified deletions, making protected evidence operationally inaccessible within the remaining decision and token budget, context overflow, exceeding the tool or interaction budget, and information leakage in no\-reference settings\. Routine token, step, and tool costs do not enter this channel; they are evaluated continuously by the global efficiency term\. Terminal budget violations are attached to the final memory decision that precedes termination\.
The scalar displayed in the All\-Returns training curve is used only for monitoring:
mkall=clip\(Rlocal\(τk\)\+rkglobal−Paux\(τk\)2,0,1\)\.m\_\{k\}^\{\\mathrm\{all\}\}=\\operatorname\{clip\}\\\!\\left\(\\frac\{R^\{\\mathrm\{local\}\}\(\\tau\_\{k\}\)\+r\_\{k\}^\{\\mathrm\{global\}\}\-P\_\{\\mathrm\{aux\}\}\(\\tau\_\{k\}\)\}\{2\},0,1\\right\)\.\(17\)The Answer\-Only curve analogously displays
mkans=clip\(rktask−Paux\(τk\),0,1\)\.m\_\{k\}^\{\\mathrm\{ans\}\}=\\operatorname\{clip\}\\\!\\left\(r\_\{k\}^\{\\mathrm\{task\}\}\-P\_\{\\mathrm\{aux\}\}\(\\tau\_\{k\}\),0,1\\right\)\.\(18\)Equations \([17](https://arxiv.org/html/2608.03137#S3.E17)\) and \([18](https://arxiv.org/html/2608.03137#S3.E18)\) are strategy\-specific monitoring values\. They are not policy advantages and are not directly comparable in absolute level\. During policy optimization, the local and global signals are normalized separately, whilect,kc\_\{t,k\}is assigned directly to the corresponding decision\. Tool\-specific rubrics, normalization details, completion thresholds, and constraint triggers are provided in Supplementary Material, Section F\.
## 4Experiments
### 4\.1Experimental Setup
Datasets\.We evaluate VerMem on ALFWorld\(Shridhar et al\.,[2021](https://arxiv.org/html/2608.03137#bib.bib19)\), SciWorld\(Wang et al\.,[2022](https://arxiv.org/html/2608.03137#bib.bib20)\), PDDL\(Ma et al\.,[2024](https://arxiv.org/html/2608.03137#bib.bib11)\), BabyAI\(Chevalier\-Boisvert et al\.,[2019](https://arxiv.org/html/2608.03137#bib.bib1)\), and HotpotQA\(Yang et al\.,[2018](https://arxiv.org/html/2608.03137#bib.bib27)\)\. VerMem is fine\-tuned only on the HotpotQA training split, whose supporting facts and distractors support LTM maintenance, STM control, and historical\-context recovery, and is then evaluated directly on all five benchmarks\. Memory states are isolated by task instance\. Each episode starts with emptyM0M\_\{0\}andH0H\_\{0\}, whileC0C\_\{0\}contains only the task instruction and initial input\.
Evaluation Metrics\.We report Success Rate \(successful episodes divided by evaluated episodes\) on ALFWorld, SciWorld, and BabyAI, Progress Rate on PDDL, and the Qwen\-Max LLM\-as\-a\-Judge scoreJjudgeJ\_\{\\mathrm\{judge\}\}on HotpotQA\. The main table reports100Jjudge100J\_\{\\mathrm\{judge\}\}\. Qwen\-Max is used only as a held\-out evaluator; it is not used for policy training, verifier scoring, or checkpoint selection\. LetS∗S^\{\*\}be a target macro\-average Success Rate andB∗B^\{\*\}an online\-token budget\.TN@S∗\\mathrm\{TN\}@S^\{\*\}is the linearly interpolated online\-token budget at which the mean success curve first reaches the targetS∗S^\{\*\}, using the two adjacent evaluated budget points that bracket the target\.SR@B∗\\mathrm\{SR\}@B^\{\*\}is the macro\-average Success Rate atB∗B^\{\*\}\. The reward ablation additionally reports Memory Quality \(MQ\), average online token number \(TN\), and average memory\-tool calls \(TC\)\.
Baselines and Backbones\.Base uses the common task solver without an explicit memory operation policy or external persistent\-memory store\. We compare it with LangMem\(LangChain Team,[2025](https://arxiv.org/html/2608.03137#bib.bib7)\), A\-Mem\(Xu et al\.,[2025](https://arxiv.org/html/2608.03137#bib.bib24)\), Mem0 and its graph\-based variant Mem0g\(Chhikara et al\.,[2025](https://arxiv.org/html/2608.03137#bib.bib2)\), and AgeMem\(Yu et al\.,[2026](https://arxiv.org/html/2608.03137#bib.bib28)\)under Qwen2\.5\-7B\-Instruct\(Qwen et al\.,[2024](https://arxiv.org/html/2608.03137#bib.bib15)\)and Qwen3\-4B\-Instruct\(Yang et al\.,[2025](https://arxiv.org/html/2608.03137#bib.bib26)\)\. VerMem\-noVerify is a matched internal control that removes semantic feedback from the local and global verifiers while retaining the same action space, SFT initialization, curriculum, task\-outcome reward, transition constraints, and GRPO configuration\. Complete implementation and baseline settings are provided in Supplementary Material, Sections C and F\.
### 4\.2Main Results
Table 2:Main results under two LLM backbones\. We report Success Rate on ALFWorld, SciWorld, and BabyAI, Progress Rate on PDDL, and100Jjudge100J\_\{\\mathrm\{judge\}\}on HotpotQA\. Each entry is the mean over three independent runs with seeds 42, 43, and 44\. Average is the unweighted arithmetic mean of the five displayed scores\. The best result in each backbone setting is shown in bold, and the second\-best result is underlined\.Overall comparison\.Table[2](https://arxiv.org/html/2608.03137#S4.T2)reports the main results under both backbones\. With Qwen2\.5\-7B\-Instruct, VerMem obtains an average score of48\.0148\.01\. This is19\.9619\.96points above Base, corresponding to a relative gain of71\.16%71\.16\\%, and6\.056\.05points above the strongest external method, AgeMem\. VerMem ranks first on ALFWorld, SciWorld, BabyAI, HotpotQA, and Average\. PDDL is the only exception, where VerMem\-noVerify is0\.350\.35points higher\. VerMem still exceeds the strongest external PDDL baseline by3\.673\.67points\. With Qwen3\-4B\-Instruct, VerMem ranks first in all six reported columns and reaches an average score of59\.8559\.85, outperforming AgeMem by5\.545\.54points\. The consistent improvement is observed with both evaluated backbones rather than only one of them\.
Performance across task types\.The improvements vary with the memory requirements of each task\. With Qwen2\.5\-7B\-Instruct, VerMem exceeds the strongest external baseline by5\.235\.23,7\.847\.84,3\.673\.67,4\.134\.13, and8\.318\.31points on ALFWorld, SciWorld, PDDL, BabyAI, and HotpotQA\. The corresponding gains with Qwen3\-4B\-Instruct are4\.764\.76,6\.896\.89,3\.683\.68,4\.214\.21, and8\.128\.12points\. The largest gains occur on HotpotQA and SciWorld\. This pattern is consistent with the greater need for multi\-hop evidence aggregation and long interaction\-state tracking in these tasks\. The smaller gain on PDDL is consistent with its more structured state representation and feedback, although the present experiments do not isolate this factor\.
Cross\-task transfer\.VerMem is fine\-tuned only on HotpotQA, yet it consistently improves over the strongest baselines on ALFWorld, SciWorld, PDDL, and BabyAI\. The gains range from3\.673\.67to7\.847\.84points with Qwen2\.5\-7B\-Instruct and from3\.683\.68to6\.896\.89points with Qwen3\-4B\-Instruct\. These environments require environment tracking, scientific\-progress maintenance, planning\-state control, and instruction execution\. The transfer results indicate that the learned memory operation policy captures reusable memory behavior rather than task\-specific document patterns\.
Figure 2:Efficiency–performance frontier under different online token budgets on ALFWorld, SciWorld, and BabyAI using Qwen2\.5\-7B\-Instruct\. Curves compare A\-Mem, Mem0, AgeMem, and VerMem, and report macro\-average Success Rate over three random seeds\. Shaded regions denote the standard deviation across seeds\. The horizontal and vertical dotted lines markS∗=40\.0S^\{\*\}=40\.0andB∗=2,500B^\{\*\}=2\{,\}500, respectively\.Table 3:Efficiency–performance comparison under controlled online\-token budgets using Qwen2\.5\-7B\-Instruct, withS∗=40\.0S^\{\*\}=40\.0andB∗=2,500B^\{\*\}=2\{,\}500\.TN@S∗\\mathrm\{TN\}@S^\{\*\}is linearly interpolated from the two adjacent evaluated budget points that bracketS∗S^\{\*\}\. LowerTN@S∗\\mathrm\{TN\}@S^\{\*\}and higherSR@B∗\\mathrm\{SR\}@B^\{\*\}are better\.Efficiency–performance frontier\.Figure[2](https://arxiv.org/html/2608.03137#S4.F2)and Table[3](https://arxiv.org/html/2608.03137#S4.T3)show that VerMem maintains the highest macro\-average SR across the evaluated budgets\. Its mean curve first reachesS∗=40\.0S^\{\*\}=40\.0at the linearly interpolated estimate of2,0802\{,\}080online tokens, reducingTN@S∗\\mathrm\{TN\}@S^\{\*\}by20\.0%20\.0\\%relative to AgeMem, and achieves anSR@B∗\\mathrm\{SR\}@B^\{\*\}of44\.3044\.30atB∗=2,500B^\{\*\}=2\{,\}500, exceeding AgeMem by5\.305\.30points\. Both verifiers are disabled during evaluation, so these measurements do not include inference\-time verification calls\.
### 4\.3Ablation Studies
Figure 3:Cumulative component ablation with Qwen2\.5\-7B\-Instruct\. ALFWorld and SciWorld report Success Rate\. HotpotQA reports100Jjudge100J\_\{\\mathrm\{judge\}\}\. Base has no explicit memory operation policy\. \+LT adds LTM operations\. \+LT/ST adds the full LTM/STM tool space\. \+noV adds stateful RL without semantic verifier feedback\. \+V adds the local and global credit branches\.Unified memory components\.Figure[3](https://arxiv.org/html/2608.03137#S4.F3)shows cumulative gains from LTM operations, the complete LTM/STM action space, stateful RL, and verifier\-guided credit\. LTM provides the largest initial improvement, STM control adds further gains on SciWorld and HotpotQA, and full VerMem improves over \+noV by6\.126\.12,9\.459\.45, and9\.679\.67points\. The progression is consistent with contributions from both the unified action space and verifier\-guided optimization\.
Table 4:Ablation of the local and global credit branches\. VerMem\-noVerify uses a normalized task\-outcome advantage and constraints\. \+Local adds the local operation advantage\. \+Global replaces the task\-only trajectory advantage with the composite global trajectory advantage\. Local\+Global uses both the local and composite global advantages and denotes full VerMem\. Architecture, SFT initialization, training data, constraints, and GRPO settings are otherwise shared\.Local and global credit branches\.Table[4](https://arxiv.org/html/2608.03137#S4.T4)shows complementary effects\. Local feedback yields larger standalone gains on ALFWorld and SciWorld, whereas the global branch yields a larger standalone gain on HotpotQA\. This pattern is consistent with, but does not by itself prove, the intended distinction between operation\-level and trajectory\-level credit\. Their combination performs best on all three evaluated tasks\.
Figure 4:GRPO training\-reward convergence on HotpotQA using Qwen2\.5\-7B\-Instruct\. Training Step indexes the 101 recorded training points over the 3,500\-update GRPO schedule, with adjacent points separated by 35 policy updates\. Curves show the unsmoothed mean over three runs, and shaded regions denote one standard deviation\. Answer\-Only and All\-Returns use the monitoring mappings defined in the reward section; their absolute levels are not directly comparable because the mappings differ\.Table 5:Reward\-function ablation on HotpotQA using Qwen2\.5\-7B\-Instruct\.JjudgeJ\_\{\\mathrm\{judge\}\}denotes the LLM\-as\-a\-Judge score on\[0,1\]\[0,1\], TN is the average online token number, MQ is Memory Quality, and TC is the average number of memory\-tool calls\.Reward function\.The monitoring curves in Figure[4](https://arxiv.org/html/2608.03137#S4.F4)suggest earlier stabilization and lower late\-stage variability for All\-Returns\. Because the two strategies use different monitoring definitions, their absolute reward levels should not be compared directly\. Table[5](https://arxiv.org/html/2608.03137#S4.T5)shows that All\-Returns raisesJjudgeJ\_\{\\mathrm\{judge\}\}from0\.5310\.531to0\.6280\.628and MQ from0\.4910\.491to0\.6730\.673, while reducing TN by7\.14%7\.14\\%\. TC increases from4\.474\.47to5\.635\.63; together with the higherJjudgeJ\_\{\\mathrm\{judge\}\}and MQ, this shows that the token reduction is not obtained merely by suppressing all memory\-tool calls\.
## 5Limitations
Our evaluation resets all memory states after each episode, so it measures within\-episode persistence and cross\-benchmark transfer of the learned policy rather than cross\-session or cross\-user memory retention\. The memory operation policy is trained on constructed HotpotQA states and evaluated with two Qwen backbones; the findings may not transfer to other model families or domains\. HotpotQAJjudgeJ\_\{\\mathrm\{judge\}\}and MQ use a fixed proprietary Qwen\-Max snapshot, which limits fully independent replication despite disclosure of the snapshot and prompts\. The frozen DeepSeek\-V3\.2 teacher and verifiers also introduce offline training cost\. Finally, learned verifier scores may inherit evaluator errors, and the present study does not establish robustness under adversarial, privacy\-sensitive, or conflicting memory content\.
## 6Conclusion
We presented Verifiable Memory \(VerMem\), a unified framework for learning LTM and STM management in long\-horizon LLM agents\. VerMem maintains persistent memory, active context, and episodic history as distinct states, while one memory operation policy coordinates seven atomic tools for persistent\-memory maintenance, active\-context control, and historical\-context recovery\. Training combines an SFT warmup with a three\-stage reinforcement learning curriculum\. The local verifier evaluates realized atomic memory transitions, while the global verifier evaluates completed trajectories and terminal memory states\. Both verifiers are removed during inference\.
Experiments across five benchmarks and two backbones show that VerMem consistently outperforms strong memory baselines and transfers from HotpotQA to interactive, planning, and instruction\-following tasks\. The budget\-controlled evaluation further demonstrates a stronger efficiency–performance trade\-off under limited online tokens\. The ablation results confirm the contributions of unified LTM/STM control, stateful reinforcement learning, complementary verifier signals, and the complete reward design\. These findings show that coordinated memory management and multi\-granularity credit assignment provide an effective basis for reliable long\-horizon agent behavior\.
## References
- Chevalier\-Boisvert et al\. \(2019\)Maxime Chevalier\-Boisvert, Dzmitry Bahdanau, Salem Lahlou, Lucas Willems, Chitwan Saharia, Thien Huu Nguyen, and Yoshua Bengio\.BabyAI: A platform to study the sample efficiency of grounded language learning\.In*International Conference on Learning Representations*, 2019\.URL[https://openreview\.net/forum?id=rJeXCo0cYX](https://openreview.net/forum?id=rJeXCo0cYX)\.
- Chhikara et al\. \(2025\)Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav\.Mem0: Building production\-ready AI agents with scalable long\-term memory, 2025\.URL[https://arxiv\.org/abs/2504\.19413](https://arxiv.org/abs/2504.19413)\.
- Gao et al\. \(2025\)Pengyu Gao, Jinming Zhao, Xinyue Chen, and Yilin Long\.An efficient context\-dependent memory framework for LLM\-centric agents\.In Weizhu Chen, Yi Yang, Mohammad Kachuee, and Xue\-Yong Fu, editors,*Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 3: Industry Track\)*, pages 1055–1069, Albuquerque, New Mexico, April 2025\. Association for Computational Linguistics\.doi:10\.18653/v1/2025\.naacl\-industry\.80\.URL[https://aclanthology\.org/2025\.naacl\-industry\.80/](https://aclanthology.org/2025.naacl-industry.80/)\.
- Goodyear et al\. \(2025\)Lyle Goodyear, Rachel Guo, and Ramesh Johari\.The effect of state representation on LLM agent behavior in dynamic routing games, 2025\.URL[https://arxiv\.org/abs/2506\.15624](https://arxiv.org/abs/2506.15624)\.
- Huo et al\. \(2026\)Yupeng Huo, Yaxi Lu, Zhong Zhang, Haotian Chen, and Yankai Lin\.Atommem: Learnable dynamic agentic memory with atomic memory operation, 2026\.URL[https://arxiv\.org/abs/2601\.08323](https://arxiv.org/abs/2601.08323)\.
- Jiang et al\. \(2024\)Xun Jiang, Feng Li, Han Zhao, Jiahao Qiu, Jiaying Wang, Jun Shao, Shihao Xu, Shu Zhang, Weiling Chen, Xavier Tang, Yize Chen, Mengyue Wu, Weizhi Ma, Mengdi Wang, and Tianqiao Chen\.Long term memory: The foundation of AI self\-evolution, 2024\.URL[https://arxiv\.org/abs/2410\.15665](https://arxiv.org/abs/2410.15665)\.
- LangChain Team \(2025\)LangChain Team\.LangMem\.[https://github\.com/langchain\-ai/langmem](https://github.com/langchain-ai/langmem), 2025\.Accessed: 2026\-07\-19\.
- Li et al\. \(2025\)Rui Li, Zeyu Zhang, Xiaohe Bo, Zihang Tian, Xu Chen, Quanyu Dai, Zhenhua Dong, and Ruiming Tang\.CAM: A constructivist view of agentic memory for LLM\-based reading comprehension\.In*Advances in Neural Information Processing Systems*, volume 38\. Curran Associates, Inc\., 2025\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2025/hash/a488aa1a0c00d76db8a922ef7815a786\-Abstract\-Conference\.html](https://proceedings.neurips.cc/paper_files/paper/2025/hash/a488aa1a0c00d76db8a922ef7815a786-Abstract-Conference.html)\.
- Liu et al\. \(2025\)Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, et al\.DeepSeek\-V3\.2: Pushing the frontier of open large language models, 2025\.URL[https://arxiv\.org/abs/2512\.02556](https://arxiv.org/abs/2512.02556)\.
- Luo et al\. \(2025\)Xufang Luo, Yuge Zhang, Zhiyuan He, Zilong Wang, Siyun Zhao, Dongsheng Li, Luna K\. Qiu, and Yuqing Yang\.Agent lightning: Train ANY AI agents with reinforcement learning, 2025\.URL[https://arxiv\.org/abs/2508\.03680](https://arxiv.org/abs/2508.03680)\.
- Ma et al\. \(2024\)Chang Ma, Junlei Zhang, Zhihao Zhu, Cheng Yang, Yujiu Yang, Yaohui Jin, Zhenzhong Lan, Lingpeng Kong, and Junxian He\.AgentBoard: An analytical evaluation board of multi\-turn LLM agents\.*Advances in Neural Information Processing Systems*, 37:74325–74362, 2024\.
- MiroMind Team \(2026\)MiroMind Team\.Mirothinker\-1\.7 & h1: Towards heavy\-duty research agents via verification, 2026\.URL[https://arxiv\.org/abs/2603\.15726](https://arxiv.org/abs/2603.15726)\.
- Peng et al\. \(2026\)Jiangweizhi Peng, Yuanxin Liu, Ruida Zhou, Charles Fleming, Zhaoran Wang, Alfredo Garcia, and Mingyi Hong\.HiPER: Hierarchical reinforcement learning with explicit credit assignment for large language model agents, 2026\.URL[https://arxiv\.org/abs/2602\.16165](https://arxiv.org/abs/2602.16165)\.
- Qian et al\. \(2026\)Hongjin Qian, Zhao Cao, and Zheng Liu\.Memobrain: Executive memory as an agentic brain for reasoning\.In*Findings of the Association for Computational Linguistics: ACL 2026*, pages 2646–2662, San Diego, California, United States, July 2026\. Association for Computational Linguistics\.doi:10\.18653/v1/2026\.findings\-acl\.127\.URL[https://aclanthology\.org/2026\.findings\-acl\.127/](https://aclanthology.org/2026.findings-acl.127/)\.
- Qwen et al\. \(2024\)Qwen, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, and Fei Huang\.Qwen2\.5 technical report\.2024\.
- Schulman et al\. \(2017\)John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov\.Proximal policy optimization algorithms\.*arXiv preprint arXiv:1707\.06347*, 2017\.
- Setlur et al\. \(2025\)Amrith Setlur, Chirag Nagpal, Adam Fisch, Xinyang Geng, Jacob Eisenstein, Rishabh Agarwal, Alekh Agarwal, Jonathan Berant, and Aviral Kumar\.Rewarding progress: Scaling automated process verifiers for LLM reasoning\.In*International Conference on Learning Representations*, 2025\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2025/hash/98711dea460bdefe0e651ca23ec98ba2\-Abstract\-Conference\.html](https://proceedings.iclr.cc/paper_files/paper/2025/hash/98711dea460bdefe0e651ca23ec98ba2-Abstract-Conference.html)\.Spotlight\.
- Shao et al\. \(2024\)Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al\.Deepseekmath: Pushing the limits of mathematical reasoning in open language models\.*arXiv preprint arXiv:2402\.03300*, 2024\.
- Shridhar et al\. \(2021\)Mohit Shridhar, Xingdi Yuan, Marc\-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht\.ALFWorld: Aligning text and embodied environments for interactive learning\.In*International Conference on Learning Representations*, 2021\.URL[https://openreview\.net/forum?id=0IOX0YcCdTn](https://openreview.net/forum?id=0IOX0YcCdTn)\.
- Wang et al\. \(2022\)Ruoyao Wang, Peter Jansen, Marc\-Alexandre Côté, and Prithviraj Ammanabrolu\.ScienceWorld: Is your agent smarter than a 5th grader?In*Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing*, pages 11279–11298, Abu Dhabi, United Arab Emirates, December 2022\. Association for Computational Linguistics\.doi:10\.18653/v1/2022\.emnlp\-main\.775\.URL[https://aclanthology\.org/2022\.emnlp\-main\.775/](https://aclanthology.org/2022.emnlp-main.775/)\.
- Wu et al\. \(2025a\)Xixi Wu, Kuan Li, Yida Zhao, Liwen Zhang, Litu Ou, Huifeng Yin, Zhongwang Zhang, Xinmiao Yu, Dingchu Zhang, Yong Jiang, Pengjun Xie, Fei Huang, Minhao Cheng, Shuai Wang, Hong Cheng, and Jingren Zhou\.Resum: Unlocking long\-horizon search intelligence via context summarization, 2025a\.URL[https://arxiv\.org/abs/2509\.13313](https://arxiv.org/abs/2509.13313)\.
- Wu et al\. \(2025b\)Yaxiong Wu, Sheng Liang, Chen Zhang, Yichao Wang, Yongyue Zhang, Huifeng Guo, Ruiming Tang, and Yong Liu\.From human memory to AI memory: A survey on memory mechanisms in the era of LLMs, 2025b\.URL[https://arxiv\.org/abs/2504\.15965](https://arxiv.org/abs/2504.15965)\.
- Xiong et al\. \(2026\)Zidi Xiong, Yuping Lin, Wenya Xie, Pengfei He, Zirui Liu, Jiliang Tang, Himabindu Lakkaraju, and Zhen Xiang\.How memory management impacts LLM agents: An empirical study of experience\-following behavior\.In*Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 623–645, San Diego, California, United States, July 2026\. Association for Computational Linguistics\.doi:10\.18653/v1/2026\.acl\-long\.27\.URL[https://aclanthology\.org/2026\.acl\-long\.27/](https://aclanthology.org/2026.acl-long.27/)\.
- Xu et al\. \(2025\)Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang\.A\-mem: Agentic memory for LLM agents\.In*Advances in Neural Information Processing Systems*, volume 38\. Curran Associates, Inc\., 2025\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2025/hash/19909c36f51abc4856b4560aff3d36d6\-Abstract\-Conference\.html](https://proceedings.neurips.cc/paper_files/paper/2025/hash/19909c36f51abc4856b4560aff3d36d6-Abstract-Conference.html)\.
- Yan et al\. \(2026\)Sikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma, Jinhe Bi, Kristian Kersting, Jeff Z\. Pan, Hinrich Schuetze, Volker Tresp, and Yunpu Ma\.Memory\-r1: Enhancing large language model agents to manage and utilize memories via reinforcement learning\.In*Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 12805–12825, San Diego, California, United States, July 2026\. Association for Computational Linguistics\.doi:10\.18653/v1/2026\.acl\-long\.583\.URL[https://aclanthology\.org/2026\.acl\-long\.583/](https://aclanthology.org/2026.acl-long.583/)\.
- Yang et al\. \(2025\)An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al\.Qwen3 technical report\.*arXiv preprint arXiv:2505\.09388*, 2025\.
- Yang et al\. \(2018\)Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D\. Manning\.HotpotQA: A dataset for diverse, explainable multi\-hop question answering\.In*Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing*, pages 2369–2380, Brussels, Belgium, October\-November 2018\. Association for Computational Linguistics\.doi:10\.18653/v1/D18\-1259\.URL[https://aclanthology\.org/D18\-1259/](https://aclanthology.org/D18-1259/)\.
- Yu et al\. \(2026\)Yi Yu, Liuyi Yao, Yuexiang Xie, Qingquan Tan, Jiaqi Feng, Yaliang Li, and Libing Wu\.Agentic memory: Learning unified long\-term and short\-term memory management for large language model agents\.In*Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 21457–21483, San Diego, California, United States, July 2026\. Association for Computational Linguistics\.doi:10\.18653/v1/2026\.acl\-long\.981\.URL[https://aclanthology\.org/2026\.acl\-long\.981/](https://aclanthology.org/2026.acl-long.981/)\.
- Zhong et al\. \(2024\)Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang\.Memorybank: Enhancing large language models with long\-term memory\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 38, pages 19724–19731\. AAAI Press, 2024\.doi:10\.1609/aaai\.v38i17\.29946\.URL[https://ojs\.aaai\.org/index\.php/AAAI/article/view/29946](https://ojs.aaai.org/index.php/AAAI/article/view/29946)\.
- Zhou et al\. \(2025\)Zijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim, Alok Prakash, Daniela Rus, Jinhua Zhao, Bryan Kian Hsiang Low, and Paul Pu Liang\.MEM1: Learning to synergize memory and reasoning for efficient long\-horizon agents, 2025\.URL[https://arxiv\.org/abs/2506\.15841](https://arxiv.org/abs/2506.15841)\.
## Appendix AMethod and Tool Specifications
The following sections provide the runtime details omitted from the main paper\. We first describe howMtM\_\{t\},CtC\_\{t\}, andHtH\_\{t\}are maintained during task execution\. We then give the tool\-specific inputs, preconditions, and state\-transition rules of the seven atomic memory tools\. The final two subsections explain historical\-context recovery and failure handling\. Training\-instance construction, verifier rubrics, and experimental configurations are provided in the subsequent appendices\.
### A\.1Runtime State and Execution Protocol
VerMem represents the memory\-decision state asst=\(q,Mt,Ct,Ht\)s\_\{t\}=\(q,M\_\{t\},C\_\{t\},H\_\{t\}\)\. The task specificationqqremains fixed within one task instance\. The long\-term memoryMtM\_\{t\}contains persistent entries that can be reused in later steps or stages\. Each entry records its content, source, current status, and revision history\. The active contextCtC\_\{t\}contains the budgeted information currently visible to the task solver\. The episodic historyHtH\_\{t\}preserves the ordered record of observations, task actions, tool outputs, intermediate conclusions, and memory decisions from the same task\.
The active context and episodic history serve different purposes\. Removing information fromCtC\_\{t\}reduces the content visible to the task solver, but does not remove the corresponding event fromHtH\_\{t\}\. This separation allows earlier evidence to be recovered after the task focus changes\.
At each memory\-decision step, the policy generates a complete commandut=\(vt,ξt\)u\_\{t\}=\(v\_\{t\},\\xi\_\{t\}\), wherevt∈𝒱∪\{∅\}v\_\{t\}\\in\\mathcal\{V\}\\cup\\\{\\varnothing\\\}\. Each non\-null command specifies one atomic tool and its structured arguments\. The policy samplesutu\_\{t\}from the bounded serializationxt=g\(st\)x\_\{t\}=g\(s\_\{t\}\)defined in Equation \([1](https://arxiv.org/html/2608.03137#S3.E1)\)\. The policy receives a bounded serialization of the task specification and budgeted views ofMtM\_\{t\},CtC\_\{t\}, andHtH\_\{t\}\. This policy\-visible serialization is distinct from the active\-context budget governingCtC\_\{t\}\. Before the command changes the real memory state, the proposed transition is checked against the tool schema, target state, required arguments, source information, and current budget\. The transition follows Equation \([2](https://arxiv.org/html/2608.03137#S3.E2)\)\. Each valid non\-null command commits its realized memory transition\. The null command leavesMtM\_\{t\}andCtC\_\{t\}unchanged\. On these states,𝒯m\\mathcal\{T\}\_\{m\}is the identity transition\. The identity decision is retained in the rollout record and does not create a memory\-state modification\. If a non\-null command is rejected,MtM\_\{t\}andCtC\_\{t\}remain unchanged, while the failed command and its reason are recorded inHtH\_\{t\}\. During training, the corresponding violation is recorded throughct,kc\_\{t,k\}\.
After the memory transition, the task solver continues with the resulting active context\. The task action and resulting observation are then incorporated through the environment interaction transition in Equation \([3](https://arxiv.org/html/2608.03137#S3.E3)\), which constructsst\+1s\_\{t\+1\}\. The following procedure summarizes this runtime process\. The local and global verifiers are used only during training and do not participate in the reported inference process\.
##### Unified runtime procedure\.
Starting fromM0M\_\{0\},C0C\_\{0\}, andH0H\_\{0\}, each memory\-decision step performs the following operations:
1. 1\.constructst=\(q,Mt,Ct,Ht\)s\_\{t\}=\(q,M\_\{t\},C\_\{t\},H\_\{t\}\)and its bounded serializationxt=g\(st\)x\_\{t\}=g\(s\_\{t\}\);
2. 2\.sample the complete commandut∼πθ\(⋅∣xt\)u\_\{t\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x\_\{t\}\);
3. 3\.computes~t=𝒯m\(st,ut\)\\widetilde\{s\}\_\{t\}=\\mathcal\{T\}\_\{m\}\(s\_\{t\},u\_\{t\}\)under the applicable execution branch\. The valid non\-null command commits its realized transition\. The valid null command leavesMtM\_\{t\}andCtC\_\{t\}unchanged, creates no memory\-state modification, and retains the identity decision in the rollout record\. The rejected non\-null command preservesMtM\_\{t\}andCtC\_\{t\}, records the failed command and reason inHtH\_\{t\}, and receives the correspondingct,kc\_\{t,k\}during training;
4. 4\.sample the task actionbtb\_\{t\}from the resulting active context, collect the next observation, and constructst\+1=𝒯e\(s~t,bt,zt\+1\)s\_\{t\+1\}=\\mathcal\{T\}\_\{e\}\(\\widetilde\{s\}\_\{t\},b\_\{t\},z\_\{t\+1\}\);
5. 5\.repeat until task termination or a configured rollout limit\.
### A\.2Atomic Tool Schemas and State Transitions
Each non\-null memory command selects one of the seven tools defined in the main paper and supplies its required arguments\. Table[6](https://arxiv.org/html/2608.03137#A1.T6)summarizes the required information, state transition, and principal validity condition of each tool\. Optional implementation fields and benchmark\-specific budgets are reported with the experimental configuration\.
Table 6:Compact specifications of the atomic memory tools in VerMem\. The table reports the information required by each operation, its state transition, and its principal validity condition\.##### Add\.
Add stores new information with persistent utility\. The new entry records the information and its source\. The operation is rejected when the content is empty, unsupported, or duplicates an existing active entry\. When an existing entry requires revision, Update is used instead\.
##### Update\.
Update revises an existing LTM entry\. It creates a new version rather than destructively overwriting the previous content\. This retains the relation between the earlier and current versions and allows later inspection of how the memory state changed\.
##### Delete\.
Delete follows soft\-deletion semantics\. The target entry is excluded from subsequent ordinary retrieval, while its content and deletion record remain available for inspection\. The retained record permits later inspection if the entry must be reconsidered\.
##### Retrieve\.
Retrieve is categorized as an STM tool because it does not changeMtM\_\{t\}\. Instead, it selects active LTM entries and introduces them intoCtC\_\{t\}\. The retrieval query is generated as part ofξt\\xi\_\{t\}\. Returned content is committed only when it satisfies the configured active\-context budget\. Deleted entries cannot be retrieved\.
##### Filter\.
Filter removes selected low\-value content fromCtC\_\{t\}\. Task instructions, the current environment state, and indispensable evidence are protected\. Filtering changes the active context but does not remove the original event fromHtH\_\{t\}\.
##### SelectEpisode\.
SelectEpisode performs episode\-level context selection over the current task history\. It identifies earlier observations, tool outputs, or intermediate conclusions that are relevant to the current task state and returns them toCtC\_\{t\}\. The operation cannot access history from another task instance\.
##### Summarize\.
Summarize replaces selected content inCtC\_\{t\}with a shorter representation\. The summary must remain faithful to the source content and preserve information required by the current task\. The original events remain available inHtH\_\{t\}\.
### A\.3Historical\-Context Recovery via Episode\-Level Context Selection
Historical\-context recovery concerns information that has left the active context but becomes useful again at a later task stage\. It operates over the episodic history of the current task\. It is therefore distinct from retrieving persistent information from LTM\.
The system first identifies same\-task historical fragments that are relevant to the current task specification and active context\. The memory operation policy then selects the fragments to be restored\. Selected content is restored toCtC\_\{t\}subject to the configured active\-context budget and retains its original source and temporal position inHtH\_\{t\}\.
The two information paths are
Mt→RetrieveCt,Ht→SelectEpisodeCt\.M\_\{t\}\\xrightarrow\{\\mathrm\{Retrieve\}\}C\_\{t\},\\qquad H\_\{t\}\\xrightarrow\{\\mathrm\{SelectEpisode\}\}C\_\{t\}\.Although both operations introduce information into the active context, their information sources and intended uses are different\. Table[7](https://arxiv.org/html/2608.03137#A1.T7)summarizes this distinction\.
Table 7:Difference between LTM retrieval and historical\-context recovery through episode\-level context selection\.This separation supports same\-task evidence recovery without treating episodic history as another persistent memory store\. It keeps Retrieve and SelectEpisode as distinct operations\.
### A\.4State Validation and Failure Handling
The non\-null memory operation is checked before it modifies the real state\. The checks cover the selected tool, required arguments, referenced entries, operation preconditions, source support, and available budget\. The transition is committed only when these conditions are satisfied\.
If a proposed command is invalid,MtM\_\{t\}andCtC\_\{t\}remain unchanged\. The failed command and its reason are retained in the task history\. During training, the corresponding structural, safety, or budget violation contributes toct,kc\_\{t,k\}\. The task solver then continues with the unchanged active context unless the task itself has reached a terminal failure\.
Update must preserve the relation between memory versions\. Delete must remain soft\. Filter must not remove indispensable evidence\. SelectEpisode must remain within the current task history\. Summarize must reduce context length without changing the supported meaning\. Retrieve, SelectEpisode, and Summarize must also satisfy the active\-context budget before their transitions are committed\.
These execution checks are separate from semantic verification\. They determine whether an operation can be executed safely\. The local verifier instead evaluates the task relevance, evidence grounding, local progress, and information fidelity of a realized operation\. The global verifier supplies only evidence\-coherence and terminal\-memory\-consistency scores after trajectory completion\. These scores enter the composite global branch together with programmatic task, supporting\-fact\-recall, and efficiency components\. Both verifiers are disabled during inference\.
## Appendix BOptimization Objective and Loss Scope
This section records the optimization settings and loss scope that supplement the objectives in the main paper\. Training\-data and rollout construction are detailed in Appendix[D](https://arxiv.org/html/2608.03137#A4); verifier, normalization, and monitoring details are provided in Appendix[F](https://arxiv.org/html/2608.03137#A6)\.
### B\.1Core Training Configuration and Loss Scope
Table[8](https://arxiv.org/html/2608.03137#A2.T8)summarizes the settings directly related to data construction and trajectory optimization\.
Table 8:Core training and rollout configuration used by VerMem\. Token limits are measured in model tokens\.During reinforcement learning, all generated tokens belonging to one atomic memory operation share the sameAt,khierA\_\{t,k\}^\{\\mathrm\{hier\}\}\. Different memory decisions retain different advantages\. The policy objective is applied only to tokens that serialize the complete commandut,ku\_\{t,k\}, including its structured arguments\. Task\-solver outputs, environment observations, input\-state tokens, and padding tokens do not receive memory\-operation credit\. Phase A, Phase B, and Phase C save separate checkpoints\. Each phase initializes the next one through policy parameters only\.
## Appendix CExperimental Configuration and Evaluation Protocol
The evaluation details below specify the splits, state initialization, baseline configurations, metric computation, Qwen\-Max evaluation, budget\-controlled analysis, and multi\-seed reporting used in the main experiments\. Training\-data construction is described in Appendix[D](https://arxiv.org/html/2608.03137#A4), and the verifier and reward implementations are described in Appendix[F](https://arxiv.org/html/2608.03137#A6)\. All methods are evaluated on the same task instances with the same backbone, active\-context budget, and decoding configuration whenever backbone substitution is supported\.
### C\.1Benchmarks and Evaluation Splits
We evaluate VerMem on ALFWorld, SciWorld, PDDL, BabyAI, and HotpotQA\. ALFWorld, SciWorld, PDDL, and BabyAI use their official test splits\. HotpotQA uses the official distractor validation split as the held\-out evaluation set because it provides both reference answers and annotated supporting facts\. The HotpotQA training split is used only for the SFT and reinforcement\-learning instances described in Appendix[D](https://arxiv.org/html/2608.03137#A4)\. Original question identifiers do not overlap between the construction\-training pool, the internal development pool, and the held\-out evaluation set\.
ALFWorld, SciWorld, and BabyAI use environment success as the task outcome\. PDDL uses the progress value returned by the benchmark\. One HotpotQA question defines one evaluation episode\. The question remains fixed asqq, while the distractor passages enter the interaction as task observations\. Supporting\-fact annotations are not exposed to the memory operation policy\. They are used for training\-instance construction, programmatic supporting\-fact recall, verifier inputs, and evaluation\. The reference answer is likewise hidden from the memory operation policy\.
Table[9](https://arxiv.org/html/2608.03137#A3.T9)summarizes the evaluation split and reported metric for each benchmark\. Every executable instance in the selected split is evaluated\. All methods receive the same task order, environment initialization, and information source\.
Table 9:Evaluation splits and primary metrics\. Supporting\-fact annotations are not exposed to the memory operation policy\. They are used for training\-instance construction, programmatic supporting\-fact recall, verifier inputs, and evaluation\.
### C\.2State Isolation and Episode Initialization
Every task instance uses an isolated memory state\. Each episode starts with emptyM0M\_\{0\}andH0H\_\{0\}, whileC0C\_\{0\}contains only the task instruction and the initial input\. The initial input for ALFWorld, SciWorld, PDDL, and BabyAI is the first observation returned by the benchmark\. For HotpotQA, it contains the question and the passage content currently exposed by the task environment\.
The memory states persist only within one episode\. After task completion, failure, or budget termination,MtM\_\{t\},CtC\_\{t\}, andHtH\_\{t\}are released\. The vector store, graph state, metadata store, and temporary context of every external baseline are also cleared before the next task instance\. Information from one evaluation example can therefore never enter another example\.
### C\.3Backbones and Inference Configuration
The main experiments use Qwen2\.5\-7B\-Instruct and Qwen3\-4B\-Instruct\. VerMem, VerMem\-noVerify, and Base directly use the corresponding backbone\. External methods use the same backbone whenever their memory interface supports backbone substitution\. The corresponding method descriptions and reported workflows are followed\.
Evaluation uses deterministic decoding\. The memory operation policy and task solver use temperature0and top\-pp1\.01\.0\. VerMem constructsxt=g\(st\)x\_\{t\}=g\(s\_\{t\}\)under the8,1928\{,\}192\-token policy\-state serialization limit\. The active\-context budget applies separately toCtC\_\{t\}\. The configured protocol determines the task\-solver and memory\-command output limits\. Retrieve, SelectEpisode, and Summarize follow the configured active\-context budget\. The valid null command does not count as a memory\-tool call or in TC, but it remains a memory decision and counts toward the memory\-decision horizon\.
Returned memory content is committed only when the resultingCtC\_\{t\}satisfies the configured active\-context budget\. Separately, VerMem uses the bounded policy\-state serialization in Equation \([1](https://arxiv.org/html/2608.03137#S3.E1)\); no additional ranking or truncation rule is introduced here\. The local and global verifiers are disabled during evaluation\.
Table[10](https://arxiv.org/html/2608.03137#A3.T10)summarizes the shared inference configuration\.
Table 10:Shared inference configuration\. Token limits are measured with the tokenizer of the corresponding backbone\.
### C\.4Baseline Configurations
Base uses the same task solver, backbone, response limit, and active\-context budget as VerMem, but it has no explicit memory operation policy and no external persistent\-memory store\.
For LangMem, A\-Mem, Mem0, Mem0g, and AgeMem, the corresponding method descriptions and reported workflows are followed\. LangMem retains modular storage, search, and consolidation\. A\-Mem retains dynamic indexing, linking, and memory evolution\. Mem0 uses its extraction, consolidation, and retrieval pipeline\. Mem0guses the graph\-based memory workflow described for the method\. AgeMem retains its unified LTM/STM tool policy\. The final assembled prompt of every method is subject to the shared evaluation context configuration\. For VerMem, the policy\-state serialization limit and the active\-context budget are separate constraints\.
VerMem\-noVerify and VerMem share the architecture, backbone, complete atomic command space, SFT initialization, Phase A/B/C curriculum, GRPO configuration, and transition constraints\. VerMem\-noVerify uses the normalized task\-outcome advantage and hard constraints\. It does not use the local operation advantage or the composite global trajectory advantage\. VerMem uses the complete local–global credit assignment\.
Table[11](https://arxiv.org/html/2608.03137#A3.T11)summarizes the role of each compared system\. No method receives a reference answer, supporting\-fact label, or future observation during inference\.
Table 11:Baseline and internal\-control configurations\. Every method uses isolated task states and the same evaluation backbone whenever backbone substitution is supported\.
### C\.5Metric Computation
Success Rate on ALFWorld, SciWorld, and BabyAI is the number of successful episodes divided by the number of evaluated episodes\. PDDL uses the normalized Progress Rate returned by the benchmark and reports it on a0–100100scale\. HotpotQA uses the Qwen\-Max LLM\-as\-a\-Judge scoreJjudgeJ\_\{\\mathrm\{judge\}\}, and the main table reports100Jjudge100J\_\{\\mathrm\{judge\}\}\.
The Average column is the arithmetic mean of ALFWorld SR, SciWorld SR, PDDL PR, BabyAI SR, and HotpotQA100Jjudge100J\_\{\\mathrm\{judge\}\}\. Every benchmark therefore receives equal weight\.
The reward\-function ablation additionally reports TN, MQ, and TC\. TN is the mean number of online input and output tokens from the memory operation policy and task solver in one HotpotQA episode\. MQ evaluates the active entries in terminalMtM\_\{t\}against the annotated supporting facts\. Soft\-deleted entries are excluded\. TC is the mean number of non\-null memory\-tool calls\. The null action∅\\varnothingis not counted in TC, but it remains part of the memory\-decision horizon\.
### C\.6Qwen\-Max Evaluation
The HotpotQAJjudgeJ\_\{\\mathrm\{judge\}\}and MQ metrics use the fixedqwen\-max\-2025\-01\-25snapshot\. The evaluator is independent of the DeepSeek\-V3\.2 teacher and the local and global verifiers\. It does not generate SFT targets, provide training rewards, or participate in a memory transition\.
Qwen\-Max follows the configured deterministic evaluation protocol\. The configured protocol determines its decoding and output limits\. Every example is evaluated independently, and the output must contain one number in\[0,1\]\[0,1\]\. Malformed outputs follow the configured retry handling\.
The answer judge compares the question, ground\-truth answer, and agent answer\. Semantically equivalent concise answers can receive full credit\. Partially correct answers receive an intermediate score, while irrelevant or contradictory answers receive a score close to zero\. The following block gives the complete answer\-judge prompt\. Bracketed fields in this and subsequent prompt blocks denote values inserted at runtime\.
##### Qwen\-Max answer\-judge prompt\.
```
You are an independent judge evaluating the correctness
of an answer to a question.
Question:
[QUESTION]
Ground-truth answer:
[GROUND_TRUTH_ANSWER]
Agent answer:
[AGENT_ANSWER]
Score the agent answer from 0.0 to 1.0.
1.0: fully correct or semantically equivalent
0.8-0.9: correct with only a minor omission or wording issue
0.6-0.7: partially correct but missing an important element
0.4-0.5: contains some correct information and a major error
0.2-0.3: mostly incorrect with limited relevant content
0.0-0.1: incorrect, contradictory, or irrelevant
Return only one number between 0.0 and 1.0.
Do not provide an explanation.
```
The MQ evaluator receives the question, reference answer, annotated supporting facts, and active terminal LTM entries\. It evaluates supporting\-fact coverage, relevance to the question, and irrelevant stored content\. The following block gives the complete MQ prompt\.
##### Qwen\-Max memory\-quality prompt\.
```
You are an independent judge evaluating the quality of
long-term memory for question answering.
Question:
[QUESTION]
Reference answer:
[REFERENCE_ANSWER]
Annotated supporting facts:
[SUPPORTING_FACTS]
Active entries in the terminal long-term memory:
[TERMINAL_LTM]
Evaluate the terminal long-term memory on three criteria:
1. Does it cover the supporting facts needed to answer the
question?
2. Are the stored entries relevant and factually consistent?
3. Is it free of irrelevant, duplicated, or misleading
content?
Score the memory from 0.0 to 1.0.
1.0: complete, relevant, consistent, and free of noise
0.8-0.9: nearly complete with only a minor omission or
irrelevant entry
0.6-0.7: useful but missing important supporting content
0.4-0.5: contains some useful evidence and substantial noise
0.2-0.3: limited relevant content with major errors
0.0-0.1: unsupported, irrelevant, or unusable memory
Return only one number between 0.0 and 1.0.
Do not provide an explanation.
```
MQ is used only for the reward\-function analysis and is not included in the Average column of the main results\.
### C\.7Budget\-Controlled Efficiency Evaluation
The efficiency experiment uses Qwen2\.5\-7B\-Instruct and compares A\-Mem, Mem0, AgeMem, and VerMem on ALFWorld, SciWorld, and BabyAI\. Success Rate is computed separately for the three benchmarks and then macro\-averaged with equal weight\.
The online token budgets are1,0001\{,\}000,1,5001\{,\}500,2,0002\{,\}000,2,5002\{,\}500,3,0003\{,\}000,3,5003\{,\}500,4,0004\{,\}000,4,5004\{,\}500, and5,0005\{,\}000\. Each value is enforced as a hard episode budget\. The unfinished episode is counted as unsuccessful when its budget is exhausted\.
Online token accounting includes every input and output token from the memory operation policy and task solver\. Repeated prompt content is counted again at every call\. Online generative memory modules used by an external baseline are also included\. Training, offline embedding, index construction, local/global verifier calls, and Qwen\-Max evaluation are excluded\.
We selectS∗=40\.0S^\{\*\}=40\.0andB∗=2,500B^\{\*\}=2\{,\}500from the validation curves and fix both values before test evaluation\.SR@B∗\\mathrm\{SR\}@B^\{\*\}is the macro\-average SR at the2,5002\{,\}500\-token budget\.TN@S∗\\mathrm\{TN\}@S^\{\*\}is the linearly interpolated online\-token budget at which the mean success curve first reachesS∗S^\{\*\}, using the two adjacent evaluated budget points that bracket the target\.
Table[12](https://arxiv.org/html/2608.03137#A3.T12)summarizes the fixed efficiency protocol\. The lines in Figure[2](https://arxiv.org/html/2608.03137#S4.F2)show the mean across three seeds, and the shaded regions denote one standard deviation\.
Table 12:Fixed settings for the budget\-controlled efficiency–performance evaluation\.
### C\.8Random Seeds and Statistical Reporting
Training and evaluation use seeds4242,4343, and4444\. Every seed independently completes SFT, Phase A, Phase B, Phase C, and downstream evaluation\. Methods use the same benchmark instances and task order within one seed\.
The main result tables and quantitative appendix tables report the mean across three seeds\. The shaded regions in Figures[2](https://arxiv.org/html/2608.03137#S4.F2)and[4](https://arxiv.org/html/2608.03137#S4.F4)denote one standard deviation\.
Phase C checkpoints are selected only with the internal development set\. Official benchmark test results, held\-out HotpotQAJjudgeJ\_\{\\mathrm\{judge\}\}, and MQ are never used for checkpoint selection\. Efficiency thresholds, evaluator prompts, and all decoding settings are fixed before test evaluation\.
## Appendix DTraining Data and Stateful Rollout Construction
The training pipeline below expands the data construction, teacher generation, quality filtering, and stateful trajectory collection procedures introduced in the Training Pipeline\. We first describe how the HotpotQA training split is converted into memory\-decision states\. We then specify the SFT data composition, the frozen teacher prompt, and the filtering protocol\. The remaining subsections present phase\-specific episode construction and stateful candidate\-trajectory collection\.
### D\.1HotpotQA Training Data and Split
All SFT and reinforcement learning instances are derived from the HotpotQA training split\(Yang et al\.,[2018](https://arxiv.org/html/2608.03137#bib.bib27)\)\. Each original example provides a questionqq, a reference answer, annotated supporting facts, and distractor passages\. The question remains fixed within one training episode\. The reference answer supports instance construction and terminal task signals, but it is not exposed to the memory operation policy\. Supporting\-fact annotations are not exposed to the memory operation policy\. They are used for training\-instance construction, programmatic supporting\-fact recall, verifier inputs, and evaluation\.
We divide the original question identifiers in the HotpotQA training split into a90%90\\%construction\-training pool and a10%10\\%internal\-development pool\. All states and episodes derived from one original question remain in the same partition\. The official HotpotQA validation and test splits are not used to construct SFT data or Phase A, B, or C episodes\.
The configured construction protocol converts supporting facts and distractor passages into source\-traceable evidence units used to validate the atomic tool schemas\. It also constructs the partial, conflicting, duplicated, and obsolete training states needed by Update and Delete\. The exact representation fields and transformation rules are determined by that protocol\. These constructed states are not treated as native HotpotQA annotations\.
The final SFT training set contains32,00032\{,\}000states, and the internal development set contains4,0004\{,\}000states\. Add, Update, Delete, Retrieve, Filter, SelectEpisode, Summarize, and the null action∅\\varnothingare balanced in both sets\. Each decision contributes4,0004\{,\}000training examples and500500development examples\. Half of each class is derived from high\-precision boundary templates, and half is retained from teacher\-generated candidates\. At least25%25\\%of each class consists of boundary or negative decision states\.
### D\.2SFT Data Generation and Teacher Prompt
Boundary templates construct memory\-decision states with a well\-defined target\. They emphasize distinctions that are easy to confuse, including Add versus Update, Retrieve versus SelectEpisode, Filter versus Summarize, and an explicit operation versus∅\\varnothing\. For example, when a fact is already represented inMtM\_\{t\}and no new evidence is available, the target is∅\\varnothingrather than a duplicate Add\. When the missing evidence appears only inHtH\_\{t\}, the target is SelectEpisode rather than Retrieve\.
The frozen DeepSeek\-V3\.2 teacher expands the diversity of operation choices and structured arguments\. For each state, the teacher receives the bounded policy inputxix\_\{i\}, the operation\-type subspace allowed by the current scenario, and the source evidence already visible in the state\. The reference answer and future observations are not provided\. Teacher\-candidate sampling, decoding, and output limits follow the configured protocol\.
The following block gives the complete teacher prompt\. The teacher must return exactly one atomic operation with its structured arguments, or∅\\varnothing\. It must not return an analysis, an explanatory paragraph, or multiple candidate operations\.
##### SFT teacher prompt\.
```
You are a teacher model for constructing supervised
examples for VerMem.
Given a task and the current memory-decision state, select
exactly one atomic memory operation, or select the null
action when no memory change is needed.
Use only information that is already visible in the state.
Do not use the reference answer, future observations, or
unsupported knowledge.
Available operations:
1. Add
Use only when new information has persistent value and
is not already represented in long-term memory.
2. Update
Use when new evidence revises an existing long-term
memory entry.
3. Delete
Use only when an existing entry is clearly incorrect,
obsolete, duplicated, or unsupported.
4. Retrieve
Use when required information exists in long-term memory
but is missing from the active context.
5. Filter
Remove irrelevant active-context items while preserving
all evidence needed by the current task.
6. SelectEpisode
Restore task-relevant fragments from the episodic history
of the current task.
7. Summarize
Replace overlong active-context content with a shorter
representation that preserves critical evidence.
8. Null action
Use when the current memory state requires no change.
Output exactly one of the following forms:
Add(content="...", source_refs=[...])
Update(memory_id="...", new_content="...",
source_refs=[...])
Delete(memory_id="...", reason="...")
Retrieve(query="...")
Filter(drop_refs=[...])
SelectEpisode(selected_refs=[...])
Summarize(source_refs=[...], summary="...")
\varnothing
Do not provide analysis, explanations,
markdown, or more than one operation.
Task:
[TASK]
Long-term memory:
[LONG_TERM_MEMORY]
Active context:
[ACTIVE_CONTEXT]
Episodic history:
[EPISODIC_HISTORY]
Allowed operation-type subspace:
[ALLOWED_ACTIONS]
Visible source evidence:
[VISIBLE_EVIDENCE]
```
### D\.3SFT Quality Filtering and Optimization
Teacher outputs are not used directly as supervision labels\. Every candidate passes seven deterministic checks:
1. 1\.the output contains exactly one operation;
2. 2\.the operation belongs to the permitted operation\-type subspace;
3. 3\.all required arguments are present;
4. 4\.every referenced memory entry, context item, or historical fragment exists;
5. 5\.all references belong to the same HotpotQA training instance;
6. 6\.the realized transition satisfies the tool semantics and preconditions in Appendix[A](https://arxiv.org/html/2608.03137#A1);
7. 7\.the output contains no answer leakage, future information, or unsupported content\.
Any candidate that fails a check is discarded without automatic repair\. Valid candidates are then deduplicated according to their state, target operation, and structured arguments\. When two examples define equivalent states and equivalent transitions, only one is retained\.
Negative decision states never use an invalid operation as the supervision target\. The generator changes information completeness, candidate location, distractor density, or memory consistency so that a superficially plausible operation becomes inappropriate\. The target is another legal operation, corrected arguments, or∅\\varnothing\. A duplicate Add may therefore be corrected to Update or∅\\varnothing, while a Retrieve tendency is corrected to SelectEpisode when the missing evidence appears only inHtH\_\{t\}\.
Table[13](https://arxiv.org/html/2608.03137#A4.T13)reports the final SFT coverage\. Each target contributes4,0004\{,\}000training examples and500500development examples\.
Target DecisionTrainingDevelopmentAdd4,000500Update4,000500Delete4,000500Retrieve4,000500Filter4,000500SelectEpisode4,000500Summarize4,000500∅\\varnothing4,000500Total32,0004,000Table 13:Balanced target\-command coverage in the SFT training and internal development sets\.SFT uses AdamW for three epochs with a learning rate of2×10−52\\times 10^\{\-5\}, a warmup ratio of0\.030\.03, a maximum gradient norm of1\.01\.0, and an effective batch size of6464\. The policy\-state serialization limit is8,1928\{,\}192tokens\. The configured protocol determines the target\-operation output limit\. The loss is applied only to the tokens that serialize the complete target commandui⋆=\(vi⋆,ξi⋆\)u\_\{i\}^\{\\star\}=\(v\_\{i\}^\{\\star\},\\xi\_\{i\}^\{\\star\}\)\. Task descriptions, input states, system instructions, and padding tokens are excluded from the target\. Checkpoint selection uses tool accuracy and executable\-call rate on the development set\. When two checkpoints have the same tool accuracy, the one with the lower invalid\-operation rate is selected\.
### D\.4Phase\-Specific Episode Construction
Reinforcement learning follows the parameter curriculum in Equation \([4](https://arxiv.org/html/2608.03137#S3.E4)\)\. Each phase initializes its task states independently and transfers only the policy parameters from the previous phase\. Table[14](https://arxiv.org/html/2608.03137#A4.T14)summarizes the data size, operation\-type subspace, and episode construction of each phase\.
Table 14:Phase\-specific episode construction\. Only policy parameters are transferred between phases; memory states are initialized independently\.##### Phase A\.
Phase A contains8,0008\{,\}000training episodes and800800development episodes\. The initialM0M\_\{0\}contains the phase\-specific persistent state,C0C\_\{0\}contains the current observation, andH0H\_\{0\}begins with the current episode\. Add, Update, Delete, and∅\\varnothingscenario families are sampled uniformly\. Supporting facts are revealed incrementally, and each episode ends with a delayed query\. The rollout terminates when the delayed query is completed, the token budget is exhausted, or another configured termination condition is reached\.
##### Phase B\.
Phase B contains8,0008\{,\}000training episodes and800800development episodes\. Each instance uses supporting documents and distractor documents from the HotpotQA distractor setting\. The initialM0M\_\{0\}contains phase\-specific entries\. The initialC0C\_\{0\}occupies part of the configured active\-context budget, andH0H\_\{0\}contains same\-task history\. Retrieve, Filter, SelectEpisode, Summarize, and∅\\varnothingscenario families are sampled uniformly\.
Historical\-context recovery cases place required evidence in earlier same\-task history while recent episodes contain related distractors\. Retrieve accessesMtM\_\{t\}, whereas SelectEpisode accessesHtH\_\{t\}\. Returned content must satisfy the configured active\-context budget\. Valid null and rejected commands count toward the memory\-decision horizon; valid null commands do not count as memory\-tool calls\.
##### Phase C\.
Phase C contains12,00012\{,\}000training episodes and1,2001\{,\}200development episodes\. Every episode requires at least one LTM operation and one STM operation\. Some episodes require historical\-context recovery so that SelectEpisode receives supervision in complete tasks\. Documents, task feedback, and observations are revealed sequentially\. Early Add or Update decisions may affect a later Retrieve, while Filter or Summarize may change the evidence available to subsequent reasoning\. Valid null and rejected commands count toward the memory\-decision horizon\.
All phases use the8,1928\{,\}192\-token bounded policy\-state serialization defined by Equation \([1](https://arxiv.org/html/2608.03137#S3.E1)\)\. The active\-context budget applies separately toCtC\_\{t\}\. The configured protocol determines the task\-solver and memory\-command output limits\.
### D\.5Stateful Candidate\-Trajectory Collection
At updatennof phasep∈\{A,B,C\}p\\in\\\{A,B,C\\\}, the current memory operation policy samples eight tasks\. For taskqq, the generator constructs the phase\-specific initial states0,p\(q\)s\_\{0,p\}^\{\(q\)\}and copies it intoK=8K=8isolated rollouts according to Equation \([8](https://arxiv.org/html/2608.03137#S3.E8)\)\. All trajectories in one candidate group share the task, initialM0M\_\{0\},C0C\_\{0\},H0H\_\{0\}, and environment condition\. They differ only in policy sampling and the resulting operation sequence\. Rollouts use temperature0\.70\.7and top\-pp0\.950\.95\.
At every step, the memory operation policy generates one complete commandut,ku\_\{t,k\}, including∅\\varnothing\. The valid non\-null command commits its transition through𝒯m\\mathcal\{T\}\_\{m\}\. The null command leavesMtM\_\{t\}andCtC\_\{t\}unchanged\. On these states,𝒯m\\mathcal\{T\}\_\{m\}is the identity transition\. The identity decision is retained in the rollout record and does not create a memory\-state modification\. If a non\-null command violates a precondition in Appendix[A](https://arxiv.org/html/2608.03137#A1),MtM\_\{t\}andCtC\_\{t\}remain unchanged, the failure record is added toHtH\_\{t\}, and the correspondingct,kc\_\{t,k\}is recorded\. The task solver then continues with the resulting active context\. Each rollout record retains the bounded policy inputxt,kx\_\{t,k\}, the complete commandut,ku\_\{t,k\}, the post\-memory states~t\\widetilde\{s\}\_\{t\}, the post\-interaction statest\+1s\_\{t\+1\}, and the command\-token log probabilities\. This preserves the transition distinction made in Equations \([2](https://arxiv.org/html/2608.03137#S3.E2)\) and \([3](https://arxiv.org/html/2608.03137#S3.E3)\) without introducing a second transition\-record symbol\.
Each trajectory terminates when the task solver produces a final answer, the environment terminates, the configured rollout limit is reached, or the token budget is exhausted\. Budget\-terminated trajectories remain in the candidate group and are processed by the composite global branch and constraint channel\. The global verifier contributes only its evidence\-coherence and terminal\-memory\-consistency components\.
##### Stateful collection procedure\.
For every task in the current update, the collector copiess0,p\(q\)s\_\{0,p\}^\{\(q\)\}intoKKisolated rollouts\. Each rollout repeatedly samplesut,ku\_\{t,k\}from the current policy conditioned onxt,kx\_\{t,k\}, realizes or rejects𝒯m\\mathcal\{T\}\_\{m\}, applies the separate task and environment transition, and stores boths~t\\widetilde\{s\}\_\{t\}andst\+1s\_\{t\+1\}\. The completed trajectories form𝒢q,p\(n\)\\mathcal\{G\}\_\{q,p\}^\{\(n\)\}for local and composite\-global credit assignment\.
Phase A and Phase B each use1,0001\{,\}000policy updates, while Phase C uses1,5001\{,\}500updates\. With eight tasks and eight trajectories per task, the three phases collect approximately64,00064\{,\}000,64,00064\{,\}000, and96,00096\{,\}000candidate trajectories, respectively\.
## Appendix EAdditional Quantitative Analyses
The analyses below provide detailed comparisons for the main results, component study, verifier ablation, efficiency evaluation, and reward\-function ablation\. Every value is reported in the main paper or computed directly from the reported results\. No additional evaluation setting is introduced\.
### E\.1Detailed Cross\-Backbone Comparison
Table[15](https://arxiv.org/html/2608.03137#A5.T15)compares VerMem with the strongest external baseline for each benchmark and backbone\. With Qwen2\.5\-7B\-Instruct, the absolute gains range from3\.673\.67points on PDDL to8\.318\.31points on HotpotQA\. With Qwen3\-4B\-Instruct, the corresponding range is3\.683\.68to8\.128\.12points\. The average gains remain close across the two backbones, reaching6\.056\.05and5\.545\.54points\. HotpotQA and SciWorld show the largest improvements in both settings\. Both benchmarks involve evidence preservation, intermediate\-state maintenance, and active\-context control\. PDDL shows the smallest gain, but VerMem remains above the strongest external planning baseline under both backbones\.
Table 15:Detailed comparison between VerMem and the strongest external baseline for each benchmark and backbone\. Gains are absolute differences on the reporting scale used in the main table\.
### E\.2Effect of Local and Composite\-Global Credit Across All Benchmarks
The main credit\-branch ablation focuses on ALFWorld, SciWorld, and HotpotQA\. Table[16](https://arxiv.org/html/2608.03137#A5.T16)compares full VerMem with the matched VerMem\-noVerify control across all five benchmarks\. VerMem\-noVerify uses the normalized task\-outcome advantage and hard constraints\. It does not use the local operation advantage or the composite global trajectory advantage\. Adding both credit branches raises four of the five task scores under Qwen2\.5\-7B\-Instruct and raises the Average by6\.256\.25points\. PDDL is the only exception, with a difference of−0\.35\-0\.35points\. With Qwen3\-4B\-Instruct, VerMem improves every task and raises the Average by6\.216\.21points\. The largest gains occur on HotpotQA and SciWorld for both backbones\. This pattern is consistent with complementary operation\-level and trajectory\-level credit\.
Table 16:Absolute difference between VerMem and VerMem\-noVerify across all tasks\. Positive values favor VerMem\. VerMem\-noVerify uses the normalized task\-outcome advantage and hard constraints; it does not use the local operation advantage or the composite global trajectory advantage\.
### E\.3Progressive Component Contributions
Table[17](https://arxiv.org/html/2608.03137#A5.T17)reports the score and marginal gain at every step of the cumulative component study\. The transition from Base to \+LT produces the largest initial improvement, especially on SciWorld\. Enabling the complete LTM/STM tool space yields further gains on all three tasks\. Stateful RL with the normalized task\-outcome advantage and hard constraints produces a larger gain on HotpotQA than on the two interactive environments\. The final transition from \+noV to \+V adds the local operation advantage and the composite global trajectory advantage\. It improves ALFWorld, SciWorld, and HotpotQA by6\.126\.12,9\.459\.45, and9\.679\.67points\. The progression is consistent with the two credit branches building on unified memory control and stateful training\.
Table 17:Numerical decomposition of the cumulative component study\. Each increment is measured relative to the preceding configuration\. \+noV uses the normalized task\-outcome advantage and hard constraints\. \+V adds the local operation advantage and the composite global trajectory advantage and denotes full VerMem\.
### E\.4Complementarity of Local and Composite\-Global Credit
Table[18](https://arxiv.org/html/2608.03137#A5.T18)separates the standalone and conditional gains of the two credit branches\. The local branch provides the larger standalone gain on ALFWorld and SciWorld, while the composite global branch provides the larger standalone gain on HotpotQA\. The conditional differences are consistent with complementary contributions after either branch has been added\. Adding composite\-global credit after local credit yields further gains of1\.581\.58,3\.583\.58, and4\.494\.49points\. Adding local credit after composite\-global credit yields2\.992\.99,5\.435\.43, and2\.572\.57points\. The gains support complementary roles for the two credit branches\.
Table 18:Standalone and conditional gains of the local and composite\-global credit branches under Qwen2\.5\-7B\-Instruct\. The composite global branch contains programmatic and global\-verifier components\. \+Local adds the local operation advantage\. \+Global replaces the task\-only trajectory advantage with the composite global trajectory advantage\.
### E\.5Pairwise Efficiency Comparison
Table[19](https://arxiv.org/html/2608.03137#A5.T19)expands the two efficiency metrics into pairwise comparisons\. VerMem reduces the token cost required to reachS∗=40\.0S^\{\*\}=40\.0by20\.0%20\.0\\%relative to AgeMem,45\.5%45\.5\\%relative to Mem0, and51\.1%51\.1\\%relative to A\-Mem\. Under the fixed budgetB∗=2,500B^\{\*\}=2\{,\}500, it improves macro\-average SR by5\.305\.30,8\.708\.70, and10\.5010\.50points\. The ordering is consistent under both metrics and supports a stronger efficiency–performance frontier for VerMem\.
Table 19:Pairwise efficiency comparison underS∗=40\.0S^\{\*\}=40\.0andB∗=2,500B^\{\*\}=2\{,\}500\. Reduction reports the relative decrease inTN@S∗\\mathrm\{TN\}@S^\{\*\}achieved by VerMem\.
### E\.6Reward\-Function Effects
Table[20](https://arxiv.org/html/2608.03137#A5.T20)reports the absolute and relative changes associated with All\-Returns\. The complete reward raisesJjudgeJ\_\{\\mathrm\{judge\}\}by0\.0970\.097and MQ by0\.1820\.182, which correspond to relative improvements of18\.3%18\.3\\%and37\.1%37\.1\\%\. TN decreases by168168tokens, or7\.14%7\.14\\%, while TC increases by1\.161\.16calls\. The simultaneous increase in TC and reduction in TN is inconsistent with uniform tool\-call suppression as the sole explanation\. These aggregate metrics do not identify operation\-specific contributions\.
Table 20:Absolute and relative effects of the complete reward on HotpotQA using Qwen2\.5\-7B\-Instruct\. Negative TN change is favorable\.The reported ablations associate the complete command space, task\-outcome stateful RL, and the complete training signal with higher scores across both evaluated backbones\. VerMem also attains the stronger reported efficiency frontier\. The results support the joint role of unified memory management and multi\-granularity credit assignment\.
## Appendix FVerifier, Credit, Reward, and Reproducibility Specifications
The verifier specification below covers the local and global verifiers, operation\-wise normalization, the constraint channel, and the command\-token loss mask and monitoring protocol introduced in the main paper\. The local verifier evaluates one realized atomic memory transition\. The global verifier supplies evidence\-coherence and terminal\-memory\-consistency scores after task completion; it does not compute task correctness, supporting\-fact recall, or efficiency\. Structural validity, state\-transition preconditions, safety requirements, and budget limits are handled by the constraint channel rather than the semantic verifier scores\. This separation keeps semantic and hard\-constraint penalties distinct\.
### F\.1Verifier Instantiation and Scoring Protocol
The local and global verifiers use frozen DeepSeek\-V3\.2 instances with separate prompts and input scopes\. The local verifier receives the taskqq, the bounded policy\-visible statext=g\(st\)x\_\{t\}=g\(s\_\{t\}\), the complete commandutu\_\{t\}, the realized memory transition, and the source evidence visible at the current decision\. It does not receive an unbounded concatenation ofHtH\_\{t\}, nor does it observe later memory decisions, the final answer, or the global\-verifier output\. The global verifier is invoked after trajectory termination\. Its input contains the task, reference answer, annotated supporting facts, final answer, ordered memory commands, their realized transitions, the terminalMtM\_\{t\},CtC\_\{t\}, andHtH\_\{t\}\. Online cost is computed programmatically and is not provided as a verifier\-produced score\.
Verifier\-input truncation, decoding, and output limits are determined by the configured protocol\. No separate ranking, removal priority, or compression rule is specified here\.
Table[21](https://arxiv.org/html/2608.03137#A6.T21)summarizes the verifier configuration\.
Table 21:Configuration of the local and global verifiers\. The two verifiers use separate prompts and are disabled during evaluation\.Each verifier dimension is scored with an integer in\{0,1,2,3,4\}\\\{0,1,2,3,4\\\}\. Score0denotes a contradicted or harmful decision; score11denotes weak support with major errors; score22denotes a partially justified decision with clear omissions; score33denotes a sound decision with minor deficiencies; and score44denotes a fully justified decision\. Scores are divided by44before reward computation\.
The implementation performs deterministic parsing and retry handling\. The exact fallback after repeated malformed verifier outputs is determined by the execution protocol\.
### F\.2Local Verifier
The local verifier evaluates one executable complete command and its realized state transition\. Executable commands include valid null decisions, which use the separate null rubric below\. It scores task relevance, evidence grounding, local progress, and information fidelity\. The four scores are divided by44and averaged to obtainrt,klocalr\_\{t,k\}^\{\\mathrm\{local\}\}\. The trajectory\-level local reward in Equation \([12](https://arxiv.org/html/2608.03137#S3.E12)\) sums the scores produced for executable decisions inℰk\\mathcal\{E\}\_\{k\}and normalizes them by the full memory\-decision horizonTkT\_\{k\}\. It is used only for monitoring\. It does not replace the operation\-wise normalization in Equation \([9](https://arxiv.org/html/2608.03137#S3.E9)\) used to computeAt,klocalA\_\{t,k\}^\{\\mathrm\{local\}\}\.
Task relevance measures whether the command addresses an actual memory need in the bounded statextx\_\{t\}\. Evidence grounding measures whether its content, target, and source are supported by the currently visible evidence\. Local progress measures whether the transition improves the current task state\. Information fidelity measures whether the operation preserves valid facts, relations, and source attribution\.
Table[22](https://arxiv.org/html/2608.03137#A6.T22)gives the tool\-specific conditions for the highest local score\. Lower scores reflect increasing departures from these conditions\.
Table 22:Operation\-specific rubric used by the local verifier\. Structural validity and hard constraints are evaluated separately\.Commands rejected by the execution checks in Appendix[A](https://arxiv.org/html/2608.03137#A1)are not sent to the local verifier\. TheirAt,klocalA\_\{t,k\}^\{\\mathrm\{local\}\}is set to zero\. They still receive the trajectory\-levelAkglobalA\_\{k\}^\{\\mathrm\{global\}\}, while the correspondingct,kc\_\{t,k\}supplies the direct decision\-level penalty\.
The following block gives the local verifier prompt\. The operation\-specific condition is selected from Table[22](https://arxiv.org/html/2608.03137#A6.T22)\.
##### Local\-verifier prompt\.
```
You are the local verifier used during VerMem training.
Evaluate exactly one realized atomic memory transition.
The transition has already passed the structural,
reference, precondition, and budget checks.
Use only the task, the state before the command, the
selected command, the realized state after the command,
and the visible source evidence.
Do not use a later trajectory state, the final answer, or
information that is absent from the supplied state.
Score four dimensions with integers from 0 to 4.
0: contradicted, harmful, or entirely unjustified
1: weakly supported with major errors
2: partially justified with clear omissions
3: sound with minor deficiencies
4: fully justified for the current state
RELEVANCE
Does the command address the current memory need?
GROUNDING
Is the command supported by visible source evidence?
LOCAL_PROGRESS
Does the realized transition improve the current task or
memory state?
INFORMATION_FIDELITY
Does the transition preserve the valid information required
from its sources?
Output exactly four lines and no additional text:
RELEVANCE: <0-4>
GROUNDING: <0-4>
LOCAL_PROGRESS: <0-4>
INFORMATION_FIDELITY: <0-4>
Task q:
[TASK]
Bounded state before the command x_t:
[STATE_BEFORE]
Complete memory command u_t:
[OPERATION]
State after the command:
[STATE_AFTER]
Visible source evidence:
[VISIBLE_EVIDENCE]
Operation-specific condition:
[OPERATION_CONDITION]
```
### F\.3Composite Global Branch and Global Verifier
The composite global branch follows Equations \([13](https://arxiv.org/html/2608.03137#S3.E13)\) and \([14](https://arxiv.org/html/2608.03137#S3.E14)\)\. It combines task correctness, supporting\-fact recall, evidence coherence, terminal\-memory consistency, and trajectory efficiency\. The global verifier produces only evidence coherence and terminal\-memory consistency; the other components are programmatic\. The four resulting reward terms in Equation \([13](https://arxiv.org/html/2608.03137#S3.E13)\) are normalized to\[0,1\]\[0,1\]and averaged with equal weight\.
VerMem\-noVerify uses the normalized task\-outcome advantage and hard constraints\. It does not use the local operation advantage or the composite global trajectory advantage\.
For HotpotQA,rktaskr\_\{k\}^\{\\mathrm\{task\}\}is the mean of exact match and token\-level F1 against the reference answer\. Both values are normalized to\[0,1\]\[0,1\]\. If no final answer is produced,rktaskr\_\{k\}^\{\\mathrm\{task\}\}is set to zero\.
The evidence component uses both annotated supporting facts and global verification\. Supporting\-fact annotations are not exposed to the memory operation policy\. They are used for training\-instance construction, programmatic supporting\-fact recall, verifier inputs, and evaluation\. The exact evidence extraction procedure follows the implementation protocol and is not redefined here\. The supporting\-fact recall is averaged with the normalized evidence\-coherence score returned by the global verifier to obtainrkevidr\_\{k\}^\{\\mathrm\{evid\}\}\. Evidence coherence measures whether the final answer follows from a complete, task\-relevant, and non\-contradictory evidence chain\.
The global verifier also returns terminal\-memory consistency\. Its normalized score formsrkstater\_\{k\}^\{\\mathrm\{state\}\}\. This dimension checks whether the terminal LTM, active context, and episodic history remain consistent with the observed evidence, whether critical supporting facts remain accessible, and whether obsolete or conflicting entries have been handled\. BecauseHtH\_\{t\}is an ordered provenance record, superseded, soft\-deleted, rejected, or failed events may remain in history when their status is explicit\.
The normalized online costC¯konline\\overline\{C\}\_\{k\}^\{\\mathrm\{online\}\}uses online tokens, task steps, and memory\-tool calls\. The three costs are independently min–max normalized within theK=8K=8candidate trajectories sampled for the same task and then averaged\. Each cost component is set to zero when all candidate trajectories have the same value for that component\. Teacher calls, verifier calls, Qwen\-Max evaluation, and offline index construction are excluded\.
The completion threshold isδ=0\.5\\delta=0\.5, as used in Equation \([15](https://arxiv.org/html/2608.03137#S3.E15)\)\. The efficiency component is available only whenrktask≥δr\_\{k\}^\{\\mathrm\{task\}\}\\geq\\delta\. Under this gate, a low\-cost failed trajectory receives no positive efficiency reward\.
The global verifier does not rescore tool syntax, required arguments, or hard preconditions\. Task completion, supporting\-fact recall, and online cost are computed programmatically\. The verifier evaluates only evidence coherence and terminal\-memory consistency\.
The following block gives the global verifier prompt\.
##### Global\-verifier prompt\.
```
You are the global verifier used during VerMem training.
Evaluate the completed task trajectory as a whole.
Answer correctness, supporting-fact recall, online cost,
tool syntax, and hard preconditions are computed
separately. Do not rescore these items.
Score two dimensions with integers from 0 to 4.
0: contradicted, incoherent, or seriously inconsistent
1: major evidence or terminal-memory failures
2: partially coherent with important omissions
3: coherent with minor deficiencies
4: complete, well supported, and internally consistent
EVIDENCE_COHERENCE
Does the final answer follow from a complete,
task-relevant, and non-contradictory evidence chain?
TERMINAL_MEMORY_CONSISTENCY
Are the terminal long-term memory, active context, and
episodic history consistent with the observed evidence?
Are critical facts preserved and accessible? Are obsolete
or conflicting entries handled appropriately? Historical
superseded, deleted, rejected, or failed events may remain
when their status is explicit.
Output exactly two lines and no additional text:
EVIDENCE_COHERENCE: <0-4>
TERMINAL_MEMORY_CONSISTENCY: <0-4>
Task q:
[TASK]
Reference answer:
[REFERENCE_ANSWER]
Annotated supporting facts:
[SUPPORTING_FACTS]
Final answer:
[FINAL_ANSWER]
Ordered memory commands and realized transitions:
[MEMORY_TRAJECTORY]
Terminal long-term memory:
[TERMINAL_LTM]
Terminal active context:
[TERMINAL_CONTEXT]
Terminal episodic history:
[TERMINAL_HISTORY]
```
### F\.4Advantage Normalization and Hierarchical Credit
The composite global score is normalized within theK=8K=8trajectories sampled from the same task\. Local scores are normalized by operation typevt,kv\_\{t,k\}, as defined in Equation \([9](https://arxiv.org/html/2608.03137#S3.E9)\)\. The seven atomic tools and the null typev=∅v=\\varnothingmaintain separate local statistics\.
When an operation type appears at least eight times in the current update, its current\-update mean and standard deviation are used\. When fewer than eight instances are available, phase\-specific running statistics are used\. Before each phase, the running statistics are initialized from128128executable development transitions for every decision type enabled in that phase\. The statistics are updated with an exponential moving\-average decay of0\.990\.99\. They are reinitialized at the start of every phase\.
Both local and global normalization use a standard\-deviation floor of0\.010\.01, and theϵ\\epsilonin the main paper is set to10−610^\{\-6\}\. If every trajectory in a candidate group has the same composite global score, all corresponding global advantages are zero\. Local and global advantages are clipped to\[−5,5\]\[\-5,5\]before they are combined\.
The three coefficients retain the unit values specified with Equation \([5](https://arxiv.org/html/2608.03137#S3.E5)\)\. Each rejected command receivesAt,klocal=0A\_\{t,k\}^\{\\mathrm\{local\}\}=0\. Its global advantage remains available because the completed trajectory still determinesAkglobalA\_\{k\}^\{\\mathrm\{global\}\}\. The constraint cost is assigned directly to the corresponding memory decision\.
Table[23](https://arxiv.org/html/2608.03137#A6.T23)summarizes the fixed normalization and optimization settings\.
Table 23:Fixed normalization, credit\-assignment, and stabilization settings used by VerMem\.
### F\.5Auxiliary Constraints
The constraint channel handles discrete and deterministic violations\. Table[24](https://arxiv.org/html/2608.03137#A6.T24)reports the fixed costs and execution behavior\. Multiple violations at one memory decision are capped atct,k=1\.0c\_\{t,k\}=1\.0\. The valid null decision counts toward the memory\-decision horizon used by the auxiliary cost, but it is not a memory\-tool call and does not incur a tool\-call cost\.
Table 24:Constraint costs and execution behavior\. Costs from multiple violations at one memory decision are capped at1\.01\.0\.Evidence grounding and information fidelity remain continuous semantic scores\. The unsupported\-operation constraint is triggered only when no valid source exists or the operation directly contradicts visible evidence\. The evidence\-loss constraint is triggered only when indispensable information is removed and cannot be recovered fromMtM\_\{t\}orHtH\_\{t\}\. Partial omissions lower the local score but do not trigger a hard constraint\.
### F\.6Token\-Level GRPO Implementation
The policy ratio is evaluated token by token asρt,k,j\(θ\)\\rho\_\{t,k,j\}\(\\theta\)in Equation \([10](https://arxiv.org/html/2608.03137#S3.E10)\)\. All generated tokens that serialize one complete commandut,ku\_\{t,k\}share the sameAt,khierA\_\{t,k\}^\{\\mathrm\{hier\}\}\. The token loss of one command is averaged over its command tokens before losses are averaged overℬop\\mathcal\{B\}\_\{\\mathrm\{op\}\}\. This gives each command equal weight before batch averaging, independent of argument length\.
The loss mask includes only the complete command tokensyt,k,1:Lt,ky\_\{t,k,1:L\_\{t,k\}\}, including the operation type and structured arguments\. The task specification,MtM\_\{t\},CtC\_\{t\},HtH\_\{t\}, system instructions, task\-solver outputs, environment observations, and padding tokens are excluded\. The null action∅\\varnothingremains a trainable memory decision, and it counts toward the memory\-decision horizon\. It is not counted as a memory\-tool call\.
The old policy is frozen before each policy update\. The reference policy remains fixed within one phase\. Phase A, Phase B, and Phase C useθSFT\\theta\_\{\\mathrm\{SFT\}\},θA\\theta\_\{A\}, andθB\\theta\_\{B\}, respectively, as their reference policies\.
The objective usesϵclip=0\.2\\epsilon\_\{\\mathrm\{clip\}\}=0\.2andβKL=0\.1\\beta\_\{\\mathrm\{KL\}\}=0\.1, as specified in Equation \([11](https://arxiv.org/html/2608.03137#S3.E11)\)\. The reinforcement\-learning rate is10−610^\{\-6\}, the maximum gradient norm is1\.01\.0, and no additional entropy bonus is used\. Local and global advantages are clipped to\[−5,5\]\[\-5,5\]\.
### F\.7Reward Monitoring and Convergence Analysis
Figure[4](https://arxiv.org/html/2608.03137#S4.F4)compares the strategy\-specific monitoring values of Answer\-Only and All\-Returns\. They use the mappingsmkansm\_\{k\}^\{\\mathrm\{ans\}\}andmkallm\_\{k\}^\{\\mathrm\{all\}\}in Equations \([18](https://arxiv.org/html/2608.03137#S3.E18)\) and \([17](https://arxiv.org/html/2608.03137#S3.E17)\)\. Since the two strategies use different monitoring definitions, their absolute values are not required to coincide at the first recorded step\.
The complete reinforcement\-learning curriculum contains3,5003\{,\}500policy updates\. We record101101training points\. Training Step0is the rollout batch collected before the first policy update, and consecutive recorded points are separated by3535cumulative updates\. At every recorded point, the monitoring value is averaged over the6464candidate trajectories in the current rollout batch\. The plotted line is the mean across seeds4242,4343, and4444, and the shaded region denotes one standard deviation across the three runs\.
For the common\[0,1\]\[0,1\]axis, Answer\-Only plotsmkansm\_\{k\}^\{\\mathrm\{ans\}\}, while All\-Returns plotsmkallm\_\{k\}^\{\\mathrm\{all\}\}, whose global term isrkglobalr\_\{k\}^\{\\mathrm\{global\}\}\. These mappings are used only for training visualization\. They are not policy advantages and do not change the policy objective or checkpoint selection\.
The curves use the raw recorded\-point statistics without a moving average, temporal smoothing, or spline interpolation\. White squares mark every tenth recorded point\. The observed curves are consistent with earlier stabilization and lower late\-stage variability for the All\-Returns configuration, but the monitoring mappings differ and their absolute levels are not directly comparable\.
### F\.8Protocol\-Level Behavior and Failure Analysis
This section organizes protocol\-level memory\-decision patterns and failure cases implied by the tool semantics and reported ablations\. It does not introduce another metric or claim an additional empirical case study\.
#### Phase\-Specific Memory Decisions
The three training phases expose different memory needs\. Phase A presents incremental facts, corrections, duplicate content, and obsolete entries\. The correct decision therefore depends on whether the visible evidence should create, revise, remove, or leave unchanged an entry inMtM\_\{t\}\. Phase B presents a prepared LTM state, a distractor\-richCtC\_\{t\}, and same\-task history inHtH\_\{t\}\. It requires the policy to distinguish Retrieve from SelectEpisode and Filter from Summarize\. Phase C combines both tool groups in one trajectory\. Early Add or Update decisions may support a later Retrieve, while Filter, Summarize, or SelectEpisode changes the evidence visible to the task solver\.
Table[25](https://arxiv.org/html/2608.03137#A6.T25)summarizes the decision patterns that define these scenarios\. The state effects follow the tool semantics in Appendix[A](https://arxiv.org/html/2608.03137#A1)\.
Table 25:Qualitative decision patterns covered by the three\-stage training curriculum\. Each row links a memory need to the corresponding atomic operation, realized state effect, and targeted failure mode\.
#### Historical\-Context Recovery
The historical\-context recovery scenarios in Phase B make the distinction betweenMtM\_\{t\}andHtH\_\{t\}explicit\. The initialC0C\_\{0\}is distractor\-rich, whileH0H\_\{0\}contains same\-task history\. Retrieve cannot recover evidence that was never written toMtM\_\{t\}\. SelectEpisode instead identifies the relevant same\-task fragment inHtH\_\{t\}and restores it toCtC\_\{t\}, subject to the configured active\-context budget\.
Filter may then remove distractors, while Summarize may compress the restored evidence before the task solver continues\. This sequence recovers earlier task evidence without treating episodic history as another persistent memory store\.
During Phase C, historical\-context recovery can occur between LTM maintenance and final reasoning\. Early Add or Update decisions may preserve persistent knowledge, while a later SelectEpisode restores an observation or tool output that was useful only after the task focus changed\. The combined state transition tests whether the policy coordinatesMtM\_\{t\},CtC\_\{t\}, andHtH\_\{t\}rather than relying on one memory source\.
#### Interpretation of Local and Composite\-Global Credit
The two credit branches answer different questions\. The local verifier evaluates whether the current executable commandutu\_\{t\}is relevant, grounded, useful, and faithful inxtx\_\{t\}\. The composite global branch combines task correctness, supporting\-fact recall, efficiency, and the global verifier’s evidence\-coherence and terminal\-memory\-consistency scores\. Table[26](https://arxiv.org/html/2608.03137#A6.T26)shows how their normalized advantages treat common trajectory patterns\.
Table 26:Qualitative interpretation of local and composite\-global credit\. The two advantages distinguish command quality from completed\-task utility\.The second and third rows explain the main role of hierarchical credit assignment\. Trajectory\-only supervision cannot separate a useful command followed by a later failure from a poor command hidden inside a successful trajectory\. Operation\-wise normalization further reduces the risk that tool\-specific score ranges dominate the update\.
#### Failure Modes
The failure modes in Table[27](https://arxiv.org/html/2608.03137#A6.T27)follow directly from the seven tool semantics and the constraint channel\. Some errors are semantic and reducert,klocalr\_\{t,k\}^\{\\mathrm\{local\}\}\. Hard failures reject the transition and contribute toct,kc\_\{t,k\}\.
Table 27:Representative failure modes and their treatment by local semantic verification, composite\-global credit, and hard constraints\.The distinction between semantic and hard failures is important\. The irrelevant Retrieve can remain structurally valid and therefore receive a low semantic score rather than automatic rejection\. Minor but executable omissions in Filter or Summarize are handled the same way\. By contrast, unsupported Update, unjustified Delete, protected\-evidence removal, and meaning\-changing Summarize violate an execution requirement\. They are rejected, receive no local semantic evaluation, retain trajectory\-level global credit, and receive a direct constraint cost\. This separation keeps the semantic and hard\-constraint penalties in distinct channels\.
#### Relation to the Quantitative Results
The qualitative patterns are consistent with the experiments in the main paper\. Figure[3](https://arxiv.org/html/2608.03137#S4.F3)shows that LTM operations provide the first major gain, while the complete LTM/STM command space adds further improvement\. This progression matches the decisions in Table[25](https://arxiv.org/html/2608.03137#A6.T25)\. Persistent storage is useful only when the policy can also retrieve, filter, summarize, or restore information at the correct time\.
Table[4](https://arxiv.org/html/2608.03137#S4.T4)shows complementary gains from the local and composite\-global credit branches\. The credit patterns in Table[26](https://arxiv.org/html/2608.03137#A6.T26)illustrate this distinction\. Local credit distinguishes the quality of one realized transition\. The global verifier evaluates evidence coherence and terminal\-memory consistency, while the composite global branch combines those scores with task correctness, supporting\-fact recall, and efficiency\. The observed gains support complementary roles for operation\-level and trajectory\-level credit\.
The reward ablation in Table[5](https://arxiv.org/html/2608.03137#S4.T5)reports higherJjudgeJ\_\{\\mathrm\{judge\}\}and MQ, lower TN, and more memory\-tool calls under All\-Returns\. This pattern is inconsistent with uniform tool\-call suppression as the sole explanation for the lower token count, but it does not identify operation\-specific contributions\. The stronger efficiency frontier in Figure[2](https://arxiv.org/html/2608.03137#S4.F2)provides the corresponding system\-level result\. Detailed numerical decompositions are reported in Appendix[E](https://arxiv.org/html/2608.03137#A5)\.Similar Articles
MemGuard: Persisting Verifier Signals for LLM-Agent Memory Governance
MemGuard introduces a system that persists verifier signals as metadata to govern memory in LLM agents, improving reliability and performance across multiple benchmarks like SWE-Bench and WebArena.
RecMem: Recurrence-based Memory Consolidation for Efficient and Effective Long-Running LLM Agents
RecMem is a recurrence-based memory consolidation method for long-running LLM agents that reduces token consumption by up to 87% while improving accuracy, by only invoking LLMs when semantically similar interactions recur.
ChronoMem: Version Control and Semantic Rollback for Large Language Model Agent Memory
ChronoMem introduces a semantic version-control layer for LLM agent memory, enabling whole-memory snapshots, natural-language rollback via hybrid retrieval, and counterfactual evaluation. It is the first open-source system and benchmark for global memory rollback in LLM agents.
AdMem: Advanced Memory for Task-solving Agents
This paper introduces AdMem, a unified memory framework for LLM-based agents that integrates semantic, episodic, and procedural memory with a bi-level short-term and long-term store, using a multi-agent architecture for automatic memory generation and adaptive retrieval. Experiments show improved robustness and success on long multi-turn tasks.
Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM Agents
This paper introduces the Memory–Clarification Boundary (MCB) benchmark to evaluate how LLM agents decide to persist, verify, or clarify memory updates, finding that models verify changing facts more reliably than they ask for clarification.