Smarter by the Moment: Environment-Driven Dynamic Policies for Continual LLM Improvement
Summary
This paper proposes Dynamic Retrieval-based Policy Generation (DRPG), a framework that uses memory retrieval and environment feedback to dynamically generate policies for continual improvement of large language models across various tasks.
View Cached Full Text
Cached at: 09/16/26, 08:53 AM
# Environment-Driven Dynamic Policies for Continual LLM Improvement Source: [https://arxiv.org/html/2609.16800](https://arxiv.org/html/2609.16800) Ting\-Wei ChangAffiliation:Department of Computer Science and Information EngineeringAffiliation:National Taiwan University, TaiwanEmail:[changtw@nlg\.csie\.ntu\.edu\.tw](mailto:[email protected])Po\-Chun ChenAffiliation:Department of Computer Science and Information EngineeringAffiliation:National Taiwan University, TaiwanEmail:[pcchen@nlg\.csie\.ntu\.edu\.tw](mailto:[email protected])Hen\-Hsen Huang & Hsin\-Hsi ChenAffiliation:Department of Computer Science and Information EngineeringAffiliation:National Taiwan University, TaiwanAffiliation:Institute of Information Science, Academia Sinica, TaiwanAffiliation:AI Research Center \(AINTU\), National Taiwan University, TaiwanEmail:[hhhuang@iis\.sinica\.edu\.tw](mailto:[email protected]) ###### Abstract Large Language Models \(LLMs\) have achieved remarkable progress across diverse domains, but continual adaptation to evolving tasks and environments remains a key challenge\. Existing memory\-augmented approaches retrieve individual past examples as direct references, but do not explicitly synthesize actionable strategies from them, causing the same types of errors to recur\. We propose Dynamic Retrieval\-based Policy Generation \(DRPG\), a framework that integrates memory\-based retrieval with a dynamic policy generator, leveraging historical data and environment feedback to produce task\-specific policies for continual LLM improvement\. We evaluate DRPG across six benchmarks spanning text\-to\-SQL, question answering, medical diagnosis, and Python programming, using seven LLMs from both proprietary and open\-weight families\. DRPG outperforms strong baselines across most datasets and models\. Further analysis demonstrates that DRPG’s policy generation is robust to retrieval strategy, operates effectively without prior policy continuity, and can leverage smaller or cross\-family models as cost\-efficient policy generators\. We also find that the benefit of policy\-level guidance depends on task characteristics, offering practical insights into when and under what conditions this mechanism is most effective\. ## 1Introduction Large language models \(LLMs\) have demonstrated impressive capabilities across a wide range of tasks, from natural language understanding to code generation\. However, as real\-world environments and requirements evolve, there is a growing need for LLMs to continuously improve their performance and adapt to new challenges[Du et al\. \(2025\)](https://arxiv.org/html/2609.16800#bib.bib4);[Shi et al\. \(2024\)](https://arxiv.org/html/2609.16800#bib.bib29);[Zheng et al\. \(2025\)](https://arxiv.org/html/2609.16800#bib.bib43)\. Among the many adaptation strategies, the online setting is particularly relevant to real\-world scenarios, as it requires models to incrementally process new tasks and data streams, reflecting how intelligent systems interact with dynamic environments\. In the online setting, several studies[Hoi et al\. \(2021\)](https://arxiv.org/html/2609.16800#bib.bib10);[Rannen\-Triki et al\. \(2024\)](https://arxiv.org/html/2609.16800#bib.bib27)update model weights incrementally as new tasks arrive\. Although this method can be effective, it is not cost\-efficient for modern LLMs due to the considerable computational overhead involved\. As a result, memory\-based methods have emerged as promising alternatives\. For example, GrowPrompt treats the lastkkprocessed instances as in\-context learning \(ICL\) examples by directly including them in the prompt, while MemPrompt stores previous questions and answers in an external memory and retrieves the most relevant cases using a retriever[Madaan et al\. \(2022\)](https://arxiv.org/html/2609.16800#bib.bib16)\. Self\-StreamICL[Wu et al\. \(2024\)](https://arxiv.org/html/2609.16800#bib.bib36)further builds on MemPrompt by using only those previous questions that were answered correctly as few\-shot references\. Multi\-Agentic\-Memory StreamICL \(MAM\-StreamICL\)[Wu et al\. \(2024\)](https://arxiv.org/html/2609.16800#bib.bib36)takes this concept further by enabling multiple models to share a collective memory, allowing the aggregation of knowledge across agents\. While these approaches effectively utilize past experiences to guide future predictions, they are limited to providing individual examples as references\. Figure 1:DRPG framework for continual adaptation of LLMs\. At each step, the agent predicts using retrieved successful examples and a policy generated by a separate policy generator based on both correct and incorrect prior cases\. The environment returns binary feedback, which is stored in memory and used to improve subsequent retrieval and policy generation\. DRPG differs from prior methods by leveraging environment feedback to explicitly synthesize policies from past experiences\.Another direction is the Dynamic Cheatsheet[Suzgun et al\. \(2025\)](https://arxiv.org/html/2609.16800#bib.bib32), which maintains an external memory that is updated during inference, attempting to extract reusable knowledge from previous answers as references for subsequent questions\. It has shown effectiveness on several datasets that require multi\-step reasoning\. However, it does not leverage environment feedback when deciding what to store, relying solely on the model’s own internal signals\. As a result, its effectiveness is limited when applied to smaller models\. Despite these efforts, none of these approaches explicitly synthesize feedback\-informed strategies that tell the agent what to do differently\. As a result, the same types of mistakes may recur even when similar errors have been encountered before\. In this paper, we propose a framework calledDynamic Retrieval\-based Policy Generation \(DRPG\)that addresses this gap by introducing an explicit policy generator that synthesizes actionable strategies, referred to as*policies*, from past experiences and environment feedback\. These policies aim to capture recurring patterns across instances, enabling the agent to avoid repeating similar mistakes\. Our main contributions are threefold\. \(1\) We propose DRPG, a framework that augments memory\-based retrieval with a dynamically generated policy for continual LLM adaptation in streaming settings\. \(2\) Experiments across six benchmarks and seven LLMs show that DRPG outperforms strong baselines across most configurations, with analysis revealing that policy generation is most effective on tasks with systematic, recurring error patterns\. \(3\) Ablation studies show that DRPG’s policy generation is robust to retrieval strategy, operates without prior policy continuity, and allows smaller or cross\-family models to serve as policy generators\. ## 2Related Work ### 2\.1LLMs Improving in Online Settings Parameter update methods involve updating model weights with each new data point in a streaming setting\.[Hoi et al\. \(2021\)](https://arxiv.org/html/2609.16800#bib.bib10)survey online learning algorithms \(e\.g\., Perceptron, Passive–Aggressive\), highlighting their iterative nature and theoretical guarantees\. For LLMs, dynamic evaluation\([Rannen\-Triki et al\., 2024](https://arxiv.org/html/2609.16800#bib.bib27)\)adapts model parameters at test time, allowing weights to act as a temporary memory that evolves during inference\. Memory\-augmented few\-shot methods like GrowPrompt and MemPrompt\([Madaan et al\., 2022](https://arxiv.org/html/2609.16800#bib.bib16)\)leverage recent examples and feedback stored in external memory to guide predictions\. StreamBench\([Wu et al\., 2024](https://arxiv.org/html/2609.16800#bib.bib36)\)benchmarks online LLM performance, showing that methods like Self\-StreamICL \(which use only correct cases as examples\) and MAM\-StreamICL \(which shares memory among models\) improve robustness and efficiency\. Dynamic Cheatsheet\([Suzgun et al\., 2025](https://arxiv.org/html/2609.16800#bib.bib32)\)introduces a memory buffer for LLMs, dynamically updating stored content based on model confidence and utility, aiming to capture useful intermediate knowledge\. Memory management relies solely on internal model signals, not external feedback\. Agent\-Pro\([Zhang et al\., 2024b](https://arxiv.org/html/2609.16800#bib.bib42)\)is an LLM agent that improves via policy\-level self\-reflection, analyzing action sequences to adjust strategies and outperform static prompt agents in game environments\. SAGE\([Liang et al\., 2024](https://arxiv.org/html/2609.16800#bib.bib15)\)enhances LLMs through multi\-agent iterative feedback, organizing agents in a loop where outputs are refined with self\-reflection and memory optimization, leading to continual performance gains\. ### 2\.2LLMs Improving in Offline or Non\-Streaming Settings Offline fine\-tuning and reinforcement learning methods\([Parthasarathy et al\., 2024](https://arxiv.org/html/2609.16800#bib.bib22);[Rafailov et al\., 2023](https://arxiv.org/html/2609.16800#bib.bib26);[Schulman et al\., 2017](https://arxiv.org/html/2609.16800#bib.bib28)\)require significant resources and risk catastrophic forgetting\([Kalajdzievski, 2024](https://arxiv.org/html/2609.16800#bib.bib11);[Song et al\., 2025](https://arxiv.org/html/2609.16800#bib.bib31);[Li et al\., 2024](https://arxiv.org/html/2609.16800#bib.bib13);[Wang et al\., 2022](https://arxiv.org/html/2609.16800#bib.bib34)\)\. LeMa\([An et al\., 2023](https://arxiv.org/html/2609.16800#bib.bib1)\)uses GPT\-4 for error analysis and correction, transferring mistake correction skills to smaller models through teacher\-generated feedback\. Approaches like LEAP\([Zhang et al\., 2024a](https://arxiv.org/html/2609.16800#bib.bib41)\)extract general principles from contrasting correct and incorrect few\-shot outputs, which then guide future inference\. SALAM\([Wang & Li, 2023](https://arxiv.org/html/2609.16800#bib.bib33)\)introduces a cooperative “study assistant” that analyzes model errors and records them in a “mistake memory”; during testing, the assistant retrieves similar errors to provide targeted guidelines that help the agent anticipate and avoid recurring mistakes\. In a related vein, Induct\-Learn\([Chen et al\., 2024](https://arxiv.org/html/2609.16800#bib.bib2)\)induces task\-level pseudo instructions from a small number of correct demonstrations and a short task phrase, then combines these instructions with demonstrations to guide the LLM’s problem\-solving process at inference time\. Similarly,[Chen et al\. \(2025\)](https://arxiv.org/html/2609.16800#bib.bib3)propose generating diverse reasoning rationales and then inducing a unified strategy, demonstrating the effectiveness of induction\-based approaches for LLM reasoning\. These methods use data\-driven reflections to generate reusable guidance\. However, they operate in offline or static settings and do not continuously adapt from environment feedback in a streaming manner\. Reflexion\([Shinn et al\., 2023](https://arxiv.org/html/2609.16800#bib.bib30)\)enables agents to learn from feedback stored in episodic and long\-term memory, but it requires repeated interactions with the same dataset\. Self\-Refine\([Madaan et al\., 2023](https://arxiv.org/html/2609.16800#bib.bib17)\)uses iterative self\-feedback to correct mistakes\. While this method enhances output quality across diverse tasks, it does not use ground truth as external feedback, which can lead to over\-refinement of correct responses or prematurely halting refinement on incorrect ones\. In contrast to these lines of work, the novelty of DRPG lies less in its individual components than in the capability their combination enables: continually synthesizing reusable, cross\-query policies from environment feedback within a stream\. Reflection methods such as Reflexion repeatedly critique a single instance and do not accumulate strategies across queries, while offline induction methods such as LEAP and Induct\-Learn produce fixed guidance once; neither continually distills feedback\-grounded, transferable strategies as the stream evolves\. ## 3Dynamic Retrieval\-based Policy Generation \(DRPG\) Framework In this section, we introduce the DRPG framework \(Figure[1](https://arxiv.org/html/2609.16800#S1.F1)\)\. At each time step, the framework operates as follows: \(1\) the agent retrieves relevant correct examples from memory; \(2\) a separate policy generator retrieves both correct and incorrect examples and synthesizes them into a concise set of actionable guidelines \(the*policy*\); \(3\) the agent produces an answer conditioned on both the retrieved examples and the policy; \(4\) the environment provides binary feedback, which is stored in memory for future use\. Below we formalize each component\. Agent\.The agent is modeled as a functionf\(⋅\)f\(\\cdot\)\. At each time steptt, given an input questionxtx\_\{t\}, the agent extracts a set of top\-kkfew\-shot examples Rt=r\(xt,\{\(xi,y^i,fbi\)∣i<t,fbi=1\},k\)R\_\{t\}=r\\left\(x\_\{t\},\\left\\\{\(x\_\{i\},\\hat\{y\}\_\{i\},fb\_\{i\}\)\\mid i<t,\\,fb\_\{i\}=1\\right\\\},\\,k\\right\)\(1\)wherefbi∈\{0,1\}fb\_\{i\}\\in\\\{0,1\\\}is the binary correctness feedback provided by the environment for theii\-th past example,RtR\_\{t\}denotes the set of top\-kkmost relevant previous examples with positive feedback \(fbi=1fb\_\{i\}=1\), retrieved from the agent’s past interactions before timett, andr\(⋅\)r\(\\cdot\)is the retriever\. The agent then produces an output as y^t=f\(xt,Rt,Pt\)\\hat\{y\}\_\{t\}=f\(x\_\{t\},R\_\{t\},P\_\{t\}\)\(2\)wherePtP\_\{t\}is the policy at timett\. Policy Generator\.The policy generatorg\(⋅\)g\(\\cdot\)is responsible for producing the policyPP, a concise set of actionable guidelines \(e\.g\., up to five bullet points\) distilled from past correct and incorrect examples\. At each time steptt, its inputs include the retrieved examplesRt′R^\{\\prime\}\_\{t\}and it produces the policy as Pt=g\(Rt′\)P\_\{t\}=g\\left\(R^\{\\prime\}\_\{t\}\\right\)\(3\) We use a retrieverr′\(⋅\)r^\{\\prime\}\(\\cdot\)to retrieve top\-kkrelevant examples from previous experiences based on the current inputxtx\_\{t\}\. The retrieverr′\(⋅\)r^\{\\prime\}\(\\cdot\)may be the same as or different from the agent’s retrieverr\(⋅\)r\(\\cdot\)\. In this work, we adopt a contrastive setting as the default configuration, where the retrieval is performed as follows: Rt,1′=r′\(xt,\{\(xi,y^i,fbi\)∣i<t,fbi=1\},k/2\)R^\{\\prime\}\_\{t,1\}=r^\{\\prime\}\\left\(x\_\{t\},\\left\\\{\(x\_\{i\},\\hat\{y\}\_\{i\},fb\_\{i\}\)\\mid i<t,\\,fb\_\{i\}=1\\right\\\},k/2\\right\)\(4\)Rt,0′=r′\(xt,\{\(xj,y^j,fbj\)∣j<t,fbj=0\},k/2\)R^\{\\prime\}\_\{t,0\}=r^\{\\prime\}\\left\(x\_\{t\},\\left\\\{\(x\_\{j\},\\hat\{y\}\_\{j\},fb\_\{j\}\)\\mid j<t,\\,fb\_\{j\}=0\\right\\\},k/2\\right\)\(5\)We then form the final retrieved set by combining these two subsets: Rt′=Rt,1′∪Rt,0′R^\{\\prime\}\_\{t\}=R^\{\\prime\}\_\{t,1\}\\cup R^\{\\prime\}\_\{t,0\}\(6\) Environment interaction\.The environment, denoted ash\(⋅\)h\(\\cdot\), produces the binary correctness feedback introduced above: after the agent outputsy^t\\hat\{y\}\_\{t\}, the environment returnsfbt=h\(xt,y^t\)∈\{0,1\}fb\_\{t\}=h\(x\_\{t\},\\hat\{y\}\_\{t\}\)\\in\\\{0,1\\\}, indicating the correctness of the answer\. Past Experiences\.During the streaming process, we maintain an updatable memory database of past experiences\. At time steptt, this database contains tuples of the form\(xi,y^i,fbi\)\(x\_\{i\},\\hat\{y\}\_\{i\},fb\_\{i\}\)fori=1,…,t−1i=1,\\ldots,t\-1\. The full procedure is summarized in Algorithm[1](https://arxiv.org/html/2609.16800#alg1)\. Algorithm 1Framework for DRPG1:Initialize agent f\(⋅\)f\(\\cdot\), retrievers r\(⋅\)r\(\\cdot\), r′\(⋅\)r^\{\\prime\}\(\\cdot\), and policy generator g\(⋅\)g\(\\cdot\) 2:for t=1t=1to TTdo 3:Receive instance xtx\_\{t\}from the data stream; 4:Retrieve Rt=r\(xt,\{\(xi,y^i,fbi\)∣i<t,fbi=1\},k\)R\_\{t\}=r\\left\(x\_\{t\},\\left\\\{\(x\_\{i\},\\hat\{y\}\_\{i\},fb\_\{i\}\)\\mid i<t,\\,fb\_\{i\}=1\\right\\\},k\\right\); 5:// Contrastive retrieval: 6: Rt,1′=r′\(xt,\{\(xi,y^i,fbi\)∣i<t,fbi=1\},k2\)R^\{\\prime\}\_\{t,1\}=r^\{\\prime\}\\left\(x\_\{t\},\\left\\\{\(x\_\{i\},\\hat\{y\}\_\{i\},fb\_\{i\}\)\\mid i<t,\\,fb\_\{i\}=1\\right\\\},\\frac\{k\}\{2\}\\right\); 7: Rt,0′=r′\(xt,\{\(xj,y^j,fbj\)∣j<t,fbj=0\},k2\)R^\{\\prime\}\_\{t,0\}=r^\{\\prime\}\\left\(x\_\{t\},\\left\\\{\(x\_\{j\},\\hat\{y\}\_\{j\},fb\_\{j\}\)\\mid j<t,\\,fb\_\{j\}=0\\right\\\},\\frac\{k\}\{2\}\\right\); 8: Rt′=Rt,1′∪Rt,0′R^\{\\prime\}\_\{t\}=R^\{\\prime\}\_\{t,1\}\\cup R^\{\\prime\}\_\{t,0\}; 9:The policy generator generates policy Pt=g\(Rt′\)P\_\{t\}=g\\left\(R^\{\\prime\}\_\{t\}\\right\); 10:The agent predicts y^t=f\(xt,Rt,Pt\)\\hat\{y\}\_\{t\}=f\(x\_\{t\},R\_\{t\},P\_\{t\}\); 11:Receive feedback fbt=h\(xt,y^t\)fb\_\{t\}=h\(x\_\{t\},\\hat\{y\}\_\{t\}\), fbt∈\{0,1\}fb\_\{t\}\\in\\\{0,1\\\}; 12:Store triplet \(xt,y^t,fbt\)\(x\_\{t\},\\hat\{y\}\_\{t\},fb\_\{t\}\)in memory; 13:endfor ### 3\.1Prompt Design We follow the prompt design of StreamBench[Wu et al\. \(2024\)](https://arxiv.org/html/2609.16800#bib.bib36)\. To minimize the impact of prompt engineering, we use the same prompt structure across different methods\. This consistent structure is applied to both the agent and the policy generator\. The main components of our prompt structure are as follows, with certain parts added or omitted depending on the specific method: \(1\)Role assignment, which specifies the role of the LLM for the given task; \(2\)Reference materials, including task\-related supporting information such as answer choices, few\-shot examples from memory, or policies; \(3\)Question, the current input \(xtx\_\{t\}\) the agent is required to answer; and \(4\)Additional instructions, covering formatting requirements, rules, or restrictions specific to the task\. For the policy generator, we use a prompt template that includes role assignment, reference materials \(retrieved correct and incorrect examples\), and additional instructions\. The policy generator analyzes past cases and produces a policy accordingly, focusing on actionable bullet points relevant to the task\. The detailed prompts for each dataset can be found in Appendix[I](https://arxiv.org/html/2609.16800#A9)\. We adapt the prompt to dataset\-specific content while retaining the same prompt structure\. ## 4Experiments ### 4\.1Datasets Following StreamBench[Wu et al\. \(2024\)](https://arxiv.org/html/2609.16800#bib.bib36), we adopt most of the datasets used in their benchmark, excluding ToolBench since it uses an LLM as a judge, which heavily relies on the LLM’s own judgment ability\. Our evaluation spans four task categories\. Text\-to\-SQL\.Spider[Yu et al\. \(2018\)](https://arxiv.org/html/2609.16800#bib.bib39), CoSQL[Yu et al\. \(2019\)](https://arxiv.org/html/2609.16800#bib.bib40), and BIRD[Li et al\. \(2023\)](https://arxiv.org/html/2609.16800#bib.bib14)are all large\-scale, cross\-domain datasets with complex SQL queries, designed to test generalization to unseen database schemas\. These three benchmarks vary in difficulty: Spider is the easiest, while BIRD is the most challenging\. Question Answering\.HotpotQA[Yang et al\. \(2018\)](https://arxiv.org/html/2609.16800#bib.bib38)is a multi\-hop question answering \(QA\) dataset that requires the LLM to locate and reason over supporting passages\. The dataset uses a distractor setting, in which the provided passages contain both useful and irrelevant information\. Medical Diagnosis\.DDXPlus is a synthetic dataset that includes patient profiles and full differential diagnoses, simulating scenarios in which a doctor diagnoses a patient’s condition\([Fansi Tchango et al\., 2022](https://arxiv.org/html/2609.16800#bib.bib6)\)\. The agent needs to select the most appropriate diagnosis from 49 candidate options, based on the patient’s background\. Python Programming\.DS\-1000[Lai et al\. \(2023\)](https://arxiv.org/html/2609.16800#bib.bib12)provides 1,000 real\-world Python programming tasks across seven libraries \(e\.g\., NumPy, pandas\), with perturbations to avoid memorization and support reliable execution\-based evaluation\. ### 4\.2Evaluation Metrics We adopt the standard evaluation metric for each dataset\. For the Text\-to\-SQL datasets, we use the common execution accuracy metric, which compares the execution results of the generated SQL queries with the ground truth\. For HotpotQA, we follow the original paper and use exact match as the primary metric, comparing the agent’s answer to the ground truth\. DDXPlus is treated as a 49\-choice multiple\-choice task, so we use accuracy as the evaluation metric\. For Python programming tasks, we use the standard pass@1 metric, in which the agent generates only one answer and the result is judged by running several predefined test cases\. ### 4\.3Baselines We compare DRPG against the following approaches, including both non\-streaming and streaming methods\. Zero\-shotevaluates the base ability of each LLM without any additional adaptation\. Self\-Refine\([Madaan et al\., 2023](https://arxiv.org/html/2609.16800#bib.bib17)\)allows the agent to generate an initial answer and then refine it using self\-generated feedback, considering only the current question and answer without external memory\. Self\-StreamICL\([Wu et al\., 2024](https://arxiv.org/html/2609.16800#bib.bib36)\)stores each question, answer, and environment feedback in an external memoryℳ\\mathcal\{M\}, then retrieves similar and correctly answered past cases as few\-shot examples for new questions\. It outperforms both GrowPrompt and MemPrompt, and serves as the main baseline in our study\. For DRPG, we use the same agent configuration as Self\-StreamICL, but additionally include a dynamically generated policy in the agent’s prompt\. The policy generator retrieveskkrelevant past examples for each update\. In our default configuration, inspired by prior work[Gao & Das \(2024\)](https://arxiv.org/html/2609.16800#bib.bib7);[An et al\. \(2023\)](https://arxiv.org/html/2609.16800#bib.bib1);[Zhang et al\. \(2024a\)](https://arxiv.org/html/2609.16800#bib.bib41), we adopt a contrastive setting that retrieves an equal number of correct and incorrect cases; however, as shown in our ablation study \(Appendix[D](https://arxiv.org/html/2609.16800#A4)\), policy generation is robust across different retrieval strategies\. ### 4\.4Models To assess the generalizability of DRPG, we evaluate across three model families spanning both proprietary API\-based and open\-weight LLMs: Gemini \(gemini\-2\.0\-flash,gemini\-2\.0\-flash\-lite\)[Gemini Team \(2023\)](https://arxiv.org/html/2609.16800#bib.bib8), Llama \(llama\-3\.3\-70b[Dubey et al\. \(2024\)](https://arxiv.org/html/2609.16800#bib.bib5),llama\-4\-maverick,llama\-4\-scout[Meta AI Team \(2025\)](https://arxiv.org/html/2609.16800#bib.bib18)\), and Mistral \(mistral\-medium[Mistral AI Team \(2025a\)](https://arxiv.org/html/2609.16800#bib.bib20),mistral\-small[Mistral AI Team \(2025b\)](https://arxiv.org/html/2609.16800#bib.bib21)\)\. This selection covers a range of model sizes and capability levels, enabling us to examine whether the benefits of policy generation are consistent across different architectures and scales\. In addition, Appendix[H](https://arxiv.org/html/2609.16800#A8)reports results on two additional open\-weight families, Qwen and Gemma\. While this spread covers multiple families and scales, it does not by itself establish generality across all modern LLM ecosystems \(e\.g\., OpenAI or Claude models\); extending the evaluation to further families is left to future work\. ### 4\.5Experimental Setup In this study, we fix several parameters to ensure experimental consistency, following many settings from StreamBench\. First, the retriever uses the BAAI/bge\-base\-en\-v1\.5[Xiao et al\. \(2023\)](https://arxiv.org/html/2609.16800#bib.bib37)embedding model\. For the number of few\-shot examples, we follow the agent settings in StreamBench[Wu et al\. \(2024\)](https://arxiv.org/html/2609.16800#bib.bib36): the retriever selectskkpast instances as ICL few\-shot examples\. Specifically,k=16k=16is used for Spider, CoSQL, BIRD, and DDXPlus, whilek=4k=4is used for DS\-1000 and HotpotQA to avoid exceeding the context window\. The policy generator uses the samekkvalue as the agent for retrieval\. For dataset ordering, we use a fixed random seed \(42\) to ensure consistent shuffling across all experiments\. Additionally, we set the temperature to 0 to maintain consistency and the maximum output tokens to 1,024 for all experiments\. We fully reuse StreamBench’s retrieval pipeline without modification, shared by DRPG and all baselines; its precise configuration is given in Appendix[C](https://arxiv.org/html/2609.16800#A3)\. ## 5Results and Analysis This section presents a comprehensive analysis of the DRPG framework\. We first establish its overall performance against strong baselines \(Section[5\.1](https://arxiv.org/html/2609.16800#S5.SS1)\), then examine the relationship between task characteristics and the effectiveness of policy\-level guidance \(Section[5\.2](https://arxiv.org/html/2609.16800#S5.SS2)\)\. Finally, we conduct a series of ablation studies examining the robustness of the policy generation mechanism \(Section[5\.3](https://arxiv.org/html/2609.16800#S5.SS3)\)\. ### 5\.1Performance of DRPG Framework Table 1:Performance comparison of DRPG \(Ours\) and baseline methods across multiple LLMs and benchmark tasks\. The best\-performing method for each model\-task pair is highlighted inbold\. The bottom rows show pairwise win counts of DRPG against each baseline across all models\.Table[1](https://arxiv.org/html/2609.16800#S5.T1)presents the main comparison between DRPG and existing baselines across six benchmarks and seven LLMs\. DRPG outperforms Self\-StreamICL with a pairwise win record of 29 out of 42 model–dataset combinations\. The margins are even more pronounced against Self\-Refine \(35/42\) and Zero\-shot \(35/42\), highlighting the overall advantage of DRPG across a wide range of tasks and model families\. The most consistent gains are observed on text\-to\-SQL benchmarks \(Spider, CoSQL, and BIRD\), where DRPG wins against Self\-StreamICL in 18 out of 21 model configurations\. DRPG also outperforms Self\-StreamICL in 5 out of 7 configurations on the multi\-hop reasoning benchmark HotpotQA\. A notable exception isllama\-4\-maverick, which shows a significant drop on HotpotQA; we hypothesize that the generated policy occasionally overrides the agent’s correct reasoning on multi\-hop questions\. On DDXPlus and DS\-1000, DRPG wins against Self\-StreamICL in only 3 out of 7 configurations for each dataset\. We observe that on these tasks, the generated policies tend to capture overly narrow, instance\-specific patterns rather than broadly applicable strategies, limiting their coverage\. Treating each model–dataset configuration as a pair, DRPG significantly outperforms Self\-StreamICL overall \(one\-sided Wilcoxon signed\-rank test,p=0\.005p=0\.005\), while on DDXPlus and DS\-1000 the differences are not significant in either direction \(two\-sidedp=0\.94p=0\.94and0\.730\.73; Appendix[G](https://arxiv.org/html/2609.16800#A7)reports all tests and visualizes the per\-configuration gains\)\. DRPG’s benefit is therefore not uniform across task types: it concentrates on tasks whose errors share recurring, generalizable structure, while on tasks requiring instance\-specific knowledge DRPG performs on par with, rather than above, Self\-StreamICL\. We examine this task\-dependent behavior in detail below\. ### 5\.2Task\-Dependent Effectiveness of Policy Generation We qualitatively examine the policies generated by DRPG to understand*why*policy generation is more effective on some tasks than others\. The key distinction lies in whether a task exhibits error patterns that can be captured by a small number of high\-level rules generalizing across instances\. Text\-to\-SQL and QA: policies capture generalizable rules\.On Spider, the generated policy forllama\-4\-scout\(Figure[2](https://arxiv.org/html/2609.16800#S5.F2)\) captures actionable, cross\-instance guidelines such as “Validate Schema and Relationships” and “Avoid Ambiguous or Redundant Results\.” Each rule applies broadly across different SQL queries, enabling an 8\.5 percentage point improvement over Self\-StreamICL\. This is possible because text\-to\-SQL errors are structurally regular: mistakes like incorrect joins, missing aggregation, or schema misuse recur regardless of the specific query\. A similar pattern holds for HotpotQA \(5/7 wins\), where multi\-hop reasoning involves recurring challenges such as identifying the correct supporting passages and avoiding distractor information\. \- Directly Address the Question: Focus on directly answering the question asked, avoiding unnecessary joins or conditions that do not contribute to answering the query\.\- Validate Schema and Relationships: Verify the schema and relationships between tables to ensure correct joins, subqueries, and conditions are used\.\- Precise Use of SQL Constructs: Choose SQL constructs that accurately handle the query requirements, such as using aggregation or grouping when necessary, and handle edge cases like empty result sets or division by zero\.\- Avoid Ambiguous or Redundant Results: Ensure that queries are clear and unambiguous, directly answering the question without providing redundant or unnecessary information\.\- Handle Multiple Results Correctly: Ensure correct usage of subqueries, aggregations, and grouping to handle complex queries with multiple results, considering all required columns and potential duplicates\. Figure 2:Example of the generated policy at time step 100 onllama\-4\-scoutfor the Spider dataset\. Each rule is broadly applicable across different SQL queries\.Medical Diagnosis and Python Programming: policies become too narrow\.In contrast, the policies generated for DDXPlus \(Figure[3](https://arxiv.org/html/2609.16800#S5.F3)\) tend to focus on specific diseases, such as “Consider Cardiovascular Diseases for Chest Pain” or “Evaluate for Anaphylaxis in Allergic Reactions\.” With 49 candidate diagnoses and diverse symptom profiles, the retrieved examples often span unrelated disease categories, making it difficult to extract a coherent diagnostic strategy\. Unlike text\-to\-SQL, where structural patterns \(e\.g\., correct use of JOINs\) generalize across queries, medical diagnosis lacks a universal procedure that applies across diverse conditions; the policy generator therefore defaults to enumerating disease\-specific heuristics, each covering only a narrow subset of future cases\. Similarly, DS\-1000 spans seven Python libraries with distinct API conventions; the precise, library\-specific knowledge required \(e\.g\., the correct parameters forpandas\.pivot\_tableornumpy\.reshape\) cannot be effectively compressed into five actionable bullet points\. For such tasks, instance\-level retrieval, which provides concrete, directly relevant examples, remains the more effective mechanism\. \- Consider Cardiovascular Diseases for Chest Pain: When patients report chest pain, especially if it’s described as tedious, heavy, or sharp, and radiates to areas like the biceps, shoulders, or under the jaw, consider cardiovascular diseases such as NSTEMI/STEMI, especially in patients with risk factors like diabetes, high cholesterol, smoking, or family history of cardiovascular diseases\.\- Evaluate for Anaphylaxis in Allergic Reactions: In cases of known severe food allergies, recent consumption of allergenic substances, symptoms like swelling, redness, itching, and widespread skin lesions, consider anaphylaxis, especially if accompanied by respiratory distress or cardiovascular symptoms\.\- Assess for Infectious Diseases Based on Exposure and Symptoms: For patients with recent travel history to high\-risk areas \(e\.g\., West Africa for Ebola\), contact with infected individuals, symptoms like fever, shortness of breath, and diffuse muscle pain, consider infectious diseases such as Ebola\. Figure 3:Partial example of the generated policy at time step 100 onllama\-4\-scoutfor DDXPlus\. Unlike Spider, rules focus on specific diseases rather than generalizable strategies\.These observations suggest that policy generation is most beneficial for tasks with*compositional structure*and*recurring error types*, and less so for tasks relying on instance\-specific knowledge or broad categorical recall\. ### 5\.3Ablation Studies We examine the sensitivity of DRPG to key design choices, focusing on the model used for the policy generator\. In DRPG, the agent and the policy generator can be instantiated using different language models\. Table[2](https://arxiv.org/html/2609.16800#S5.T2)presents the results under various model pairings, including both within\-series and cross\-series configurations\. Overall, the agent’s own capability remains the dominant factor determining final performance, while the policy generator provides auxiliary guidance that leads to performance improvements in most cases\. Notably, even when using a smaller LLM as the policy generator, the induced strategies still benefit a stronger agent\. An interesting case ismaverick: when it generates its own policy, performance drops sharply on HotpotQA \(50\.60\), but whenscoutormistral\-smallserves as the policy generator, this drop does not occur \(63\.73 and 63\.07, respectively\)\. This suggests that cross\-model policy generation can sometimes avoid failure modes present in same\-model generation, though we note that forscoutandmistral\-smallas agents, same\-model and cross\-model configurations perform comparably\. These results suggest that lightweight models can serve as cost\-efficient policy generators without sacrificing downstream performance, and that practitioners may benefit from experimenting with cross\-model configurations\. Table 2:Performance when using different models for the main agent and the policy generator\. For Self\-StreamICL, there is no separate policy generator\. We abbreviate Llama\-4\-Maverick asmaverick, Llama\-4\-Scout asscout, and Mistral\-Small\-2503 asmistral\-small\. The best\-performing combination is highlighted inbold\.We further verify two additional design choices\. First, replacing the contrastive retrieval strategy with correct\-only or wrong\-only retrieval yields comparable performance, confirming that the policy generator can reliably distill strategies regardless of the composition of its input \(Appendix[D](https://arxiv.org/html/2609.16800#A4)\)\. Second, providing the previous policyPt−1P\_\{t\-1\}as input to the policy generator does not consistently improve results, indicating that the mechanism operates effectively in a stateless manner without requiring policy continuity \(Appendix[E](https://arxiv.org/html/2609.16800#A5)\)\. Together with the cross\-model result above, these findings demonstrate that DRPG’s policy generation is a modular component with minimal sensitivity to design choices\. In addition, a no\-feedback ablation in Appendix[F](https://arxiv.org/html/2609.16800#A6)isolates the contribution of environment feedback: generating policies from the retrieved examples without correctness labels consistently underperforms DRPG across all six benchmarks\. ### 5\.4Cost Analysis Relative to Self\-StreamICL, DRPG adds one retrieval \(fetching the successful and failed cases for policy generation\) and one LLM call \(the policy generator\) per query\. Table[3](https://arxiv.org/html/2609.16800#S5.T3)compares the per\-query computational cost of the three methods for the representative configurationmistral\-medium×\\timesSpider \(N=2,147\); token counts include all of a method’s LLM calls\. Because most runs were executed on free\-tier APIs, the latency should be read only as a relative comparison; its ordering is consistent with the call structure: DRPG \(≈\\approx2 calls\)\>\>Self\-StreamICL\>\>Zero\-shot\. Per\-query token counts for all 42 \(model×\\timesdataset\) configurations are provided in Appendix[B](https://arxiv.org/html/2609.16800#A2)\. Table 3:Per\-query cost comparison for the representative configurationmistral\-medium×\\timesSpider \(N=2,147\)\. Token counts include all LLM calls of each method; latency is measured on free\-tier APIs and should be read as a relative comparison\. ## 6Conclusion We present DRPG, a framework that augments memory\-based retrieval with a dynamic policy generator for continual LLM improvement in online settings\. Evaluation across six benchmarks and seven LLMs shows that DRPG outperforms strong baselines on the majority of model–dataset configurations, significantly so overall against Self\-StreamICL, with the largest gains on structured prediction tasks, while performing on par with Self\-StreamICL on tasks that require instance\-specific knowledge\. Ablation studies further demonstrate that policy generation is robust to retrieval strategy, operates statelessly, and allows the use of a separate, potentially smaller model as the policy generator\. Our analysis also reveals that the benefit of policy\-level guidance varies with task characteristics: it is most effective for tasks with systematic, recurring error patterns capturable by high\-level rules, and less so for tasks requiring instance\-specific knowledge\. We believe this finding offers practical guidance for practitioners and future work on when to deploy policy\-level adaptation mechanisms in streaming settings\. ## Acknowledgments This work was supported by the National Science and Technology Council, Taiwan, under grants NSTC 114\-2221\-E\-002\-070\-MY3 and NSTC 115\-2634\-F\-002\-012, as well as by financial support from the Featured Area Research Center Program within the framework of the Higher Education Sprout Project by the Ministry of Education \(115L900901\)\. ## References - An et al\. \(2023\)Shengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng, Jian\-Guang Lou, and Weizhu Chen\.Learning from mistakes makes llm better reasoner\.*arXiv preprint arXiv:2310\.20689*, 2023\. - Chen et al\. \(2024\)Po\-Chun Chen, Sheng\-Lun Wei, Hen\-Hsen Huang, and Hsin\-Hsi Chen\.Induct\-learn: Short phrase prompting with instruction induction\.In Yaser Al\-Onaizan, Mohit Bansal, and Yun\-Nung Chen \(eds\.\),*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, pp\. 5204–5231, Miami, Florida, USA, November 2024\. Association for Computational Linguistics\.doi:10\.18653/v1/2024\.emnlp\-main\.297\.URL[https://aclanthology\.org/2024\.emnlp\-main\.297/](https://aclanthology.org/2024.emnlp-main.297/)\. - Chen et al\. \(2025\)Po\-Chun Chen, Hen\-Hsen Huang, and Hsin\-Hsi Chen\.Diverge to induce prompting: Multi\-rationale induction for zero\-shot reasoning\.In Kentaro Inui, Sakriani Sakti, Haofen Wang, Derek F\. Wong, Pushpak Bhattacharyya, Biplab Banerjee, Asif Ekbal, Tanmoy Chakraborty, and Dhirendra Pratap Singh \(eds\.\),*Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics*, pp\. 102–115, Mumbai, India, December 2025\. The Asian Federation of Natural Language Processing and The Association for Computational Linguistics\.ISBN 979\-8\-89176\-303\-6\.doi:10\.18653/v1/2025\.findings\-ijcnlp\.6\.URL[https://aclanthology\.org/2025\.findings\-ijcnlp\.6/](https://aclanthology.org/2025.findings-ijcnlp.6/)\. - Du et al\. \(2025\)Shangheng Du, Jiabao Zhao, Jinxin Shi, Zhentao Xie, Xin Jiang, Yanhong Bai, and Liang He\.A survey on the optimization of large language model\-based agents\.*arXiv preprint arXiv:2503\.12434*, 2025\. - Dubey et al\. \(2024\)Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al\-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al\.The llama 3 herd of models\.*arXiv preprint arXiv:2407\.21783*, 2024\. - Fansi Tchango et al\. \(2022\)Arsene Fansi Tchango, Rishab Goel, Zhi Wen, Julien Martel, and Joumana Ghosn\.Ddxplus: A new dataset for automatic medical diagnosis\.*Advances in neural information processing systems*, 35:31306–31318, 2022\. - Gao & Das \(2024\)Xiang Gao and Kamalika Das\.Customizing language model responses with contrastive in\-context learning\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 38, pp\. 18039–18046, 2024\. - Gemini Team \(2023\)Google Gemini Team\.Gemini: A family of highly capable multimodal models\.*ArXiv*, abs/2312\.11805, 2023\.URL[https://arxiv\.org/pdf/2312\.11805](https://arxiv.org/pdf/2312.11805)\. - Gemma Team \(2026\)Gemma Team\.Gemma 4 technical report, 2026\.URL[https://arxiv\.org/abs/2607\.02770](https://arxiv.org/abs/2607.02770)\.arXiv:2607\.02770\. - Hoi et al\. \(2021\)Steven CH Hoi, Doyen Sahoo, Jing Lu, and Peilin Zhao\.Online learning: A comprehensive survey\.*Neurocomputing*, 459:249–289, 2021\. - Kalajdzievski \(2024\)Damjan Kalajdzievski\.Scaling laws for forgetting when fine\-tuning large language models\.*arXiv preprint arXiv:2401\.05605*, 2024\. - Lai et al\. \(2023\)Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Wen\-tau Yih, Daniel Fried, Sida Wang, and Tao Yu\.Ds\-1000: A natural and reliable benchmark for data science code generation\.In*International Conference on Machine Learning*, pp\. 18319–18345\. PMLR, 2023\. - Li et al\. \(2024\)Hongyu Li, Liang Ding, Meng Fang, and Dacheng Tao\.Revisiting catastrophic forgetting in large language model tuning\.In Yaser Al\-Onaizan, Mohit Bansal, and Yun\-Nung Chen \(eds\.\),*Findings of the Association for Computational Linguistics: EMNLP 2024*, pp\. 4297–4308, Miami, Florida, USA, November 2024\. Association for Computational Linguistics\.doi:10\.18653/v1/2024\.findings\-emnlp\.249\.URL[https://aclanthology\.org/2024\.findings\-emnlp\.249/](https://aclanthology.org/2024.findings-emnlp.249/)\. - Li et al\. \(2023\)Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua Li, Bowen Li, Bailin Wang, Bowen Qin, Ruiying Geng, Nan Huo, et al\.Can llm already serve as a database interface? a big bench for large\-scale database grounded text\-to\-sqls\.*Advances in Neural Information Processing Systems*, 36:42330–42357, 2023\. - Liang et al\. \(2024\)Xuechen Liang, Yangfan He, Yinghui Xia, Xinyuan Song, Jianhui Wang, Meiling Tao, Li Sun, Xinhang Yuan, Jiayi Su, Keqin Li, et al\.Self\-evolving agents with reflective and memory\-augmented abilities\.*arXiv preprint arXiv:2409\.00872*, 2024\. - Madaan et al\. \(2022\)Aman Madaan, Niket Tandon, Peter Clark, and Yiming Yang\.Memory\-assisted prompt editing to improve GPT\-3 after deployment\.In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang \(eds\.\),*Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing*, pp\. 2833–2861, Abu Dhabi, United Arab Emirates, December 2022\. Association for Computational Linguistics\.doi:10\.18653/v1/2022\.emnlp\-main\.183\.URL[https://aclanthology\.org/2022\.emnlp\-main\.183/](https://aclanthology.org/2022.emnlp-main.183/)\. - Madaan et al\. \(2023\)Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al\.Self\-refine: Iterative refinement with self\-feedback\.*Advances in Neural Information Processing Systems*, 36:46534–46594, 2023\. - Meta AI Team \(2025\)Meta AI Team\.The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, 2025\.URL[https://ai\.meta\.com/blog/llama\-4\-multimodal\-intelligence](https://ai.meta.com/blog/llama-4-multimodal-intelligence)\. - Min et al\. \(2022\)Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer\.Rethinking the role of demonstrations: What makes in\-context learning work?*arXiv preprint arXiv:2202\.12837*, 2022\. - Mistral AI Team \(2025a\)Mistral AI Team\.Medium is the new large\., 2025a\.URL[https://mistral\.ai/news/mistral\-medium\-3](https://mistral.ai/news/mistral-medium-3)\.Section: news\. - Mistral AI Team \(2025b\)Mistral AI Team\.Mistral small 3\.1, 2025b\.URL[https://mistral\.ai/news/mistral\-small\-3\-1](https://mistral.ai/news/mistral-small-3-1)\.Section: news\. - Parthasarathy et al\. \(2024\)Venkatesh Balavadhani Parthasarathy, Ahtsham Zafar, Aafaq Khan, and Arsalan Shahid\.The ultimate guide to fine\-tuning llms from basics to breakthroughs: An exhaustive review of technologies, research, best practices, applied research challenges and opportunities\.*arXiv preprint arXiv:2408\.13296*, 2024\. - Patel et al\. \(2024\)Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini\.Splitwise: Efficient generative llm inference using phase splitting\.In*2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture \(ISCA\)*\. IEEE, 2024\. - Pope et al\. \(2023\)Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Anselm Levskaya, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean\.Efficiently scaling transformer inference\.In*Proceedings of Machine Learning and Systems*, volume 5, 2023\. - Qwen Team \(2026\)Qwen Team\.Qwen3\.5: Towards native multimodal agents, February 2026\.URL[https://qwen\.ai/blog?id=qwen3\.5](https://qwen.ai/blog?id=qwen3.5)\.Accessed: 2026\-08\-05\. - Rafailov et al\. \(2023\)Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn\.Direct preference optimization: Your language model is secretly a reward model\.*Advances in Neural Information Processing Systems*, 36:53728–53741, 2023\. - Rannen\-Triki et al\. \(2024\)Amal Rannen\-Triki, Jorg Bornschein, Razvan Pascanu, Marcus Hutter, Andras György, Alexandre Galashov, Yee Whye Teh, and Michalis K Titsias\.Revisiting dynamic evaluation: Online adaptation for large language models\.*arXiv preprint arXiv:2403\.01518*, 2024\. - Schulman et al\. \(2017\)John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov\.Proximal policy optimization algorithms\.*arXiv preprint arXiv:1707\.06347*, 2017\. - Shi et al\. \(2024\)Haizhou Shi, Zihao Xu, Hengyi Wang, Weiyi Qin, Wenyuan Wang, Yibin Wang, Zifeng Wang, Sayna Ebrahimi, and Hao Wang\.Continual learning of large language models: A comprehensive survey\.*ACM Computing Surveys*, 2024\. - Shinn et al\. \(2023\)Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao\.Reflexion: Language agents with verbal reinforcement learning\.*Advances in Neural Information Processing Systems*, 36:8634–8652, 2023\. - Song et al\. \(2025\)Shezheng Song, Hao Xu, Jun Ma, Shasha Li, Long Peng, Qian Wan, Xiaodong Liu, and Jie Yu\.How to alleviate catastrophic forgetting in llms finetuning? hierarchical layer\-wise and element\-wise regularization\.*arXiv preprint arXiv:2501\.13669*, 2025\. - Suzgun et al\. \(2025\)Mirac Suzgun, Mert Yuksekgonul, Federico Bianchi, Dan Jurafsky, and James Zou\.Dynamic cheatsheet: Test\-time learning with adaptive memory\.*arXiv preprint arXiv:2504\.07952*, 2025\. - Wang & Li \(2023\)Danqing Wang and Lei Li\.Learning from mistakes via cooperative study assistant for large language models\.In Houda Bouamor, Juan Pino, and Kalika Bali \(eds\.\),*Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing*, pp\. 10667–10685, Singapore, December 2023\. Association for Computational Linguistics\.doi:10\.18653/v1/2023\.emnlp\-main\.659\.URL[https://aclanthology\.org/2023\.emnlp\-main\.659/](https://aclanthology.org/2023.emnlp-main.659/)\. - Wang et al\. \(2022\)Yihan Wang, Si Si, Daliang Li, Michal Lukasik, Felix Yu, Cho\-Jui Hsieh, Inderjit S Dhillon, and Sanjiv Kumar\.Two\-stage llm fine\-tuning with less specialization and more generalization\.*arXiv preprint arXiv:2211\.00635*, 2022\. - Wei et al\. \(2023\)Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, et al\.Larger language models do in\-context learning differently\.*arXiv preprint arXiv:2303\.03846*, 2023\. - Wu et al\. \(2024\)Cheng\-Kuang Wu, Zhi Rui Tam, Chieh\-Yen Lin, Yun\-Nung Vivian Chen, and Hung\-yi Lee\.Streambench: Towards benchmarking continuous improvement of language agents\.*Advances in Neural Information Processing Systems*, 37:107039–107063, 2024\. - Xiao et al\. \(2023\)Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff\.C\-pack: Packed resources for general chinese embeddings, 2023\. - Yang et al\. \(2018\)Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning\.Hotpotqa: A dataset for diverse, explainable multi\-hop question answering\.*arXiv preprint arXiv:1809\.09600*, 2018\. - Yu et al\. \(2018\)Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, et al\.Spider: A large\-scale human\-labeled dataset for complex and cross\-domain semantic parsing and text\-to\-sql task\.*arXiv preprint arXiv:1809\.08887*, 2018\. - Yu et al\. \(2019\)Tao Yu, Rui Zhang, He Yang Er, Suyi Li, Eric Xue, Bo Pang, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Zihan Li, et al\.Cosql: A conversational text\-to\-sql challenge towards cross\-domain natural language interfaces to databases\.*arXiv preprint arXiv:1909\.05378*, 2019\. - Zhang et al\. \(2024a\)Tianjun Zhang, Aman Madaan, Luyu Gao, Steven Zheng, Swaroop Mishra, Yiming Yang, Niket Tandon, and Uri Alon\.In\-context principle learning from mistakes\.*arXiv preprint arXiv:2402\.05403*, 2024a\. - Zhang et al\. \(2024b\)Wenqi Zhang, Ke Tang, Hai Wu, Mengna Wang, Yongliang Shen, Guiyang Hou, Zeqi Tan, Peng Li, Yueting Zhuang, and Weiming Lu\.Agent\-pro: Learning to evolve via policy\-level reflection and optimization\.*arXiv preprint arXiv:2402\.17574*, 2024b\. - Zheng et al\. \(2025\)Junhao Zheng, Shengjie Qiu, Chengming Shi, and Qianli Ma\.Towards lifelong learning of large language models: A survey\.*ACM Computing Surveys*, 57\(8\):1–35, 2025\. ## Limitations Our evaluation adopts a controlled online setting with clean binary correctness feedback at every step, identical to the StreamBench protocol used by Self\-StreamICL and the other baselines, so that all methods are compared under the same conditions and the observed behavior can be attributed to policy generation itself\. We note that per\-step binary correctness feedback is an idealization: even in text\-to\-SQL, execution can reveal syntactic errors but does not by itself confirm that the produced query is semantically correct, and fully reliable correctness signals generally require additional verification such as user confirmation or downstream checks\. In real deployments we expect such feedback to be sparse rather than available at every step; a natural extension of DRPG is to update memory and generate policies only from the subset of interactions that do receive feedback, and how its advantage over feedback\-free methods depends on feedback density is an open question\. Our conclusions hold for this controlled setting and should not be extrapolated to real\-deployment conditions such as non\-stationary streams or noisy, delayed, partial, non\-binary, or sparse feedback; studying robustness under such degraded\-feedback environments is left to future work\. In addition, while we use a consistent prompt structure across all methods, model families, and datasets, we have not conducted a systematic sensitivity study over prompt paraphrases, and we leave establishing prompt robustness to future work\. Due to budget limitations, we did not evaluate DRPG with reasoning LLMs such as OpenAI GPT\-o3 or DeepSeek R1, which we leave to future work\. Although our benchmarks span four diverse task categories, the effectiveness of DRPG may vary on other tasks\. All experiments use a single fixed data ordering \(seed=42\) with temperature set to 0; sensitivity to data order in streaming settings remains unexplored, and we do not report variance across multiple runs\. We also observe that the generated policy can occasionally hurt performance, as seen withllama\-4\-maverickon HotpotQA; understanding when policy\-level guidance may conflict with the agent’s reasoning is an important direction for future investigation\. We do not include Dynamic Cheatsheet[Suzgun et al\. \(2025\)](https://arxiv.org/html/2609.16800#bib.bib32)as a baseline, since its implementation supports only specific datasets and does not cover the benchmarks we use; as a closer controlled alternative, Appendix[F](https://arxiv.org/html/2609.16800#A6)compares DRPG with a no\-environment\-feedback variant of policy generation\. ## Appendix AUse of Large Language Models ChatGPT was used for grammar refinement and occasional assistance in programming tasks\. All language model outputs served only as references; the final text and code were written entirely by the authors\. ## Appendix BExperimental Cost All experiments conducted in this study were performed within the free\-tier allocations provided by the respective API providers\. Specifically, we utilized the complimentary access quotas offered by Google AI Studio, Nvidia NIM \(NVIDIA Inference Microservices\), and Meta Llama API\. These free\-tier limits were sufficient for our experimental requirements\. Consequently, no additional computational expenses or API usage fees were incurred during the course of this research\. Table[4](https://arxiv.org/html/2609.16800#A2.T4)complements the representative comparison in Section[5\.4](https://arxiv.org/html/2609.16800#S5.SS4)by reporting the per\-query input/output token counts for all 42 \(model×\\timesdataset\) configurations\. Figure[4](https://arxiv.org/html/2609.16800#A2.F4)further visualizes this cost–performance trade\-off, plotting the score of each configuration against its per\-query*output*\(completion\) tokens\. We focus on output tokens because autoregressive decoding makes them the dominant factor in wall\-clock latency, whereas input tokens are consumed in a single parallel prefill pass\([Pope et al\., 2023](https://arxiv.org/html/2609.16800#bib.bib24);[Patel et al\., 2024](https://arxiv.org/html/2609.16800#bib.bib23)\)\. Relative to Self\-Refine, DRPG generally sits toward the upper left: higher scores at comparable or lower output cost\. Relative to Self\-StreamICL, DRPG spends additional output tokens \(mainly for policy generation\), and whether this cost translates into gains depends on the task type, consistent with the analysis in Section 5: clear improvements on text\-to\-SQL, comparable performance on tasks relying on instance\-specific knowledge\. Figure 4:Score versus per\-query output \(completion\) tokens for all 42 model–dataset configurations and four methods \(scores from Table[1](https://arxiv.org/html/2609.16800#S5.T1), token counts from Table[4](https://arxiv.org/html/2609.16800#A2.T4)\)\. Each point is one model under one method; token counts include all LLM calls of the method\. Up and left is better\.Table 4:Per\-query computational cost \(input/output tokens per query, i\.e\. total tokens divided by the number of questions\)\. Each DRPG value already includes both of its LLM calls \(policy generation and answer\); each Self\-Refine value includes all refinement rounds\. ## Appendix CRetrieval Pipeline Details We fully reuse StreamBench’s retrieval pipeline without any modification; DRPG and Self\-StreamICL share this exact, unmodified retriever, so retrieval is not a component tuned for DRPG\. Table[5](https://arxiv.org/html/2609.16800#A3.T5)summarizes the precise configuration\. Table 5:Precise configuration of the retrieval pipeline, fully reused from StreamBench and shared by DRPG and all baselines\. ## Appendix DRobustness to Retrieval Strategy To examine the sensitivity of DRPG to the retrieval strategy used for policy generation, we compare three configurations: contrastive \(our default\), which retrieves an equal number of correct and incorrect examples; correct\-only, which retrieves only successful cases; and wrong\-only, which uses only failed cases\. As shown in Table[6](https://arxiv.org/html/2609.16800#A4.T6), all three strategies achieve comparable performance across benchmarks and models\. This indicates that the policy generation mechanism can reliably distill useful strategies from past experiences, regardless of whether the input examples are successes, failures, or a mixture of both\. The policy generator is not sensitive to the composition of its input, suggesting that the act of synthesizing high\-level rules is itself the key driver of improvement, rather than the specific examples provided\. This also challenges the common practice of discarding failure cases in memory\-based methods[Wu et al\. \(2024\)](https://arxiv.org/html/2609.16800#bib.bib36);[Min et al\. \(2022\)](https://arxiv.org/html/2609.16800#bib.bib19);[Wei et al\. \(2023\)](https://arxiv.org/html/2609.16800#bib.bib35), consistent with recent findings that analyzing errors can yield generalizable insights[Zhang et al\. \(2024a\)](https://arxiv.org/html/2609.16800#bib.bib41)\. We note the scope of this ablation: it varies the*composition*of the retrieved cases \(contrastive, correct\-only, wrong\-only\), not the relevance computation itself\. We did not ablate the embedding model, similarity metric, embedded fields, orkk, since the retriever \(Appendix[C](https://arxiv.org/html/2609.16800#A3)\) is shared and fixed across DRPG and all baselines and is not a component we optimized; a systematic relevance\-computation ablation is left to future work\. Table 6:Performance of different retrieval strategies in policy generation\. The “Contrastive” method uses an equal number of past correct and incorrect few\-shot examples\. The best\-performing method for each model\-task pair is highlighted inbold\. ## Appendix EEffect of Referencing the Previous Policy In our default DRPG design, the policy generator constructs each policy independently at every time step\. To examine whether maintaining policy continuity across steps is beneficial, we introduce a variant in which the policy generator receivesPt−1P\_\{t\-1\}as additional input when generatingPtP\_\{t\}\. Table[7](https://arxiv.org/html/2609.16800#A5.T7)reports the full results\. Neither configuration consistently outperforms the other across models and tasks, indicating that the policy generator can effectively synthesize useful strategies from retrieved examples alone, without requiring access to prior policies\. This stateless property simplifies system design and avoids the risk of error accumulation from inheriting outdated or overly specific rules, while the agent’s growing memory of past experiences \(Eq\. 1\) still ensures continual adaptation at the framework level\. Table 7:Ablation study on the impact of using the previous turn’s policy\. “w/” denotes with previous policy, “w/o” denotes without\. The best\-performing setting for each model\-task pair is highlighted inbold\. ## Appendix FEffect of Environment Feedback: A No\-Feedback Ablation To isolate the contribution of environment feedback, we construct a no\-feedback variant of DRPG, in which the policy is generated entirely from the retrieved few\-shot examples without any correct/incorrect labels; following Dynamic Cheatsheet, this variant uses no environment feedback\. All other components \(the policy generator, prompts, retriever,kk, and seed\) are identical to the DRPG configuration in Table[1](https://arxiv.org/html/2609.16800#S5.T1)\. Due to cost constraints, we run this ablation withllama\-3\.3\-70bon all six benchmarks\. As shown in Table[8](https://arxiv.org/html/2609.16800#A6.T8), DRPG outperforms its no\-feedback variant on all six datasets, indicating that the gains come specifically from feedback\-conditioned policy generation rather than from adding policy text or summarizing retrieved examples\. The gap is largest on DDXPlus and DS\-1000, where policies generated without feedback fall far below Self\-StreamICL; consistent with the task\-dependent analysis in Section 5, DRPG does not surpass Self\-StreamICL on these two datasets even with feedback, while feedback keeps its performance close to that baseline\. Table 8:No\-feedback ablation withllama\-3\.3\-70b\(same configuration as Table[1](https://arxiv.org/html/2609.16800#S5.T1)\)\. “DRPG w/o feedback” generates the policy from the retrieved examples without correctness labels\. The best result per dataset is highlighted inbold\. ## Appendix GSignificance Tests and Per\-Configuration Gains Treating each model–dataset configuration in Table[1](https://arxiv.org/html/2609.16800#S5.T1)as a paired observation \(DRPG vs\. Self\-StreamICL;n=7n=7models per dataset, 42 configurations overall\), we run paired Wilcoxon signed\-rank tests\. Across all 42 configurations, DRPG significantly outperforms Self\-StreamICL \(one\-sidedp=0\.005p=0\.005\)\. Table[9](https://arxiv.org/html/2609.16800#A7.T9)reports the per\-dataset two\-sided tests: the improvements are significant on Spider and CoSQL, while on DDXPlus and DS\-1000 the differences are not significant in either direction, with mean differences of only−0\.2\-0\.2and−0\.3\-0\.3points\. In other words, DRPG does not significantly degrade performance on any task type while improving significantly overall\. Table 9:Paired Wilcoxon signed\-rank tests of DRPG vs\. Self\-StreamICL per dataset \(n=7n=7models each\), computed from Table[1](https://arxiv.org/html/2609.16800#S5.T1)\. Overall, across all 42 configurations, DRPG is significantly better \(one\-sidedp=0\.005p=0\.005\)\.Figure[5](https://arxiv.org/html/2609.16800#A7.F5)visualizes the per\-configuration gainsΔ\\Delta\(DRPG−\-Self\-StreamICL\), grouped by model family\. The gains concentrate on the text\-to\-SQL benchmarks and HotpotQA, consistent with the task\-dependent analysis in Section 5\. Figure 5:Δ\\Delta\(DRPG−\-Self\-StreamICL\), in percentage points, for every model–dataset configuration, grouped by model family \(computed from Table[1](https://arxiv.org/html/2609.16800#S5.T1); shown to one decimal\)\. ## Appendix HResults on Additional Open\-Weight Model Families: Qwen and Gemma To further examine generality beyond the three model families in Table[1](https://arxiv.org/html/2609.16800#S5.T1), we evaluateqwen3\.5\-122b\-a10b\(Qwen\)[Qwen Team \(2026\)](https://arxiv.org/html/2609.16800#bib.bib25)andgemma\-4\-31b\-it\(Gemma\)[Gemma Team \(2026\)](https://arxiv.org/html/2609.16800#bib.bib9)on all six benchmarks, comparing Zero\-shot, the strongest baseline Self\-StreamICL, and DRPG under the same experimental setup as Table[1](https://arxiv.org/html/2609.16800#S5.T1); the results are shown in Table[10](https://arxiv.org/html/2609.16800#A8.T10)\. Both models are hybrid reasoning models; we disable their thinking mode in all runs \(thinking budget set to 0, so no reasoning traces are generated\), consistent with our scope of evaluating standard, non\-reasoning inference as stated in the Limitations\. For Qwen, DRPG outperforms Self\-StreamICL on Spider, BIRD, HotpotQA, and DS\-1000, performs comparably on CoSQL, and falls below Self\-StreamICL on DDXPlus\. This is consistent with the task\-dependent pattern analyzed in Section 5: policy\-level guidance helps most on tasks whose errors share recurring, generalizable patterns, whereas on DDXPlus the generated policies tend to capture overly narrow, instance\-specific patterns rather than broadly applicable strategies, limiting their benefit\. For Gemma, DRPG outperforms Self\-StreamICL on five of the six datasets and is comparable on HotpotQA\. Notably, Self\-StreamICL falls well below Zero\-shot on BIRD and DS\-1000 for this model, while DRPG does not exhibit the same degradation: it stays above Zero\-shot on BIRD and recovers most of the gap on DS\-1000\. This suggests that policy\-level guidance can be more robust than accumulating raw exemplars for some models\. Table 10:Performance ofqwen3\.5\-122b\-a10bandgemma\-4\-31b\-itacross the six benchmarks, evaluated under the same protocol as Table[1](https://arxiv.org/html/2609.16800#S5.T1)with the thinking mode of both models disabled \(thinking budget set to 0\)\. The best result per model–dataset pair is highlighted inbold\. ## Appendix IPrompt Templates for the Agent and Policy Generator This section presents the prompt templates used in our DRPG framework\. We provide agent prompts \(Figures[6](https://arxiv.org/html/2609.16800#A9.F6)–[9](https://arxiv.org/html/2609.16800#A9.F9)\) that guide task execution with few\-shot examples and policies, and policy generator prompts \(Figures[10](https://arxiv.org/html/2609.16800#A9.F10)–[13](https://arxiv.org/html/2609.16800#A9.F13)\) that refine policies from error analysis\. The templates cover text\-to\-SQL, question answering, medical diagnosis, and code generation tasks\. \[Role assignment\]You are performing the text\-to\-SQL task\.\[Reference materials\]Here are some examples:\{few\-shot examples\}Please pay special attention to the following points, which are derived from previous cases\.\{policy\}\[Question\]Now it’s your turn\.\- SQL schema:\{schema\}\- Using valid SQLite, answer the following question for the SQL schema provided above\.\- Question:\{question\}\[Additional instructions\]Now, generate the correct SQL code directly \(Do NOT generate other text except the SQL code\):’’’sql\\n<your SQL code\>\\n’’’Figure 6:The prompt template for the DRPG for text\-to\-SQL\. The\{schema\}would be replaced by database schema based on dataset\.\[Role assignment\]You are doing a question\-answering task\.\[Reference materials\]Here are some example cases:\{few\-shot examples\}Please pay special attention to the following points, which are derived from previous cases\.\{policy\}\[Question\]Now you are given the following context, which might help you answer the question:Context:\{context\}Question:\{question\}\[Additional instructions\]Note that you only need to answer with a short text span without explanation\. Now, provide your answer in the following JSON format:\{"answer": "<your answer text span\>"\}Figure 7:The prompt template for the DRPG for HotpotQA\. The\{context\}would be replaced by provided passages\.\[Role assignment\]Act as a medical doctor and diagnose the patient based on the provided patient profile\.\[Reference materials\]All possible diagnoses for you to choose from are as follows \(one diagnosis per line, in the format of <number\>\. <diagnosis\>\):\{option text\}Here are some example cases\.\{few\-shot examples\}Please pay special attention to the following points, which are derived from previous cases\.\{policy\}\[Question\]Now it’s your turn\.\{profile\}\[Additional instructions\]Now, directly provide the diagnosis for the patient in the following format:<number\>\. <diagnosis\>Figure 8:The prompt template for the DRPG used in DDXPlus\. The\{profile\}is replaced with the provided patient profile, and the\{option text\}contains all possible diagnoses\.\[Role assignment\]You are performing a python programming task to satisfy the user’s requirements\.\[Reference materials\]Here are some examples:\{few\-shot examples\}Please pay special attention to the following points, which are derived from previous cases\.\{policy\}\[Question\]Now it’s your turn\.The user’s requirements \(enclosed in "”\):\{question\}\[Additional instructions\]You need to provide your solution in python code to satisfy the user’s requirements\. Your code will be tested as follows \(enclosed in "”\):\{test cases\}Now, generate your code directly in the following format:’’’python<your code\>’’’Figure 9:The prompt template for the DRPG used in DS\-1000\. The\{question\}is replaced with the user’s requirements, and\{test cases\}is the provided execution context for code testing\.\[Role assignment\]You are optimizing the policy for a text\-to\-SQL agent to reduce errors in generating correct SQL queries\.\[Reference materials\]Here are several error cases that have occurred:"’’\{previous wrong cases\}"’’Here are several correct cases that have occurred:"’’\{previous correct cases\}"’’\[Additional instructions\]Please analyze all the cases and revise the policy accordingly\. Focus on text\-to\-SQL correctness and how to pass the required tests\.Remove redundant or obsolete points\. If certain previous policy points are still valid, keep them\.Output the revised policy, starting with POLICY: on a new line, followed immediately by a bulleted list \(each point on a new line, starting with \- \)\.Limit the revised policy to at most 5 concise and actionable bullet points relevant to common text\-to\-SQL mistakes\.Example output format:POLICY:\- \[first point\]\- \[second point\]Figure 10:The prompt template for the policy generator in text\-to\-SQL\.\[Role assignment\]You are optimizing a QA agent’s policy to reduce errors\.\[Reference materials\]Here are several error cases that have occurred:"’’\{previous wrong cases\}"’’Here are several correct cases that have occurred:"’’\{previous correct cases\}"’’\[Additional instructions\]Please analyze all the cases and revise the policy accordingly\. Focus on question answering correctness and how to pass the required tests\.Remove redundant or obsolete points\. If certain previous policies are still valid, keep them\.Output the revised policy, starting with POLICY: on a new line, followed immediately by a bulleted list \(each point on a new line, starting with \- \)\.Limit the revised policy to at most 5 concise, actionable bullet points relevant to common QA mistakes\.Example output format:POLICY:\- \[first point\]\- \[second point\]Figure 11:The prompt template for the policy generator in HotpotQA\.\[Role assignment\]You are optimizing the policy for a medical QA agent to reduce diagnostic errors\.\[Reference materials\]Here are several error cases that have occurred:"’’\{previous wrong cases\}"’’Here are several correct cases that have occurred:"’’\{previous correct cases\}"’’\[Additional instructions\]Please analyze all the cases and revise the policy accordingly\. Focus on diagnosis correctness and how to pass the required tests\.Remove redundant or obsolete points\. If certain previous policies are still valid, keep them\.Output the revised policy, starting with POLICY: on a new line, followed immediately by a bulleted list \(each point on a new line, starting with \- \)\.Limit the revised policy to at most 5 concise, actionable bullet points relevant to common diagnostic mistakes\.Example output format:POLICY:\- \[first point\]\- \[second point\]Figure 12:The prompt template for the policy generator in DDXPlus\.\[Role assignment\]You are optimizing the policy for a Python programming agent to reduce errors in solving user requirements\.\[Reference materials\]Here are several error cases that have occurred:"’’\{previous wrong cases\}"’’Here are several correct cases that have occurred:"’’\{previous correct cases\}"’’\[Additional instructions\]Please analyze all the cases and revise the policy accordingly\. Focus on Python code correctness and how to pass the required tests\.Remove redundant or obsolete points\. If certain previous policies are still valid, keep them\.Output the revised policy, starting with POLICY: on a new line, followed immediately by a bulleted list \(each point on a new line, starting with \- \)\.Limit the revised policy to at most 5 concise, actionable bullet points relevant to common Python coding and testing mistakes\.Example output format:POLICY:\- \[first point\]\- \[second point\]Figure 13:The prompt template for the policy generator in DS\-1000\.
Similar Articles
Dynamic Latent Routing
Dynamic Latent Routing (DLR) lets LLMs learn their own inner monologue by composing sub-policies via search, inspired by language compositionality. In low-data fine-tuning, DLR matches or outperforms standard supervised fine-tuning.
PolicyBank: Evolving Policy Understanding for LLM Agents
PolicyBank proposes a memory mechanism that enables LLM agents to autonomously refine their understanding of organizational policies through iterative interaction and corrective feedback, closing specification gaps that cause systematic behavioral divergence from true requirements. The work introduces a systematic testbed and demonstrates PolicyBank can close up to 82% of policy-gap alignment failures, significantly outperforming existing memory mechanisms.
From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning
This paper proposes the LLM-as-Environment-Engineer framework, where a policy model analyzes failures to automatically redesign the training environment for reinforcement learning, and introduces MAPF-FrozenLake as a controllable testbed. The framework, using Qwen3-4B, outperforms larger models like GPT and Gemini, showing that policy learning improves the model's ability to diagnose weaknesses.
Towards Reliable and Robust LLM Planning: Symbolic Feedback-Driven Iterative Self-Refinement Framework
This paper proposes a symbolic feedback-driven iterative self-refinement framework to improve the robustness and reliability of large language models in long-horizon planning tasks. The method uses natural language prompting, a symbolic verifier, and a plan recognizer to enhance feasibility and correctness.
@dair_ai: New paper on giving LLM agents experience that improves the weights and stays readable at the same time. Agent-experien…
JERP introduces a method for LLM agents to jointly learn interpretable natural-language rules and update policy parameters from the same interaction trajectories, improving performance on AlfWorld and WebShop while maintaining inspectability.