Mitigating Context Interference for Reliable and Efficient Search Agents

arXiv cs.CL Papers

Summary

This paper systematically studies context interference in multi-turn LLM-based search agents, finding that interference primarily arises from the latest retrieved documents, and introduces a distill-based context refiner to mitigate it. Incorporating context refinement into RL training pipelines significantly improves reliability and efficiency.

arXiv:2608.10743v1 Announce Type: new Abstract: Recent research empowers Large Language Models (LLMs) as multi-turn search agents to iteratively retrieve and generate outputs until complex tasks are solved. However, the contexts of multi-turn search agents are lengthy and complex. For example, the retrieved set of documents in each turn would inevitably introduce irrelevant information that distracts LLMs, referring to \textit{context interference}, potentially hindering the reliability and efficiency of search agents. Therefore, we conduct a systematic study on context interference in multi-turn search agents, focusing on investigating i) which parts of the context of search agents will contribute to the context interference, ii) how to refine the contexts of search agents to mitigate the interference, and iii) can incorporating context refinement into search agent training yield further improvements. We reveal that interference primarily arises from the latest retrieved documents. Based on the explored findings, we then introduce a distill-based context refiner to dynamically mitigate context interference for multi-turn search agents. Finally, we validate that incorporating context refinement into RL training pipelines of search agents can significantly enhance both reliability and efficiency. This study highlights the importance of mitigating context interference of search agents, inspiring a novel paradigm of ``refine context and then generate'' for AI agents.
Original Article
View Cached Full Text

Cached at: 08/12/26, 08:37 AM

# Mitigating Context Interference for Reliable and Efficient Search Agents
Source: [https://arxiv.org/html/2608.10743](https://arxiv.org/html/2608.10743)
Boyang Xue1,6, Bin Wu2, Shuofei Qiao3, Sheng Wang4, Rui Wang1,6, Yiming Du1,6, Hongru Wang5, Jeff Z\. Pan5, Emine Yilmaz2∗, Kam\-Fai Wong1,6∗, Aldo Lipani2 1The Chinese University of Hong Kong,2University College London 3Zhejiang University,4The University of Hong Kong,5The University of Edinburgh 6MoE Key Laboratory of High Confidence Software Technologies \{byxue, kfwong\}@se\.cuhk\.edu\.hk

###### Abstract

Recent research empowers Large Language Models \(LLMs\) as multi\-turn search agents to iteratively retrieve and generate outputs until complex tasks are solved\. However, the contexts of multi\-turn search agents are lengthy and complex\. For example, the retrieved set of documents in each turn would inevitably introduce irrelevant information that distracts LLMs, referring tocontext interference, potentially hindering the reliability and efficiency of search agents\. Therefore, we conduct a systematic study on context interference in multi\-turn search agents, focusing on investigating i\) which parts of the context of search agents will contribute to the context interference, ii\) how to refine the contexts of search agents to mitigate the interference, and iii\) can incorporating context refinement into search agent training yield further improvements\. We reveal that interference primarily arises from the latest retrieved documents\. Based on the explored findings, we then introduce a distill\-based context refiner to dynamically mitigate context interference for multi\-turn search agents\. Finally, we validate that incorporating context refinement into RL training pipelines of search agents can significantly enhance both reliability and efficiency\. This study highlights the importance of mitigating context interference of search agents, inspiring a novel paradigm of “refine context and then generate” for AI agents\.

Mitigating Context Interference for Reliable and Efficient Search Agents

Boyang Xue1,6, Bin Wu2, Shuofei Qiao3, Sheng Wang4, Rui Wang1,6, Yiming Du1,6,Hongru Wang5††thanks:Co\-corresponding authors\., Jeff Z\. Pan5, Emine Yilmaz2∗, Kam\-Fai Wong1,6∗, Aldo Lipani21The Chinese University of Hong Kong,2University College London3Zhejiang University,4The University of Hong Kong,5The University of Edinburgh6MoE Key Laboratory of High Confidence Software Technologies\{byxue, kfwong\}@se\.cuhk\.edu\.hk

## 1Introduction

Large Language Models \(LLMs\) have demonstrated strong performance in tackling complex tasks using their pretrained knowledge\(gpt5;deepseekai2025deepseekr1incentivizingreasoningcapability\)\. Recent work has further empowered them to invoke search engines to retrieve external knowledge, essentially training them as multi\-turn search agents that iterate on retrieval and generation until tasks are solved\(wang2025theoryagentstoolusedecisionmakers;jia2025fastslowtoolaugmentedthinking;huang2025reinforcedinternalexternalknowledgesynergistic\)\.

![Refer to caption](https://arxiv.org/html/2608.10743v1/x1.png)Figure 1:Demonstration of how context interference affects the performance of LLM\-based search agents\. \(“recall rate”=Nr/NN\_\{\\text\{r\}\}/Ndenotes the proportion of questions for which the retrieved documents contain the correct answer \(NrN\_\{\\text\{r\}\}\) among all questions \(NN\)\. “recall accuracy”=Nrc/NN\_\{\\text\{rc\}\}/Nrefers to the proportion of correctly answered questions \(NrcN\_\{\\text\{rc\}\}\) inNrN\_\{\\text\{r\}\}, relative to all questions \(NN\)\. “accuracy” represents the proportion of all correctly answered questionsNcN\_\{\\text\{c\}\}out of the questions \(NN\)\.\)However, the contexts of multi\-turn search agents are lengthy and complex, encompassing the question, multi\-round search queries, retrieved documents, and reasoning steps\(jin2025search\), which may include irrelevant information\. For instance, the retriever always returns a set of documents to ensure coverage of the search query, which also introduces noisy or irrelevant documents into the context\(dong2025understand\)\. This refers tocontext interference, which indicates “feeding too much irrelevant context may confuse LLMs to focus on wrong information\(coleman2023incontextinterferencechatbasedlarge;haseeb2025contextengineeringmultiagentllm;gupta\-etal\-2024\-llm;jiang2025enhancingrobustnesslargelanguage\)\.” The context interference presented in each round may distract LLMs from irrelevant information and persistently degrade subsequent generation quality\(li2025singleturnsurveymultiturninteractions;laban2025llmslostmultiturnconversation\), thereby hindering the efficiency and reliability of search agents\. As in Figure[1](https://arxiv.org/html/2608.10743#S1.F1), the gaps between “recall rate” and “recall accuracy” demonstrate that search agents always retrieve the documents containing the useful information but fail to generate the correct answer, highlighting the context interference effect of accurate knowledge expression of search agents\.

Prior studies to mitigate context interference have predominantly centered on dialogue systems\(jacqmin\-etal\-2022\-follow\)and retrieval\-augmented generation \(RAG\)\(glass\-etal\-2022\-re2g;nguyen2025maragmultiagentretrievalaugmentedgeneration;yu2024rankragunifyingcontextranking\), while largely overlooking techniques on multi\-turn search agent settings\. Therefore, we systematically study thecontext interferenceissue on search agents in this work with three research questions \(RQ\):

i\)Which parts of contexts will contribute to context interference for multi\-turn search agents?

ii\)How to refine contexts of search agents to mitigate such interference?

iii\)Can leveraging context refinement in RL training pipelines of search agents yield further performance improvements?

![Refer to caption](https://arxiv.org/html/2608.10743v1/x2.png)Figure 2:Examples of the search agent with \(a\) contextual interference in the retrieved documents and \(b\) refined contexts with the most critical relevant information\. The question derives from PopQA\(mallen2022not\), and we employ Qwen2\.5\-7b\-Instruct\(yang2024qwen2\)as the foundation LLM of the search agent\. We present more multi\-turn QA examples of context interference mitigation in the Appendix\.In light of the above questions, we investigate the context interference effect on multi\-turn search agents with respect toreliabilityandefficiencyacross a series of closed\-book QA benchmarks\. ForRQ i, we compare the performances of multi\-turn search agents with different inputs by masking specific parts of history contexts\. Results identify that context interference primarily derives from the latest retrieved document of search agents and slightly arises from previous search queries and documents\. ForRQ ii, we first explore a series of context interference mitigation strategies based on the conclusion fromRQ i\. Based on the exploration, we propose to distill a context refinement dataset comprising both retrieved documents and refined texts, which contains the most critical information in documents to the search query\. Then we train a context refiner using the dataset, which can mitigate context interference for multi\-turn search agents as exemplified in Figure[2](https://arxiv.org/html/2608.10743#S1.F2)and improve both efficiency and reliability\. ForRQ iii, we further incorporate the context refiner into the reinforcement learning \(RL\) training pipelines of search agents\. Experiments demonstrate that training multi\-turn search agents with dynamically refined context achieves significant performance improvements regarding both the reliability and efficiency over other training baselines\.

The contributions of this work are as follows:

\(1\) This work first investigates the context interference issue on multi\-turn search agents, highlighting the necessity of context refinement to improve both efficiency and reliability of search agents, inspiring a novel paradigm of “refine context and then generate” for AI agents111We have released the codes of this work on[https://github\.com/AmourWaltz/CRRL\.git](https://github.com/AmourWaltz/CRRL.git)\.\.

\(2\) This work reveals that context interference primarily derives from the latest documents in multi\-turn search agents, and therefore introduces a distill\-based context refiner to dynamically eliminate context interference for search agents, which can be applicable to mitigate contextual interference in other search agent scenarios\.

\(3\) This work further incorporates context refinement into RL training pipelines of search agents, which can further enhance both reliability and efficiency, providing insight into refining contexts during search agent training for future work\.

## 2Preliminary of Search Agent

To establish a theoretical foundation to analyze the context interference issue, we introduce the concepts of the search agent’s internal/external knowledge, a Markov Decision Process \(MDP\), and practical settings of multi\-turn search agents\. Notation definitions of this work can be found in Appendix[A](https://arxiv.org/html/2608.10743#A1)\. Related works are in Appendix[B](https://arxiv.org/html/2608.10743#A2)\.

#### Internal/External Knowledge of Search Agent

Previous works identify the concept ofInternal/External Knowledgefor LLM agent\(wang2025theoryagentstoolusedecisionmakers;jia2025fastslowtoolaugmentedthinking\), where internal knowledge𝓚I\\boldsymbol\{\\mathcal\{K\}\}\_\{I\}is learned from the pretrained corpus and encoded within the model parameters, and external knowledge𝓚E\\boldsymbol\{\\mathcal\{K\}\}\_\{E\}is accessed through external tools \(e\.g\., search engines\)\. For an LLM\-based search agent𝓜\\boldsymbol\{\\mathcal\{M\}\}with parametersθ\\theta, the output𝒚\\boldsymbol\{y\}is jointly determined by its internal parametric knowledge𝓚I∈θ\\boldsymbol\{\\mathcal\{K\}\}\_\{I\}\\in\\theta, the externally retrieved documents𝒅∈𝓚E\\boldsymbol\{d\}\\in\\boldsymbol\{\\mathcal\{K\}\}\_\{E\}, and the input task𝒙\\boldsymbol\{x\}as𝒚=𝓜θ​\(𝒙,𝒅\)\\boldsymbol\{y\}=\\boldsymbol\{\\mathcal\{M\}\}\_\{\\theta\}\(\\boldsymbol\{x\},\\boldsymbol\{d\}\), where𝒙\\boldsymbol\{x\}generally refers to the query and sequentially concatenated history outputs\.

As𝓚I\\boldsymbol\{\\mathcal\{K\}\}\_\{I\}embedded in parametersθ\\thetacannot be directly accessed, the manifestation of𝓚I\\boldsymbol\{\\mathcal\{K\}\}\_\{I\}is conditioned on the context of𝒙\\boldsymbol\{x\}and𝒅\\boldsymbol\{d\}\. The retriever generally returns a set of documents in𝒅∈𝓚E\\boldsymbol\{d\}\\in\\boldsymbol\{\\mathcal\{K\}\}\_\{E\}intended to comprehensively cover the search query, but this inevitably introduces redundancy and noise\. Such extraneous content can in turn distort the utilization of both𝓚I\\boldsymbol\{\\mathcal\{K\}\}\_\{I\}and𝓚E\\boldsymbol\{\\mathcal\{K\}\}\_\{E\}, and undermine the reliability of search agents\. In Figure[1](https://arxiv.org/html/2608.10743#S1.F1), “recall rate” denotes the questions calling𝓚E\\boldsymbol\{\\mathcal\{K\}\}\_\{E\}containing helpful information, while “recall accuracy” refers to actually solved questions using𝓚E\\boldsymbol\{\\mathcal\{K\}\}\_\{E\}, and “accuracy” represents successfully answered questions through the combined use of𝓚I\\boldsymbol\{\\mathcal\{K\}\}\_\{I\}and𝓚E\\boldsymbol\{\\mathcal\{K\}\}\_\{E\}\. The gaps between “recall rate” and “recall accuracy” / “accuracy” exemplify the negative impact of context interference\. Moreover, distractions from irrelevant information will incur extra retrievals and inference costs, degrading the efficiency of search agents\.

#### Markov Decision Process of Agent

The process of an LLM agent𝓜θ\\boldsymbol\{\\mathcal\{M\}\}\_\{\\theta\}with𝓚I∈θ\\boldsymbol\{\\mathcal\{K\}\}\_\{I\}\\in\\thetainteracting with an environment𝑬\\boldsymbol\{E\}\(referred to𝓚E\\boldsymbol\{\\mathcal\{K\}\}\_\{E\}\) to complete a task can be regarded as a Markov Decision Process:\(𝓢,𝓐,𝓣,𝓞,𝓡\)\\left\(\\boldsymbol\{\\mathcal\{S\}\},\\boldsymbol\{\\mathcal\{A\}\},\\boldsymbol\{\\mathcal\{T\}\},\\boldsymbol\{\\mathcal\{O\}\},\\boldsymbol\{\\mathcal\{R\}\}\\right\)\. Initially, a specific task𝒙\\boldsymbol\{x\}is provided as the initial environmental state\. Assuming the interaction proceeds inNNrounds and in roundnn\(n<Nn<N\), the LLM agent receives observation𝒐n∈𝓞\\boldsymbol\{o\}\_\{n\}\\in\\boldsymbol\{\\mathcal\{O\}\}and takes action𝒂n∈𝓐\\boldsymbol\{a\}\_\{n\}\\in\\boldsymbol\{\\mathcal\{A\}\}\. The state𝒔n\\boldsymbol\{s\}\_\{n\}at roundnnis the history context of all preceding concatenated sequences𝒔n=\(𝒙,𝒂0,𝒐1,𝒂1,…,𝒂n−1,𝒐n\)∈𝓢\\boldsymbol\{s\}\_\{n\}=\(\\boldsymbol\{x\},\\boldsymbol\{a\}\_\{0\},\\boldsymbol\{o\}\_\{1\},\\boldsymbol\{a\}\_\{1\},\\dots,\\boldsymbol\{a\}\_\{n\-1\},\\boldsymbol\{o\}\_\{n\}\)\\in\\boldsymbol\{\\mathcal\{S\}\}\.𝓜θ\\boldsymbol\{\\mathcal\{M\}\}\_\{\\theta\}is responsible for deciding𝒂n\\boldsymbol\{a\}\_\{n\}based on𝒔n\\boldsymbol\{s\}\_\{n\}:𝒂n∼𝓜θ\(⋅\|𝒔n\)\\boldsymbol\{a\}\_\{n\}\\sim\\boldsymbol\{\\mathcal\{M\}\}\_\{\\theta\}\(\\cdot\|\\boldsymbol\{s\}\_\{n\}\), and a retriever𝓔\\boldsymbol\{\\mathcal\{E\}\}interacts with𝑬\\boldsymbol\{E\}to determines the state transition𝓣\\boldsymbol\{\\mathcal\{T\}\}\. Upon task completion afterNNrounds, the trajectory is characterized as𝝉=\(𝒙,𝒔0,𝒂0,𝒐1,…,𝒂N−1,𝒐N,𝒂N\)\\boldsymbol\{\\tau\}=\(\\boldsymbol\{x\},\\boldsymbol\{s\}\_\{0\},\\boldsymbol\{a\}\_\{0\},\\boldsymbol\{o\}\_\{1\},\\dots,\\boldsymbol\{a\}\_\{N\-1\},\\boldsymbol\{o\}\_\{N\},\\boldsymbol\{a\}\_\{N\}\)where the final response𝒚\\boldsymbol\{y\}is parsed from𝒂N\\boldsymbol\{a\}\_\{N\}\. A final rewardrris provided by the reward model𝓡\\boldsymbol\{\\mathcal\{R\}\}by comparing𝒚\\boldsymbol\{y\}and the reference𝒚^\\boldsymbol\{\\hat\{y\}\}asr=𝓡​\(𝒚,𝒚^\)r=\\boldsymbol\{\\mathcal\{R\}\}\(\\boldsymbol\{y\},\\boldsymbol\{\\hat\{y\}\}\)\.

#### Multi\-turn Search Agent

For multi\-turn search agents, theii\-th observation𝒐i\\boldsymbol\{o\}\_\{i\}is generally a list of retrieved Top\-KKdocuments𝒅i=\[𝒅i,1,𝒅i,2,…,𝒅i,K\]\\boldsymbol\{d\}\_\{i\}=\[\\boldsymbol\{d\}\_\{i,1\},\\boldsymbol\{d\}\_\{i,2\},\\dots,\\boldsymbol\{d\}\_\{i,K\}\]returned by the retriever𝓔​\(𝒒i−1\|𝑲b\)\\boldsymbol\{\\mathcal\{E\}\}\(\\boldsymbol\{q\}\_\{i\-1\}\|\\boldsymbol\{K\}\_\{b\}\)where𝒒i−1\\boldsymbol\{q\}\_\{i\-1\}is the search query generated in the previous round and𝑲b\\boldsymbol\{K\}\_\{b\}is the external knowledge base as𝓚E\\boldsymbol\{\\mathcal\{K\}\}\_\{E\}\. The action𝒂i\\boldsymbol\{a\}\_\{i\}includes a thinking step𝒕i\\boldsymbol\{t\}\_\{i\}and a search query𝒒i\\boldsymbol\{q\}\_\{i\}\. Followingjin2025search, LLMs are instructed to encapsulate their search queries, retrieved documents, and final answer between specially designated tokens respectively\. Both𝒒i\\boldsymbol\{q\}\_\{i\}and𝒅i\\boldsymbol\{d\}\_\{i\}are appended to the context in each turn\. When generating𝒕i\\boldsymbol\{t\}\_\{i\}, all the preceding sequences in𝒔i\\boldsymbol\{s\}\_\{i\}are fed into𝓜θ\\boldsymbol\{\\mathcal\{M\}\}\_\{\\theta\}and contribute to𝒂i\\boldsymbol\{a\}\_\{i\}\. However, not all previous documents and search queries are pertinent to the current thinking step, and the inclusion of irrelevant content can introduce context interference, impairing the LLMs’ efficiency and reliability to accurately express knowledge\.

## 3Context Interference in Search Agent

MethodsSingle\-Hop QAMulti\-Hop QAAvg\.NQTriviaQAPopQAHotpotQA2wikiMusiqueBamboogleQwen2\.5\-7b\-InstructDirect17\.7 / 0\.044\.2 / 0\.015\.5 / 0\.017\.9 / 0\.024\.0/ 0\.03\.9 / 0\.08\.0 / 0\.018\.7 / 0\.0CoT17\.7 / 0\.047\.0 / 0\.013\.4 / 0\.021\.0 / 0\.023\.6 / 0\.04\.7 / 0\.030\.4 / 0\.022\.5 / 0\.0IRCoT30\.6 / 2\.051\.2 / 1\.831\.7 / 2\.124\.6 / 3\.018\.0 / 3\.49\.8 / 3\.128\.8 / 2\.527\.5 / 2\.6IRCoT\-oo29\.8 /1\.851\.3 /1\.732\.7 /1\.725\.3 / 2\.722\.7 / 3\.111\.2 / 2\.934\.4 / 1\.929\.6 /2\.3IRCoT\-o​qoq32\.8 / 2\.051\.7/ 2\.033\.0/ 2\.026\.0/ 3\.023\.6 /3\.111\.4/ 3\.134\.4 /1\.730\.4/ 2\.4IRCoT\-o​q​poqp32\.9/ 2\.651\.5 / 2\.732\.7 / 2\.825\.4 /2\.622\.8 / 3\.110\.6 /2\.836\.0/ 2\.530\.3 / 2\.7Qwen2\.5\-3b\-InstructDirect9\.7 / 0\.025\.7 / 0\.07\.7 / 0\.013\.5 / 0\.017\.1 / 0\.01\.7 / 0\.03\.2 / 0\.011\.2 / 0\.0CoT12\.4 / 0\.035\.4 / 0\.09\.4 / 0\.015\.5 / 0\.016\.4 / 0\.02\.6 / 0\.020\.0/ 0\.016\.0 / 0\.0IRCoT21\.6 /1\.245\.2 / 1\.129\.6 / 1\.124\.1 / 1\.623\.5 / 1\.96\.8 / 1\.619\.2 / 1\.424\.3 / 1\.4IRCoT\-oo22\.4 / 1\.247\.0 /1\.131\.8/1\.124\.7 /1\.521\.3 /1\.78\.0/ 1\.519\.4 /1\.224\.9 /1\.3IRCoT\-o​qoq23\.5 / 1\.347\.2 / 1\.229\.9 / 1\.424\.9/ 1\.624\.2/ 1\.77\.0 / 1\.619\.4 / 1\.225\.2/ 1\.4IRCoT\-o​q​poqp24\.6/ 1\.547\.7/ 1\.330\.5 / 1\.524\.5 / 1\.822\.5 / 1\.75\.1 /1\.519\.0 / 1\.224\.8 / 1\.5Table 1:Performance \(EM/ART\) of different inference methods for search agents across QA test sets, measured by Exact Match \(EM, for reliability\) and Average Retrieval Times \(ART, for efficiency\)\.In this section, we first detail the evaluation setting in Sec\.[3\.1](https://arxiv.org/html/2608.10743#S3.SS1)\. Then we analyze context interference effects in different parts of contexts of search agents in Sec\.[3\.2](https://arxiv.org/html/2608.10743#S3.SS2), and finally propose the context refiner to mitigate interference in Sec[3\.3](https://arxiv.org/html/2608.10743#S3.SS3)\.

### 3\.1Evaluation Settings

#### Dataset

Various closed\-book QA datasets are employed to evaluate the performance of search agents, which necessitate extra retrieval to address, encompassing both single\- and multi\-hop scenarios\.Single\-hop QAincludes:Natural Questions \(NQ\)\(kwiatkowski2019natural\),TriviaQA\(joshi2017triviaqa\), andPopQA\(mallen2022not\)\.Multi\-hop QAincludes:HotpotQA\(yang2018hotpotqa\),2WikiMultiHopQA \(2Wiki\)\(ho2020constructing\),MuSiQue\(trivedi2022musique\), andBamboogle\(press2022measuring\)\. Dataset details are presented in Appendix[C](https://arxiv.org/html/2608.10743#A3)\.

#### Search Agent

We employ two foundation LLMs𝓜\\boldsymbol\{\\mathcal\{M\}\}for search agents:Qwen\-2\.5\-7b\-InstructandQwen\-2\.5\-3b\-Instruct\(yang2024qwen2\)\. For retrieval, we leverage E5\(wang2022text\)as the retriever𝓔\\boldsymbol\{\\mathcal\{E\}\}and 2018 Wikipedia dump\(karpukhin2020dense\)as the knowledge base𝑲b\\boldsymbol\{K\}\_\{b\}respectively\. The number of retrieved passagesKKis set to 3 across all retrieval\-based methods\.

#### Evaluation Metrics

We employ several metrics to evaluate both thereliabilityandefficiencyof search agents\. For reliability, we assess the correctness of the generated answer𝒚\\boldsymbol\{y\}with the reference𝒚^\\boldsymbol\{\\hat\{y\}\}usingExact Match \(EM\)\(song2025r1searcherincentivizingsearchcapability;jin2025search\), a standard string\-matching metric that checks whether𝒚≡𝒚^\\boldsymbol\{y\}\\equiv\\boldsymbol\{\\hat\{y\}\}, which is a percentage that represents the proportion of correctly answered questions out of all questions\. Efficiency is evaluated regarding different aspects\. Since the time cost of search agents is primarily determined by the number of retrieval operations, theaverage retrieval times \(ART\), which denotes the average number of retrievals required per question, is employed to intuitively evaluate the efficiency of the generation processes\. In addition, theaverage context length \(Len\.\)in multi\-turn generations and theaverage inference time per question \(AIT\)\(/seconds\) are also employed to assist efficiency assessments for search agents\.

### 3\.2Context Interference in Different Parts of Context of Search Agents

#### Background

To figure out context interference effects of different parts in the contexts of search agents \(asRQ i\), we compare the performance of multi\-turn search agents with different input contexts by masking specific segments \(actions and observations in preceding rounds\) of the history state\. The history state𝒔i\\boldsymbol\{s\}\_\{i\}includes input question𝒙\\boldsymbol\{x\}and a series of actions𝒂0:i−1\\boldsymbol\{a\}\_\{0:i\-1\}and observations𝒐1:i\\boldsymbol\{o\}\_\{1:i\}\. Specifically,𝒙\\boldsymbol\{x\}representing the initial state is fixed\.𝒑i\\boldsymbol\{p\}\_\{i\}in𝒂i\\boldsymbol\{a\}\_\{i\}generally involves summarizing and reasoning from𝒐i\\boldsymbol\{o\}\_\{i\}\.𝒒i\\boldsymbol\{q\}\_\{i\}is to interact with𝑲b\\boldsymbol\{K\}\_\{b\}to get retrieved documents in𝒐i\+1\\boldsymbol\{o\}\_\{i\+1\}\. We present several inference methods for search agents as follows\.

#### Inference Methods

We employ direct inference \(Direct\) and Chain\-of\-Thought \(CoT\) reasoning\(wei2022chain\)as two retrieval\-free baselines, which represent LLMs’ internal knowledge𝓚I\\boldsymbol\{\\mathcal\{K\}\}\_\{I\}to answer questions\. The prompt templates are in Appendix[F](https://arxiv.org/html/2608.10743#A6)\. For retrieval\-based settings of search agents, we employ Information Retrieval with CoT \(IRCoT\)\(trivedi2022interleaving\), which enables LLMs to actively call the retriever for questions beyond their knowledge scope after thinking\.

To understand the effect of different parts of context, several variants based onIRCoTare developed\. Since search agents typically decompose a complex question into a set of sub\-questions\(jin2025search;song2025r1searcherincentivizingsearchcapability\), the generated search queries and retrieved documents in a history state are generally mutually independent, exhibiting rarely sequential dependencies\. Therefore, we can specifically mask different parts in the states as follows\. 1\) To investigate interference in previous documents,IRCoT\-oo\(w/oo:−1\\boldsymbol\{o\}\_\{:\-1\}\)only incorporates the latest observation𝒐i\\boldsymbol\{o\}\_\{i\}of retrieved documents in context\[𝒙,𝒑0,𝒒1,𝒑1,𝒒1,…,𝒒i−1,𝒐i\]\[\\boldsymbol\{x\},\\boldsymbol\{p\}\_\{0\},\\boldsymbol\{q\}\_\{1\},\\boldsymbol\{p\}\_\{1\},\\boldsymbol\{q\}\_\{1\},\\dots,\\boldsymbol\{q\}\_\{i\-1\},\\boldsymbol\{o\}\_\{i\}\]when generating𝒂i\\boldsymbol\{a\}\_\{i\}\. 2\) For interference in search queries,IRCoT\-o​qoq\(w/oo:−1,q:−1\\boldsymbol\{o\}\_\{:\-1\},\\boldsymbol\{q\}\_\{:\-1\}\)with context\[𝒙,𝒑0,𝒑1,…,𝒑i−1,𝒒i−1,𝒐i\]\[\\boldsymbol\{x\},\\boldsymbol\{p\}\_\{0\},\\boldsymbol\{p\}\_\{1\},\\dots,\\boldsymbol\{p\}\_\{i\-1\},\\boldsymbol\{q\}\_\{i\-1\},\\boldsymbol\{o\}\_\{i\}\]is also employed\. 3\) For previous thinking steps, we utilizeIRCoT\-o​q​p\{oqp\}\(w/oo:−1,q:−1,p:−1\\boldsymbol\{o\}\_\{:\-1\},\\boldsymbol\{q\}\_\{:\-1\},\\boldsymbol\{p\}\_\{:\-1\}\)relies exclusively on the latest thinking, search query, and observation as context\[𝒙,𝒑i−1,𝒒i−1,𝒐i\]\[\\boldsymbol\{x\},\\boldsymbol\{p\}\_\{i\-1\},\\boldsymbol\{q\}\_\{i\-1\},\\boldsymbol\{o\}\_\{i\}\]when generating𝒂i\\boldsymbol\{a\}\_\{i\}\.

#### Analysis and Findings

As presented in Table[1](https://arxiv.org/html/2608.10743#S3.T1), benefiting from𝑲b\\boldsymbol\{K\}\_\{b\}, all retrieval\-based methods outperform retrieval\-free baselines\. Moreover,IRCoT\-oooutperformsIRCoTin both reliability and efficiency, suggesting that previous retrieved documents before roundiicontain context interference for generating𝒂i\\boldsymbol\{a\}\_\{i\}\.IRCoT\-o​q\{oq\}marginally outperformsIRCoT\-oo, indicating that previous search queries𝒒:−1\\boldsymbol\{q\}\_\{:\-1\}also carry slight interference, although removing them may incur little extra retrieval\.IRCoT\-o​q​poqpexhibits a slight drop in reliability and a notable decrease in efficiency compared with others, implying that previous thinking steps store key information for future steps\. LLMs may repeat previous search after masking𝒑:−1\\boldsymbol\{p\}\_\{:\-1\}and thus add retrieval costs\. Consequently,context interference when generatingai\\boldsymbol\{a\}\_\{i\}may arise from both previous search queries and documents\.

However, as presented in Figure[3](https://arxiv.org/html/2608.10743#S3.F3), although the aboveIRCoTvariants gain slight improvements in accuracy, the gap between “recall rate” and “recall accuracy” remains considerable, indicating context interference in previous rounds is not the dominant factor\. Since the latest thinking𝒑i−1\\boldsymbol\{p\}\_\{i\-1\}and search query𝒒i−1\\boldsymbol\{q\}\_\{i\-1\}mostly encapsulate summaries and reasoning of previous context, it can be inferred thatthe latest observationoi\\boldsymbol\{o\}\_\{i\}is subject to the primary cause of context interference\. To mitigate the interference, further context refinement methods for the latest observed documents are required on multi\-turn search agents\.

![Refer to caption](https://arxiv.org/html/2608.10743v1/figures/recall_inter.png)Figure 3:Demonstrations of context interference effects on four IRCoT variants of search agents\. Results are averaged on all QA test sets; metrics follow the definitions in Figure[1](https://arxiv.org/html/2608.10743#S1.F1)\.![Refer to caption](https://arxiv.org/html/2608.10743v1/x3.png)Figure 4:Training pipeline for the distill\-based context refiner\.

### 3\.3Context Refiner for Search Agent

Table 2:Performance \(EM/ART\) of context refinement methods for mitigating interference, measured by Exact Match \(EM\) and Average Retrieval Times \(ART\) across QA test sets\.#### Background

Our preliminary findings and analysis in Sec\.[3\.2](https://arxiv.org/html/2608.10743#S3.SS2)demonstrate that we can marginally mitigate context interference by removing irrelevant information like previous documents and search queries in multi\-turn search agents but not enough, which highlights the necessity of further capturing critical information and filtering noise in the latest retrieved documents \(asRQ ii\)\. More generally, we desire to develop a context refiner𝓕\\boldsymbol\{\\mathcal\{F\}\}that, in each roundii, refine the context to preserve the most relevant information𝒅~i=𝓕​\(𝒒i−1,𝒅i\)\\boldsymbol\{\\tilde\{d\}\}\_\{i\}=\\boldsymbol\{\\mathcal\{F\}\}\(\\boldsymbol\{q\}\_\{i\-1\},\\boldsymbol\{d\}\_\{i\}\)to search query𝒒i−1\\boldsymbol\{q\}\_\{i\-1\}from the latest retrieved documents𝒅i\\boldsymbol\{d\}\_\{i\}\.𝒅~i\\boldsymbol\{\\tilde\{d\}\}\_\{i\}is served as theii\-th observation𝒐i\\boldsymbol\{o\}\_\{i\}and then appended to𝒔i\\boldsymbol\{s\}\_\{i\}to generate𝒂i\\boldsymbol\{a\}\_\{i\}\.

Prior works on mitigating context interference are demonstrated in Appendix[B\.2](https://arxiv.org/html/2608.10743#A2.SS2)\. However, directly designing precise schemes to capture key information from numerous documents is challenging\(glass\-etal\-2022\-re2g\)\. A compression model may focus on summarizing global contents, potentially leading to information loss or the introduction of extraneous knowledge\(li2025singleturnsurveymultiturninteractions\)\. Relatively small LLMs exhibit limited capability for key information extraction \(in Table[2](https://arxiv.org/html/2608.10743#S3.T2)\)\. Therefore, we propose to distill a dataset for context refinement from advanced LLMs, enabling relatively weak models to refine context to mitigate interference\.

![Refer to caption](https://arxiv.org/html/2608.10743v1/figures/context_len.png)Figure 5:Averaged context lengths of search agents across all QA test sets, comparing IRCoT \(baseline\), three context refinement methods, and the proposed context refiner for efficiency assessments\.
#### Context Refiner

Given a distill dataset𝓓d=\{𝒙i,𝒚^i\}i=1N\\boldsymbol\{\\mathcal\{D\}\}\_\{d\}=\\\{\\boldsymbol\{x\}\_\{i\},\\boldsymbol\{\\hat\{y\}\}\_\{i\}\\\}\_\{i=1\}^\{N\}, a foundation LLM𝓜π\\boldsymbol\{\\mathcal\{M\}\}\_\{\\pi\}, an advanced teacher LLM𝓜T\\boldsymbol\{\\mathcal\{M\}\}\_\{T\}, a retriever𝓔\\boldsymbol\{\\mathcal\{E\}\}, and a knowledge base𝑲b\\boldsymbol\{K\}\_\{b\}, we infer each query𝒙\\boldsymbol\{x\}on𝓜T\\boldsymbol\{\\mathcal\{M\}\}\_\{T\}using IRCoT\(trivedi2023interleaving\)\. Inii\-th round of inference for𝒙\\boldsymbol\{x\}, the retriever return documents𝒅i=𝓔​\(𝒒i−1\|𝑲b\)\\boldsymbol\{d\}\_\{i\}=\\boldsymbol\{\\mathcal\{E\}\}\(\\boldsymbol\{q\}\_\{i\-1\}\|\\boldsymbol\{K\}\_\{b\}\)\. Then we instruct the teacher model𝓜T\\boldsymbol\{\\mathcal\{M\}\}\_\{T\}to specifically extract only critical information related to𝒒i−1\\boldsymbol\{q\}\_\{i\-1\}from𝒅i\\boldsymbol\{d\}\_\{i\}\. Extracted information𝒅~i=𝓜T​\(𝒒i−1,𝒅i\)\\boldsymbol\{\\tilde\{d\}\}\_\{i\}=\\boldsymbol\{\\mathcal\{M\}\}\_\{T\}\(\\boldsymbol\{q\}\_\{i\-1\},\\boldsymbol\{d\}\_\{i\}\)is then appended into context𝒔i\\boldsymbol\{s\}\_\{i\}to generate the next\-step action𝒂i\\boldsymbol\{a\}\_\{i\}\. For each correct trajectory with𝒚≡𝒚~\\boldsymbol\{y\}\\equiv\\boldsymbol\{\\tilde\{y\}\}, we retain all step\-wise pairs of extracted data\(𝒅~i,<​𝒅i,𝒒i−1​\>\)\(\\boldsymbol\{\\tilde\{d\}\}\_\{i\},\\text\{<\}\\boldsymbol\{d\}\_\{i\},\\boldsymbol\{q\}\_\{i\-1\}\\text\{\>\}\)and employ an entailment model to verify that𝒅~i\\boldsymbol\{\\tilde\{d\}\}\_\{i\}is entirely encompassed within𝒅i\\boldsymbol\{d\}\_\{i\}and does not introduce extra knowledge\. Finally, all qualified data points are re\-formatted and incorporated into the context refinement dataset𝓓c=\{𝒅~j,𝒅i,𝒒i\}i=1M\\boldsymbol\{\\mathcal\{D\}\}\_\{\\text\{c\}\}=\\\{\\boldsymbol\{\\tilde\{d\}\}\_\{j\},\\boldsymbol\{d\}\_\{i\},\\boldsymbol\{q\}\_\{i\}\\\}\_\{i=1\}^\{M\}\.

Given𝓓c\\boldsymbol\{\\mathcal\{D\}\}\_\{\\text\{c\}\}, we train the model𝓜π\\boldsymbol\{\\mathcal\{M\}\}\_\{\\pi\}to enable its ability of context refinement using supervised fine\-tuning \(SFT\) as follows\.

π∗\\displaystyle\\pi^\{\*\}=arg⁡minπ⁡ℒπSFT\\displaystyle=\\arg\\min\_\{\\pi\}\\mathcal\{L\}^\{\\text\{SFT\}\}\_\{\\pi\}\(1\)ℒπSFT\\displaystyle\\mathcal\{L\}^\{\\text\{SFT\}\}\_\{\\pi\}=−1M​∑i=1M𝔼\(𝒅~i,𝒅i,𝒒i\)∼𝓓c​ℒπ\(i\)\\displaystyle=\-\\frac\{1\}\{M\}\\sum\_\{i=1\}^\{M\}\\mathbb\{E\}\_\{\(\\boldsymbol\{\\tilde\{d\}\}\_\{i\},\\boldsymbol\{d\}\_\{i\},\\boldsymbol\{q\}\_\{i\}\)\\sim\\boldsymbol\{\\mathcal\{D\}\}\_\{c\}\}\\mathcal\{L\}^\{\(i\)\}\_\{\\pi\}\(2\)ℒπ\(i\)\\displaystyle\\mathcal\{L\}^\{\(i\)\}\_\{\\pi\}=log⁡𝓜π​\(𝒅~i∣𝒅i,𝒒i\)\\displaystyle=\\log\\boldsymbol\{\\mathcal\{M\}\}\_\{\\pi\}\(\\boldsymbol\{\\tilde\{d\}\}\_\{i\}\\mid\\boldsymbol\{d\}\_\{i\},\\boldsymbol\{q\}\_\{i\}\)\(3\)𝓜π\\boldsymbol\{\\mathcal\{M\}\}\_\{\\pi\}is trained to generate refined documents𝒅~i\\boldsymbol\{\\tilde\{d\}\}\_\{i\}, yielding a context refiner𝓕=𝓜π∗\\boldsymbol\{\\mathcal\{F\}\}=\\boldsymbol\{\\mathcal\{M\}\}\_\{\\pi^\{\*\}\}for dynamic context refinement in multi\-turn search agents222The teacher model𝓜T\\boldsymbol\{\\mathcal\{M\}\}\_\{T\}in this work is GPT\-4 and the base model𝓜π\\boldsymbol\{\\mathcal\{M\}\}\_\{\\pi\}of context refiner𝓕\\boldsymbol\{\\mathcal\{F\}\}is Qwen2\.5\-7b\-Instruct or Qwen2\.5\-3b\-Instruct, which is the same as their respective inference models\.\.

#### Context Refinement Methods

We employ several existing context refinement methods as comparisons: 1\) We utilize the compression method by introducing GPT\-4\(gpt4\)to summarize the previous contexts \(GPT\-Compress\); 2\) We employ GPT\-4 to dynamically refine the latest Top\-KKretrieved documents𝒅i\\boldsymbol\{d\}\_\{i\}in each round based on the search query𝒒i\\boldsymbol\{q\}\_\{i\}\(GPT\-Refine\)\. 3\) We also employ the foundation LLM itself𝓜π\\boldsymbol\{\\mathcal\{M\}\}\_\{\\pi\}to dynamically refine the latest documents𝒅i\\boldsymbol\{d\}\_\{i\}using𝒒i\\boldsymbol\{q\}\_\{i\}\(Self\-Refine\)\.

#### Results and Analysis

As in Table[2](https://arxiv.org/html/2608.10743#S3.T2)and Figure[5](https://arxiv.org/html/2608.10743#S3.F5),IRCoTbaseline is also presented for intuitive comparisons of context interference mitigation\.GPT\-Refineconsistently outperformsGPT\-Compressionin terms of reliability, while being slightly inferior in search times and context length, indicating that extracting and preserving search query\-relevant key information is more effective in mitigating contextual interference than general summarization\. These two baselines can effectively reduce the retrieval wastes of search agents on irrelevant information in context overIRCoT\.Self\-Refine, by contrast, does not yield performance improvements and even fails to reduce context length, which can be attributed to the lack of information extraction capability of the foundation model𝓜π\\boldsymbol\{\\mathcal\{M\}\}\_\{\\pi\}\. In comparison,Context Refinertrained on the context refinement dataset achieves close performance toGPT\-Refine, suggesting that even relatively weak LLMs can also acquire the ability to refine context through fine\-tuning, without relying on external models for search agents\. In addition, Context Refiner achieves marginally competitive performance to prompt\-driven GPT\-Refine which represents the performance of the teacher model\. This suggests that Context Refiner can refine context to mitigate context interference and outperforms other compression and self\-refine baselines\.

## 4Context Refinement for Search Agent Training

![Refer to caption](https://arxiv.org/html/2608.10743v1/x4.png)Figure 6:Demonstration of our proposed CRRL method\.To further explore the potential of mitigating context interference in search agent training pipelines \(asRQ iii\), we extend our context refiner to the RL training pipeline of search agents, proposing a novelContext\-RefinedReinforcementLearning \(CRRL\) framework to dynamically refine context during training rollouts, reducing context interference in trajectory quality as follows\.

### 4\.1Context\-Refined Reinforcement Learning

Reinforcement learning \(RL\)\(kaelbling1996reinforcement\)has emerged as a paradigm for search agent training\(song2025r1searcherincentivizingsearchcapability;jin2025search;chen2025learning\)using PPO\(schulman2017proximal\)and GRPO\(shao2024deepseekmath\)\. During rollouts, these RL algorithms enable LLMs to iteratively interact with search engines and append the retrieved documents to contexts, potentially introducing interference in RL training and resulting in suboptimal performance of RL\. Nonetheless, current RL methods overlook the impact of context interference of rollouts\. Derived from Sec\.[3](https://arxiv.org/html/2608.10743#S3), we have figured out that removing previous documents and search queries, as well as leveraging context refiner for the latest documents in contexts, can eliminate context interference for multi\-turn search agents, which can also improve the quality of rollouts\. Therefore, we propose the CRRL algorithm\.

Current RL pipelines for search agents are mainly based on Proximal Policy Optimization \(PPO\)\(schulman2017proximal\)and Group Relative Policy Optimization \(GRPO\)\(shao2024deepseekmath\)\. GRPO performs multiple rollouts per task and calculates the relative reward within the group as the advantage, which is more lightweight without the value model and demonstrates comparable performance with PPO\(jin2025search\)\. Hence, this paper adopts GRPO as the default RL algorithm, and the proposed CRRL is also based on GRPO\.

To mitigate context interference of GRPO rollouts and obtain high\-quality trajectories, our CRRL dynamically refines the input context during rollout of multi\-turn search agents with the context refiner𝓕\\boldsymbol\{\\mathcal\{F\}\}\. As shown in Figure[6](https://arxiv.org/html/2608.10743#S4.F6), during rollout, the trajectories of CRRL only contain thinking steps and refined documents, which indicates that the action𝒂ij=\[𝒑ij,𝒒ij\]\\boldsymbol\{a\}^\{j\}\_\{i\}=\[\\boldsymbol\{p\}^\{j\}\_\{i\},\\boldsymbol\{q\}^\{j\}\_\{i\}\]is produced with the context𝒔ij=\[𝒙,𝒑0:i−1j,𝒒i−1j,𝒅~ij\]\\boldsymbol\{s\}\_\{i\}^\{j\}=\[\\boldsymbol\{x\},\\boldsymbol\{p\}^\{j\}\_\{0:i\-1\},\\boldsymbol\{q\}^\{j\}\_\{i\-1\},\\boldsymbol\{\\tilde\{d\}\}^\{j\}\_\{i\}\]where𝒅~ij=𝓕​\(𝒅ij\),𝒅ij=𝓔​\(𝒒i−1j\)\\boldsymbol\{\\tilde\{d\}\}^\{j\}\_\{i\}=\\boldsymbol\{\\mathcal\{F\}\}\(\\boldsymbol\{d\}\_\{i\}^\{j\}\),\\boldsymbol\{d\}\_\{i\}^\{j\}=\\boldsymbol\{\\mathcal\{E\}\}\(\\boldsymbol\{q\}\_\{i\-1\}^\{j\}\)\. RL pipelines for search agents explicitly incorporate retrieval interleaved reasoning, and the token\-level losses are only computed over the rollouts of LLM\-generated tokens, including both search queries and thinking steps, where loss masking is introduced for retrieved tokens, ensuring the stabilization of training while preserving the ability to adaptively retrieve\. When optimizing the policy𝓜π\\boldsymbol\{\\mathcal\{M\}\}\_\{\\pi\}on𝒂i=\[𝒑i,𝒒i\]\\boldsymbol\{a\}\_\{i\}=\[\\boldsymbol\{p\}\_\{i\},\\boldsymbol\{q\}\_\{i\}\], the CRRL algorithm can be represented as follows:

ℒ𝓜πCRRL=−𝔼\{𝒂ij\}j=1G∼𝓜ref\(⋅\|𝒔ij\)​\[𝒢π−β​KL\]\\displaystyle\\mathcal\{L\}^\{\\textrm\{CRRL\}\}\_\{\\boldsymbol\{\\mathcal\{M\}\}\_\{\\pi\}\}=\-\\mathbb\{E\}\_\{\\\{\\boldsymbol\{a\}\_\{i\}^\{j\}\\\}\_\{j=1\}^\{G\}\\sim\{\\boldsymbol\{\\mathcal\{M\}\}\}\_\{\\text\{ref\}\}\(\\cdot\|\\boldsymbol\{s\}\_\{i\}^\{j\}\)\}\\left\[\\mathcal\{G\}\_\{\\pi\}\-\\beta\\text\{KL\}\\right\]\(4\)𝒢π=1G​∑j=1G1∑k=1N−1\|ai,kj\|​\[ℛπj\]\\displaystyle\\mathcal\{G\}\_\{\\pi\}=\\frac\{1\}\{G\}\\sum\_\{j=1\}^\{G\}\\frac\{1\}\{\\sum\_\{k=1\}^\{N\-1\}\|\{a\}\_\{i,k\}^\{j\}\|\}\\left\[\\mathcal\{R\}^\{j\}\_\{\\pi\}\\right\]\(5\)ℛπ,i,kj=\\displaystyle\\mathcal\{R\}^\{j\}\_\{\\pi,i,k\}=∑kN−1min⁡\(rπ,i,kj​Aj,clip​\(rπ,i,kj,1−ϵ,1\+ϵ\)​Aj\)\\displaystyle\\sum\_\{k\}^\{N\-1\}\\min\\left\(r\_\{\\pi,i,k\}^\{j\}A\_\{j\},\\text\{clip\}\\left\(r\_\{\\pi,i,k\}^\{j\},1\-\\epsilon,1\+\\epsilon\\right\)A\_\{j\}\\right\)\(6\)rπ,i,kj=𝓜π​\(ai,kj\|𝒙,𝒂i,<kj,𝒔ij\)𝓜ref​\(ai,kj\|𝒙,𝒂i,<kj,𝒔ij\)\\displaystyle r\_\{\\pi,i,k\}^\{j\}=\\frac\{\{\\boldsymbol\{\\mathcal\{M\}\}\}\_\{\\pi\}\(a^\{j\}\_\{i,k\}\|\\boldsymbol\{x\},\\boldsymbol\{a\}^\{j\}\_\{i,<k\},\\boldsymbol\{s\}\_\{i\}^\{j\}\)\}\{\\boldsymbol\{\\mathcal\{M\}\}\_\{\{\\mathrm\{ref\}\}\}\(a^\{j\}\_\{i,k\}\|\\boldsymbol\{x\},\\boldsymbol\{a\}^\{j\}\_\{i,<k\},\\boldsymbol\{s\}\_\{i\}^\{j\}\)\}\(7\)where𝓜ref\\boldsymbol\{\\mathcal\{M\}\}\_\{\{\\mathrm\{ref\}\}\}represents reference model\. The termϵ\\epsilonis a clipping ratio\.β\\betais the coefficient for the KL divergence\. The advantage estimate is computed on the group\-relative rewards of trajectory𝝉j\\boldsymbol\{\\tau\}^\{j\}as

Aj=rj−μjσj\\displaystyle A\_\{j\}=\\frac\{r^\{j\}\-\\mu^\{j\}\}\{\\sigma^\{j\}\}\(8\)whererj=𝓡​\(𝒚j\)r^\{j\}=\\boldsymbol\{\\mathcal\{R\}\}\(\\boldsymbol\{y\}^\{j\}\)is the final reward, andμj\\mu^\{j\}andσj\\sigma^\{j\}denote the mean and standard deviation of the rewards within the group\.

Table 3:Performance results of EM/ART across QA test sets on several baselines as well as our proposed CRRL method, measured by Exact Match \(EM\) and Average Retrieval Times \(ART\)\.
### 4\.2Training Setup

We sample 40k data from the training sets of NQ and HotpotQA and merge them into𝓓t\\boldsymbol\{\\mathcal\{D\}\}\_\{t\}where𝓓t∩𝓓d=∅\\boldsymbol\{\\mathcal\{D\}\}\_\{t\}\\cap\\ \\boldsymbol\{\\mathcal\{D\}\}\_\{d\}=\\varnothing\. Evaluation is conducted on seven QA datasets to assess both in\-domain and out\-of\-domain performance\. The GRPO implementation is based on Verl\(sheng2025verl\)\. More implementation details can be found in Appendix[D](https://arxiv.org/html/2608.10743#A4)\.

We employ two types of baselines as comparisons\. 1\)Retrieval\-free Fine\-Tuning: We train LLMs using both supervised fine\-tuning\(SFT\)and\(GRPO\)\-based RL\(shao2024deepseekmath\)methods which only contain reasoning and answer steps\. 2\)Retrieval\-based Fine\-Tuning: To obtain the trajectories with the retriever, we utilize rejection sampling to generate several candidate responses by interacting with the search engine for each question from the training set𝓓t\\boldsymbol\{\\mathcal\{D\}\}\_\{t\}\. Then we collect paths including correct answers for rejection fine\-tuning\(RFT\)\(yuan2023scalingrelationshiplearningmathematical\)\. We also employ the RL pipeline for multi\-turn search agents inSearch\-GRPOfollowing\(jin2025search\)andSearch\-o1following\(li\-etal\-2025\-search\), which optimizes LLM rollouts to autonomously call the retriever when they lack relevant knowledge\.

### 4\.3Results and Analysis

Table 4:Efficiency comparisons of baselines and our proposed CRRL during inference on test sets of average retrieval times \(ART\), average context length \(Len\.\), and average inference time per question \(AIT\)\.In Table[3](https://arxiv.org/html/2608.10743#S4.T3), experimental results demonstrate that our proposed CRRL method outperforms other baselines in EM scores, suggesting that incorporating context refinement into the RL training pipeline of search agents can effectively mitigate context interference and improve the reliability of generations\. In Table[4](https://arxiv.org/html/2608.10743#S4.T4), during the inference phase, our proposed CRRL achieves a significant reduction in both ART and Len\. compared to IRCoT and Search\-GRPO, which compensate for context refiner overhead arising from the extra inference costs, and result in lower AIT, suggesting that introducing a context refiner during RL training to relax contextual interference can finally improve the efficiency of the search agent in inference\.

## 5Conclusion

This work investigates the context interference issue on search agents across a variety of QA benchmarks\. We first demonstrate that context interference, largely stemming from the latest retrieved documents, poses a key challenge to multi\-turn search agents\. Then, we present a distill\-based context refiner to dynamically mitigate interference in multi\-turn search agents and thus significantly boost both reliability and efficiency\. Furthermore, we introduce context refinement into the RL training pipelines of search agents, which can further yield performance improvements\. These findings highlight the importance of context refinement to mitigate context interference to construct reliable and efficient search agents, paving the way for a new paradigm of dynamic “refine context and then generate” for AI agents in future work\.

## Limitations

The limitations and future work of this study are listed as follows:

#### Task Settings

This study mainly focuses on mitigating the context interference issue on search agent tasks\. However, similar problems also arise in other agent settings, such astool useandplanning\. The principal factors of context interference in various tasks may differ, necessitating specific mitigation strategies to further improve reliability and efficiency of AI agents\.

#### Paradigm Design

The context refiner proposed in this study is implemented as an auxiliary module to the search agent, rather than being integrated into the training pipelines of agents\. In future work, we plan to develop a dedicated training algorithm that internalizes the context refinement capability within agents themselves\. This may inspire a new paradigm of “receiving observation→\\rightarrowrefining context→\\rightarrowgenerating action” for agents, which can dynamically eliminate the context interference, achieving more efficient and reliable AI agents\.

## Acknowledgments

This work is partially supported by Hong Kong RGC GRF No\. 14206324\.

## References

## Appendix AProtocols

The definitions of the notations are summarized in Table[6](https://arxiv.org/html/2608.10743#A3.T6)\.

Table 5:Summarized notations in this work\.
## Appendix BRelated Work

### B\.1LLM\-based Search Agent

Although LLMs exhibit impressive capabilities, they often lack updated or domain\-specific knowledge\(peng2023study;li2023large\), which undermines LLMs’ reliability\. Therefore, search engines\(zhao2024dense\)are widely integrated to provide external evidence\. A common paradigm is retrieval\-augmented generation \(RAG\)\(gao2024retrievalaugmentedgenerationlargelanguage;lewis2020retrieval\), in which a search engine retrieves documents relevant to the search query and feeds them into LLMs\. More recent work treats search engines as interactive tools\(schick2023toolformer\), prompting or fine\-tuning LLMs to act as search agents\(xiong2025rag;cogkernal;webaggregator;li2026cso\)\. Approaches such as IRCoT\(trivedi2023interleaving\)and ReAct\(yaoreact\)employ prompting to interleave reasoning with iterative search calls\. Search\-R1\(jin2025search\)optimizes LLMs to produce high\-quality trajectories through real\-time, multi\-turn search interactions using RL\. Despite these advances, prior studies have largely focused on eliciting LLMs to output accurate knowledge in reasoning paths despite given lengthy contexts with noise, while overlooking the contextual interference introduced by multi\-turn search interactions, which may cascade across subsequent actions, degrading the reliability and efficiency of the final answers for search agents\.

### B\.2Contextual Interference

LLMs are highly sensitive to input contexts, leaving them vulnerable to noise or irrelevant content that can degrade output quality\(xie\-etal\-2024\-ask;prompt2025amirhossein;webaggregator\)\. Existing strategies to mitigate this issue fall into three main categories\.\(1\) Key Information Extractionidentifies and preserves the most critical or relevant information in contents as in dialogue state tracking\(jacqmin\-etal\-2022\-follow\)or RAG reranking\(glass\-etal\-2022\-re2g;nguyen2025maragmultiagentretrievalaugmentedgeneration;yu2024rankragunifyingcontextranking\);\(2\) Compression methodcondenses lengthy input sequences into summaries or latent states to weaken noise\(jiang2023llmlingua;yi2025surveyrecentadvancesllmbased;li2025singleturnsurveymultiturninteractions\)\.\(3\) Prompt\-based methodexplicitly instructs LLMs to disregard irrelevant content\(rajeev2025catsconfusereasoningllm\)\. However, Key Information Extraction relies on task\-specific schema, Compression may distort key information and incur extra compression cost, and Prompt\-based methods are sensitive to prompt design\. More importantly, these issues are exacerbated in multi\-turn search\-agent scenarios, where the large volume of retrieved documents after several turns introduces more noise and irrelevant information\.

### B\.3Reinforcement Learning for Agent

Reinforcement Learning \(RL\)\(kaelbling1996reinforcement\)has emerged as a paradigm for LLM post\-training or alignment\(ouyang2022training\)\. A variety of RL algorithms have been introduced, like PPO\(schulman2017proximal\)and GRPO\(shao2024deepseekmath\)\. With specific environments and reward designs, LLMs can evolve as autonomous agents capable of adaptive decision\-making and interactions with the environment\. One representative application is Search Agent\(song2025r1searcherincentivizingsearchcapability;jin2025search;chen2025learning\), which interacts with search engines to iteratively gather external knowledge into its reasoning, and thereby performs knowledge\-intensive tasks more effectively\. However, current RL research focuses on deriving the optimal action from rollout trajectories while overlooking the potential influence of context interference within trajectories—particularly in search agents—thereby constraining the achievable performance of RL algorithms\.

## Appendix CDataset Details

Experiments are conducted to evaluate the performance of search agents on various closed\-book QA datasets, which necessitate extra retrieval to address, encompassing both single\- and multi\-hop scenarios\.Single\-hop QAincludes: 1\)Natural Questions \(NQ\)\(kwiatkowski2019natural\), which is constructed by Google Search queries along with annotated short answers; 2\)TriviaQA\(joshi2017triviaqa\), which contains closed\-book trivia QA pairs to gauge models’ factual knowledge; and 3\)PopQA\(mallen2022not\), which consists of entity\-centric QA pairs converted from a knowledge tuple retrieved in Wikidata\.Multi\-hop QAincludes: 1\)HotpotQA\(yang2018hotpotqa\), the first large\-scale dataset requiring reasoning across multiple Wikipedia paragraphs; 2\)2WikiMultiHopQA \(2Wiki\)\(ho2020constructing\), which provides evidence information containing a reasoning path for multi\-hop questions; 3\)MuSiQue\(trivedi2022musique\), which features more difficult 2\-4 hop questions; and 4\)Bamboogle\(press2022measuring\), which is made up only of complex questions that Google answers incorrectly\. Dataset statistics of seven test sets are presented in Table[6](https://arxiv.org/html/2608.10743#A3.T6)\.

Table 6:Data statistics of questions in seven test sets\.![Refer to caption](https://arxiv.org/html/2608.10743v1/figures/recall_com.png)Figure 7:Demonstration of how contextual interference affects the performance of LLM search agents\. “recall rate”=Nr/NN\_\{\\text\{r\}\}/Ndenotes the proportion of questions for which the retrieved documents contain the correct answer \(NrN\_\{\\text\{r\}\}\) among all questions \(NN\)\. “recall accuracy”=Nrc/NN\_\{\\text\{rc\}\}/Nrefers to the proportion of correctly answered questions \(NrcN\_\{\\text\{rc\}\}\) inNrN\_\{\\text\{r\}\}, relative to all questions \(NN\)\. “accuracy” represents the proportion of all correctly answered questionsNcN\_\{\\text\{c\}\}out of the questions \(NN\)\.### C\.1Concerns about Data Contamination

We have carefully considered the concern about Data contamination during the experiments and verified that the external knowledge base𝑲b\\boldsymbol\{K\}\_\{b\}will not contaminate the internal knowledge𝑲I\\boldsymbol\{K\}\_\{I\}as follows\.

In our experiments, the timeline of test sets is synchronized with the wiki dump, so𝑲b\\boldsymbol\{K\}\_\{b\}is regarded as the ground\-truth knowledge\. Although there may be partial knowledge conflicts between𝑲b\\boldsymbol\{K\}\_\{b\}and𝑲I\\boldsymbol\{K\}\_\{I\}due to knowledge updates, questions in test sets usually have clear timeline information as presented below\.

> HotpotQA: The 1895/96 Football League season was the eighth in Football League history with Everton, their Goodison Park home, is a football stadium located in Walton, Liverpool, in which country?

> PopQA: In the 80s who wrote the novel Empire of The Sun?

Therefore,𝑲b\\boldsymbol\{K\}\_\{b\}will serve as a ground truth and will not contaminate𝑲I\\boldsymbol\{K\}\_\{I\}\. A very small number of ambiguous questions can be ignored and will not affect the performance evaluation\. The search agent framework and evaluation setting of this work are implemented based on Search\-R1, which is reliable and widely adopted by a series of studies\.

## Appendix DTraining Details

For GRPO training, we set the policy LLM learning rate to 1e\-6 and sample 4 responses per prompt, following the GRPO implementation in Verl\(sheng2025verl\)\. The batch size is set at 32, with a mini\-batch size of 8 and a micro\-batch size of 4\. The maximum input sequence length and generation length are set to 2048 and 500 respectively\. We enable gradient checkpointing and use Fully Sharded Data Parallel \(FSDP\) with CPU offloading\. For efficient LLM rollouts, we adopt vLLM\(kwon2023efficient\)with a tensor parallel size of 1 and GPU memory utilization ratio of 0\.7\. The rollout sampling temperature is set to 0\.7 and the top\-p value to 1\.0\. The KL divergence regularization coefficientβ\\betaand clip ratioϵ\\epsilonare set to 0\.001 and 0\.2\. The maximum action budgetBBis set to 8\. In cases where training diverges, we evaluate at the most recent stable checkpoint according to the training reward curve; otherwise, the final checkpoint is used for evaluation\.

For training\-based baselines, due to the computational resource limitation with only 4×40G A100 GPU cards, fine\-tuning on the all corpus of 160k samples of the training corpus is expensive\. Therefore, we employ a subset with 60k samples randomly sampled from the original training corpus\. All other training settings are maintained\. Experiments conducted using CRRL are to validate the effectiveness of incorporating context refinement in RL pipelines for search agent training, which is regardless of the training corpus quantity\. We will further add and clarify details of the training setting differences in the final version of the manuscript\.

## Appendix EExperiments

### E\.1Context Interference Effects

We have presented the results of context interference effects on “recall rate” and “recall accuracy” of context refinement and baselines in Table[7](https://arxiv.org/html/2608.10743#A5.T7)as follows\.

Table 7:Context interference effects on “recall rate” and “recall accuracy” of context refinement and baseline methods, measured by Exact Match \(EM\)\.
### E\.2Alternative Ranking Baselines

#### Baseline Setting

To avoid the interference derived from the irrelevant/distracting search results, we employ alternative ranking baselines\. In our experiments, the retriever will consistently return the top\-3 highest\-scoring documents, which may contain irrelevant documents with relatively lower retrieval scores\. The retrieval score ranges from \[0, 1\]\. Therefore, we set different thresholds for retrieval scores to filter out irrelevant documents\. We report the performance \(EM/ART\) of the ranking baseline with other context refinement methods in Table[8](https://arxiv.org/html/2608.10743#A5.T8)for interference mitigation as follows\. EM and ART denote the Exact Match \(EM\) and Average Retrieval Times \(ART\), respectively\.

Table 8:Performance results of EM/ART across QA test sets on several baselines as well as our proposed CRRL method, measured by Exact Match \(EM\) and Average Retrieval Times \(ART\)\.
#### Analysis

The ranking baselines marginally outperform the IRCoT but underperform other context refinement methods, which can be attributed that

1\. Ranking baselines can effectively filter out irrelevant documents to mitigate context interference but can also remove the ground\-truth documents\. Therefore, they can not lead to consistent performance improvements in both reliability and efficiency\.

2\. The performance of ranking baselines on different cases rely on the threshold setting, which lacks flexibility compared with other model\-based context refinement methods\.

## Appendix FPrompt Template

> Direct Inference PromptYou are an excellent Question\-Answering assistant\. Please answer the following question based on your knowledge\. You can directly provide the answer inside<answer\>and</answer\>, without detailed illustrations\. For example,<answer\>North America</answer\>\. Question: \{question\}CoT PromptYou are an excellent Question\-Answering assistant\. Please answer the following question based on your knowledge\. You must conduct reasoning inside<think\>and</think\>to think step by step first\. You can directly provide the answer inside<answer\>and</answer\>, without detailed illustrations\. For example,<answer\>North America</answer\>\. Question: \{question\}IRCoT PromptYou are an excellent Question\-Answering assistant\. Please answer the following question based on your knowledge\. You must conduct reasoning inside<think\>and</think\>to think step by step first\. After reasoning, if you find you lack some knowledge, you can call a search engine by<search\>and</search\>and it will return the top searched results between<information\>and</information\>\. You can search as many times as your want\. If you find no further external knowledge needed, you can directly provide the answer inside <answer\> and </answer\>, without detailed illustrations\. For example,<answer\>North America</answer\>\. Question: \{question\}

Table 9:Demonstrations of one generation of a search agent given original retrieved documents \(ID 1\) and extracted information \(ID 2\) respectively\.Table 10:Demonstrations of one generation of a search agent given original retrieved documents \(ID 1\) and extracted information \(ID 2\) respectively\.

Similar Articles

Context-Aware RL for Agentic and Multimodal LLMs

Hugging Face Daily Papers

Introduces ContextRL, a reinforcement learning approach that teaches LLMs to identify which context supports an answer, achieving gains on agentic and multimodal benchmarks.

Multi-Turn Reasoning When Context Arrives in Pieces: Scalable Sharding and Memory-Augmented RL

arXiv cs.CL

This paper addresses the 'Lost in Conversation' problem where LLMs struggle with information revealed across multiple turns. It proposes a scalable sharding pipeline to create multi-turn training data from single-turn QA datasets and uses reinforcement learning with verifiable rewards to train a memory-augmented policy that maintains a compact rolling memory, improving multi-turn reasoning accuracy and generalizing zero-shot to harder tasks.

Learning Agent-Compatible Context Management for Long-Horizon Tasks

arXiv cs.AI

Introduces AdaCoM, an external LLM-based context manager for frozen agents, using reinforcement learning to improve long-horizon task performance by preserving task constraints and pruning stale content, with experiments on web search and deep research benchmarks.

Why Retrying Fails: Context Contamination in LLM Agent Pipelines

arXiv cs.AI

This paper introduces the Context-Contaminated Restart Model (CCRM) to formally analyze how failed attempts in LLM agent pipelines contaminate context and increase error rates during retries. It provides theoretical proofs and validates the model against SWE-bench data, showing significant discrepancies with standard independent models.