参数与上下文:面向鲁棒检索增强生成的TRACE微调

arXiv cs.LG 论文

摘要

本文提出TRACE,这是一个用于检索增强生成(RAG)的微调框架,它使用多智能体辩论轨迹和答案完整性正则化来处理知识冲突,从而提高对误导性上下文的鲁棒性并减少不完整答案。

arXiv:2609.30337v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) mitigates knowledge obsolescence and factual hallucination in large language models by introducing external context. However, when retrieved knowledge conflicts with the model's internal parametric knowledge, the model may either blindly follow misleading context or incorrectly rely on parametric knowledge, leading to unreliable responses. To address this issue, this paper proposes TRACE (Debate-TRace and Answer-Completeness rEgularized fine-tuning), a robust fine-tuning framework for RAG under knowledge conflicts. First, we propose a fine-tuning method that leverages multi-agent debate traces to extract correct candidates, incorrect candidates, and answer-shift patterns, providing fine-grained supervision for reliable knowledge-source selection. In addition, we design an answer completeness regularization mechanism to alleviate empty, overly short, and prematurely terminated responses via answer-tail token reinforcement and premature termination suppression. The fine-tuning objective combines correct-answer supervision, incorrect-candidate suppression, answer-tail token reinforcement, and premature termination suppression, enabling the model to use reliable external context, resist misleading or irrelevant retrieved content, and fall back to parametric knowledge when retrieved evidence is unreliable. Experiments across multiple knowledge-conflict scenarios and datasets show that TRACE improves robustness against misleading retrieved knowledge and reduces incomplete answers. These results demonstrate that multi-agent debate traces and answer completeness regularization jointly enhance knowledge-source selection, conflict robustness, and answer quality in RAG models. Our code is available at https://github.com/PHD-lanyu/TRACE.
查看原文
查看缓存全文

缓存时间: 2026/09/29 09:34

# Parameters vs. Context: TRACE Fine-Tuning for Robust Retrieval-Augmented Generation
Source: [https://arxiv.org/html/2609.30337](https://arxiv.org/html/2609.30337)
Zhengchen Huang1, Yundong Sun1,\*, Minrui Song2, Shuanglong Yao1, Ye Liu1, Ji Chen1, Xing Wang1Affiliation:Affiliation:1School of Information Science & Engineering, LinYi University, ChinaAffiliation:Affiliation:2School of Computer Science and Technology, Harbin Institute of Technology, ChinaAffiliation:Affiliation:\*Corresponding author: Yundong Sun, hitffmy@163\.com

###### Abstract

Retrieval\-Augmented Generation \(RAG\) mitigates knowledge obsolescence and factual hallucination in large language models by introducing external context\. However, when retrieved knowledge conflicts with the model’s internal parametric knowledge, the model may either blindly follow misleading context or incorrectly rely on parametric knowledge, leading to unreliable responses\. To address this issue, this paper proposes TRACE \(Debate\-TRace andAnswer\-Completeness rEgularized fine\-tuning\), a robust fine\-tuning framework for RAG under knowledge conflicts\. First, we propose a fine\-tuning method that leverages multi\-agent debate traces to extract correct candidates, incorrect candidates, and answer\-shift patterns, providing fine\-grained supervision for reliable knowledge\-source selection\. In addition, we design an answer completeness regularization mechanism to alleviate empty, overly short, and prematurely terminated responses via answer\-tail token reinforcement and premature termination suppression\. The fine\-tuning objective combines correct\-answer supervision, incorrect\-candidate suppression, answer\-tail token reinforcement, and premature termination suppression, enabling the model to use reliable external context, resist misleading or irrelevant retrieved content, and fall back to parametric knowledge when retrieved evidence is unreliable\. Experiments across multiple knowledge\-conflict scenarios and datasets show that TRACE improves robustness against misleading retrieved knowledge and reduces incomplete answers\. These results demonstrate that multi\-agent debate traces and answer completeness regularization jointly enhance knowledge\-source selection, conflict robustness, and answer quality in RAG models\. Our code is available at[TRACE](https://github.com/PHD-lanyu/TRACE)\.

###### Index Terms:

Retrieval\-Augmented Generation; knowledge conflict; multi\-agent debate; fine\-tuning; large language models

## IIntroduction

Large Language Models \(LLMs\), such as GPT\[[1](https://arxiv.org/html/2609.30337#bib.bib1)\]and Llama\[[2](https://arxiv.org/html/2609.30337#bib.bib3)\], have become strong backbones for open\-domain question answering, knowledge reasoning, and text generation\. Despite this progress, their responses still depend on parametric knowledge that can be outdated, incomplete, or incorrect\[[3](https://arxiv.org/html/2609.30337#bib.bib4)\], which may lead to factual hallucinations\[[4](https://arxiv.org/html/2609.30337#bib.bib5)\]and unreliable answers\. Retrieval\-Augmented Generation \(RAG\)\[[5](https://arxiv.org/html/2609.30337#bib.bib6)\]mitigates this problem by conditioning generation on retrieved documents or knowledge\-base content, thereby allowing models to use external evidence beyond their fixed parameters\[[6](https://arxiv.org/html/2609.30337#bib.bib7),[7](https://arxiv.org/html/2609.30337#bib.bib8)\]\.

The central difficulty in RAG is that retrieved context can conflict with the model’s parametric knowledge\. In practical retrieval pipelines, external context may be correct, misleading, self\-conflicting, or irrelevant\[[8](https://arxiv.org/html/2609.30337#bib.bib9)\]\. As shown in Fig\.[1](https://arxiv.org/html/2609.30337#S1.F1), conventional RAG prompting becomes unstable when the retrieved context is unreliable: incorrect or irrelevant context can mislead the model and even make it perform worse than query\-only prompting, while self\-conflicting context provides only limited gains\. These observations indicate that LLMs do not inherently know when to trust retrieved evidence, when to rely on parametric knowledge, or how to revise an answer when the two sources disagree\[[9](https://arxiv.org/html/2609.30337#bib.bib10)\]\. Therefore, robust RAG requires source\-aware generation: the model must use reliable retrieved evidence, resist misleading context, and preserve correct parametric knowledge when retrieval is unreliable\.

![Refer to caption](https://arxiv.org/html/2609.30337v1/fig1.png)

Fig\. 1:Performance under different knowledge\-conflict contexts on ConFiQA\-MR, ConFiQA\-SC, and ExplainPE\.Existing work improves RAG under knowledge conflicts from three main directions: prompt\-based knowledge integration\[[10](https://arxiv.org/html/2609.30337#bib.bib11)\], retrieval\-quality assessment\[[11](https://arxiv.org/html/2609.30337#bib.bib13)\], and learnable knowledge\-selection strategies\[[12](https://arxiv.org/html/2609.30337#bib.bib16)\]\. Prompt\-based methods guide models to compare internal knowledge with external context\[[13](https://arxiv.org/html/2609.30337#bib.bib12)\]; retrieval\-quality assessment methods identify, filter, or correct low\-quality retrieved content\[[14](https://arxiv.org/html/2609.30337#bib.bib14)\]; and learnable selection methods use decoding control, preference optimization, or reinforcement learning to choose among information sources under conflicting conditions\[[15](https://arxiv.org/html/2609.30337#bib.bib35)\]\. However, these methods still face two limitations\. First, many methods introduce additional prompts, reflection stages, retrieval evaluation modules, or multi\-step reasoning at test time, which increases inference overhead\. Second, their training signals are usually constructed from final answers, decoding distributions, or sampled rewards, so they rarely preserve the intermediate candidate answers, erroneous paths, and answer revisions that appear during conflict reasoning\. As a result, the supervision is often too coarse to teach how a model should move from a misleading answer to a reliable one\.

Multi\-agent debate provides a natural source of process\-level supervision for this problem\. Prior studies show that interaction, questioning, and revision among multiple model instances can improve factual judgment and complex reasoning\[[16](https://arxiv.org/html/2609.30337#bib.bib2),[17](https://arxiv.org/html/2609.30337#bib.bib17)\]\. Inspired by this observation, we propose TRACE \(Debate\-TRace andAnswer\-Completeness rEgularized fine\-tuning\), a robust fine\-tuning method for RAG under knowledge conflicts\. TRACE uses multi\-agent debate only during training, rather than retaining debate at inference time\. It extracts correct candidate answers, incorrect candidate answers, and answer\-shift traces from debate trajectories, then converts these traces into positive and negative fine\-tuning signals\. In this way, the model learns not only which answer is correct, but also which misleading candidates should be suppressed and how answer revisions occur when parametric and retrieved knowledge conflict\.

Reliable source selection alone is not sufficient, because knowledge\-conflict fine\-tuning can also introduce answer\-incompleteness errors\. In our empirical analysis, some models identify the correct answer direction but stop too early, producing empty answers, overly short answers, incomplete entities, or unclear answer boundaries\. TRACE therefore adds answer completeness regularization, which reinforces answer\-tail tokens and suppresses premature termination inside the answer span\. This auxiliary constraint encourages the model to produce complete and parsable final answers after selecting the appropriate knowledge source\.

The main contributions of this paper are as follows:

1. \(1\)We propose a debate\-trace fine\-tuning method that converts correct candidates, incorrect candidates, and answer\-shift traces into fine\-grained supervision signals, enabling the model to select more reliable knowledge when parametric and retrieved knowledge conflict\.
2. \(2\)We identify an answer\-incompleteness issue that can emerge during knowledge\-conflict fine\-tuning, and design answer completeness regularization to reduce empty, overly short, and incomplete outputs through answer\-tail token reinforcement and premature termination suppression\.
3. \(3\)We conduct experiments across correct, wrong, self\-conflicting, irrelevant, and partially irrelevant retrieval scenarios\. The results show that TRACE improves robustness in explicit conflict settings while preserving the use of correct retrieved knowledge, and ablation studies validate the proposed debate\-trace supervision and answer completeness regularization modules\.

## IIRELATED WORK

### II\-ARAG under Knowledge Conflicts

RAG under knowledge conflicts aims to decide when retrieved context should override, complement, or be ignored in favor of parametric knowledge\. Prompt\-based methods address this problem by explicitly eliciting source comparison or self\-critique during inference\. Astute RAG\[[10](https://arxiv.org/html/2609.30337#bib.bib11)\]elicits internal knowledge and integrates it with retrieved content in a source\-aware manner, Self\-RAG\[[18](https://arxiv.org/html/2609.30337#bib.bib18)\]controls retrieval and critique through reflection tokens, and InstructRAG\[[13](https://arxiv.org/html/2609.30337#bib.bib12)\]uses self\-generated rationales to extract valid evidence from noisy retrieval\. These methods make the reasoning process more interpretable, but they typically depend on additional prompts, reflection steps, or multi\-stage reasoning at test time\.

Retrieval\-quality assessment methods improve RAG by judging or repairing the external context before generation\. CRAG\[[11](https://arxiv.org/html/2609.30337#bib.bib13)\]estimates retrieval reliability and invokes corrective retrieval when the context is unreliable, RobustRAG\[[14](https://arxiv.org/html/2609.30337#bib.bib14)\]studies robustness against retrieval corruption, and TruthfulRAG\[[19](https://arxiv.org/html/2609.30337#bib.bib19)\]converts retrieved text into a knowledge graph to identify factual\-level conflicts\. These methods reduce the risk of misleading retrieval, but their main focus is obtaining or filtering better context rather than training the generator to choose between parametric and retrieved knowledge\. In contrast, TRACE assumes that the retrieved context may remain correct, wrong, self\-conflicting, or irrelevant, and trains the model to make the source\-selection decision inside generation\.

### II\-BLearnable Knowledge\-Source Selection

Learnable knowledge\-selection methods are the closest line of work to TRACE because they directly optimize how a model uses parametric and contextual knowledge\. CK\-PlUG\[[12](https://arxiv.org/html/2609.30337#bib.bib16)\]controls reliance on parametric and contextual knowledge during decoding, Context\-DPO\[[20](https://arxiv.org/html/2609.30337#bib.bib20)\]improves context faithfulness through preference pairs, KnowPO\[[21](https://arxiv.org/html/2609.30337#bib.bib21)\]formulates knowledge\-aware preference optimization, and Knowledgeable\-R1\[[22](https://arxiv.org/html/2609.30337#bib.bib15)\]strengthens resistance to contextual interference through parametric\-knowledge reinforcement\. These methods provide important baselines for source\-aware RAG, but their supervision is mainly derived from final answers, sampled responses, reward signals, or decoding\-level control\. TRACE differs by using debate trajectories to preserve correct candidates, incorrect candidates, and answer shifts, so the fine\-tuning data contains process\-level evidence about how answers become wrong or get corrected under conflict\.

Existing conflict\-oriented RAG work also pays limited attention to answer completeness after source selection\. Most prior methods evaluate whether the model follows the appropriate knowledge source or produces a factually correct final answer, whereas our observations show that knowledge\-conflict fine\-tuning can produce empty answers, overly short answers, incomplete entities, and unclear answer boundaries\. TRACE therefore treats final\-answer completeness as a separate generation\-side objective, complementing source selection with answer\-tail token reinforcement and premature\-termination suppression\.

### II\-CMulti\-agent Debate as Process\-level Supervision

Multi\-agent debate exposes intermediate errors, corrections, and answer revisions that are difficult to observe from a single final response\. Liang et al\.\[[23](https://arxiv.org/html/2609.30337#bib.bib22)\]introduced a debate framework in which multiple agents discuss a problem and a judge determines the final answer\. Subsequent studies show that debate can improve factual judgment and reasoning\[[17](https://arxiv.org/html/2609.30337#bib.bib17),[16](https://arxiv.org/html/2609.30337#bib.bib2)\], support consensus formation among diverse models\[[24](https://arxiv.org/html/2609.30337#bib.bib30)\], and provide trajectories for post\-training or preference optimization\[[25](https://arxiv.org/html/2609.30337#bib.bib31),[26](https://arxiv.org/html/2609.30337#bib.bib32)\]\. These findings suggest that debate records are not only inference\-time reasoning traces, but also potential supervision sources\.

The connection between debate traces and RAG knowledge conflicts remains underexplored\. Knowledge conflicts are common in LLM applications\[[27](https://arxiv.org/html/2609.30337#bib.bib33)\], and recent work studies how to verify or select knowledge under inconsistent evidence\[[28](https://arxiv.org/html/2609.30337#bib.bib34)\]\. However, existing RAG conflict methods usually do not mine debate trajectories for source labels, positive candidates, negative candidates, and answer\-shift patterns\. TRACE fills this gap by using multi\-agent debate only during data construction, then converting the resulting trajectories into conflict\-aware supervision for a single target model\. This design preserves the training value of debate while avoiding debate\-time overhead during deployment\.

## IIIMETHODOLOGY

### III\-ATask Definition

Given a questionQQand a retrieved contextCC, a RAG model generates an answerA=fθ​\(Q,C\)A=f\_\{\\theta\}\(Q,C\)\. The retrieved context may be reliable, misleading, irrelevant, or internally conflicting, and the model may also rely on its parametric knowledge\. This paper focuses on answerable knowledge\-conflict settings in which at least one source can support the golden answer\. When parametric and retrieved knowledge disagree, the model should select the more reliable source; when both sources are reliable, it should maintain stable answer generation rather than over\-suspecting either source\.

### III\-BOverview

TRACE addresses the two methodological gaps identified in the Introduction and Related Work: coarse conflict supervision and incomplete final answers after knowledge\-conflict fine\-tuning\. As shown in Fig\.[2](https://arxiv.org/html/2609.30337#S3.F2), TRACE first uses multi\-agent debate to expose correct candidates, incorrect candidates, and answer\-shift trajectories under knowledge conflicts\. It then converts these traces into source\-aware positive and negative fine\-tuning records, allowing the target model to learn which knowledge source to trust without running debate at inference time\. Finally, TRACE adds answer completeness regularization to reduce empty, overly short, and prematurely terminated answers\.

![Refer to caption](https://arxiv.org/html/2609.30337v1/overview3.png)

Fig\. 2:Overview of TRACE\. Debate is used only during data construction to expose source\-selection errors and answer shifts\. The target model is then fine\-tuned with debate\-trace supervision and answer completeness regularization\.
### III\-CFine\-tuning Method Based on Multi\-agent Debate Traces

The first component of TRACE constructs process\-level supervision from multi\-agent debate traces\. This component contains three steps: collecting debate traces under retrieved context, converting the traces into source\-labeled training records, and optimizing the target model with positive supervision for reliable answers and unlikelihood suppression for misleading candidates\. This design turns debate from an inference\-time procedure into a training\-time supervision source\.

#### III\-C1Debate Trace Collection

For each training sample, TRACE first queries the base model withQQonly to obtain a query\-only response, which provides an estimate of the model’s parametric\-knowledge behavior\. TRACE then conducts a multi\-round debate with the retrieved contextCC\. The debate has three roles: the response model produces candidate answers, the critic model checks factual consistency, context reliability, and knowledge\-source selection, and the judge model decides whether the debate should stop and selects key information from the final candidates\.

TRACE uses a strong\-to\-weak debate setting to make the trace informative for fine\-tuning\. The response model is the weak target model to be fine\-tuned, the critic model is stronger, and the judge model has intermediate capability\. This configuration encourages the debate to expose misleading candidates, corrections, and answer shifts that are difficult to obtain from a single final response\. Because debate is used only to construct training data, TRACE does not introduce multi\-agent interaction or additional reasoning rounds during deployment\.

#### III\-C2Construction of Source\-Aware Training Records

TRACE converts each debate trace into compact training records instead of directly using the full debate transcript as model input\. Each record is built from the query\-only response, intermediate debate candidates, the final debate answer, and the golden answer\. This representation avoids overfitting to surface debate language and focuses the training signal on source selection, candidate correction, and answer stability\.

Each training record receives a knowledge\-source labelsi∈\{internal,external,both\}s\_\{i\}\\in\\\{\\mathrm\{internal\},\\mathrm\{external\},\\mathrm\{both\}\\\}\. The*internal*label indicates that parametric knowledge supports the golden answer while the retrieved context is unreliable; the*external*label indicates that the retrieved context supports the golden answer while the query\-only response is incorrect; and the*both*label indicates that both sources support the golden answer\. The first two labels correspond to conflict settings, while*both*serves as a non\-conflict anchor that preserves ordinary RAG answering ability\. If neither source supports the golden answer, TRACE excludes the sample because it is closer to unanswerability detection than source\-aware answer generation\.

On top of source labels, TRACE assigns functional type labels according to each constructed record’s training role\. These roles include reliable supervision, context usage, context resistance, correct\-to\-wrong prevention, dual\-source error suppression, and boundary correction\. Intermediate and final debate answers are divided into correct and incorrect candidates according to the golden answer, with at most two representative incorrect candidates retained from each debate trace to reduce noisy negatives\. Table[I](https://arxiv.org/html/2609.30337#S3.T1)summarizes the sample taxonomy\.

TABLE I:Taxonomy of fine\-tuning samples for TRACE\.Knowledge\-source labelFunctional typeConstruction conditionTraining roleinternal / external / both

Reliable supervisionGolden answers or reliable correct candidate answers observed during debate\.Supervises correct answers and preserves basic answering ability\.Correct\-to\-wrong preventionA correct candidate appears in an intermediate round, but the later or final response drifts to an incorrect answer\.Suppresses shifts from correct candidates to incorrect answers and improves answer stability\.Boundary correctionThe candidate is close to the correct answer but is overly short, boundary\-unclear, or difficult to judge\.Improves answer completeness, parsability, and final evaluation determinability\.externalContext usageThe retrieved context supports the golden answer, but the query\-only response is incorrect\.Trains the model to actively use external evidence when it is reliable\.internalContext resistanceThe query\-only response is correct, but the retrieved context contains misleading or irrelevant\.Trains the model to avoid blindly following erroneous retrieved content\.bothDual\-source hard negativeBoth parametric knowledge and retrieved context support the correct answer, but incorrect candidates still appear during debate\.Reduces the probability of abnormal incorrect candidates under dual\-source reliable conditions\.A single debate trace can be expanded into multiple training records because different candidates in the same trace may serve different training roles\. During training, the source labelsis\_\{i\}determines the source\-balance coefficientγsi\\gamma\_\{s\_\{i\}\}, and the functional type labelτi\\tau\_\{i\}determines the fixed sample\-type weightwτ⁡\(i\)w\_\{\\tau\(i\)\}\. The main tunable parameters in this component are the incorrect\-candidate suppression weightλneg\\lambda\_\{\\mathrm\{neg\}\}and the source\-balance coefficientsγint\\gamma\_\{\\mathrm\{int\}\},γext\\gamma\_\{\\mathrm\{ext\}\}, andγboth\\gamma\_\{\\mathrm\{both\}\}\.

#### III\-C3Source\-Aware Positive\-Negative Objective

Letx=\(Q,C\)x=\(Q,C\)denote the input,y\+y^\{\+\}denote a correct answer, andy−y^\{\-\}denote an incorrect candidate answer\. For correct answers, TRACE uses cross\-entropy \(CE\) supervision:

ℒCE\(x,y\+\)=−1T\+∑t=1T\+logpθ\(yt\+∣x,y<t\+\),\\mathcal\{L\}\_\{\\mathrm\{CE\}\}\(x,y^\{\+\}\)=\-\\frac\{1\}\{T^\{\+\}\}\\sum\_\{t=1\}^\{T^\{\+\}\}\\log p\_\{\\theta\}\(y\_\{t\}^\{\+\}\\mid x,y\_\{<t\}^\{\+\}\),\(1\)whereT\+T^\{\+\}is the length of the correct answer\. This term teaches the model to generate reliable answers under each source condition\.

For incorrect candidates, TRACE uses an unlikelihood \(UL\) objective\[[29](https://arxiv.org/html/2609.30337#bib.bib23)\]to reduce the probability of misleading answer content\. To avoid damaging the output template, this loss is computed only inside the answer span:

ℒUL​\(x,y−\)\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{UL\}\}\(x,y^\{\-\}\)=−1∑tmt\+ϵ∑t=1T−mt\\displaystyle=\-\\frac\{1\}\{\\sum\_\{t\}m\_\{t\}\+\\epsilon\}\\sum\_\{t=1\}^\{T^\{\-\}\}m\_\{t\}\(2\)log⁡\(1−pθ​\(yt−∣x,y<t−\)\+ϵ\),\\displaystyle\\log\\left\(1\-p\_\{\\theta\}\(y\_\{t\}^\{\-\}\\mid x,y\_\{<t\}^\{\-\}\)\+\\epsilon\\right\),wheremt∈\{0,1\}m\_\{t\}\\in\\\{0,1\\\}is the answer\-content mask\. The mask equals 1 only for tokens inside the answer span, so structural tags and answer\-boundary tokens are not treated as negative samples\. This design suppresses wrong answer content without encouraging malformed or incomplete answer boundaries\.

At the same time, TRACE also uses lightweight balance modulation as an auxiliary stabilization strategy for cases where parametric knowledge is correct but external retrieval is wrong\. This component follows the motivation of Knowledge Balance Modulation in Knowledgeable\-R1\[[22](https://arxiv.org/html/2609.30337#bib.bib15)\], but it is not treated as the main contribution of this paper\.

During the early stage of fine\-tuning, the loss can be large and later gradually decreases\. Therefore, we maintain an exponential moving average \(EMA\)μt\\mu\_\{t\}to represent the historical loss level of the current training stage and to reduce the influence of large early losses:

μt=\(1−η\)​μt−1\+η​ℓt,\\mu\_\{t\}=\(1\-\\eta\)\\mu\_\{t\-1\}\+\\eta\\ell\_\{t\},\(3\)whereη\\etais the EMA update rate andℓt\\ell\_\{t\}is the training loss of the current sample\. When computing the scaling factor, the historical baseline before the update,μt−1\\mu\_\{t\-1\}, is compared with the current loss:

ri=\{max⁡\(β,μt−1ℓt\+ϵ\),si=internal,1,otherwise\.r\_\{i\}=\\begin\{cases\}\\max\\left\(\\beta,\\dfrac\{\\mu\_\{t\-1\}\}\{\\ell\_\{t\}\+\\epsilon\}\\right\),&\\begin\{array\}\[\]\{l\}s\_\{i\}=\\mathrm\{internal\},\\\\\[\-2\.84526pt\] \\end\{array\}\\\\\[8\.53581pt\] 1,&\\text\{otherwise\}\.\\end\{cases\}\(4\)For samples where internal knowledge is correct but external retrieval is wrong, this scaling moderately reduces overly large penalties and protects the model’s ability to answer with reliable parametric knowledge\. The lower boundβ\\betaprevents such samples from being completely ignored\. For other source labels,ri=1r\_\{i\}=1, so no adjustment is applied\.

Reliable\-supervision records use only the CE term, while conflict\-oriented records use joint positive\-negative fine\-tuning\. The main source\-aware objective is:

ℒmain\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{main\}\}=1N​∑i=1Nωi​\[ℒCE​\(xi,yi\+\)\+λneg​𝕀​\(yi−\)​ℒUL​\(xi,yi−\)\],\\displaystyle=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\omega\_\{i\}\\big\[\\mathcal\{L\}\_\{\\mathrm\{CE\}\}\(x\_\{i\},y\_\{i\}^\{\+\}\)\+\\lambda\_\{\\mathrm\{neg\}\}\\mathbb\{I\}\(y\_\{i\}^\{\-\}\)\\mathcal\{L\}\_\{\\mathrm\{UL\}\}\(x\_\{i\},y\_\{i\}^\{\-\}\)\\big\],\(5\)where𝕀⁡\(yi−\)\\mathbb\{I\}\(y\_\{i\}^\{\-\}\)indicates whether theii\-th record contains an incorrect candidate,λneg\\lambda\_\{\\mathrm\{neg\}\}controls the suppression strength, andωi\\omega\_\{i\}is the sample\-level weight:

ωi=wτ⁡\(i\)⋅γsi⋅ri,\\omega\_\{i\}=w\_\{\\tau\(i\)\}\\cdot\\gamma\_\{s\_\{i\}\}\\cdot r\_\{i\},\(6\)wherewτ⁡\(i\)w\_\{\\tau\(i\)\}is determined by the functional type,γsi\\gamma\_\{s\_\{i\}\}is determined by the knowledge\-source label, andrir\_\{i\}is the stabilization factor for internal\-correct/external\-wrong samples\.

### III\-DAnswer Completeness Regularization

#### III\-D1Motivation: Incomplete Answer Generation

Source\-aware fine\-tuning improves knowledge selection, but it may also exacerbate answer\-incompleteness errors\. In these cases, the model often starts to generate the correct answer but stops at a partial word, a partial entity, or the first fragment of a multi\-word answer\. Such outputs are especially harmful under exact\-match evaluation and also reduce semantic clarity for downstream users\.

TABLE II:Examples of incomplete answers without answer completeness regularization\.QuestionIncomplete outputComplete answerWhat position does the spouse of the head of state of the country where Heath Ledger is a citizen hold?ConsorConsort of the United KingdomWhat position is held by the spouse of the head of government of the country where John F\. Kelly is a citizen?FirstFirst Lady of the United StatesWhat is the historical significance of Fort Dearborn in relation to the creator of*The Mandalorian*?Fort DearFort DearbornTable[II](https://arxiv.org/html/2609.30337#S3.T2)shows that these errors are not always caused by choosing the wrong knowledge source\. Outputs such as “Consor”, “Fort Dear”, and “First” suggest that the model has identified the correct answer direction but terminates before completing the answer\. Therefore, TRACE augments source\-aware fine\-tuning with an auxiliary objective that explicitly encourages complete final\-answer generation\.

#### III\-D2Design of Answer Completeness Regularization

Answer completeness regularization has two terms: answer\-tail token reinforcement and premature termination suppression\. The first term gives stronger supervision to later answer tokens, while the second term discourages termination tokens before the answer span is complete\.

First, answer\-tail token reinforcement strengthens the learning signal inside the answer span\. Let the target output sequence bey=\{y1,y2,…,yT\}y=\\\{y\_\{1\},y\_\{2\},\\ldots,y\_\{T\}\\\}, where the answer content is located in<answer\>\.\.\.</answer\>\. Letmt∈\{0,1\}m\_\{t\}\\in\\\{0,1\\\}indicate whether thett\-th token is inside the answer span, and letat∈\[0,1\]a\_\{t\}\\in\[0,1\]denote the relative position of this token inside the answer span\. A largerata\_\{t\}means that the token is closer to the end of the answer\. The coefficientρ\\rhocontrols the magnitude of answer\-tail reinforcement\. For a single answer, the answer\-span weighted supervision loss is defined as:

ℒans\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{ans\}\}=−1∑t=1Tmt\+ϵ∑t=1Tmt\(1\+ρat\)logpθ\(yt∣x,y<t\)\.\\displaystyle=\-\\frac\{1\}\{\\sum\_\{t=1\}^\{T\}m\_\{t\}\+\\epsilon\}\\sum\_\{t=1\}^\{T\}m\_\{t\}\(1\+\\rho a\_\{t\}\)\\log p\_\{\\theta\}\(y\_\{t\}\\mid x,y\_\{<t\}\)\.\(7\)This means that in a long answer, later tokens receive relatively larger weights than earlier tokens\. The design prevents the model from generating only the answer prefix while ignoring the latter part of multi\-word entities, restrictive phrases, or long proper names, thereby reducing cases such as generating “Consor” instead of “Consort of the United Kingdom”\.

Second, premature termination suppression reduces early stopping before the answer is complete\. LetEEdenote the set of termination tokens, including EOS and</answer\>related tokens\. Inside the answer\-content span, the model should continue generating answer content, so the probability of termination should be reduced\. Letptendp\_\{t\}^\{\\mathrm\{end\}\}be the termination probability at thett\-th position:

ptend=∑e∈Epθ​\(e∣x,y<t\)\.p\_\{t\}^\{\\mathrm\{end\}\}=\\sum\_\{e\\in E\}p\_\{\\theta\}\(e\\mid x,y\_\{<t\}\)\.\(8\)The premature termination suppression loss is:

ℒend=−1∑t=1Tmt\+ϵ∑t=1Tmtlog\(1−ptend\)\.\\mathcal\{L\}\_\{\\mathrm\{end\}\}=\-\\frac\{1\}\{\\sum\_\{t=1\}^\{T\}m\_\{t\}\+\\epsilon\}\\sum\_\{t=1\}^\{T\}m\_\{t\}\\log\(1\-p\_\{t\}^\{\\mathrm\{end\}\}\)\.\(9\)This term is essentially an unlikelihood constraint on termination tokens\. It suppresses early generation of ending tags before the answer body is complete, reducing empty answers, overly short answers, truncated entities, and incomplete answer boundaries\.

The answer completeness regularization for a single answer is:

ℒcomp=ℒans\+λend​ℒend\.\\mathcal\{L\}\_\{\\mathrm\{comp\}\}=\\mathcal\{L\}\_\{\\mathrm\{ans\}\}\+\\lambda\_\{\\mathrm\{end\}\}\\mathcal\{L\}\_\{\\mathrm\{end\}\}\.\(10\)
The final training objective combines source\-aware fine\-tuning and answer completeness regularization:

ℒtotal=ℒmain\+λcomp​ℒcomp,\\mathcal\{L\}\_\{\\mathrm\{total\}\}=\\mathcal\{L\}\_\{\\mathrm\{main\}\}\+\\lambda\_\{\\mathrm\{comp\}\}\\mathcal\{L\}\_\{\\mathrm\{comp\}\},\(11\)whereλcomp\\lambda\_\{\\mathrm\{comp\}\}controls the contribution of answer completeness regularization\.

#### III\-D3Difference from Format Constraints

Answer completeness regularization is different from a pure format constraint\. A format constraint asks whether the answer is placed in the required template and whether tags are well formed\. For example,<answer\>\.\.\.</answis a malformed tag and belongs to the scope of format constraints\. In contrast, answer completeness regularization asks whether the content inside a valid answer span is semantically complete\.

For example, the output “First” can be placed inside a valid answer tag, but it is incomplete when the golden answer is “First Lady of the United States\.” TRACE therefore treats answer completeness as an independent generation objective\. It does not require the model to generate longer answers in general; rather, it encourages the model to preserve complete entities, restrictive phrases, and answer boundaries once the final knowledge source has been selected\.

## IVExperiments and Results

We evaluate around four questions that directly correspond to the claims in the preceding sections\.RQ1asks whether TRACE can preserve the ability to use correct retrieved context while resisting misleading retrieved context\.RQ2asks whether the same source\-aware training remains effective under more complex retrieval conditions, including self\-conflicting, irrelevant, and partially relevant contexts\.RQ3asks whether debate traces, strong\-to\-weak supervision, and answer completeness regularization are necessary\.RQ4asks whether the debate\-trace training framework can improve a substantially weaker target model\.

### IV\-AExperimental Setup

The evaluation covers five retrieved\-context scenarios\.Scenario I \(S1\)uses correct contextual knowledge, where retrieved passages support the golden answer\.Scenario II \(S2\)uses adversarial contextual knowledge, where retrieved passages contradict the golden answer\. We further evaluate three harder retrieval settings:Scenario III \(S3\), self\-conflicting contextual knowledge;Scenario IV \(S4\), irrelevant contextual knowledge; andScenario V \(S5\), partially relevant contextual knowledge mixed with distractors\.

#### IV\-A1Models and Baselines

We use Qwen2\.5–7B\-Instruct and Llama3\.1–8B\-Instruct as the main backbone models to evaluate whether TRACE is stable across model families\. In the debate stage, unless otherwise specified, the critic model is Qwen3–32B and the judge model is DeepSeek\-R1–8B\. The maximum number of debate rounds is 5\. However, the debate can terminate earlier based on the judge model’s stopping decision, resulting in fewer than 5 rounds per question on average\.

We compare TRACE with query\-only prompting, RAG prompting, Astute\-RAG\[[10](https://arxiv.org/html/2609.30337#bib.bib11)\], CK\-PlUG\[[12](https://arxiv.org/html/2609.30337#bib.bib16)\], supervised fine\-tuning \(SFT\), GRPO with RAG\[[30](https://arxiv.org/html/2609.30337#bib.bib25)\], and Knowledgeable\-R1\[[22](https://arxiv.org/html/2609.30337#bib.bib15)\]\. These baselines cover no\-retrieval prompting, standard retrieval prompting, prompt\-based conflict handling, learnable knowledge\-source control, supervised fine\-tuning, and reinforcement\-learning\-based source selection\. We use exact match \(EM\) as the primary metric, following the standard protocol of these benchmarks, and report all scores as percentages where higher is better\. All fine\-tuned variants are trained for one epoch\. During inference, all methods use deterministic decoding with temperature set to 0 to reduce sampling variance\.

TRACE is implemented with LoRA fine\-tuning\[[31](https://arxiv.org/html/2609.30337#bib.bib26)\], while some compared methods use full\-parameter reinforcement\-learning fine\-tuning and report results on H100 GPUs\. Our experiments are conducted on RTX 4090 GPUs\. This setting emphasizes whether debate\-trace supervision can provide practical gains under a parameter\-efficient training regime\.

#### IV\-A2Datasets

The datasets cover correct, wrong, self\-conflicting, irrelevant, and partially relevant retrieval scenarios\. S1 and S2 are mainly constructed from ConFiQA\[[20](https://arxiv.org/html/2609.30337#bib.bib20)\]and include three subtasks: QA, MR, and MC\. QA denotes single\-hop question answering\. MR denotes multi\-hop reasoning where only one evidence step is wrong\. MC denotes multi\-hop reasoning where multiple steps in the evidence chain are wrong\. PC\-QA, PC\-MR, and PC\-MC evaluate the model’s ability to use correct retrieved knowledge, whereas NC\-QA, NC\-MR, and NC\-MC evaluate robustness when the retrieved context contradicts the correct answer\.

For extended experiments, S3 uses the SC dataset constructed from ConFiQA, where mutually conflicting evidence is provided to evaluate the model’s ability to handle internal contradictions within the context\. S4 uses the ExplainPE medical question\-answering dataset\[[32](https://arxiv.org/html/2609.30337#bib.bib24)\], where the retrieved content is largely irrelevant to the question, evaluating the model’s resistance to irrelevant\-context interference\. S5 uses HotPotQA\[[33](https://arxiv.org/html/2609.30337#bib.bib27)\], 2WikiMultiHopQA\[[34](https://arxiv.org/html/2609.30337#bib.bib28)\], and MuSiQue\[[35](https://arxiv.org/html/2609.30337#bib.bib29)\], where useful evidence and distracting passages appear together, evaluating evidence selection under partially relevant retrieval results\.

#### IV\-A3Source of Retrieved Contexts

For all datasets, we use the contexts provided by the benchmarks and do not introduce an additional retriever\. Correct and wrong passages are directly taken from the benchmark structure rather than re\-indexed or re\-retrieved from a larger corpus\. We use the first five passages provided by each benchmark and keep their original order\. Truncation is applied only when the input exceeds the maximum length supported by the backbone model\.

### IV\-BResults on Standard Conflict Scenarios

The standard conflict scenarios evaluate the central requirement of source\-aware RAG: using reliable context in S1 while resisting wrong context in S2\. Table[III](https://arxiv.org/html/2609.30337#S4.T3)reports EM results on both backbone models\.

##### Scenario I: Correct contextual knowledge \(S1\)

TRACE preserves strong context utilization when retrieved evidence is reliable\. On Qwen2\.5–7B\-Instruct, TRACE achieves 90\.92%, 91\.05%, and 86\.91% on PC\-MR, PC\-MC, and PC\-QA, respectively, outperforming Knowledgeable\-R1 by 15\.84, 15\.54, and 6\.01 percentage points\. On Llama3\.1–8B\-Instruct, TRACE reaches 91\.41%, 84\.63%, and 91\.91%, improving over Knowledgeable\-R1 by 17\.65, 4\.39, and 11\.88 percentage points\. These results show that TRACE does not gain robustness by simply rejecting retrieved content; it continues to exploit correct external evidence\.

##### Scenario II: Adversarial contextual knowledge \(S2\)

TRACE also improves robustness when retrieved evidence is wrong, although the margin depends on the backbone and subtask\. Conventional RAG prompting suffers a clear drop under S2; for example, on Qwen2\.5–7B\-Instruct, it obtains only 13\.47%, 8\.06%, and 11\.31% on NC\-MR, NC\-MC, and NC\-QA, all lower than query\-only prompting\. In contrast, TRACE reaches 47\.64%, 39\.04%, and 37\.48%, improving over GRPO w/ RAG by 20\.70, 19\.30, and 11\.47 percentage points and over Knowledgeable\-R1 by 3\.70, 1\.70, and 8\.08 percentage points\. On Llama3\.1–8B\-Instruct, TRACE reaches 50\.17%, 43\.42%, and 44\.91%\. It is best on NC\-MC and NC\-QA, while Knowledgeable\-R1 remains stronger on NC\-MR\. Overall, the S2 results support the main claim that debate\-trace supervision improves resistance to misleading retrieved content without removing the ability to use correct context\.

TABLE III:EM \(%\) in standard conflict scenarios\. The best result is in bold, and the second\-best result is underlined\.S1: CorrectS2: WrongMethodPC\-MRPC\-MCPC\-QANC\-MRNC\-MCNC\-QAQwen2\.5–7BQuery\-only prompting27\.72%24\.66%31\.67%25\.93%25\.82%32\.31%RAG prompting65\.68%66\.39%74\.35%13\.47%8\.06%11\.31%CK\-PlUG64\.69%66\.55%78\.66%11\.62%8\.06%7\.92%Astute\-RAG65\.51%66\.05%77\.62%12\.79%7\.07%10\.34%SFT71\.95%77\.70%74\.70%24\.92%21\.05%21\.97%GRPO w/ RAG77\.56%77\.36%80\.03%26\.94%19\.74%26\.01%Knowledgeable\-R175\.08%75\.51%80\.90%43\.94%37\.34%29\.40%TRACE90\.92%91\.05%86\.91%47\.64%39\.04%37\.48%Llama3\.1–8BQuery\-only prompting29\.37%26\.18%39\.93%27\.10%27\.63%42\.65%RAG prompting64\.85%61\.99%76\.42%22\.90%16\.28%24\.88%CK\-PlUG54\.79%58\.45%69\.71%12\.12%9\.05%17\.29%Astute\-RAG65\.84%64\.86%77\.97%17\.00%9\.05%17\.29%SFT72\.88%79\.22%73\.84%42\.59%35\.53%35\.86%GRPO w/ RAG78\.05%79\.73%82\.62%41\.58%35\.69%39\.26%Knowledgeable\-R173\.76%80\.24%80\.03%55\.39%41\.12%44\.59%TRACE91\.41%84\.63%91\.91%50\.17%43\.42%44\.91%

### IV\-CExtended Evaluation under More Complex Retrieval Scenarios

The extended evaluation tests whether TRACE generalizes beyond the standard correct\-context and wrong\-context setting\. Table[IV](https://arxiv.org/html/2609.30337#S4.T4)reports EM results under self\-conflicting, irrelevant, and partially relevant retrieval\.

##### Scenario III: Self\-conflicting contextual knowledge \(S3\)

TRACE is particularly effective when retrieved evidence is internally contradictory\. In S3, the model must choose among mutually conflicting contextual claims rather than merely decide whether to use retrieval\. TRACE reaches 90\.58% on Qwen2\.5–7B\-Instruct, outperforming Knowledgeable\-R1 by 14\.25 percentage points, and reaches 92\.58% on Llama3\.1–8B\-Instruct, outperforming GRPO w/ RAG by 16\.00 percentage points\. This result is consistent with the method design: debate traces expose competing candidate answers and answer shifts, which provide direct supervision for resolving self\-conflicting evidence\.

##### Scenario IV: Irrelevant contextual knowledge \(S4\)

TRACE remains competitive when the retrieved context is largely irrelevant, but the strength of the result is backbone\-dependent\. On ExplainPE with Qwen2\.5–7B\-Instruct, query\-only prompting obtains 64\.45%, while RAG prompting drops to 62\.21%, showing that irrelevant retrieval can hurt answer accuracy\. TRACE reaches 69\.14%, outperforming Knowledgeable\-R1 and indicating stronger resistance to irrelevant\-context interference\. On Llama3\.1–8B\-Instruct, TRACE reaches 47\.95%, improving over RAG prompting by 8\.79 percentage points and slightly outperforming GRPO w/ RAG, but remaining below Knowledgeable\-R1\.

##### Scenario V: Partially relevant contextual knowledge \(S5\)

S5 reveals the main boundary of TRACE\. On Qwen2\.5–7B\-Instruct, TRACE is best on MuSiQue and second\-best on 2Wiki, but it is below Knowledgeable\-R1 on HotPotQA\. On Llama3\.1–8B\-Instruct, TRACE is second\-best on MuSiQue but lags behind the strongest baseline on HotPotQA and 2Wiki\. This pattern suggests that TRACE is strongest when the central challenge is source conflict or misleading candidates, while partially relevant long\-context multi\-hop settings require finer paragraph\-level evidence localization and evidence\-chain composition\. These results motivate future extensions that combine debate\-trace supervision with explicit evidence\-chain supervision\.

TABLE IV:EM \(%\) in more complex retrieval scenarios\. The best result is in bold, and the second\-best result is underlined\.S3: ConflictS4: IrrelevantS5: Partly IrrelevantMethodSCExplainPEHotPotQA2WikiMuSiQueQwen2\.5–7BQuery\-only prompting29\.67%64\.45%20\.90%25\.54%4\.36%RAG prompting59\.50%62\.21%20\.36%22\.53%6\.41%CK\-PlUG55\.00%55\.00%22\.74%24\.76%6\.25%Astute\-RAG54\.20%56\.74%17\.87%20\.35%6\.29%SFT68\.50%66\.60%30\.14%32\.20%11\.75%GRPO w/ RAG75\.33%66\.50%27\.93%33\.95%11\.79%Knowledgeable\-R176\.33%67\.57%31\.45%37\.52%12\.04%TRACE90\.58%69\.14%29\.29%37\.44%12\.37%Llama3\.1–8BQuery\-only prompting32\.08%43\.26%20\.69%21\.02%6\.16%RAG prompting61\.17%39\.16%24\.44%23\.50%8\.19%CK\-PlUG42\.00%31\.54%22\.35%24\.63%5\.25%Astute\-RAG59\.83%40\.14%1\.65%30\.26%9\.64%SFT70\.12%47\.17%33\.59%38\.24%13\.36%GRPO w/ RAG76\.58%47\.56%34\.84%41\.22%16\.59%Knowledgeable\-R173\.67%49\.61%37\.06%45\.37%14\.69%TRACE92\.58%47\.95%31\.17%38\.88%15\.18%

### IV\-DAblation Study

The ablation study evaluates whether the main modules in TRACE are responsible for the observed gains\. We evaluate three variants on the QA, MC, and MR subsets of ConFiQA\.*Without debate*removes multi\-agent debate and directly uses query\-only and RAG responses from the base LLM\.*Without strong\-teaching\-weak*keeps the debate format but replaces all debate roles with the base LLM\.*Without answer completeness regularization*sets the weight of answer completeness regularization to 0\.

TABLE V:Ablation results on ConFiQA EM \(%\)\. The best result in each column is in bold\.VariantPC\-MCNC\-MCPC\-MRNC\-MRPC\-QANC\-QAFull TRACE91\.05%39\.04%90\.92%47\.64%86\.91%37\.48%w/o debate85\.14%35\.86%86\.30%44\.94%87\.26%36\.99%w/o strong\-teaching\-weak92\.06%38\.15%87\.45%45\.45%85\.54%35\.86%w/o answer completeness regularization85\.97%38\.65%83\.00%42\.59%85\.88%32\.14%

Removing debate reduces the average EM from 65\.51% to 62\.75%, with drops on all wrong\-context metrics\. The only local improvement is PC\-QA, which increases from 86\.91% to 87\.26%\. This pattern indicates that simple single\-hop questions under correct context can be solved without debate traces, but robustness under misleading retrieval benefits from intermediate candidate answers and positive\-negative samples exposed by debate\.

Replacing strong\-to\-weak debate with same\-model debate also weakens overall performance\. PC\-MC increases from 91\.05% to 92\.06%, but NC\-MC, PC\-MR, NC\-MR, PC\-QA, and NC\-QA all decrease, and average EM drops from 65\.51% to 64\.09%\. This result suggests that stronger critic and judge models provide useful correction and counterfactual\-discrimination signals, especially under wrong\-context settings\.

Removing answer completeness regularization causes the largest average drop, from 65\.51% to 61\.37%\. All six metrics decrease, with the largest drop on NC\-QA\. This supports the claim that source selection alone is insufficient: after the model identifies a reliable answer direction, answer\-tail reinforcement and premature\-termination suppression help preserve complete answer boundaries and reduce incomplete outputs\.

To complement these quantitative ablations, Appendix[A](https://arxiv.org/html/2609.30337#A1)provides two representative cases that illustrate how TRACE resists misleading retrieved evidence and preserves complete answer boundaries\.

### IV\-ERole Conversion between Strong and Weak Models

The role\-conversion experiment tests whether TRACE can transfer supervision from relatively stronger models to a substantially weaker target model\. In the main setting, the strong model is Qwen3–32B, the intermediate model is DeepSeek\-R1–8B, and the weak target model is Qwen2\.5–7B\. In the role\-conversion setting, Qwen2\.5–7B is used as the strongest model, Qwen2\.5–3B is used as the intermediate model, and Qwen2\.5–1\.5B is used as the weak target model for fine\-tuning\.

Fig\. 3:Role\-conversion results between stronger debate models and a weaker target model, with comparison to the strongest baseline\.Fig\.[3](https://arxiv.org/html/2609.30337#S4.F3)shows that TRACE can still improve a weak target model in several knowledge\-conflict settings\. Although Qwen2\.5–1\.5B has only 21\.3% of the parameters of Qwen2\.5–7B, the fine\-tuned model remains competitive and surpasses the 7B\-level Knowledgeable\-R1 baseline on some metrics\. This result suggests that debate\-trace supervision can transfer useful source\-selection behavior to weaker models, although the effect is not uniform across all scenarios\.

## VConclusion and Future Work

This paper addresses knowledge conflicts in RAG by proposing TRACE, a source\-aware fine\-tuning framework that learns when to use retrieved context and when to rely on parametric knowledge\. The key idea is to use multi\-agent debate only during data construction: debate trajectories expose correct candidates, misleading candidates, and answer\-shift patterns, which are converted into positive and negative fine\-tuning signals for a single target model\. TRACE further adds answer completeness regularization so that source selection is paired with complete final\-answer generation\.

Experiments across correct, wrong, self\-conflicting, irrelevant, and partially relevant retrieval settings show that TRACE preserves the use of reliable context while improving robustness to misleading or conflicting retrieval\. Ablation results confirm the contribution of debate traces, strong\-to\-weak supervision, and answer completeness regularization, and the role\-conversion study suggests that the supervision can also benefit a weaker target model\. These findings indicate that debate trajectories can serve as process\-level supervision for source\-aware RAG without requiring debate at inference time\.

A current limitation is that TRACE is strongest in explicit source\-conflict settings and less uniformly effective in partially relevant long\-context multi\-hop scenarios, where answer correctness depends on fine\-grained evidence localization and evidence\-chain composition\. Future work will therefore combine debate\-trace supervision with explicit evidence\-chain supervision, reduce the cost and noise of debate\-based data construction, and evaluate source\-aware training under more realistic retrieval pipelines beyond benchmark\-provided contexts\.

## Acknowledgments

This research was supported by the Key R&D Program of Shandong Province, China \(Project No\. 2024CXGC010109\), the Shandong Provincial Natural Science Foundation \(Project No\. ZR2026QC1574\), and the 2026 Linyi University High\-Level Talent \(Doctoral\) Research Start\-up Fund \(Natural Science\), Grant No\. Z6126063\.

## References

- \[1\]A\. Singh, A\. Fry, A\. Perelman, A\. Tart, A\. Ganesh, A\. El\-Kishky, A\. McLaughlin, A\. Low, A\. Ostrow, A\. Ananthram,et al\.\(2025\)Openai gpt\-5 system card\.arXiv preprint arXiv:2601\.03267\.Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p1.1)\.
- \[2\]H\. Touvron, T\. Lavril, G\. Izacard, X\. Martinet, M\. Lachaux, T\. Lacroix, B\. Rozière, N\. Goyal, E\. Hambro, F\. Azhar,et al\.\(2023\)Llama: open and efficient foundation language models\.arXiv preprint arXiv:2302\.13971\.Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p1.1)\.
- \[3\]\(2022\)Truthfulqa: measuring how models mimic human falsehoods\.InProceedings of the 60th annual meeting of the association for computational linguistics \(volume 1: long papers\),pp\. 3214–3252\.Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p1.1)\.
- \[4\]Z\. Ji, N\. Lee, R\. Frieske, T\. Yu, D\. Su, Y\. Xu, E\. Ishii, Y\. J\. Bang, A\. Madotto, and P\. Fung\(2023\)Survey of hallucination in natural language generation\.ACM computing surveys55\(12\),pp\. 1–38\.Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p1.1)\.
- \[5\]P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel,et al\.\(2020\)Retrieval\-augmented generation for knowledge\-intensive nlp tasks\.Advances in neural information processing systems33,pp\. 9459–9474\.Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p1.1)\.
- \[6\]Y\. Gao, Y\. Xiong, X\. Gao, K\. Jia, J\. Pan, Y\. Bi, Y\. Dai, J\. Sun, H\. Wang, H\. Wang,et al\.\(2023\)Retrieval\-augmented generation for large language models: a survey\.arXiv preprint arXiv:2312\.10997\.Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p1.1)\.
- \[7\]K\. Guu, K\. Lee, Z\. Tung, P\. Pasupat, and M\. Chang\(2020\)Retrieval augmented language model pre\-training\.InInternational conference on machine learning,pp\. 3929–3938\.Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p1.1)\.
- \[8\]S\. Longpre, K\. Perisetla, A\. Chen, N\. Ramesh, C\. DuBois, and S\. Singh\(2021\)Entity\-based knowledge conflicts in question answering\.InProceedings of the 2021 conference on empirical methods in natural language processing,pp\. 7052–7063\.Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p2.1)\.
- \[9\]J\. Xie, K\. Zhang, J\. Chen, R\. Lou, and Y\. Su\(2024\)Adaptive chameleon or stubborn sloth: revealing the behavior of large language models in knowledge conflicts\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 35623–35646\.Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p2.1)\.
- \[10\]F\. Wang, X\. Wan, R\. Sun, J\. Chen, and S\. O\. Arik\(2025\)Astute rag: overcoming imperfect retrieval augmentation and knowledge conflicts for large language models\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 30553–30571\.Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p3.1),[§II\-A](https://arxiv.org/html/2609.30337#S2.SS1.p1.1),[§IV\-A1](https://arxiv.org/html/2609.30337#S4.SS1.SSS1.p2.1)\.
- \[11\]S\. Yan, J\. Gu, Y\. Zhu, and Z\. Ling\(2024\)Corrective retrieval augmented generation\.CoRRabs/2401\.15884\.External Links:[Link](https://doi.org/10.48550/arXiv.2401.15884)Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p3.1),[§II\-A](https://arxiv.org/html/2609.30337#S2.SS1.p2.1)\.
- \[12\]B\. Bi, S\. Liu, Y\. Wang, Y\. Xu, J\. Fang, L\. Mei, and X\. Cheng\(2026\)Parameters vs\. context: fine\-grained control of knowledge reliance in language models\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=TJ3DqFiGau)Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p3.1),[§II\-B](https://arxiv.org/html/2609.30337#S2.SS2.p1.1),[§IV\-A1](https://arxiv.org/html/2609.30337#S4.SS1.SSS1.p2.1)\.
- \[13\]Z\. Wei, W\. Chen, and Y\. Meng\(2025\)InstructRAG: instructing retrieval\-augmented generation via self\-synthesized rationales\.InInternational Conference on Learning Representations,pp\. 82731–82754\.Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p3.1),[§II\-A](https://arxiv.org/html/2609.30337#S2.SS1.p1.1)\.
- \[14\]C\. Xiang, T\. Wu, Z\. Zhong, D\. Wagner, D\. Chen, and P\. Mittal\(2024\)Certifiably robust rag against retrieval corruption\.arXiv preprint arXiv:2405\.15556\.Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p3.1),[§II\-A](https://arxiv.org/html/2609.30337#S2.SS1.p2.1)\.
- \[15\]E\. Choi, J\. Park, H\. Lee, and J\. Lee\(2025\)Conflict\-aware soft prompting for retrieval\-augmented generation\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 26969–26983\.Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p3.1)\.
- \[16\]H\. K\. Choi, J\. Zhu, and S\. Li\(2025\)Debate or vote: which yields better decisions in multi\-agent large language models?\.Advances in Neural Information Processing Systems38,pp\. 101732–101764\.Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p4.1),[§II\-C](https://arxiv.org/html/2609.30337#S2.SS3.p1.1)\.
- \[17\]Y\. Du, S\. Li, A\. Torralba, J\. B\. Tenenbaum, and I\. Mordatch\(2024\)Improving factuality and reasoning in language models through multiagent debate\.InForty\-first international conference on machine learning,Cited by:[§I](https://arxiv.org/html/2609.30337#S1.p4.1),[§II\-C](https://arxiv.org/html/2609.30337#S2.SS3.p1.1)\.
- \[18\]A\. Asai, Z\. Wu, Y\. Wang, A\. Sil, and H\. Hajishirzi\(2024\)Self\-rag: learning to retrieve, generate, and critique through self\-reflection\.InInternational conference on learning representations,Vol\.2024,pp\. 9112–9141\.Cited by:[§II\-A](https://arxiv.org/html/2609.30337#S2.SS1.p1.1)\.
- \[19\]S\. Liu, Y\. Shang, and X\. Zhang\(2026\)Truthfulrag: resolving factual\-level conflicts in retrieval\-augmented generation with knowledge graphs\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 32168–32176\.Cited by:[§II\-A](https://arxiv.org/html/2609.30337#S2.SS1.p2.1)\.
- \[20\]B\. Bi, S\. Huang, Y\. Wang, T\. Yang, Z\. Zhang, H\. Huang, L\. Mei, J\. Fang, Z\. Li, F\. Wei,et al\.\(2025\)Context\-dpo: aligning language models for context\-faithfulness\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 10280–10300\.Cited by:[§II\-B](https://arxiv.org/html/2609.30337#S2.SS2.p1.1),[§IV\-A2](https://arxiv.org/html/2609.30337#S4.SS1.SSS2.p1.1)\.
- \[21\]R\. Zhang, Y\. Xu, Y\. Xiao, R\. Zhu, X\. Jiang, X\. Chu, J\. Zhao, and Y\. Wang\(2025\)Knowpo: knowledge\-aware preference optimization for controllable knowledge selection in retrieval\-augmented language models\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 25895–25903\.Cited by:[§II\-B](https://arxiv.org/html/2609.30337#S2.SS2.p1.1)\.
- \[22\]C\. Lin, Y\. Wen, D\. Su, H\. Tan, F\. Sun, M\. Chen, C\. Bao, and Z\. Lv\(2026\)Resisting contextual interference in RAG via parametric\-knowledge reinforcement\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=6Qc6sO1jh9)Cited by:[Appendix A](https://arxiv.org/html/2609.30337#A1.p1.1),[§II\-B](https://arxiv.org/html/2609.30337#S2.SS2.p1.1),[§III\-C3](https://arxiv.org/html/2609.30337#S3.SS3.SSS3.p3.1),[§IV\-A1](https://arxiv.org/html/2609.30337#S4.SS1.SSS1.p2.1)\.
- \[23\]T\. Liang, Z\. He, W\. Jiao, X\. Wang, Y\. Wang, R\. Wang, Y\. Yang, S\. Shi, and Z\. Tu\(2024\)Encouraging divergent thinking in large language models through multi\-agent debate\.InProceedings of the 2024 conference on empirical methods in natural language processing,pp\. 17889–17904\.Cited by:[§II\-C](https://arxiv.org/html/2609.30337#S2.SS3.p1.1)\.
- \[24\]J\. Chen, S\. Saha, and M\. Bansal\(2024\)Reconcile: round\-table conference improves reasoning via consensus among diverse llms\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 7066–7085\.Cited by:[§II\-C](https://arxiv.org/html/2609.30337#S2.SS3.p1.1)\.
- \[25\]S\. R\. Motwani, C\. Smith, R\. J\. Das, R\. Rafailov, I\. Laptev, P\. Torr, F\. Pizzati, R\. Clark, and C\. S\. de Witt\(2025\)MALT: improving reasoning with multi\-agent LLM training\.InWorkshop on Reasoning and Planning for Large Language Models,External Links:[Link](https://openreview.net/forum?id=lIf7grAC7n)Cited by:[§II\-C](https://arxiv.org/html/2609.30337#S2.SS3.p1.1)\.
- \[26\]X\. Zhou, H\. Huang, and L\. Liao\(2025\)Debate, reflect, and distill: multi\-agent feedback with tree\-structured preference optimization for efficient language model enhancement\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 9122–9137\.Cited by:[§II\-C](https://arxiv.org/html/2609.30337#S2.SS3.p1.1)\.
- \[27\]R\. Xu, Z\. Qi, Z\. Guo, C\. Wang, H\. Wang, Y\. Zhang, and W\. Xu\(2024\)Knowledge conflicts for llms: a survey\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 8541–8565\.Cited by:[§II\-C](https://arxiv.org/html/2609.30337#S2.SS3.p2.1)\.
- \[28\]S\. Zeng, J\. Zhang, B\. Li, Y\. Lin, T\. Zheng, D\. Everaert, H\. Lu, H\. Liu, Y\. Xing, M\. X\. Cheng,et al\.\(2025\)Towards knowledge checking in retrieval\-augmented generation: a representation perspective\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 2952–2969\.Cited by:[§II\-C](https://arxiv.org/html/2609.30337#S2.SS3.p2.1)\.
- \[29\]S\. Welleck, I\. Kulikov, S\. Roller, E\. Dinan, K\. Cho, and J\. Weston\(2020\)Neural text generation with unlikelihood training\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=SJeYe0NtvH)Cited by:[§III\-C3](https://arxiv.org/html/2609.30337#S3.SS3.SSS3.p2.2)\.
- \[30\]D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi,et al\.\(2025\)DeepSeek\-r1 incentivizes reasoning in llms through reinforcement learning\.Nature645\(8081\),pp\. 633–638\.Cited by:[§IV\-A1](https://arxiv.org/html/2609.30337#S4.SS1.SSS1.p2.1)\.
- \[31\]E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, W\. Chen,et al\.\(2022\)Lora: low\-rank adaptation of large language models\.\.Iclr\.Cited by:[§IV\-A1](https://arxiv.org/html/2609.30337#S4.SS1.SSS1.p3.1)\.
- \[32\]Y\. Wen, Z\. Wang, and J\. Sun\(2024\)Mindmap: knowledge graph prompting sparks graph of thoughts in large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 10370–10388\.Cited by:[§IV\-A2](https://arxiv.org/html/2609.30337#S4.SS1.SSS2.p2.1)\.
- \[33\]Z\. Yang, P\. Qi, S\. Zhang, Y\. Bengio, W\. Cohen, R\. Salakhutdinov, and C\. D\. Manning\(2018\)HotpotQA: a dataset for diverse, explainable multi\-hop question answering\.InProceedings of the 2018 conference on empirical methods in natural language processing,pp\. 2369–2380\.Cited by:[§IV\-A2](https://arxiv.org/html/2609.30337#S4.SS1.SSS2.p2.1)\.
- \[34\]X\. Ho, A\. D\. Nguyen, S\. Sugawara, and A\. Aizawa\(2020\)Constructing a multi\-hop qa dataset for comprehensive evaluation of reasoning steps\.InProceedings of the 28th International Conference on Computational Linguistics,pp\. 6609–6625\.Cited by:[§IV\-A2](https://arxiv.org/html/2609.30337#S4.SS1.SSS2.p2.1)\.
- \[35\]H\. Trivedi, N\. Balasubramanian, T\. Khot, and A\. Sabharwal\(2022\)♫ MuSiQue: multihop questions via single\-hop question composition\.Transactions of the Association for Computational Linguistics10,pp\. 539–554\.Cited by:[§IV\-A2](https://arxiv.org/html/2609.30337#S4.SS1.SSS2.p2.1)\.

## Appendix ACase Analysis

This appendix provides two representative cases to illustrate the behavior of TRACE under misleading retrieved contexts\. We summarize the key misleading evidence in each case and compare the outputs of Knowledgeable\-R1\[[22](https://arxiv.org/html/2609.30337#bib.bib15)\]and TRACE\. Case 1 mainly reflects context resistance under erroneous retrieval, and case 2 shows the importance of answer completeness regularization mechanism\.

### A\-ACase 1: Resisting a Misleading Citizenship Chain

Question\.

> What is the native language of the head of government of the country where Robert Evans has citizenship?

Misleading evidence\.The retrieved passage incorrectly describes Robert Evans as a citizen of Belgium\. It then builds a misleading reasoning chain from Belgium to its head of government and further associates the final answer with Dutch\. The passage also contains noisy and internally inconsistent statements, such as connecting Belgium with Narendra Modi and describing him as a native Puerto Rican Spanish speaker\.

Knowledgeable\-R1 output\.

> <think\> Robert Evans is a citizen of Belgium\. The head of government in Belgium is the Prime Minister\. The native language of the Prime Minister of Belgium is Dutch\. </think\> <answer\> Dutch </answer\>

TRACE output\.

> <think\> The information provided about Robert Evans is incorrect, as he is a U\.S\. citizen, not a Belgian\. However, the question asks about his citizenship, which is the United States\. The head of government of the United States is the President, and the native language of the President is English\. </think\> <answer\> English </answer\>

Analysis\.This case illustrates the context\-resistance ability of TRACE\. The thinking process of Knowledgeable\-R1 shows that it accepts the incorrect retrieved premise that Robert Evans is a Belgian citizen, and then follows the misleading chain “Belgium, Prime Minister,Dutch\.” As a result, it produces an incorrect final answer\. By contrast, the thinking process of TRACE explicitly identifies the retrieved citizenship statement as unreliable and falls back to parametric knowledge\. It then maps the country to the United States, the head of government to the U\.S\. President, and the native language to English\. This example provides qualitative evidence that debate\-trace supervision can help the model resist misleading retrieved content when parametric knowledge is more reliable\.

### A\-BCase 2: Preserving the Complete Answer Boundary

Question\.

> What position is held by the spouse of the head of state of the country where Bam Adebayo is a citizen?

Misleading evidence\.The retrieved passage contains noisy and irrelevant descriptions\. It incorrectly introduces Bam Adebayo as being from the Russian Empire, then shifts to unrelated statements about Charles’s Empire, Charles III, and Ronald Reagan\.

Model outputs\.Knowledgeable\-R1 produces an answer in the correct general direction but with an incomplete answer:

> <answer\> First Lady </answer\>

TRACE gives a more complete answer:

> <answer\> First Lady of the United States </answer\>

Analysis\.The core of this analysis is not mainly about choosing between parametric knowledge or retrieved knowledge\. The key difference lies in answer completeness\. Knowledgeable\-R1 outputs only the head noun phrase, while TRACE preserves the complete answer boundary by including the country\-specific qualifier “of the United States\.” For exact\-match style evaluation and for semantic clarity, this qualifier is important because it makes the final answer more specific and less ambiguous\. This example therefore supports the role of answer\-completeness regularization: answer\-tail token reinforcement and premature termination suppression help LLM generate a complete final answer rather than stopping at a shorter answer fragment\.

### A\-CSummary

The first case shows that TRACE can resist some misleading retrieved chain and recover the answer from reliable parametric knowledge more effectively during some situations\. The second case shows that TRACE can produce a more complete answer\. Together, they provide qualitative evidence for the two central components of TRACE: the fine\-tuning method based on multi\-agent debate traces helps LLM choose more reliable knowledge and answer completeness regularization mechanism helps LLMs reduce the generation of incomplete answer\.

相似文章

上下文优化下的检索增强生成:从梯度下降视角

arXiv cs.CL

本文研究检索增强生成作为上下文优化过程,表明线性自注意力可以在统一的RAG目标上实现梯度下降。它提出了一种轻量级方法,适用于冻结的RAG大语言模型,通过预测上下文条件的更新,在多个问答基准上提升了性能。

语境之代价:在多模态检索增强生成中缓解文本偏差

arXiv cs.CL

本文识别并形式化了多模态RAG中的“再污染”现象,即添加准确上下文会导致模型因注意力崩溃(视觉盲区和位置偏差)而放弃正确预测。作者提出BAIR,一种无参数的推理时框架,能恢复视觉显著性并惩罚文本干扰因素,从而在医学、公平性和地理空间基准上提高可靠性。

TRACE:可信检索增强对话引擎

arXiv cs.AI

TRACE 是一种面向公共服务聊天机器人的检索增强对话引擎,通过增强对嘈杂目录的检索质量来改进约束感知推荐,同时减少幻觉响应。

为什么检索增强生成会失败:图视角

arXiv cs.CL

本文探讨了检索增强生成(RAG)系统即使在获取到正确证据的情况下仍然失败的原因。通过电路追踪和归因图,作者发现正确的预测展现出更深的推理路径和更分散的证据流,而失败则表现为浅层、碎片化的模式。他们提出了一个基于图的错误检测框架和有针对性的干预措施,以提高RAG的可靠性。