EAVer:基于端到端代理策略的长文本事实性验证
摘要
EAVer 提出了一种端到端的代理策略,用于通过分组断言、高效搜索和重用证据来验证长文本中的事实,在诸如 VeriFastScore 和 FaStFact-Bench 等基准测试上以更少的搜索次数实现更优性能。
arXiv:2609.22223v1 Announce Type: new
Abstract: Long-form factuality verification is commonly implemented as a static decompose-search-verify pipeline, with separately prompted modules processing claims and invoking external search. Treating claims independently makes LLM and search calls scale with claim count and causes repeated searches for overlapping evidence about related claims. We introduce EAVer, an End-to-end Agentic Verifier that learns to control the complete response-level verification workflow as a unified policy. EAVer groups semantically related claims, routes each group to direct verification or targeted search based on confidence, and keeps evidence returned by search in compact in-context memos for cross-claim reuse. To train this policy, we develop a privileged-teacher synthesis pipeline that converts gold claim annotations into executable multi-turn tool-interaction trajectories with live search rather than post-hoc rationales. Structural, label-alignment, tool-use, search-budget, and leakage checks yield 1,447 quality-controlled trajectories. We further construct 794 bidirectional same-trajectory preference pairs that keep claim grouping, search, and evidence fixed, enabling decision-focused Direct Preference Optimization (DPO) over factuality-decision tokens. The results with Qwen3-8B show that EAVer outperforms the strongest search-based baseline on each benchmark by 2.88 Macro-F1 points on VeriFastScore and 4.73 points on the out-of-distribution FaStFact-Bench, while using about 80% fewer searches than the most search-efficient baseline. Moreover, EAVer consistently improves performance across models ranging from 4B to 32B parameters, demonstrating its strong generalizability.
查看缓存全文
缓存时间: 2026/09/22 09:10
# EAVer: Long-Form Factuality Verification as an End-to-End Agentic Policy
Source: [https://arxiv.org/html/2609.22223](https://arxiv.org/html/2609.22223)
Kening Zheng1,2, Aoying Zheng2,3, Zhigang Chang2, Yazhi Guo2, Miaotian Guo2, Qingwei Zong2Xianhai Xie2, Weiqiang Jin4, Chengze Li1, Hanrong Zhang1, Jie Yang1, Wei\-Chieh Huang1Lingzhe Zhang1,5, Liancheng Fang1, Xin Zou6, Hanqian Li6, Jiahao Huo6, Yibo Yan6Zizhuang Deng3, Lei Miao2, Wei Guo2, Haihong Tang2, Bo Zheng2, Philip S\. Yu1\\corresponding††thanks:Work done during an internship\.
###### Abstract
Long\-form factuality verification is commonly implemented as a static decompose\-search\-verify pipeline, with separately prompted modules processing claims and invoking external search\. Treating claims independently makes LLM and search calls scale with claim count and causes repeated searches for overlapping evidence about related claims\. We introduceEAVer, anEnd\-to\-endAgenticVerifier that learns to control the complete response\-level verification workflow as a unified policy\.EAVergroups semantically related claims, routes each group to direct verification or targeted search based on confidence, and keeps evidence returned by search in compact in\-context memos for cross\-claim reuse\. To train this policy, we develop a privileged\-teacher synthesis pipeline that converts gold claim annotations into executable multi\-turn tool\-interaction trajectories with live search rather than post\-hoc rationales\. Structural, label\-alignment, tool\-use, search\-budget, and leakage checks yield 1,447 quality\-controlled trajectories\. We further construct 794 bidirectional same\-trajectory preference pairs that keep claim grouping, search, and evidence fixed, enabling decision\-focused Direct Preference Optimization \(DPO\) over factuality\-decision tokens\. The results with Qwen3\-8B show thatEAVeroutperforms the strongest search\-based baseline on each benchmark by 2\.88 Macro\-F1 points on VeriFastScore and 4\.73 points on the out\-of\-distribution FaStFact\-Bench, while using about 80% fewer searches than the most search\-efficient baseline\. Moreover,EAVerconsistently improves performance across models ranging from 4B to 32B parameters, demonstrating its strong generalizability\.
1University of Illinois Chicago2Taobao & Tmall Group3Shandong University4Xi’an Jiaotong University5Peking University6The Hong Kong University of Science and Technology \(Guangzhou\)
## Introduction
Large language models \(LLMs\) can produce fluent, extended responses across open\-ended question answering and instruction\-following settings\([Brown et al\. 2020](https://arxiv.org/html/2609.22223#bib.bib17);[Ouyang et al\. 2022](https://arxiv.org/html/2609.22223#bib.bib18)\)\. Yet fluency does not guarantee factual reliability: generated answers may contain plausible but unsupported statements that are difficult to detect at the response level\([Lin et al\. 2022](https://arxiv.org/html/2609.22223#bib.bib19);[Manakul et al\. 2023](https://arxiv.org/html/2609.22223#bib.bib20)\)\. This challenge is amplified in long\-form generation, where factual statements are distributed across sentences and paragraphs and may connect entities, dates, quantities, events, and causal relations\. Reliable evaluation therefore requires fine\-grained verification that identifies individual claims and grounds each judgment in the relevant evidence for that claim\.
Figure 1:Search efficiency of EAVer\.*Left:*Existing per\-claim evaluators repeatedly search for overlapping evidence;EAVeradaptively searches and reuses evidence across related claims\.*Right:*Mean searches per sample and link redundancy on an independent VeriFastScore validation split\. Link redundancy is the fraction of links repeated within the same evaluation sample\.Existing long\-form factuality evaluators generally follow a decompose\-search\-verify recipe: extract factual claims, search for external evidence, and predict whether each claim is supported\([Min et al\. 2023](https://arxiv.org/html/2609.22223#bib.bib1);[Wei et al\. 2024](https://arxiv.org/html/2609.22223#bib.bib3);[Song et al\. 2024](https://arxiv.org/html/2609.22223#bib.bib4)\)\. Subsequent work broadens the range of verifiable content, uses iterative search for difficult claims, or batches extraction and verification for greater efficiency\([Xie et al\. 2025](https://arxiv.org/html/2609.22223#bib.bib5);[Rajendhran et al\. 2025](https://arxiv.org/html/2609.22223#bib.bib6)\)\. Nevertheless, search control and verification remain hand\-orchestrated or separately optimized rather than jointly learned as a unified policy for response\-level evaluation\.
Despite their different implementations, existing methods share two structural limitations\. Most hard\-code verification as an orchestration of separately prompted LLM and search\-API calls, and process each claim largely in isolation\. Consequently, the number of calls grows with the number of claims, while related claims repeatedly search for overlapping evidence about the same entities and events\. To quantify this inefficiency, we evaluate each method on an independent VeriFastScore validation split and record searches per sample and link redundancy\. Figure[1](https://arxiv.org/html/2609.22223#Sx1.F1)illustrates this structural contrast and quantifies its cost: prior per\-claim pipelines issue 35\.9–103\.7 searches per sample with 23\.5–66\.5% repeated links, whereasEAVeruses 3\.42 searches with only 0\.3% redundancy\. The high redundancy of prior pipelines persists even when exact query strings differ, showing that per\-claim control and string\-level caching do not fully address cross\-claim evidence reuse\. More fundamentally, a fixed controller cannot be optimized as a response\-level policy that learns when to search, what to search for, and when available evidence is already sufficient\. Although VeriFastScore trains its verifier, it consumes evidence collected by a separate search stage and therefore does not learn an interactive policy for search and evidence reuse\([Rajendhran et al\. 2025](https://arxiv.org/html/2609.22223#bib.bib6)\)\. Motivated by language\-model agents that interleave reasoning with external actions\([Yao et al\. 2023](https://arxiv.org/html/2609.22223#bib.bib21);[Schick et al\. 2023](https://arxiv.org/html/2609.22223#bib.bib22)\), we introduceEAVer, an end\-to\-end agentic verifier that learns to control the complete response\-level verification workflow\. Rather than executing a fixed sequence of modules,EAVerrepresents verification as a single trajectory that jointly organizes claims, acquires evidence, and produces factuality judgments\. It groups semantically related claims, routes each group to direct verification or targeted search based on model confidence, and compresses evidence returned by search into compact in\-context memos that later claims can reuse directly from the shared interaction history\. Learning this structured behavior requires supervision beyond final factuality labels\. Building on synthetic supervision for instruction following\([Wang et al\. 2023](https://arxiv.org/html/2609.22223#bib.bib23)\), we develop a policy\-aware synthesis pipeline in which a privileged teacher uses gold atomic claims and labels to generate executable multi\-turn tool\-interaction trajectories with live search rather than post\-hoc explanations; the student receives none of these gold annotations at inference time\. We apply quality\-control filters for structural validity, label alignment, proper tool use, search\-budget compliance, and the absence of label leakage, retaining only demonstrations that faithfully instantiate the complete verification policy\.
Base verifiers are strongly biased toward the supported class, making false support particularly consequential: it allows unsupported content to pass as factual\. We therefore first learn the core policy from executable verification demonstrations and study decision\-focused Direct Preference Optimization \(DPO\)\([Rafailov et al\. 2023](https://arxiv.org/html/2609.22223#bib.bib24)\)only as a conservative extension that restricts preference learning to factuality decisions within otherwise matched trajectories, without directly optimizing the learned search and evidence\-use behavior\. Experiments show thatEAVerachieves the strongest Macro\-F1 among search\-based systems on both benchmarks while using about one fifth as many searches as the most search\-efficient baseline\. Moreover, the proposed policy\-learning framework yields consistent gains across backbone generations, parameter scales, and model families\. Our contributions are:
- ❶Empirical finding\.We systematically analyze the search efficiency of representative long\-form factuality verifiers, uncovering substantial per\-sample search overhead and cross\-claim evidence redundancy\.
- ❷Policy\-learning contribution\.We develop a privileged\-teacher synthesis and quality\-control framework that converts claim supervision into executable verification behavior, together with a controlled preference\-construction procedure that isolates factuality\-decision errors\.
- ❸Methodological contribution\.We proposeEAVer, an end\-to\-end agentic long\-form verifier that unifies claim grouping, adaptive search, in\-context evidence reuse, and factuality prediction within a single trainable policy\.
- ❹Empirical validation\.EAVerachieves the strongest Macro\-F1 among search\-based systems on both benchmarks while using about one fifth as many searches as the most search\-efficient baseline, with consistent improvements across backbone generations, parameter scales, and model families\.
Figure 2:Fixed verification pipelines versus EAVer\.*Left:*FActScore, SAFE, and VeriScore verify claims independently; VeriFastScore batches sentence\-level search\.*Center:*Fixed pipelines have four limitations: high API and LLM costs, no end\-to\-end training, cross\-claim search redundancy, and static claim processing\.*Right:*EAVergroups related claims, routes them to direct verification or search by confidence, and stores evidence for reuse\.
## Related Work
#### Long\-form factuality evaluation\.
FActScore established fine\-grained evaluation by decomposing a generation into atomic facts and measuring the fraction supported by a reliable knowledge source\([Min et al\. 2023](https://arxiv.org/html/2609.22223#bib.bib1)\)\. FacTool generalized tool\-augmented factuality detection across tasks through claim extraction, query generation, evidence collection, and agreement checking\([Chern et al\. 2023](https://arxiv.org/html/2609.22223#bib.bib2)\), while SAFE operationalized long\-form evaluation with LLM\-based decomposition and web\-search verification\([Wei et al\. 2024](https://arxiv.org/html/2609.22223#bib.bib3)\)\. VeriScore further distinguishes verifiable claims from subjective or otherwise uncheckable content\([Song et al\. 2024](https://arxiv.org/html/2609.22223#bib.bib4)\)\. More recent systems improve either search adaptivity or throughput\. FIRE iteratively searches for and verifies evidence until an individual claim can be decided\([Xie et al\. 2025](https://arxiv.org/html/2609.22223#bib.bib5)\), whereas VeriFastScore uses sentence\-level batch search and a fine\-tuned model to jointly extract and verify claims\([Rajendhran et al\. 2025](https://arxiv.org/html/2609.22223#bib.bib6)\)\. However, their overall control remains claim\-local or stage\-wise, with search and verification either prompt\-orchestrated or optimized separately\. Our work instead studies a joint policy for response\-level search and verification\.
#### Learned search agents\.
Recent work trains language models to treat search as part of the reasoning policy rather than a fixed preprocessing step\. R1\-Searcher and Search\-R1 use outcome\-based reinforcement learning to elicit autonomous, multi\-turn search\([Song et al\. 2025a](https://arxiv.org/html/2609.22223#bib.bib8);[Jin et al\. 2025](https://arxiv.org/html/2609.22223#bib.bib7)\), while ReSearch learns interleaved reasoning and search without supervised reasoning traces\([Chen et al\. 2025](https://arxiv.org/html/2609.22223#bib.bib9)\)\. Smart\-Searcher combines an SFT cold start with reinforcement learning to encourage dynamic use of internal and external knowledge\([Song et al\. 2025b](https://arxiv.org/html/2609.22223#bib.bib10)\)\. These methods primarily target knowledge\-intensive question answering, where search supports a single final answer\. Long\-form factuality verification instead requires exhaustive coverage of claims across an entire response, aligned claim\-level judgments, and control of cumulative search cost; our work studies learned search under this multi\-claim objective\.
## Task Definition
The input is a questionqqand a model responseyy\. For evaluation, the response is annotated with a gold set of factual claimsC=\{ci\}i=1nC=\\\{c\_\{i\}\\\}\_\{i=1\}^\{n\}, where each gold claim has a binary label
zi∈\{supported,unsupported\}\.z\_\{i\}\\in\\\{\\textsc\{supported\},\\textsc\{unsupported\}\\\}\.The system must predict a claim setC^=\{c^j\}j=1n^\\widehat\{C\}=\\\{\\widehat\{c\}\_\{j\}\\\}\_\{j=1\}^\{\\widehat\{n\}\}and assign each predicted claim a labelz^j∈\{supported,unsupported\}\\widehat\{z\}\_\{j\}\\in\\\{\\textsc\{supported\},\\textsc\{unsupported\}\\\}\. In our current evaluation, predicted claims are greedily matched one\-to\-one to gold claims using string similarity with threshold 0\.5\. For matched claims, predicted labels are scored against the corresponding gold labels\. We report claim\-extraction F1, macro\-F1, and False Support Rate \(FSR\)\.
#### Claim extraction F1\.
LetMMbe the set of matched predicted–gold claim pairs\. We compute extraction precision, recall, and their harmonic mean as
Pext=\|M\|\|C^\|,Rext=\|M\|\|C\|,F1ext=2PextRextPext\+Rext\.P\_\{\\mathrm\{ext\}\}=\\frac\{\|M\|\}\{\|\\widehat\{C\}\|\},\\qquad R\_\{\\mathrm\{ext\}\}=\\frac\{\|M\|\}\{\|C\|\},\\qquad F\_\{1\}^\{\\mathrm\{ext\}\}=\\frac\{2P\_\{\\mathrm\{ext\}\}R\_\{\\mathrm\{ext\}\}\}\{P\_\{\\mathrm\{ext\}\}\+R\_\{\\mathrm\{ext\}\}\}\.The last quantity is reported asclaim\_ext\_F1in Table[2](https://arxiv.org/html/2609.22223#Sx4.T2)\. The combined metric evaluates decomposition independently of factuality labels and penalizes both omitted gold claims and over\-decomposition with unmatched predictions\.
#### Macro\-F1\.
Macro\-F1 is the average of F1 forsupportedand F1 forunsupported\. It is the primary score because the supported class dominates the validation set\.
#### False Support Rate \(FSR\)\.
FSR is the fraction of goldunsupportedclaims incorrectly labeledsupported:
FSR=\#\(gold=unsupported,pred=supported\)\#\(gold=unsupported\)\.\\mathrm\{FSR\}=\\frac\{\\\#\(\\mathrm\{gold\}=\\textsc\{unsupported\},\\ \\mathrm\{pred\}=\\textsc\{supported\}\)\}\{\\\#\(\\mathrm\{gold\}=\\textsc\{unsupported\}\)\}\.Lower FSR is better because false support allows unsupported content to be accepted as factual\. FSR therefore isolates the asymmetric error that is most consequential for a factuality verifier\. Macro\-F1 summarizes performance across both classes, so reporting the two metrics together captures both balanced class performance and false\-support risk\.
## EAVer
EAVeris designed around three principles: \(1\) claim verification should be optimized as a single trajectory rather than a hand\-written pipeline, \(2\) search should be conditional on model confidence, and \(3\) evidence should be reusable across claims that share entities or events\.
### Policy\-Aware Verification Trajectory Synthesis
Final\-label supervision offers little guidance for the behavior preceding a factuality decision\. For example, the zero\-shot Qwen3\-8B verifier in Table[2](https://arxiv.org/html/2609.22223#Sx4.T2)has FSRs of 89\.18% on VeriFastScore and 89\.72% on FaStFact\-Bench, demonstrating a consistent tendency to accept unsupported claims as factual before policy training\. Search access alone does not teach the model when to doubt a plausible claim, seek counter\-evidence, or overturn the response\. Rather than distilling only final labels, we synthesize executable demonstrations that jointly supervise claim decomposition, confidence\-conditioned search, query formulation, evidence summarization, and negative verification\.
Figure[3](https://arxiv.org/html/2609.22223#Sx4.F3)illustrates one such real trajectory from the in\-domain VeriFastScore dataset\.
Figure 3:A representative EAVer inference trajectory\.Five claims form two entity groups\. The high\-confidence group is verified directly; the medium\-confidence group uses two searches and shared evidence before all five verdicts are returned\.We use Claude Opus 4\.8 as a privileged teacher during data construction\. It privately receives gold atomic claims and labels to construct label\-consistent trajectories; the student sees neither at inference\. The teacher copies each claim verbatim into a<claims\>JSON array and assigns two policy fields:confidence\(*high*,*medium*, or*low*\) and anentity\_key\. Claims sharing an entity key are processed together\. An all\-high\-confidence group is verified directly in a batch, whereas any medium\- or low\-confidence claim triggers one focused search for the group\. The trajectory uses five structured tags:<claims\>,<search\>,<memo\>,<verify\>, and<answer\>\. When the teacher ends a turn with<search\>query</search\>, generation pauses; the runtime parses the query, calls the search tool, injects its results as an<information\>environment message, and resumes with full history\. This yields multi\-turn tool\-interaction transcripts rather than post\-hoc textual rationales alone\.
After search, the teacher compresses the long, noisy observation into a compact<memo\>evidence summary before emitting one<verify\>decision per claim and an aligned<answer\>JSON\. The memo is an explicit evidence\-summary scaffold, not an external cache or lookup store: it remains in the shared dialogue context so that claims in the same entity group can implicitly reuse the evidence without a separate recall action\. We retain trajectories that pass structural, label\-alignment, tool\-use, search\-budget, and private\-key leakage checks\. Twenty human experts additionally conducted random spot checks of the retained trajectories, with 98% of inspected trajectories passing the audit\. The retained demonstrations provide process\-level supervision for claim organization, evidence acquisition and reuse, and aligned factuality decisions\.
Table 1:Synthetic training\-data overview\.Trajectory length counts generated text and injected search\-result observations after the initial prompt, giving a tokenizer\-independent measure\. Label shares are computed over 19,334 SFT decisions and 15,686 decisions in the DPO chosen completions\. Each rejected DPO completion shares the chosen completion’s searches and evidence and differs only at corrected decision labels; the 794 pairs comprise 405 false\-support\-to\-unsupported and 389 false\-unsupported\-to\-supported pairs, respectively\.Table 2:Main results\.Results onVeriFastScore\(in\-domain\) andFaStFact\-Bench\(out\-of\-domain, human\-annotated\)\. Searches/sample averages search\-API requests over both datasets; LLM calls are excluded, and “–” denotes no external search\. Evaluation metrics are percentages\. Best reported values in each column arebold\.
### Decision\-Focused Preference Post\-training
After supervised training on our policy\-aware synthetic trajectories, the model has learned the complete verification policy, including claim organization, search, evidence use, and factuality prediction\. Preference post\-training is therefore used only as an optional correction for residual verdict errors\. Applying trajectory\-level preference optimization to independently sampled completions would entangle the desired label correction with incidental differences in search and evidence\-use behavior\. Among models with comparable macro\-F1, we prefer the one with lower FSR: false support allows unsupported content to pass as factual, and identifying such negative cases is therefore more consequential for a verifier than confirming additional supported ones\. We use preference post\-training to seek an operating point with lower FSR while constraining changes to overall macro\-F1, and localize the update through both controlled preference construction and a decision\-focused loss\.
Specifically, we analyze rollouts produced by the policy trained on the synthetic trajectory corpus and select trajectories with incorrect verdicts for which the collected evidence is sufficient to support the gold decision\. In these cases, the model has already obtained the evidence needed for verification but still emits an incorrect label, indicating a decision failure rather than a search failure\. Rather than regenerate the full reasoning trajectory, we keep the trajectory fixed and correct only the erroneous verdict\. For each selected case, the original erroneous completion serves as the rejected response, while the chosen response is created by correcting the affected<verify\>labels and synchronized<answer\>entries to the gold decisions\. Claim grouping, search queries, search observations, evidence summaries, and all non\-decision text remain identical\. The resulting corpus contains 794 pairs spanning both false\-support\-to\-unsupported and false\-unsupported\-to\-supported corrections\. Because the surrounding completion is matched, the preference signal isolates the factuality decision rather than differences in tool use or language realization\.
We further restrict the DPO log\-probability calculation to the decision tokens expressingsupportedorunsupportedinside<verify\>and<answer\>\. Letmtm\_\{t\}be a binary mask that is one only for these decision tokens\. The resulting masked sequence score is
sθM\(y∣x\)=∑t=1\|y\|mtlogπθ\(yt∣x,y<t\)\.s\_\{\\theta\}^\{M\}\(y\\mid x\)=\\sum\_\{t=1\}^\{\|y\|\}m\_\{t\}\\log\\pi\_\{\\theta\}\(y\_\{t\}\\mid x,y\_\{<t\}\)\.\(1\)Defining the reference\-relative score asrθM\(y,x\)=sθM\(y∣x\)−srefM\(y∣x\)r\_\{\\theta\}^\{M\}\(y,x\)=s\_\{\\theta\}^\{M\}\(y\\mid x\)\-s\_\{\\mathrm\{ref\}\}^\{M\}\(y\\mid x\), and the preference margin asΔrθM=rθM\(y\+,x\)−rθM\(y−,x\)\\Delta r\_\{\\theta\}^\{M\}=r\_\{\\theta\}^\{M\}\(y^\{\+\},x\)\-r\_\{\\theta\}^\{M\}\(y^\{\-\},x\), our decision\-focused objective is
ℒDF\-DPO=−𝔼\(x,y\+,y−\)logσ\(βΔrθM\)\.\\mathcal\{L\}\_\{\\mathrm\{DF\\text\{\-\}DPO\}\}=\-\\mathbb\{E\}\_\{\(x,y^\{\+\},y^\{\-\}\)\}\\log\\sigma\\\!\\left\(\\beta\\Delta r\_\{\\theta\}^\{M\}\\right\)\.\(2\)All claim text, search, evidence, memo, and formatting tokens therefore havemt=0m\_\{t\}=0and are excluded from the preference loss; conventional DPO is recovered by settingmt=1m\_\{t\}=1for every completion token\. Spanning both correction directions prevents the training signal from reducing FSR simply by pushing every claim towardunsupported\. This optional stage directly supervises only the verdict tokens and does not reward changes to the preceding search trajectory\. Section[Main Results](https://arxiv.org/html/2609.22223#Sx6)evaluates whether this localized update reduces false support without degrading balanced factuality performance\.
## Experimental Setup
#### Backbone\.
We conduct our main evaluation ofEAVeron Qwen2\.5\-7B\-Instruct\([Qwen Team 2025](https://arxiv.org/html/2609.22223#bib.bib11)\)and Qwen3\-8B\([Yang and others 2025](https://arxiv.org/html/2609.22223#bib.bib12)\), providing comprehensive validation across two strong open\-weight models\. Additional ablations span model scales from 4B to 32B and multiple model families, including Llama 3\.1\([Grattafiori et al\. 2024](https://arxiv.org/html/2609.22223#bib.bib13)\), to assess the generalizability of our synthesized trajectories\.
#### Search setting\.
Following Search\-R1\([Jin et al\. 2025](https://arxiv.org/html/2609.22223#bib.bib7)\), we implement the search actions used during training with a local dense\-search backend, enabling stable and reproducible interaction trajectories\. Queries are encoded with E5\-base\-v2\([Wang et al\. 2022](https://arxiv.org/html/2609.22223#bib.bib14)\), and a FAISS index\([Johnson et al\. 2017](https://arxiv.org/html/2609.22223#bib.bib15)\)returns the top three search results\. This fixed local search backend avoids fluctuations in the latency, availability, and outputs of external search APIs\. For evaluation, we keep the model checkpoint,<search\>action protocol, and decoding configuration fixed, but replace the local search backend with the Google Search API without retraining the model\.
#### Data\.
We evaluate on two benchmarks\. For in\-domain evaluation, we construct a clean test set from the official 5,900\-example test split of VeriFastScore\([Rajendhran et al\. 2025](https://arxiv.org/html/2609.22223#bib.bib6)\)\. We remove persona and creative\-writing sources, abstentions and invalid responses, examples without valid binary claims, and any sample whose question–response fingerprint matches data used for training, trajectory synthesis, rollout generation, or development\. The resulting set contains 22,532 claims, with zero sample\-level fingerprint overlap with all touched pools\. For out\-of\-distribution evaluation, we use the independently released, human\-annotated FaStFact\-Bench\([Wan et al\. 2025](https://arxiv.org/html/2609.22223#bib.bib16)\)\. We map its fine\-grained labels to our binary scheme and discard non\-factual housekeeping annotations, abstentions, and samples without verifiable claims, yielding 6,954 claims\.
#### Baselines\.
We compare against two baseline families shown in Table[2](https://arxiv.org/html/2609.22223#Sx4.T2)\. The search\-free group, which uses no external search, includes LLM\-only Qwen2\.5\-7B\-Instruct\([Qwen Team 2025](https://arxiv.org/html/2609.22223#bib.bib11)\), LLM\-only Qwen3\-8B\([Yang and others 2025](https://arxiv.org/html/2609.22223#bib.bib12)\), Qwen3\-Max, GPT\-4\.1, Claude Sonnet 4\.6, and Claude Opus 4\.8, each using one model invocation per sample\. The search\-augmented group comprises FActScore\([Min et al\. 2023](https://arxiv.org/html/2609.22223#bib.bib1)\), SAFE\([Wei et al\. 2024](https://arxiv.org/html/2609.22223#bib.bib3)\), VeriScore\([Song et al\. 2024](https://arxiv.org/html/2609.22223#bib.bib4)\), FIRE\([Xie et al\. 2025](https://arxiv.org/html/2609.22223#bib.bib5)\), and VeriFastScore\([Rajendhran et al\. 2025](https://arxiv.org/html/2609.22223#bib.bib6)\)\. All experiments are run three times, and we report the mean over the three runs\.
## Main Results
Table[2](https://arxiv.org/html/2609.22223#Sx4.T2)reports the main comparison on two held\-out benchmarks under a unified evaluation protocol\. We group methods by category: zero\-shot LLM verifiers without search, search\-augmented pipelines with GPT\-4\.1 backends, andEAVer\.
#### Effectiveness\.
We select Qwen2\.5\-7B\-Instruct, Qwen3\-8B, and Llama\-3\.1\-8B as three representative backbones from the Qwen and Llama families to evaluate the effectiveness ofEAVeracross model architectures\. Among them,EAVer\(Qwen3\-8B\) achieves the best overall Macro\-F1 on both the in\-domain VeriFastScore benchmark and the out\-of\-domain FaStFact\-Bench, reaching 73\.43 and 74\.05, respectively\. Compared with the search\-free Qwen3\-8B verifier, training on our agentic verification trajectories reduces FSR from 89\.18 to 41\.57 on VeriFastScore and from 89\.72 to 35\.87 on FaStFact\-Bench\. This pronounced reduction shows that the gain is not inherited from the backbone alone: the trained policy learns to use evidence to identify unsupported claims rather than defaulting tosupported\.EAVer\(Qwen3\-8B\) also obtains the highest claim extraction F1 on both benchmarks\. By contrast, most prior search\-augmented pipelines exhibit a substantial drop in claim extraction F1 out of domain\. Our case inspection suggests that their static, LLM\-driven decomposition stages tend to over\-extract fine\-grained or redundant claims, inflating\|C^\|\|\\widehat\{C\}\|without a commensurate increase in matched gold claims and thereby lowering extraction precision and F1\. VeriFastScore attains the lowest in\-domain FSR, but this result should be interpreted in light of its two\-stage fine\-tuning on roughly 9K synthetic prompt–response pairs and its collection of sentence\-level evidence before claim decomposition\([Rajendhran et al\. 2025](https://arxiv.org/html/2609.22223#bib.bib6)\)\. Finally, Claude Opus 4\.8 achieves a strong 72\.21 Macro\-F1 on FaStFact\-Bench as a search\-free verifier, providing empirical support for its use as our privileged teacher\. Although neither trajectory synthesis nor post\-training uses FaStFact\-Bench examples or annotations, the resulting Qwen3\-8B student reaches 74\.05 Macro\-F1, suggesting that the teacher\-guided agentic supervision transfers beyond its source data\.
#### Cost\-efficient verification\.
For a like\-for\-like comparison, Table[2](https://arxiv.org/html/2609.22223#Sx4.T2)counts external search requests only\.EAVer\(Qwen3\-8B\) averages 3\.14 searches per sample, about 80% fewer than the 15\.62 used by VeriFastScore, the most search\-efficient baseline\. This measure does not include the LLM calls incurred by multi\-stage orchestration\. For example, SAFE can require up to2\+5×\(1\+1\)\+1=132\+5\\times\(1\+1\)\+1=13calls for a single relevant claim: 2 LLM calls for self\-containment and relevance, 5 query\-generation calls paired with 5 Google Search calls, and 1 final verdict call\([Wei et al\. 2024](https://arxiv.org/html/2609.22223#bib.bib3)\)\. In contrast,EAVergroups related claims, directly verifies high\-confidence groups, searches only for uncertain groups, and reuses evidence returned by search through in\-context memos, allowing one search to support decisions for several related claims at once\.
Figure 4:Decision\-focused DPO lowers FSR for EAVer on VeriFastScore\.Results are reported on VeriFastScore across three backbones\. Circles showEAVerbefore DPO, and diamonds showEAVerafter DPO\. Arrows connect matched checkpoints; leftward movement indicates lower FSR, while upward movement indicates higher Macro\-F1\. Colors identify backbones\.Figure 5:Mechanism\-level component analysis for Qwen3\-8B\.\(a\) On FaStFact\-Bench, entity grouping reduces token\-level trajectory length per sample\. \(b\) On FaStFact\-Bench, disabling memo\-based evidence reuse increases searches per sample and repeated\-link redundancy\. \(c\) On VeriFastScore, bars show claim share and final verification accuracy by confidence level\.
#### Decision\-focused preference post\-training\.
Figure[4](https://arxiv.org/html/2609.22223#Sx6.F4)evaluates decision\-focused DPO applied toEAVeron VeriFastScore across three backbones under fixed search and decoding settings\. Across all three backbones, applying DPO toEAVerlowers FSR while preserving or modestly improving Macro\-F1\. For Qwen2\.5\-7B, which exhibits the largest FSR reduction, FSR drops from 38\.51 to 34\.16, while Macro\-F1 rises from 71\.77 to 72\.31\. The consistent upper\-left shifts show that preference post\-training strengthens the ability ofEAVerto reject unsupported claims rather than merely shifting the prediction prior\.
## Analysis and Extensions
We organize the remaining experiments around two questions: whetherEAVer’s core components jointly enable efficient and reliable verification, and whether the learned agentic policy generalizes across model generations, parameter scales, and families alike\.
### Component Ablations
Table 3:Component ablations\.Entity grouping shortens trajectories, confidence routing avoids unnecessary search, and memo reuse limits redundant searches\.Figure[5](https://arxiv.org/html/2609.22223#Sx6.F5)\(a,b\) and Table[3](https://arxiv.org/html/2609.22223#Sx7.T3)report FaStFact\-Bench ablations, while panel \(c\) diagnoses confidence behavior on VeriFastScore\. In panel \(a\), grouping cuts trajectory length from 4,821 to 2,474 tokens \(49%\) through shared claim context\. In panel \(b\), disabling memos increases searches per sample from 4\.10 to 5\.72 and repeated links from 0\.3% to 7\.4%, confirming that evidence reuse avoids redundant retrieval\. Panel \(c\) shows a non\-monotonic relation between confidence and accuracy: high\-, medium\-, and low\-confidence claims reach 85\.1%, 67\.3%, and 80\.4%, respectively\. Medium\-confidence claims can be deceptively familiar, prompting broad yet non\-discriminative queries\. For example, verifying “Einstein won the Nobel Prize in 1922” may search pages that mention both the 1921 prize and its 1922 presentation, making the award year easy to misread\. Low\-confidence claims often contain niche names or technical terms that support narrower queries and better\-matched evidence\.
### Generalization of Model Families and Sizes
Using the same training and evaluation configuration, we testEAVeron six Qwen and Llama models ranging from 4B to 32B on the out\-of\-domain FaStFact\-Bench\.
Figure 6:Cross\-backbone generalization in Macro\-F1 and FSR\.Base andEAVer\-trained models are evaluated on the out\-of\-distribution FaStFact\-Bench across six Qwen and Llama backbones\. Higher Macro\-F1 and lower FSR indicate better performance;EAVerimproves both metrics for every tested model and scale\.Figure[6](https://arxiv.org/html/2609.22223#Sx7.F6)shows thatEAVerconsistently improves long\-form verification across both model families and at every tested scale, increasing Macro\-F1 and reducing FSR for all six backbones\. The improvement magnitude varies across backbones, while EAVer improves Macro\-F1 and reduces FSR for all six models\. Qwen2\.5\-7B, the weakest base model in our evaluation, achieves a 17\.4\-point Macro\-F1 improvement after policy training\. This demonstrates policy transfer across model generations, families, and scales without backbone\-specific synthesis\.
## Conclusion
Privileged\-teacher trajectories trainEAVerto jointly learn claim grouping, selective search, and evidence reuse, improving factuality with about 80% fewer searches and transferring across model families and scales\. This supports jointly learning search, evidence retention, and factuality decisions\.
## References
- T\. B\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell, S\. Agarwal, A\. Herbert\-Voss, G\. Krueger, T\. Henighan, R\. Child, A\. Ramesh, D\. M\. Ziegler, J\. Wu, C\. Winter, C\. Hesse, M\. Chen, E\. Sigler, M\. Litwin, S\. Gray, B\. Chess, J\. Clark, C\. Berner, S\. McCandlish, A\. Radford, I\. Sutskever, and D\. AmodeiLanguage models are few\-shot learners\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 1877–1901\.External Links:[Link](https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html)Cited by:[Introduction](https://arxiv.org/html/2609.22223#Sx1.p1.1)\.
- Chenet al\.\(2025\)M\. Chen, L\. Sun, T\. Li, H\. Sun, Y\. Zhou, C\. Zhu, H\. Wang, J\. Z\. Pan, W\. Zhang, H\. Chen, F\. Yang, Z\. Zhou, and W\. ChenReSearch: learning to reason with search for LLMs via reinforcement learning\.InAdvances in Neural Information Processing Systems,External Links:2503\.19470,[Link](https://arxiv.org/abs/2503.19470)Cited by:[Learned search agents\.](https://arxiv.org/html/2609.22223#Sx2.SS0.SSS0.Px2.p1.1)\.
- Chernet al\.\(2023\)I\. Chern, S\. Chern, S\. Chen, W\. Yuan, K\. Feng, C\. Zhou, J\. He, G\. Neubig, and P\. LiuFacTool: factuality detection in generative ai – a tool augmented framework for multi\-task and multi\-domain scenarios\.External Links:2307\.13528,[Link](https://arxiv.org/abs/2307.13528)Cited by:[Long\-form factuality evaluation\.](https://arxiv.org/html/2609.22223#Sx2.SS0.SSS0.Px1.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.The Llama 3 herd of models\.External Links:2407\.21783,[Document](https://dx.doi.org/10.48550/arXiv.2407.21783),[Link](https://arxiv.org/abs/2407.21783)Cited by:[Backbone\.](https://arxiv.org/html/2609.22223#Sx5.SSx2.SSS0.Px1.p1.1)\.
- Jinet al\.\(2025\)B\. Jin, H\. Zeng, Z\. Yue, J\. Yoon, S\. O\. Arik, D\. Wang, H\. Zamani, and J\. HanSearch\-R1: training LLMs to reason and leverage search engines with reinforcement learning\.InSecond Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=Rwhi91ideu)Cited by:[Learned search agents\.](https://arxiv.org/html/2609.22223#Sx2.SS0.SSS0.Px2.p1.1),[Search setting\.](https://arxiv.org/html/2609.22223#Sx5.SSx2.SSS0.Px2.p1.1)\.
- Johnsonet al\.\(2017\)J\. Johnson, M\. Douze, and H\. JegouBillion\-scale similarity search with GPUs\.External Links:1702\.08734,[Link](https://arxiv.org/abs/1702.08734)Cited by:[Search setting\.](https://arxiv.org/html/2609.22223#Sx5.SSx2.SSS0.Px2.p1.1)\.
- Linet al\.\(2022\)S\. Lin, J\. Hilton, and O\. EvansTruthfulQA: measuring how models mimic human falsehoods\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Dublin, Ireland,pp\. 3214–3252\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.229),[Link](https://aclanthology.org/2022.acl-long.229/)Cited by:[Introduction](https://arxiv.org/html/2609.22223#Sx1.p1.1)\.
- Manakulet al\.\(2023\)P\. Manakul, A\. Liusie, and M\. GalesSelfCheckGPT: zero\-resource black\-box hallucination detection for generative large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,Singapore,pp\. 9004–9017\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.557),[Link](https://aclanthology.org/2023.emnlp-main.557/)Cited by:[Introduction](https://arxiv.org/html/2609.22223#Sx1.p1.1)\.
- Minet al\.\(2023\)S\. Min, K\. Krishna, X\. Lyu, M\. Lewis, W\. Yih, P\. Koh, M\. Iyyer, L\. Zettlemoyer, and H\. HajishirziFActScore: fine\-grained atomic evaluation of factual precision in long form text generation\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,Singapore,pp\. 12076–12100\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.741),[Link](https://aclanthology.org/2023.emnlp-main.741/)Cited by:[Introduction](https://arxiv.org/html/2609.22223#Sx1.p2.1),[Long\-form factuality evaluation\.](https://arxiv.org/html/2609.22223#Sx2.SS0.SSS0.Px1.p1.1),[Table 2](https://arxiv.org/html/2609.22223#Sx4.T2.1.11.1),[Baselines\.](https://arxiv.org/html/2609.22223#Sx5.SSx2.SSS0.Px4.p1.1)\.
- Ouyanget al\.\(2022\)L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray, J\. Schulman, J\. Hilton, F\. Kelton, L\. Miller, M\. Simens, A\. Askell, P\. Welinder, P\. F\. Christiano, J\. Leike, and R\. LoweTraining language models to follow instructions with human feedback\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 27730–27744\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/hash/b1efde53be364a73914f58805a001731-Abstract.html)Cited by:[Introduction](https://arxiv.org/html/2609.22223#Sx1.p1.1)\.
- Qwen Team \(2025\)Qwen TeamQwen2\.5 technical report\.External Links:2412\.15115,[Link](https://arxiv.org/abs/2412.15115)Cited by:[Backbone\.](https://arxiv.org/html/2609.22223#Sx5.SSx2.SSS0.Px1.p1.1),[Baselines\.](https://arxiv.org/html/2609.22223#Sx5.SSx2.SSS0.Px4.p1.1)\.
- Rafailovet al\.\(2023\)R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. FinnDirect preference optimization: your language model is secretly a reward model\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 53728–53741\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/a85b405ed65c6477a4fe8302b5e06ce7-Abstract-Conference.html)Cited by:[Introduction](https://arxiv.org/html/2609.22223#Sx1.p4.1)\.
- Rajendhranet al\.\(2025\)R\. Rajendhran, A\. Zadeh, M\. Sarte, C\. Li, and M\. IyyerVeriFastScore: speeding up long\-form factuality evaluation\.InFindings of the Association for Computational Linguistics: EMNLP 2025,Suzhou, China,pp\. 9234–9259\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.491),[Link](https://aclanthology.org/2025.findings-emnlp.491/)Cited by:[Introduction](https://arxiv.org/html/2609.22223#Sx1.p2.1),[Introduction](https://arxiv.org/html/2609.22223#Sx1.p3.1),[Long\-form factuality evaluation\.](https://arxiv.org/html/2609.22223#Sx2.SS0.SSS0.Px1.p1.1),[Table 2](https://arxiv.org/html/2609.22223#Sx4.T2.1.15.1),[Data\.](https://arxiv.org/html/2609.22223#Sx5.SSx2.SSS0.Px3.p1.1),[Baselines\.](https://arxiv.org/html/2609.22223#Sx5.SSx2.SSS0.Px4.p1.1),[Effectiveness\.](https://arxiv.org/html/2609.22223#Sx6.SSx2.SSS0.Px1.p1.1)\.
- Schicket al\.\(2023\)T\. Schick, J\. Dwivedi\-Yu, R\. Dessi, R\. Raileanu, M\. Lomeli, E\. Hambro, L\. Zettlemoyer, N\. Cancedda, and T\. ScialomToolformer: language models can teach themselves to use tools\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 68539–68551\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/d842425e4bf79ba039352da0f658a906-Abstract-Conference.html)Cited by:[Introduction](https://arxiv.org/html/2609.22223#Sx1.p3.1)\.
- Songet al\.\(2025a\)H\. Song, J\. Jiang, Y\. Min, J\. Chen, Z\. Chen, W\. X\. Zhao, L\. Fang, and J\. WenR1\-Searcher: incentivizing the search capability in LLMs via reinforcement learning\.External Links:2503\.05592,[Link](https://arxiv.org/abs/2503.05592)Cited by:[Learned search agents\.](https://arxiv.org/html/2609.22223#Sx2.SS0.SSS0.Px2.p1.1)\.
- Songet al\.\(2025b\)H\. Song, J\. Jiang, W\. Tian, Z\. Chen, Y\. Wu, J\. Zhao, Y\. Min, W\. X\. Zhao, L\. Fang, and J\. WenSmart\-Searcher: incentivizing the dynamic knowledge acquisition of LLMs via reinforcement learning\.InFindings of the Association for Computational Linguistics: EMNLP 2025,Suzhou, China,pp\. 13572–13586\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.731),[Link](https://aclanthology.org/2025.findings-emnlp.731/)Cited by:[Learned search agents\.](https://arxiv.org/html/2609.22223#Sx2.SS0.SSS0.Px2.p1.1)\.
- Songet al\.\(2024\)Y\. Song, Y\. Kim, and M\. IyyerVeriScore: evaluating the factuality of verifiable claims in long\-form text generation\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Miami, Florida, USA,pp\. 9447–9474\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.552),[Link](https://aclanthology.org/2024.findings-emnlp.552/)Cited by:[Introduction](https://arxiv.org/html/2609.22223#Sx1.p2.1),[Long\-form factuality evaluation\.](https://arxiv.org/html/2609.22223#Sx2.SS0.SSS0.Px1.p1.1),[Table 2](https://arxiv.org/html/2609.22223#Sx4.T2.1.13.1),[Baselines\.](https://arxiv.org/html/2609.22223#Sx5.SSx2.SSS0.Px4.p1.1)\.
- Wanet al\.\(2025\)Y\. Wan, H\. Tan, X\. Zhu, X\. Zhou, Z\. Li, Q\. Lv, C\. Sun, J\. Zeng, Y\. Xu, J\. Lu, Y\. Liu, and Z\. GuoFaStFact: faster, stronger long\-form factuality evaluations in LLMs\.InFindings of the Association for Computational Linguistics: EMNLP 2025,Suzhou, China,pp\. 23814–23854\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1295),[Link](https://aclanthology.org/2025.findings-emnlp.1295/)Cited by:[Data\.](https://arxiv.org/html/2609.22223#Sx5.SSx2.SSS0.Px3.p1.1)\.
- Wanget al\.\(2022\)L\. Wang, N\. Yang, X\. Huang, B\. Jiao, L\. Yang, D\. Jiang, R\. Majumder, and F\. WeiText embeddings by weakly\-supervised contrastive pre\-training\.External Links:2212\.03533,[Link](https://arxiv.org/abs/2212.03533)Cited by:[Search setting\.](https://arxiv.org/html/2609.22223#Sx5.SSx2.SSS0.Px2.p1.1)\.
- Wanget al\.\(2023\)Y\. Wang, Y\. Kordi, S\. Mishra, A\. Liu, N\. A\. Smith, D\. Khashabi, and H\. HajishirziSelf\-Instruct: aligning language models with self\-generated instructions\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Toronto, Canada,pp\. 13484–13508\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.754),[Link](https://aclanthology.org/2023.acl-long.754/)Cited by:[Introduction](https://arxiv.org/html/2609.22223#Sx1.p3.1)\.
- Weiet al\.\(2024\)J\. Wei, C\. Yang, X\. Song, Y\. Lu, N\. Hu, J\. Huang, D\. Tran, D\. Peng, R\. Liu, D\. Huang, C\. Du, and Q\. V\. LeLong\-form factuality in large language models\.External Links:2403\.18802,[Link](https://arxiv.org/abs/2403.18802)Cited by:[Introduction](https://arxiv.org/html/2609.22223#Sx1.p2.1),[Long\-form factuality evaluation\.](https://arxiv.org/html/2609.22223#Sx2.SS0.SSS0.Px1.p1.1),[Table 2](https://arxiv.org/html/2609.22223#Sx4.T2.1.12.1),[Baselines\.](https://arxiv.org/html/2609.22223#Sx5.SSx2.SSS0.Px4.p1.1),[Cost\-efficient verification\.](https://arxiv.org/html/2609.22223#Sx6.SSx2.SSS0.Px2.p1.1)\.
- Xieet al\.\(2025\)Z\. Xie, R\. Xing, Y\. Wang, J\. Geng, H\. Iqbal, D\. Sahnan, I\. Gurevych, and P\. NakovFIRE: fact\-checking with iterative retrieval and verification\.InFindings of the Association for Computational Linguistics: NAACL 2025,Albuquerque, New Mexico,pp\. 2901–2914\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-naacl.158),[Link](https://aclanthology.org/2025.findings-naacl.158/)Cited by:[Introduction](https://arxiv.org/html/2609.22223#Sx1.p2.1),[Long\-form factuality evaluation\.](https://arxiv.org/html/2609.22223#Sx2.SS0.SSS0.Px1.p1.1),[Table 2](https://arxiv.org/html/2609.22223#Sx4.T2.1.14.1),[Baselines\.](https://arxiv.org/html/2609.22223#Sx5.SSx2.SSS0.Px4.p1.1)\.
- Yanget al\.\(2025\)A\. Yanget al\.Qwen3 technical report\.External Links:2505\.09388,[Document](https://dx.doi.org/10.48550/arXiv.2505.09388),[Link](https://arxiv.org/abs/2505.09388)Cited by:[Backbone\.](https://arxiv.org/html/2609.22223#Sx5.SSx2.SSS0.Px1.p1.1),[Baselines\.](https://arxiv.org/html/2609.22223#Sx5.SSx2.SSS0.Px4.p1.1)\.
- Yaoet al\.\(2023\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. R\. Narasimhan, and Y\. CaoReAct: synergizing reasoning and acting in language models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=WE_vluYUL-X)Cited by:[Introduction](https://arxiv.org/html/2609.22223#Sx1.p3.1)\.
## Appendix AAppendix: Evaluation and Reproducibility Details
#### Evaluation protocol\.
The agent generates a complete trajectory containing<claims\>,<search\>,<memo\>, and<verify\>blocks, and concludes with an<answer\>JSON listing all predicted claims and labels\. Predicted claims are matched to gold claims by string similarity \(SequenceMatcher, threshold 0\.5\); matched claims contribute to confusion\-matrix metrics\. We report*claim coverage*\(fraction of gold claims matched\),*Macro F1*\(mean of supported\-F1 and unsupported\-F1, our primary metric\),*FSR*\(the fraction of gold\-unsupportedclaims that are predicted assupported\),*searches per sample*, and*searches per claim*\(SPC\)\. Agentic evaluation usesmax\_turns=20=20andmax\_new\_tokens=8192=8192\.
## Appendix BTraining Design and Hyperparameters
This section provides the optimization, tokenization, and systems details for the two policy\-training stages\. The first stage teaches the complete verification behavior from executable demonstrations\. The second stage starts from the resulting policy and applies a conservative preference update to factuality decisions\. All backbone comparisons use the same 1,447 demonstrations, action schema, and optimization settings unless an architecture\-specific memory constraint is stated explicitly\.
### Policy Demonstration Training
#### Training objective\.
Each training item is a complete multi\-turn interaction consisting of a system instruction, a question and long\-form model response, assistant actions, retrieval observations, evidence memos, claim\-level decisions, and the final structured answer\. We optimize causal language\-model likelihood only on assistant outputs\. Tokens belonging to the system, user, and retrieval\-environment turns are assigned the ignore index, as is the assistant role prefix\. The supervised target contains the assistant content and its end\-of\-turn marker\. Consequently, the policy learns to generate<claims\>,<search\>,<memo\>,<verify\>, and<answer\>outputs while conditioning on, but not imitating, the question, source response, or tool observation\. This assistant\-only objective also prevents long retrieved passages from dominating the loss\.
#### Sequence construction\.
We serialize each sample with the native conversation markers of its backbone family and preserve all turns in chronological order\. For Qwen2\.5 and Qwen3, each message is enclosed by<\|im\_start\|\>and<\|im\_end\|\>markers\. For Llama 3\.1, we use the corresponding header and end\-of\-turn tokens and omit the auxiliary knowledge\-date preamble so that training and evaluation use identical templates\. We do not insert a separate reasoning marker and do not pack multiple samples into one sequence\. Sequences are right\-truncated to 8,192 tokens and padded to a multiple of eight; if a tokenizer has no dedicated padding token, its end\-of\-sequence token is used for padding\.
Table 4:Shared policy demonstration training configuration\.The same semantic training corpus and objective are used for all backbones\. Only the memory\-sensitive settings listed in Table[5](https://arxiv.org/html/2609.22223#A2.T5)vary\.
#### Optimization and precision\.
Table[4](https://arxiv.org/html/2609.22223#A2.T4)summarizes the shared configuration\. Training uses the TRL supervised fine\-tuning trainer, bfloat16 model and optimizer computation, FlashAttention\-2, and gradient checkpointing with non\-reentrant recomputation\. We use AdamW with a peak learning rate of2×10−52\\times 10^\{\-5\}, cosine decay, a 0\.05 warmup ratio, weight decay 0\.01, and gradient clipping at 1\.0\. The per\-device batch size is one\. On eight GPUs, two gradient\-accumulation steps give an effective global batch of 16 for the models up to 14B parameters\. Each configuration is trained independently with seeds 42, 123, and 2027\. Logging is performed every five optimizer steps\. Checkpoints are written at epoch boundaries and contain model weights only, which avoids retaining optimizer states solely for evaluation\.
#### Distributed execution\.
All reported policies are trained on a single node with eight NVIDIA H20 GPUs, each with approximately 95–96 GB of memory\. Models up to 14B parameters use DeepSpeed ZeRO\-2 with optimizer\-state CPU offload\. The dense 32B model uses ZeRO\-3, which shards parameters across all eight GPUs and offloads optimizer states to host memory\. For the 32B setting, we additionally reduce gradient accumulation from two to one and disable data\-loader worker processes to keep peak host memory below the node limit\. These changes affect systems memory only; the demonstration corpus, loss mask, learning rate, schedule, and sequence length remain unchanged\.
Table 5:Backbone\-specific training settings\.The selected epoch is used for the reported evaluations\. The 32B changes are imposed by host and device memory constraints rather than backbone\-specific tuning\.
### Decision\-Focused Preference Training
The input is the 794\-pair bidirectional corpus described in the preference\-pair construction below\. Each pair contains a chosen and rejected completion with identical claim decomposition, tool interaction, retrieved evidence, memo content, and formatting\. Only synchronized factuality labels differ\.
The model attends to the full prompt and completion, including all search observations, but the DPO sequence score is accumulated only over thesupportedandunsupporteddecision tokens inside<verify\>and<answer\>\. The decision\-token mask therefore changes which positions contribute to the preference objective without deleting contextual evidence from the forward pass\. This distinction is essential: the policy can use the complete trajectory to evaluate a label while the preference gradient cannot reward incidental changes to search queries, evidence wording, or memo style\.
Table 6:Decision\-focused preference training configuration\.Table[6](https://arxiv.org/html/2609.22223#A2.T6)gives the complete preference\-training configuration\. We train for three epochs with a per\-device batch of one and effective global batch of 32 across eight GPUs\. The learning rate is5×10−75\\times 10^\{\-7\}with cosine decay and a 0\.10 warmup ratio\. We use sigmoid DPO withβ=0\.1\\beta=0\.1, zero weight decay, no auxiliary negative\-log\-likelihood term, bfloat16 precision, gradient checkpointing, and a paged 8\-bit AdamW optimizer\. The maximum combined sequence length is 5,120 tokens and the maximum prompt length is 2,048 tokens\. Reference\-model log probabilities are precomputed with the same decision mask\. Models are saved once per epoch; final checkpoint selection uses the disjoint development set and the constrained criterion described in the main paper\.
## Appendix CPolicy Demonstration and Preference Data Construction
This section describes the two training corpora at the level of their statistical construction rather than implementation details\. We refer to the first corpus as*policy demonstrations*because it supervises the complete behavior of the verifier, not only its final factuality labels\. The second corpus contains controlled*preference pairs*that refine the decision rule without changing the learned search policy\.
### Policy Demonstration Synthesis
#### Source pool and sampling\.
The source examples are drawn from the VeriFastScore training pool and contain a question, a long\-form model response, and an ordered list of atomic claims with binary gold labels\. We retain only examples with at least one valid binary claim and no more than 40 claims\. Invalid responses, abstentions, and examples dominated by persona or creative writing are excluded\. Sampling covers the eight source categories that occur in the corresponding development distribution: ELI5, AskHistorian, FreshQA, FActScore, new books, LongFacts, ShareGPT, and WritingPrompts\. Source quotas follow the development distribution rather than oversampling examples with unsupported claims\. This preserves a realistic mixture of supported, unsupported, and mixed\-label samples\.
Before synthesis, normalized question–response fingerprints are compared against every development and test pool\. Any overlap is removed\. The same fingerprint audit is repeated after merging synthesis batches and before model fitting\. Because the claim annotations are used as privileged synthesis inputs, they are never used to select examples from the held\-out evaluation benchmarks\.
#### Controlled retrieval environment\.
Teacher trajectories interact with the same retrieval protocol later exposed to the student policy\. The corpus is formed by deduplicating the title and content of retrieval snippets associated with VeriFastScore claims, yielding 1,098,517 unique passages\. An E5\-base\-v2 encoder maps queries and passages to 768\-dimensional vectors\. A flat FAISS inner\-product index returns the top three passages for every search\. The retrieved passages are injected into the next environment turn inside an<information\>block\. Using a fixed local index keeps evidence stable across synthesis shards, training, and controlled evaluation, and ensures that the teacher cannot rely on capabilities unavailable to the student\.
#### Privileged\-teacher principle\.
A strong teacher receives the question, source response, exact ordered gold claims, and gold binary labels\. This private key is used for three purposes only: \(1\) the<claims\>array must reproduce the annotated atomic claims verbatim and in order; \(2\) confidence values are calibrated so that ambiguous details are routed to retrieval; and \(3\)<verify\>and<answer\>decisions must match the annotated labels\. The teacher must nevertheless reach those decisions through common knowledge or retrieved evidence\. It is forbidden to mention the private key, labels, hints, supervision, or any equivalent source of privileged information in its generated trajectory\.
Listing[1](https://arxiv.org/html/2609.22223#LST1)shows the central portion of the teacher prompt\. The omitted portions give detailed confidence examples, entity\-key naming rules, and the same structural constraints that are checked automatically after generation\.
Youareanexpertlong\-formfactualityverificationagentproducinga
high\-qualityreasoningdemonstrationforastudentmodeltolearnfrom\.
\[PRIVILEGEDANSWERKEY:neverrevealorreferenceinoutput\]
Youaregiven:
\-theexactatomicclaimstoextract,and
\-thecorrectsupportedorunsupportedlabelforeachclaim\.
Thekeyhasonlythreeuses:
\(1\)reproducetheclaimsverbatimandinorderin<claims\>;
\(2\)calibrateconfidenceforsearchrouting;
\(3\)reachthesamefinallabelsthroughgenuineevidencereasoning\.
Neverrefertothekey,goldlabels,hints,groundtruth,privileged
information,orprovidedsupervisioninanyoutputtoken\.
Evidencerule:
supported=evidencesubstantiatesthespecificclaim\.
unsupported=evidencecontradictstheclaimordoesnotconfirmits
specificdate,number,name,orcausallink\.
Hardconstraints:
\-<claims\>isparseableJSONwithconsecutiveids\.
\-<search\>isthelasttaginanactionturn\.
\-Everysearchisfollowedbyone<memo\>beforeverification\.
\-Claimssharinganentity\_keyareprocessedinonegroup\.
\-Everyclaimreceivesexactlyone<verify\>\.
\-<answer\>containsallclaimsinorder\.
\-Atmosteightsearchesareissued\.
Listing 1:Privileged\-teacher instruction excerpt used during policy demonstration synthesis\.The corresponding teacher user message is instantiated separately for each sample:
Question:
\{question\}
Responsetoverify:
\{response\}
\[PRIVATESUPERVISION:neverrevealorreferenceinanyoutputtoken\]
DecomposetheresponseintoEXACTLYthese\{N\}atomicclaims,verbatimand
inthegivenorder,andreachEXACTLYtheselabelsin<verify\>and<answer\>:
1\.\[supported\]\{claim\_1\}
2\.\[unsupported\]\{claim\_2\}
\.\.\.
N\.\[supported\|unsupported\]\{claim\_N\}
Alloutputmustreadasagenuineinvestigation\.Startwith<claims\>\.
Listing 2:Teacher user\-message template\. Braced fields are filled per sample\.
#### Five\-tag trajectory schema\.
Every retained demonstration follows one compact action language:<claims\>specifies claim decomposition and routing,<search\>invokes the only external tool,<memo\>stores a short evidence summary in the shared context,<verify\>records one claim decision, and<answer\>aggregates all claims and labels\. We intentionally omit separate route, recall, and free\-form reasoning tags\. The reduced schema exposes the decisions that matter for control while keeping evidence reuse implicit in the conversation history\.
The<claims\>array contains four fields per claim\.idis a consecutive integer,claimis the verbatim atomic claim,confidenceis*high*,*medium*, or*low*, andentity\_keyis a normalized central entity or topic\. Claims with the same entity key form a group\. If all members of a group have high confidence, the teacher verifies them directly in a single batch\. If any member has medium or low confidence, the group issues one focused search, writes one memo, and then verifies every member, including any high\-confidence members, against the shared context\. This makes search count depend on uncertain entity groups rather than raw claim count\.
<claims\>
\[
\{"id":1,
"claim":"EinsteinwontheNobelPrizeinPhysics\.",
"confidence":"high",
"entity\_key":"einstein"\},
\{"id":2,
"claim":"EinsteinwontheNobelPrizein1922\.",
"confidence":"medium",
"entity\_key":"einstein"\}
\]
</claims\>
<search\>EinsteinNobelPrizeinPhysicsawardyear</search\>
<information\>
Doc1statesthatthe1921NobelPrizeinPhysicswasawardedto
AlbertEinsteinandpresentedin1922\.
</information\>
<memo\>Einsteinreceivedthe1921NobelPrizeinPhysics\(Doc1\)\.</memo\>
<verifyid=1\>supported</verify\>
<verifyid=2\>unsupported</verify\>
<answer\>\{"claims":\[
\{"claim":"EinsteinwontheNobelPrizeinPhysics\.",
"label":"supported"\},
\{"claim":"EinsteinwontheNobelPrizein1922\.",
"label":"unsupported"\}
\]\}</answer\>
Listing 3:Illustrative five\-tag policy trajectory\. Retrieval observations are generated by the environment\.
#### Confidence calibration\.
High confidence is reserved for stable, textbook\-level facts that the teacher can recall specifically\. Medium confidence denotes familiar subject matter with uncertainty about an exact date, quantity, name, or causal relation\. Low confidence denotes niche, recent, or technical content that requires retrieval\. The teacher is explicitly instructed not to mark every claim low, which would teach wasteful search, or every claim high, which would suppress tool use\. Calibration is evaluated indirectly through schema consistency, search placement, and the empirical distribution of searches rather than by treating confidence as an independently supervised target\.
#### Entity grouping and evidence reuse\.
Entity keys are lowercase, snake\-case identifiers of the central subject\. They are intentionally broader than an individual claim\. For example, claims about Einstein’s award, birthplace, and nationality share the keyeinstein; a key such aseinstein\_birth\_year\_1879would be rejected as unnecessarily narrow\. Groups are processed in first\-appearance order\. A memo is a one\-sentence summary that retains the decisive names, dates, numbers, and a compact document reference\. Because the memo remains in the dialogue history, later decisions can attend to it directly\. There is no external memory store and no separate recall operation\.
#### Interactive rollout\.
Generation is paused whenever the teacher completes a<search\>action\. The environment parses the query, returns three passages inside an<information\>turn, and resumes generation with the complete history\. The search closing tag is used as a generation stop condition, ensuring that the tool action is the final tag of its turn\. At most eight searches and 30 assistant\-generation calls are allowed for one sample\. A generation that reaches its length limit without a complete<answer\>is discarded rather than truncated into a nominally valid demonstration\.
#### Quality\-control gates\.
Every synthesized sample passes all of the following checks before entering the training corpus:
1. 1\.Schema validity\.The<claims\>payload is parseable JSON; claim identifiers are consecutive; confidence values belong to the three\-level vocabulary; and entity keys satisfy the normalized naming rule\.
2. 2\.Decomposition alignment\.The number, order, and stripped text of generated claims match the annotated claims exactly\. This prevents teacher paraphrases from changing the unit scored by the evaluator\.
3. 3\.Tool protocol\.Search actions occur at turn boundaries, stay within the eight\-search budget, and are followed by a memo before the corresponding decisions\.
4. 4\.Decision alignment\.Every claim identifier occurs in exactly one<verify\>; its label equals the annotation; and the final<answer\>repeats all claims and labels in order\.
5. 5\.Leakage rejection\.Every assistant turn is scanned case\-insensitively for references to an answer key, gold labels, privileged information, supervision, hints, or equivalent phrasing\. A match rejects the complete sample rather than editing the offending phrase\.
6. 6\.Student\-view reconstruction\.The privileged system and user turns are removed and replaced with the deployment prompt\. Assistant actions and retrieval observations are preserved verbatim\. The reconstructed conversation is scanned again\.
7. 7\.Dataset\-level audit\.Duplicate identifiers and normalized question–response fingerprints are removed, and overlap with all development and held\-out sets is required to be zero\.
After quality control and deduplication, the final corpus contains 1,447 trajectories\. The multi\-stage rejection procedure is deliberately conservative: a malformed, misaligned, truncated, or potentially leaked sample is dropped rather than repaired\. Random human spot checks provide an additional semantic audit of tool use and evidence\-label consistency\.
#### Student view\.
The student receives no gold claims or labels\. Its system prompt states the action schema, confidence\-conditioned routing rule, evidence semantics, and search budget\. Its user message contains only the question, long\-form response, and an approximate claim\-count cue\. Listing[4](https://arxiv.org/html/2609.22223#LST4)provides the complete policy portion of the prompt used to serialize the final demonstrations\.
Youarealong\-formfactualityverificationagent\.
Youwillreceiveaquestionandamodel’sresponse\.Verifytheresponseby:
decomposingitintoatomicclaimswithconfidencecalibration,groupingby
entity,andverifyingeachgroupwithsearchwhenneeded\.
Process:
1\.Emit<claims\>asaJSONarray\.Eachobjecthas:
"id":int1\.\.N
"claim":atomicclaimtext,verbatimfromtheresponse
"confidence":"high"\|"medium"\|"low"
"entity\_key":lowercasesnake\_casecentralentity
2\.Processeachuniqueentity\_keyinfirst\-appearanceorder:
\-Ifallclaimsinthegrouparehighconfidence,outputbatched
<verify\>tagsonly\.
\-Ifanyclaimismediumorlowconfidence,outputone<search\>
coveringthegroup,waitfor<information\>,writeone<memo\>,
andthenverifyallclaimsinthegroup\.
3\.Finishwith:
<answer\>\{"claims":\[
\{"claim":"\.\.\.","label":"supported\|unsupported"\},\.\.\.
\]\}</answer\>
Rules:
\-Searchonlywhenagroupcontainsamediumorlowconfidenceclaim\.
\-Aftereveryusefulsearch,writeamemobeforeverification\.
\-Processallclaimssharinganentity\_keyinonegroupturn\.
\-"unsupported"includescaseswhereevidencedoesnotconfirmthe
specificdate,number,name,orcausallink\.
\-Useatmosteightsearches\.
Listing 4:Student policy prompt stored in the demonstration corpus\.Question:
\{question\}
ModelResponse:
\{response\}
\(Thisresponsecontainsapproximately\{num\_claims\}verifiableclaims\.
Verifythemasdescribedabove\.\)
Youhaveaccesstoonetool:
<search\>query</search\>
Theretrievalenginerepliesinside<information\>\.\.\.</information\>\.
Aftereveryusefulsearch,write:
<memo\>one\-linedistilledevidence\(DocX\)</memo\>
Outputthefinalaggregatein:
<answer\>\{"claims":\[\.\.\.\]\}</answer\>
Listing 5:Student user\-message and tool\-description template\.
#### Final corpus statistics\.
The 1,447 retained trajectories contain 19,334 claim decisions\. Of these, 70\.4% aresupportedand 29\.6% areunsupported\. The mean trajectory issues 2\.24 searches\. Mean generated trajectory length, including injected retrieval observations after the initial prompt, is 9\.16 thousand characters\. Character length is reported because it is independent of the tokenizer used for each backbone\. The fact that search count is much smaller than mean claim count reflects both direct verification of high\-confidence groups and reuse of one retrieval result across related claims\.
### Preference\-Pair Construction
#### Motivation and candidate rollouts\.
Policy demonstrations teach decomposition, routing, search, memo formation, and factuality decisions jointly\. Preference data are constructed only after this policy has converged, and are used to correct residual decision errors\. We generate policy rollouts on a candidate pool that is disjoint from the development and held\-out evaluation sets\. Only structurally valid trajectories with at least one incorrect verdict are considered\. We further require the existing context to contain sufficient evidence for the gold decision so that the selected case represents a decision failure, not a missing\-search failure\.
#### Same\-trajectory counterfactuals\.
For each eligible trajectory, the original model completion becomes the rejected member of a preference pair\. The chosen member is produced by copying the complete trajectory and correcting only erroneous labels in two synchronized locations: the corresponding<verify\>content and the matchinglabelvalue inside<answer\>\. Claim text, identifiers, confidence, entity keys, group order, search queries, retrieval observations, memos, turn boundaries, and all other tokens are kept fixed\. This construction turns each pair into a controlled counterfactual in which the factuality decision is the only changed variable\.
Sharedtrajectoryprefix:
<claims\>\.\.\.</claims\>
<search\>\.\.\.</search\>
<information\>\.\.\.</information\>
<memo\>\.\.\.</memo\>
Rejected:
<verifyid=7\>supported</verify\>
\.\.\.
\{"claim":"\.\.\.","label":"supported"\}
Chosen:
<verifyid=7\>unsupported</verify\>
\.\.\.
\{"claim":"\.\.\.","label":"unsupported"\}
Listing 6:Schematic same\-trajectory preference pair\. All omitted context is byte\-identical\.Conventional preference construction from independently sampled completions would allow the chosen and rejected members to differ in search count, retrieved documents, decomposition style, memo content, length, and factuality labels\. Those differences make it difficult to identify which behavior the optimizer is rewarding\. In a same\-trajectory pair, identical tokens cancel in the chosen\-versus\-rejected comparison, while the decision\-token mask further excludes every unchanged position from the explicit sequence score\. The full evidence remains visible as context at the active label positions\.
#### Bidirectional corrections\.
A corpus containing only false\-support corrections would lower FSR but could teach a degenerate policy that predictsunsupportedfor every uncertain claim\. We therefore retain both correction directions\. There are 405 false\-support\-to\-unsupported pairs and 389 false\-unsupported\-to\-supported pairs, giving 794 pairs in total\. The near\-balanced direction count allows the preference signal to target specific evidence\-conditioned errors rather than shift the global class prior\. The chosen completions contain 15,686 claim decisions, of which 61\.7% aresupportedand 38\.3% areunsupported\. They average 3\.69 searches and 13\.80 thousand characters per trajectory\.
#### Preference\-pair validation\.
Each pair is admitted only after the following invariants are checked:
1. 1\.The rejected rollout is structurally valid and its erroneous claim identifiers are matched to gold annotations\.
2. 2\.Every corrected<verify\>label is synchronized with the corresponding<answer\>label\.
3. 3\.Chosen labels match the gold decisions, while the rejected completion retains at least one identified false\-support or false\-unsupported error\.
4. 4\.After replacing decision values with placeholders, chosen and rejected trajectories are identical\. This checks the same\-trajectory property\.
5. 5\.Both correction directions are retained during balancing; no unsupported\-only label policy can satisfy all pairs\.
6. 6\.Duplicate pairs and question–response fingerprint overlap with development or held\-out evaluation pools are rejected\.
#### Relationship to the training loss\.
The pair construction and decision\-focused loss enforce the same locality at two different levels\. Data construction controls what differs betweeny\+y^\{\+\}andy−y^\{\-\}; the loss mask controls which token log probabilities contribute to the optimization objective\. The active positions are only label values inside<verify\>and<answer\>\. Search and evidence tokens remain in the causal prefix, so a decision can still depend on retrieved information, but they receive no direct preference reward\.相似文章
基于证据链评估的校准式选择性事实核查
本文介绍了一种名为证据链评估(ECE)的选择性事实核查框架,该框架允许基于LLM的验证代理在证据薄弱、稀疏或不一致时放弃给出判断。在ECE-Bench上,ECE实现了93.7%覆盖率下的97.8%选择性准确率,展示了处理认知薄弱证据时面向安全性的权衡。
EVE-Agent: 可验证证据的自我进化智能体
EVE-Agent 提出了一个自我进化搜索智能体框架,通过生成问题、答案和证据片段,并基于证据的边际准确性增益进行训练,确保证据可验证性。这提高了基于依据的正确性,且无需人工标注。
从片段到语义:重新思考多语言事实核查的证据粒度
本文介绍了SEEK,一个用于多语言事实核查中语义证据提取的框架,该框架从完整文章中构建连贯的证据块,并使用LoRA微调多语言大语言模型,在宏观F1分数上相比基线提升了高达20%。
ElementCheck:基于句子元素的复杂性感知长文本事实性评估
ElementCheck是一个复杂性感知框架,通过提取句子元素并将其组织成元素图来提高长文本事实性评估,并伴随新的基准FastFact-Sent以增强验证准确性。
基于策略性标注的人类锚定事实性评估
本文引入了一种使用失败空间分析的特定于事实性的标注策略,以在有限的标注预算下改善针对大语言模型的人类锚定事实性评估,并在基准系统上取得了显著的效率提升。