When Absence Is Evidence: Evaluating Completeness-Sensitive Negative Reasoning in Large Language Models

arXiv cs.CL Papers

Summary

Introduces CROWN-QA, a benchmark for completeness-sensitive negative reasoning in LLMs, showing models struggle to distinguish justified negative answers from insufficient evidence, often over-closing.

arXiv:2608.04591v1 Announce Type: new Abstract: Large language models (LLMs) are often asked whether something is absent from a record, list, or retrieved context. Yet non-observation licenses a negative answer only when evidence completely covers the query scope; otherwise, the answer should remain unknown. We call this completeness-sensitive negative reasoning. We introduce CROWN-QA, comprising CROWN-Synth, a controlled paired core that fixes the question and observed facts while varying only query-relative coverage, and CROWN-Real, a real-document contrast-set evaluation with controlled coverage variants. Across three LLM families, models show unstable closure judgments and substantial over-closure, failing to reliably distinguish a justified negative answer (Certified-Negative) from insufficient evidence (Unknown). The dominant CROWN-Synth failure is asymmetric: models often recognize implicitly complete evidence yet treat implicitly partial evidence as query-covering. Prompting redistributes errors between over- and under-closure rather than consistently resolving them. Structured certificate elicitation traces many errors to evidence-coverage mischaracterization. CROWN-Real shows that the core partial-coverage asymmetry persists on real-document content, while its strength and the balance between over- and under-closure vary by model, prompt, and source.
Original Article
View Cached Full Text

Cached at: 08/06/26, 07:50 AM

# When Absence Is Evidence: Evaluating Completeness-Sensitive Negative Reasoning in Large Language Models
Source: [https://arxiv.org/html/2608.04591](https://arxiv.org/html/2608.04591)
###### Abstract

Large language models \(LLMs\) are often asked whether something is absent from a record, list, or retrieved context\. Yet non\-observation licenses a negative answer only when evidence completely covers the query scope; otherwise, the answer should remain unknown\. We call this completeness\-sensitive negative reasoning\. We introduce CROWN\-QA, comprising CROWN\-Synth, a controlled paired core that fixes the question and observed facts while varying only query\-relative coverage, and CROWN\-Real, a real\-document contrast\-set evaluation with controlled coverage variants\. Across three LLM families, models show unstable closure judgments and substantial over\-closure, failing to reliably distinguish a justified negative answer \(Certified\-Negative\) from insufficient evidence \(Unknown\)\. The dominant CROWN\-Synth failure is asymmetric: models often recognize implicitly complete evidence yet treat implicitly partial evidence as query\-covering\. Prompting redistributes errors between over\- and under\-closure rather than consistently resolving them\. Structured certificate elicitation traces many errors to evidence\-coverage mischaracterization\. CROWN\-Real shows that the core partial\-coverage asymmetry persists on real\-document content, while its strength and the balance between over\- and under\-closure vary by model, prompt, and source\.

## Introduction

Large language models \(LLMs\) are commonly evaluated by their ability to find and use positive evidence\. In retrieval\-augmented generation \(RAG\), for example, evaluation often asks whether retrieved context is relevant, sufficient, and faithfully used to answer the user query\(Lewiset al\.[2020](https://arxiv.org/html/2608.04591#bib.bib1); Eset al\.[2024](https://arxiv.org/html/2608.04591#bib.bib29); Saad\-Falconet al\.[2024](https://arxiv.org/html/2608.04591#bib.bib17); Niuet al\.[2024](https://arxiv.org/html/2608.04591#bib.bib18); Chenet al\.[2024](https://arxiv.org/html/2608.04591#bib.bib30); Yanget al\.[2024](https://arxiv.org/html/2608.04591#bib.bib31); Jorenet al\.[2025](https://arxiv.org/html/2608.04591#bib.bib2)\)\. In real settings, however, many questions—including high\-stakes ones—require reasoning about absence: whether a condition is absent from a contraindication list, an applicant is absent from an exclusion list, or a proceedings index contains no matching title\.

The key challenge is that absence is evidential only when coverage is both complete and query\-covering\. Figure[1](https://arxiv.org/html/2608.04591#Sx1.F1)holds the query and listed titles fixed while varying coverage status and scope\. The complete conference\-wide index \(center\) licensesCertified\-Negative, whereas both the partial search results \(left\) and the complete main\-track index \(right\) remainUnknown\. Thus, the core problem is not detecting absence, but deciding when absence licenses negation\.

Query:Did any MLConf 2024 paper, across all tracks, have a title containing “calibrated retrieval”?

Observed titles:The listed titles are identical across all three variants; none contains the queried phrase\.

Partialall tracksCompleteall tracksCompletemain track onlyTop\-20 keyword\-search results\.Official conference\-wide title index\.Official main\-track title index\.UnknownCertified\-NegativeUnknown

Figure 1:Matched variants with identical titles\. Only the complete conference\-wide index covers the query; the partial results and the complete main\-track index remainUnknown\.Completeness is query\-relative: evidence may be complete for its own scope yet fail to cover the query scope\. As Figure[2](https://arxiv.org/html/2608.04591#Sx1.F2)shows, support and query\-scope coverage are separate axes: a justified “no” requires absent support and complete query\-covering evidence\. CROWN\-QA fixes the former by construction and evaluates only the latter\. This scope\-containment view is related to query–knowledge relevance in RAG\(Liet al\.[2024](https://arxiv.org/html/2608.04591#bib.bib20)\)and gives a query\-relative form of the open\- versus closed\-world distinction: non\-closing evidence leaves an unobserved fact unknown, whereas complete query\-covering evidence licenses a negative answer\(Razniewskiet al\.[2024](https://arxiv.org/html/2608.04591#bib.bib7)\)\. This distinction is rarely isolated in natural\-language LLM evaluation\.

Prior work studies abstention, unanswerability, ambiguous information needs, and knowledge boundaries\(Penget al\.[2025](https://arxiv.org/html/2608.04591#bib.bib3); Kirichenkoet al\.[2025](https://arxiv.org/html/2608.04591#bib.bib4); Madhusudhanet al\.[2025](https://arxiv.org/html/2608.04591#bib.bib22); Zhang and others[2024](https://arxiv.org/html/2608.04591#bib.bib24); Chenet al\.[2025](https://arxiv.org/html/2608.04591#bib.bib23)\), retrieval robustness, context sufficiency, evidence\-based QA, and factuality\(Yuet al\.[2024](https://arxiv.org/html/2608.04591#bib.bib19); Jorenet al\.[2025](https://arxiv.org/html/2608.04591#bib.bib2); Glockneret al\.[2025](https://arxiv.org/html/2608.04591#bib.bib21); Fatahi Bayatet al\.[2025](https://arxiv.org/html/2608.04591#bib.bib25)\), as well as omitted information\(Fuet al\.[2025](https://arxiv.org/html/2608.04591#bib.bib5)\), false premises\(Shafieiet al\.[2025](https://arxiv.org/html/2608.04591#bib.bib26)\), and negation\(García\-Ferreroet al\.[2023](https://arxiv.org/html/2608.04591#bib.bib6)\)\. These settings ask whether to answer, refuse, or identify missing or misleading information; they do not hold the query and observed facts fixed while varying whether coverage contains the query scope\. CROWN\-Synth isolates this factor after fixing support as absent: the same unsupported fact remainsUnknownunder non\-closing evidence but becomesCertified\-Negativeunder complete query\-covering evidence\. Thus, always abstaining fails on the latter, whereas always answering “no” fails on the former\.

![Refer to caption](https://arxiv.org/html/2608.04591v1/x1.png)Figure 2:Conceptual outcome space\. Positive\-support cases are ordinary QA and are outside CROWN\-QA; CROWN\-QA evaluates the no\-support branch, where query\-scope coverage determinesCertified\-NegativeversusUnknown\.We introduce CROWN\-QA \(Completeness Reasoning Over What is Not in Question Answering\), combining a controlled paired core, real\-document contrast sets, and diagnostic certificates for query scope, evidence scope, and closure judgment\. The contributions of this paper are as follows\.

- •Problem and formalization\.We define completeness\-sensitive negative reasoning as an absence\-conditioned QA task:Certified\-Negativeis justified only when evidence completely covers the query scope; otherwise, the answer isUnknown\.
- •Benchmark and diagnostics\.We introduce CROWN\-QA: CROWN\-Synth, a controlled paired core with L1–L4 regimes, Scope\-Mismatch cases, and paired completeness sensitivity; and CROWN\-Real, a real\-document contrast\-set evaluation\. Directional metrics and structured certificates support diagnosis\. To our knowledge, CROWN\-Synth is the first controlled natural\-language QA benchmark to isolate query\-relative coverage as the label\-changing intervention\.
- •Empirical findings and diagnostic analysis\.Across three LLM families, CROWN\-Synth shows strong over\-closure, especially on implicitly partial evidence, while prompting shifts errors between over\- and under\-closure\. CROWN\-Real reproduces the partial\-coverage gap on real\-document content, but its size varies by model, prompt, and source\. Certificate errors most often first appear in the reported evidence\-coverage field\.

## Related Work

Work\(a\)\(b\)\(c\)\(d\)\(e\)Sufficient Ctx\.\(Jorenet al\.[2025](https://arxiv.org/html/2608.04591#bib.bib2)\)✗✗✗✗✗UAEval4RAG\(Penget al\.[2025](https://arxiv.org/html/2608.04591#bib.bib3)\)✗✗✗✗✗AbstentionBench\(Kirichenkoet al\.[2025](https://arxiv.org/html/2608.04591#bib.bib4)\)✗✗✗✗✗AbsenceBench\(Fuet al\.[2025](https://arxiv.org/html/2608.04591#bib.bib5)\)✓∗✗✗✗✗Negation\(García\-Ferreroet al\.[2023](https://arxiv.org/html/2608.04591#bib.bib6)\)✓∗✗✗✗✗KB completeness/negation\(Razniewskiet al\.[2024](https://arxiv.org/html/2608.04591#bib.bib7)\)✓✓✗✓†—CROWN\-QA✓✓✓✓✓

Table 1:Positioning of CROWN\-QA\. Dimensions: \(a\) absence or explicit negation as the primary semantic target; \(b\) query\-relative set coverage or completeness explicitly modeled; \(c\) same\-question, same\-fact paired coverage interventions; \(d\) query–coverage Scope\-Mismatch; and \(e\) negative licensing distinguished from uncertainty in NL QA\.∗These works evaluate omission or explicit negation, but not whether missing support licenses a negative answer\.†KB completeness represents scope symbolically rather than inferring it from language;*—*indicates that \(e\) does not apply to symbolic KB semantics\.RAG sufficiency and unanswerability\.RAG grounds generation in retrieved context\(Lewiset al\.[2020](https://arxiv.org/html/2608.04591#bib.bib1)\), and evaluations measure relevance, faithfulness, hallucination, sufficiency, and unanswerability\(Saad\-Falconet al\.[2024](https://arxiv.org/html/2608.04591#bib.bib17); Niuet al\.[2024](https://arxiv.org/html/2608.04591#bib.bib18); Jorenet al\.[2025](https://arxiv.org/html/2608.04591#bib.bib2); Penget al\.[2025](https://arxiv.org/html/2608.04591#bib.bib3)\)\. Related benchmarks examine missing or misleading evidence\(Glockneret al\.[2025](https://arxiv.org/html/2608.04591#bib.bib21)\), query–knowledge relevance\(Liet al\.[2024](https://arxiv.org/html/2608.04591#bib.bib20)\), and broader abstention cases\(Kirichenkoet al\.[2025](https://arxiv.org/html/2608.04591#bib.bib4)\)\. These works primarily evaluate answer support, context sufficiency, or abstention\.

Absence and negation in LLMs\.Recent work shows that LLMs struggle with omitted information and explicit negation\(Fuet al\.[2025](https://arxiv.org/html/2608.04591#bib.bib5); García\-Ferreroet al\.[2023](https://arxiv.org/html/2608.04591#bib.bib6)\)\. Omitted\-content benchmarks ask what is missing, while negation benchmarks test linguistic negation\.

Completeness in knowledge bases\.Database and knowledge\-base research distinguishes closed\-world assumptions, where missing facts are treated as false, from open\-world assumptions, where they remain unknown\(Razniewskiet al\.[2024](https://arxiv.org/html/2608.04591#bib.bib7)\)\. Partial completeness and completeness assertions encode coverage through formal semantics or metadata\(Razniewskiet al\.[2015](https://arxiv.org/html/2608.04591#bib.bib9); Darariet al\.[2013](https://arxiv.org/html/2608.04591#bib.bib8)\)\.

Selective prediction and knowledge boundaries\.Selective prediction trades coverage for risk, while conformal methods provide uncertainty sets or risk\-control guarantees\(El\-Yaniv and Wiener[2010](https://arxiv.org/html/2608.04591#bib.bib10); Geifman and El\-Yaniv[2017](https://arxiv.org/html/2608.04591#bib.bib11); Angelopouloset al\.[2021](https://arxiv.org/html/2608.04591#bib.bib16)\); related LLM work studies model self\-knowledge, unanswerability, and clarification under ambiguity\(Kadavathet al\.[2022](https://arxiv.org/html/2608.04591#bib.bib12); Yinet al\.[2023](https://arxiv.org/html/2608.04591#bib.bib13); Rajpurkaret al\.[2018](https://arxiv.org/html/2608.04591#bib.bib14); Coleet al\.[2023](https://arxiv.org/html/2608.04591#bib.bib15); Madhusudhanet al\.[2025](https://arxiv.org/html/2608.04591#bib.bib22); Zhang and others[2024](https://arxiv.org/html/2608.04591#bib.bib24)\)\. These lines manage unreliable answers through abstention, uncertainty sets, or clarification\.

Positioning\.Across these lines, the absence of support is handled as a reason to abstain, hedge, or clarify; none of them tests when absence itself licenses a negative answer\. CROWN\-QA targets exactly this distinction \(Table[1](https://arxiv.org/html/2608.04591#Sx2.T1)\): CROWN\-Synth isolates it by fixing the question and observed facts and varying query\-relative coverage alone, and CROWN\-Real tests whether the same distinction transfers to real documents\.

## Task Definition

### Absence\-Conditioned Completeness Judgment

Letqqbe a natural\-language question and letE=\{e1,…,en\}E=\\\{e\_\{1\},\\ldots,e\_\{n\}\\\}denote the evidence context\. The context may contain factual statements and coverage information, expressed explicitly or through source semantics, that specifies the scope described byEEand whether its coverage is complete, partial, or unspecified\. Such a scope may involve a time period, entity, attribute, list, database field, or document collection\. We denote the scope requested by the question asS​\(q\)S\(q\)and the coverage scope asserted by the evidence asS​\(E\)S\(E\)\.

CROWN\-QA is conditioned on the absence of positive support: the queried fact is not supported by the observed factual content ofEE\. The task is therefore not to find a positive answer, but to decide whether this observed absence is licensed as a negative answer by the coverage information\. The model must output one of two labels:

𝒴=\{Certified\-Negative,Unknown\}\.\\mathcal\{Y\}=\\\{\\textsc\{Certified\-Negative\},\\textsc\{Unknown\}\\\}\.Certified\-Negativeis correct when the observed absence is licensed by complete coverage of the query scope; “certified” is limited to the stated evidence coverage\.Unknownis correct when coverage is partial, sampled, unspecified, ambiguous, or complete only for a scope that does not coverS​\(q\)S\(q\)\.

Formally, letc=Comp​\(E,S​\(q\)\)∈\{0,1\}c=\\mathrm\{Comp\}\(E,S\(q\)\)\\in\\\{0,1\\\}indicate whetherEEclosesS​\(q\)S\(q\)\. Herec=1c=1only whenEEestablishes complete coverage ofS​\(E\)S\(E\)andS​\(q\)⊆S​\(E\)S\(q\)\\subseteq S\(E\)\. Since positive support is absent by construction, the target label is

y∗​\(q,E\)=\{Certified\-Negative,c=1,Unknown,c=0\.y^\{\*\}\(q,E\)=\\begin\{cases\}\\textsc\{Certified\-Negative\},&c=1,\\\\ \\textsc\{Unknown\},&c=0\.\\end\{cases\}Herec=1c=1only whenEEestablishes complete coverage of its asserted scopeS​\(E\)S\(E\)andS​\(q\)⊆S​\(E\)S\(q\)\\subseteq S\(E\)\. This containment relation is semantic rather than lexical; a complete source for a different time, entity, attribute, population, or collection hasc=0c=0\.

CROWN\-QA excludes positive\-evidence retrieval, already studied in QA and RAG, and isolates the downstream licensing step: once support is absent, does coverage justify a negative answer?

### Failure Modes

Although the gold label is deterministic, the model must infer query\-relative coverage from natural language\. Lety^∈𝒴\\hat\{y\}\\in\\mathcal\{Y\}denote the model prediction and lety∗y^\{\*\}denote the gold label\. We focus on two directional errors that aggregate accuracy can obscure\.

Over\-Closure\.Over\-closure occurs when a model outputsCertified\-Negativeeven though the evidence is non\-closing for the query scope:

y^=Certified\-Negative,y∗=Unknown\.\\hat\{y\}=\\textsc\{Certified\-Negative\},\\qquad y^\{\*\}=\\textsc\{Unknown\}\.This corresponds to treating observed absence as evidence of absence without licensed coverage\.

Under\-Closure\.Under\-closure occurs when a model outputsUnknowneven though complete query\-covering evidence licenses a certified negative answer:

y^=Unknown,y∗=Certified\-Negative\.\\hat\{y\}=\\textsc\{Unknown\},\\qquad y^\{\*\}=\\textsc\{Certified\-Negative\}\.This corresponds to unnecessary abstention despite licensed negative evidence\.

### Metrics

We first report the over\-closure rate \(OCR\) and under\-closure rate \(UCR\):

OCR\\displaystyle\\mathrm\{OCR\}=Pr⁡\[y^=Certified\-Negative∣y∗=Unknown\],\\displaystyle=\\Pr\[\\hat\{y\}=\\textsc\{Certified\-Negative\}\\mid y^\{\*\}=\\textsc\{Unknown\}\],UCR\\displaystyle\\mathrm\{UCR\}=Pr⁡\[y^=Unknown∣y∗=Certified\-Negative\]\.\\displaystyle=\\Pr\[\\hat\{y\}=\\textsc\{Unknown\}\\mid y^\{\*\}=\\textsc\{Certified\-Negative\}\]\.
OCR measures false certification under non\-closing evidence, whereas UCR measures false abstention under complete query\-covering evidence\.

We report class\-balanced accuracy, denotedAcc\\mathrm\{Acc\}, which averages the correct rates for the two gold labels:

Acc=12\(\\displaystyle\\mathrm\{Acc\}=\\tfrac\{1\}\{2\}\\big\(Pr⁡\[y^=Certified\-Negative∣y∗=Certified\-Negative\]\\displaystyle\\Pr\[\\hat\{y\}=\\textsc\{Certified\-Negative\}\\mid y^\{\*\}=\\textsc\{Certified\-Negative\}\]\+Pr\[y^=Unknown∣y∗=Unknown\]\)\.\\displaystyle\+\\Pr\[\\hat\{y\}=\\textsc\{Unknown\}\\mid y^\{\*\}=\\textsc\{Unknown\}\]\\big\)\.
We use this definition for both CROWN\-Synth and CROWN\-Real, so the two gold labels receive equal weight despite their different label ratios\.

For CROWN\-Synth, we additionally report paired completeness sensitivity \(CS\) on same\-question, same\-fact pairs\. Letxc=\(q,Ec\)x\_\{c\}=\(q,E\_\{c\}\)denote the complete query\-covering member with gold labelCertified\-Negative, and letxo=\(q,Eo\)x\_\{o\}=\(q,E\_\{o\}\)denote its matched non\-closing member with gold labelUnknown\. CS is the fraction of pairs answered correctly in both directions:

CS=Pr⁡\[y^​\(xc\)=Certified\-Negative∧y^​\(xo\)=Unknown\]\.\\mathrm\{CS\}=\\Pr\[\\hat\{y\}\(x\_\{c\}\)=\\textsc\{Certified\-Negative\}\\ \\wedge\\ \\hat\{y\}\(x\_\{o\}\)=\\textsc\{Unknown\}\]\.
High CS indicates that the model changes its answer in the intended direction when only query\-relative coverage changes\.

## CROWN\-QA Benchmark

CROWN\-QA combines two forms of coverage control\. CROWN\-Synth uses exact same\-question, same\-fact pairs, whereas CROWN\-Real uses A/B/C contrast sets grounded in real documents\. All examples are absence\-conditioned; Table[2](https://arxiv.org/html/2608.04591#Sx4.T2)summarizes both components\.

PropertyCROWN\-SynthCROWN\-RealSourceSynthetic worldsACL Anthology proceedings; DailyMed drug labelsDomains52UnitMatched pair \(2 variants\)Contrast set \(A/B/C\)Scale5,000 examples \(2,500 pairs\)1,599 examples \(533 sets\)CN:UNK1:11:2Non\-closingPartial / Scope\-Mismatch \(1:1\)B: narrower complete; C: non\-exhaustive \(1:1\)ControlSame question and observed factsSame question and target phrase; source content variesCoverageL1–L4 coverage\-expressionregimesA: query\-covering complete;B: narrower complete;C: non\-exhaustiveTable 2:CROWN\-QA design summary\. All examples are absence\-conditioned\. CROWN\-Synth provides exact paired control, whereas CROWN\-Real tests transfer on real\-document content with controlled coverage variants\.### CROWN\-Synth: Controlled Paired Core

CROWN\-Synth is the controlled core of CROWN\-QA\. It is generated from synthetic worlds in which observed factual content, query scope, evidence coverage scope, and coverage status can be independently controlled\. This design isolates whether a model changes its answer because evidence closes the query scope, rather than because the queried item or surrounding factual content changes\.

Paired construction\.For each base world, we generate a queried factrrand an observed evidence set in whichrris absent\. We then create matched members with the same question and identical observed factual content\. The query\-covering member establishes complete coverage ofS​\(q\)S\(q\), yieldingCertified\-Negative\. Its matched non\-closing member is partial, sampled, unspecified, or complete only for a scope that does not containS​\(q\)S\(q\), yieldingUnknown\. A worked pair is provided in Appendix A\.

Coverage regimes\.A benchmark that always states “this is the complete record” can reduce to keyword spotting\. We therefore vary how coverage is expressed across four regimes: \(L1\) explicit—“this is the complete list”; \(L2\) paraphrased—“the official registry of allXX”; \(L3\) implicit—coverage is conveyed by a balanced inventory of four complete and four partial source\-type families, crossed with domain and realized in a shared sentence frame, without explicit completeness language; and \(L4\) adversarial—a complete source is saliently mentioned but does not establish that the displayed evidence closes the query scope\. L1–L3 vary coverage explicitness, whereas L4 tests robustness to a non\-licensing completeness statement\. Coverage regime and coverage relation are annotated independently, so Scope\-Mismatch cases occur at every L1–L4 regime\. We report the two axes separately to distinguish cue matching from scope reasoning\.

Scope\-Mismatch cases\.Complete–partial pairs, particularly under explicit coverage regimes, may reward detecting whether a closure cue is present\. We therefore add cases in which the evidence is complete for a scopeS​\(E\)S\(E\)that does not contain the query scopeS​\(q\)S\(q\)\. For example, a question may ask whether any MLConf 2024 paper across all tracks has a matching title, while the evidence provides a complete index of main\-track papers only\. Although support is absent and the source is complete for its own scope,S​\(q\)⊈S​\(E\)S\(q\)\\not\\subseteq S\(E\), so the correct label isUnknown\. These cases expose policies that treat any completeness cue as sufficient and test semantic scope containment rather than shallow cue matching\. Each item is annotated as a period\-overlap mismatch, hierarchy mismatch, population mismatch, or attribute mismatch\.

Domain templates\.The benchmark covers domains in which negative answers are practically meaningful\. Academic records test whether a requirement is absent from a complete student record\. Grant eligibility tests whether an applicant is excluded under a complete list of exclusion criteria\. Medical contraindications test whether a condition appears in a contraindication list declared complete for the relevant scope\. Legal and policy rules test whether an exception or prohibition applies\. Proceedings and catalogs test whether an item is absent from an official list declared complete for the relevant scope\.

### CROWN\-Real: Real\-Document Transfer Set

CROWN\-Real tests whether the failure patterns isolated on CROWN\-Synth persist when coverage must be interpreted from real\-document content and source structure\. It is grounded in two source families: NLP proceedings from the ACL Anthology\(Association for Computational Linguistics[2026](https://arxiv.org/html/2608.04591#bib.bib27)\)and public drug labels from DailyMed\(U\.S\. National Library of Medicine[2026](https://arxiv.org/html/2608.04591#bib.bib28)\)\. Unlike the strict same\-fact pairs in CROWN\-Synth, CROWN\-Real uses three\-member contrast sets that share a question and an exact target phrase, while the displayed source content varies across members\. The target phrase is absent from every displayed context\. Variant A provides complete query\-covering evidence, Variant B is complete for a narrower scope, and Variant C is non\-exhaustive\.

For proceedings, A combines the full title indexes of two official collections, B retains one complete collection, and C contains a deterministic strict subset of titles from the same two collections\. For drug labels, A contains the full Warnings and Precautions section with its source\-native boundaries, B contains one complete numbered subsection with its scope and closing boundary, and C uses the Highlights material as an explicitly non\-exhaustive control\. Thus, titles, label text, headings, and document boundaries come from real sources; coverage is varied through controlled extraction and selection rather than generated factual claims\.

We verify by normalized exact matching that the queried phrase is absent from every variant and exclude any item containing positive support\. Source identity and section or collection boundaries are retained so that closure is licensed by the evidence shown to the model rather than by hidden metadata\. Medical C is analyzed only as an explicit\-partial control; implicit\-partial transfer is evaluated on proceedings C\. Because CROWN\-Real uses contrast sets rather than strict pairs, we report class\-balancedAcc\\mathrm\{Acc\}and variant\-specific recalls instead of paired completeness sensitivity\.

### Quality Control

Each CROWN\-Synth pair must satisfy three constraints: \(i\) the queried fact is absent from the observed factual content; \(ii\) the question and observed factual content are identical across the paired members; and \(iii\) only coverage status or scope changes the gold label\. Gold labels follow controlled metadata throughc=Comp​\(E,S​\(q\)\)c=\\mathrm\{Comp\}\(E,S\(q\)\), rather than human annotation\. We manually inspect a stratified sample across domains, coverage regimes, and coverage relations, using an LLM judge as a secondary consistency check\. We reject implicit\-partial items that inadvertently establish exhaustive coverage and Scope\-Mismatch items solvable by an obvious non\-overlapping token difference, such as disjoint entities or years\.

For CROWN\-Real, we verify by normalized exact matching that the target is absent from every A/B/C member\. We also check source identity, section or collection boundaries, and whether the displayed evidence itself establishes the intended A/B/C coverage relation\. Items containing positive support or requiring hidden metadata to justify the gold label are excluded\.

## Structured Scope\-and\-Coverage Elicitation

To examine where closure errors arise, we evaluate a structured elicitation condition using the same base LLM\. Given a questionqqand evidence contextEE, the model returns a completeness certificateC=\(S^q,S^E,c^\)C=\(\\hat\{S\}\_\{q\},\\hat\{S\}\_\{E\},\\hat\{c\}\), whereS^q\\hat\{S\}\_\{q\}is the model\-reported query scope,S^E\\hat\{S\}\_\{E\}describes the model\-reported evidence coverage, including its scope and completeness status, andc^∈\{0,1\}\\hat\{c\}\\in\\\{0,1\\\}is the final coverage judgment\. All three fields are elicited jointly in one structured response\. Onlyc^\\hat\{c\}is mapped to the scored label; the two scope fields are retained for diagnostic analysis\.

The model should setc^=1\\hat\{c\}=1only when the coverage described byS^E\\hat\{S\}\_\{E\}is complete and contains the full scope described byS^q\\hat\{S\}\_\{q\}\. Accordingly,c^=0\\hat\{c\}=0for partial, sampled, incomplete, unspecified, ambiguous, or scope\-mismatched coverage\. This is a semantic judgment rather than a literal comparison of natural\-language strings\. Reporting the two scopes separately makes the query–coverage relation explicit, particularly in Scope\-Mismatch cases, rather than collapsing the decision into a single sufficiency judgment\.

For scoring, the Boolean field is mapped to the common two\-label space:

y^​\(C\)=\{Certified\-Negative,c^=1,Unknown,c^=0\.\\hat\{y\}\(C\)=\\begin\{cases\}\\textsc\{Certified\-Negative\},&\\hat\{c\}=1,\\\\ \\textsc\{Unknown\},&\\hat\{c\}=0\.\\end\{cases\}Becauseccis the gold coverage judgment, the directional error rates for this condition are

OCR=Pr⁡\[c^=1∣c=0\],UCR=Pr⁡\[c^=0∣c=1\]\.\\mathrm\{OCR\}=\\Pr\[\\hat\{c\}=1\\mid c=0\],\\qquad\\mathrm\{UCR\}=\\Pr\[\\hat\{c\}=0\\mid c=1\]\.We compareS^q\\hat\{S\}\_\{q\},S^E\\hat\{S\}\_\{E\}, andc^\\hat\{c\}with the benchmark metadata to identify whether an incorrect certificate first diverges in query\-scope extraction, evidence\-coverage characterization, or the final Boolean judgment\. This analysis diagnoses errors in the model\-reported fields rather than internal reasoning\.

## Experiments

Our experiments address five questions\.

- •RQ1: Paired closure judgment\.How reliably do LLMs distinguishCertified\-NegativefromUnknownwhen only query\-relative coverage changes, with the question and observed facts fixed?
- •RQ2: Failure conditions\.Are errors concentrated in partial\-coverage or Scope\-Mismatch cases, and which coverage regimes yield the largest complete–partial gaps across model families?
- •RQ3: Prompting effects\.Do explicit rules, chain\-of\-thought, abstention\-aware prompting, self\-checking, and certificate elicitation improve paired discrimination, or merely shift errors between over\- and under\-closure?
- •RQ4: Diagnostic decomposition\.At which certificate field does an incorrect output first differ from the benchmark annotation: query scope, evidence coverage, or the final coverage judgment?
- •RQ5: Real\-document transfer\.Which CROWN\-Synth failure patterns, particularly partial\-coverage over\-closure, persist on CROWN\-Real, and how do they vary across models, prompts, and source structures?

Qwen3\.5\-9BClaude Haiku 4\.5Gemma\-4\-12BConditionAcc↑\\uparrowOCR↓\\downarrowUCR↓\\downarrowCS↑\\uparrowAcc↑\\uparrowOCR↓\\downarrowUCR↓\\downarrowCS↑\\uparrowAcc↑\\uparrowOCR↓\\downarrowUCR↓\\downarrowCS↑\\uparrowNaive70\.637\.221\.642\.873\.740\.811\.847\.661\.576\.00\.923\.2Def\.79\.926\.913\.260\.083\.513\.319\.667\.175\.044\.65\.450\.5Def\.\+CoT83\.927\.94\.468\.189\.912\.47\.879\.885\.224\.84\.970\.9Def\.\+Abstain78\.213\.030\.656\.683\.112\.121\.766\.281\.131\.06\.862\.3Def\.\+Self\-check74\.58\.842\.349\.485\.216\.613\.170\.473\.151\.82\.046\.6Cert\.82\.820\.414\.066\.187\.913\.311\.075\.871\.255\.61\.942\.6Cert\.\+CoT85\.724\.64\.071\.987\.917\.66\.675\.979\.140\.31\.658\.3Table 3:Overall results on CROWN\-Synth \(%\)\.### Models and Evaluation Conditions

We evaluate three models from two open\-weight families and one API\-based family: Qwen3\.5\-9B, Gemma\-4\-12B, and Claude Haiku 4\.5\.

The evaluation conditions are summarized in Appendix B\. We evaluate seven LLM conditions\. Naive measures the model’s default treatment of missing support; Definition\-aware states the task rule explicitly\. Its CoT, Abstain, and Self\-check variants test step\-by\-step reasoning, a conservative default toUnknown, and second\-pass revision\. Certificate elicits\(S^q,S^E,c^\)\(\\hat\{S\}\_\{q\},\\hat\{S\}\_\{E\},\\hat\{c\}\); onlyc^\\hat\{c\}is mapped to the scored label, while the two scope fields are retained for analysis\. Certificate\+CoT adds reasoning before the same structured output\.

### Experimental Protocol

For each example, all LLM conditions receive the same question and evidence context, use no few\-shot demonstrations, and use greedy decoding at temperature zero\. Outputs are scored in the common two\-label space\. Metric definitions use proportions in\[0,1\]\[0,1\]; for readability, tables report rates as percentages and rate differences in percentage points\. We report percentile 95% confidence intervals from 10,000 bootstrap replicates, resampling matched pairs for CROWN\-Synth and A/B/C contrast sets for CROWN\-Real\. Appendix D reports the full condition tables and bootstrap intervals\.

### Overall Closure Profiles

Table[3](https://arxiv.org/html/2608.04591#Sx6.T3)provides the aggregate view for RQ1 and previews the prompting effects examined in RQ3\. Under Naive prompting, OCR exceeds UCR for all three models, revealing a default tendency to license negative answers from non\-closing evidence rather than remain uncertain\. Explicit task rules improveAcc\\mathrm\{Acc\}and CS for all three models, and adding CoT further improves both metrics\. However, these aggregate gains do not consistently reduce OCR and UCR together: abstention lowers OCR by increasing UCR, while self\-checking and certificate elicitation have model\-dependent effects\.

Closure judgments remain unstable \(RQ1\)\.The evaluated LLMs exhibit some completeness\-sensitive reasoning, but do not reliably distinguishCertified\-NegativefromUnknown\. Explicit rules improve performance, yet the distinction remains sensitive to the model and prompting condition\. The central limitation is therefore an unstable query\-relative closure judgment\.

### Failure Conditions

Pair: Completevs\. PartialPair: Completevs\. Scope\-MismatchModelConditionCS↑\\uparrowBoth\-CN↓\\downarrowBoth\-UNK↓\\downarrowCS↑\\uparrowBoth\-CN↓\\downarrowBoth\-UNK↓\\downarrowQwenNaive28\.056\.613\.457\.614\.626\.6Def\.44\.643\.811\.675\.59\.614\.5Def\.\+CoT58\.634\.56\.677\.620\.31\.3Cert\.59\.330\.210\.072\.99\.817\.1HaikuNaive26\.167\.96\.069\.013\.317\.3Def\.66\.218\.615\.067\.97\.924\.2Def\.\+CoT72\.620\.57\.087\.04\.38\.6Cert\.70\.421\.97\.681\.24\.614\.2GemmaNaive2\.197\.90\.044\.253\.91\.6Def\.33\.461\.84\.267\.726\.35\.6Def\.\+CoT66\.731\.80\.675\.016\.78\.1Cert\.17\.082\.70\.268\.228\.23\.4Table 4:Branch\-specific CS and same\-label pair outcomes on CROWN\-Synth \(%\)\. The full outcome decomposition appears in Table[12](https://arxiv.org/html/2608.04591#A4.T12)in Appendix D\.Table[4](https://arxiv.org/html/2608.04591#Sx6.T4)compares complete–partial pairs with complete–Scope\-Mismatch pairs\. In both pair types, the complete query\-covering member has gold labelCertified\-Negative, and the matched non\-closing member has gold labelUnknown\. CS requires both members to be correct\. Both\-CN indicates that both members are predicted asCertified\-Negative, producing over\-closure on the non\-closing member\. Both\-UNK indicates that both members are predicted asUnknown, producing under\-closure on the complete member\. Across every model, CS is lower and Both\-CN is higher for complete–partial pairs than for complete–Scope\-Mismatch pairs\. This indicates that, when coverage changes from complete to partial, models often fail to switch their prediction fromCertified\-NegativetoUnknownand instead predictCertified\-Negativefor both members\.

Having established that complete–partial pairs produce more errors than complete–Scope\-Mismatch pairs, we next examine which coverage regimes account for these failures\. Table[5](https://arxiv.org/html/2608.04591#Sx6.T5)divides the 1,250 complete–partial pairs across L1–L4\. The first four columns report regime\-specific CS, while the final two columns report the separate correct rates for the complete and partial members within L3 complete–partial pairs\.

As shown in Table[5](https://arxiv.org/html/2608.04591#Sx6.T5), L3 has the lowest or tied\-lowest CS in every model–condition row\. Within L3, the correct rate for the complete member ranges from 82\.4 to 100\.0%, whereas that for the partial member ranges from 0\.0 to 27\.9%\. Thus, the low L3 CS is driven primarily by predictingCertified\-Negativefor the partial member, rather than by errors on the complete member\. This complete–partial asymmetry holds across all four complete and four partial source\-type families in the domain\-crossed L3 inventory \(Appendix D, Table[14](https://arxiv.org/html/2608.04591#A4.T14)\)\.

Errors concentrate in complete–partial pairs, especially L3 \(RQ2\)\.Compared with Scope\-Mismatch pairs, complete–partial pairs have lower CS and more Both\-CN outcomes across every model and condition shown\. Within the complete–partial branch, L3 has the lowest or tied\-lowest CS: models usually answer the implicit\-complete member correctly but often also answer the implicit\-partial member asCertified\-Negative\.

### Prompting Effects

Table[3](https://arxiv.org/html/2608.04591#Sx6.T3)shows that explicit rules and CoT often improve aggregate performance, but RQ3 asks whether these gains repair the failure modes identified above or merely shift errors across items\. We therefore compare predictions item by item\. Table[6](https://arxiv.org/html/2608.04591#Sx6.T6)reports the accuracy change after adding CoT within each item group\. Positive values indicate more corrections than regressions\.

Regime\-SpecificCSL3 MemberCorrect RateModelCond\.L1L2L3L4Comp\.\-CN↑\\uparrowPart\.\-UNK↑\\uparrowQwenNaive45\.436\.712\.517\.386\.524\.7Def\.58\.870\.09\.639\.789\.719\.9Def\.\+CoT94\.997\.18\.733\.391\.716\.0Cert\.85\.079\.27\.765\.186\.918\.6HaikuNaive38\.030\.49\.326\.696\.213\.1Def\.97\.899\.410\.657\.182\.427\.9Def\.\+CoT100\.099\.415\.175\.692\.922\.1Cert\.94\.699\.717\.669\.699\.418\.3GemmaNaive3\.80\.00\.04\.5100\.00\.0Def\.49\.842\.80\.040\.7100\.00\.0Def\.\+CoT99\.499\.46\.461\.598\.76\.7Cert\.22\.416\.90\.028\.8100\.00\.0Table 5:Regime\-specific CS on 1,250 complete–partial pairs \(%\)\. L1–L4 report CS within each coverage regime\. Comp\.\-CN and Part\.\-UNK report the correct rates for the complete and partial members of L3, respectively\. Full results appear in Table[13](https://arxiv.org/html/2608.04591#A4.T13)in Appendix D\.ModelTransitionCompleteL3\-Part\.L4\-Part\.SMQwenDef\.→\\rightarrowDef\.\+CoT\+8\.9\-3\.8\-7\.7\-11\.1Cert\.→\\rightarrowCert\.\+CoT\+9\.9\-4\.8\-29\.8\-5\.7HaikuDef\.→\\rightarrowDef\.\+CoT\+11\.8\-5\.8\-4\.2\+3\.6Cert\.→\\rightarrowCert\.\+CoT\+4\.4\+1\.3\-37\.2\-0\.6GemmaDef\.→\\rightarrowDef\.\+CoT\+0\.5\+6\.7\+6\.1\+9\.8Cert\.→\\rightarrowCert\.\+CoT\+0\.3\+0\.6\+19\.9\+9\.9Table 6:Item\-level accuracy change after adding CoT \(percentage points\)\. L3\-Part\. and L4\-Part\. denote partial members in L3 and L4, respectively; SM denotes Scope\-Mismatch members\. Positive values indicate accuracy gains, and negative values indicate losses\. Full transition results appear in Table[15](https://arxiv.org/html/2608.04591#A4.T15)in Appendix D\.CoT does not provide a uniform correction\. Under Definition\-aware prompting, it repairs complete cases for Qwen and Haiku but regresses on L3 and L4 partial evidence\. With certificates, the regression concentrates on L4 for both models, while the L3 change is small and differs in sign\. Gemma instead shows broader repairs on non\-closing cases\. Thus, the same instruction can move models in opposite directions\.

As shown in Table[15](https://arxiv.org/html/2608.04591#A4.T15)in Appendix D, abstention\-aware prompting produces a more predictable shift towardUnknown: it repairs many non\-closing cases but creates new under\-closure on complete evidence\. Self\-checking is less consistent and does not yield a common improvement pattern across models\.

Prompting redistributes rather than removes closure errors \(RQ3\)\.Explicit rules and reasoning can improve aggregate performance, but no intervention consistently repairs partial\-coverage over\-closure while preserving correct judgments on complete evidence across models\. Prompting often redistributes errors between over\- and under\-closure and does not consistently improve discrimination between complete query\-covering evidence and non\-closing evidence\.

### Diagnostic Decomposition

To answer RQ4, we analyze the three fields of each incorrect completeness certificateC=\(S^q,S^E,c^\)C=\(\\hat\{S\}\_\{q\},\\hat\{S\}\_\{E\},\\hat\{c\}\)\. We compare the reported query scopeS^q\\hat\{S\}\_\{q\}withS​\(q\)S\(q\), the reported evidence coverageS^E\\hat\{S\}\_\{E\}with the gold evidence scopeS​\(E\)S\(E\)and its completeness status, and the final judgmentc^\\hat\{c\}withcc\. Each error is assigned to its earliest erroneous field: query\-scope whenS^q\\hat\{S\}\_\{q\}is incorrect, evidence\-coverage whenS^E\\hat\{S\}\_\{E\}is incorrect after an adequate query scope, and Boolean when both scope fields are adequate butc^≠c\\hat\{c\}\\neq c\.

Model\#Err\.QueryS^q\\hat\{S\}\_\{q\}EvidenceS^E\\hat\{S\}\_\{E\}Booleanc^\\hat\{c\}Qwen85824\.450\.125\.5Haiku60623\.945\.430\.7Gemma143826\.772\.01\.3Table 7:Earliest erroneous field among incorrect Certificate outputs\. \#Err\. is the error count; the remaining columns report percentages\.Evidence\-coverage errors inS^E\\hat\{S\}\_\{E\}form the largest category for all three models and are especially concentrated for Gemma\. Qwen and Haiku show more distributed profiles, including more errors in the final judgmentc^\\hat\{c\}\. Thus, the earliest reported error most often lies in the characterization of evidence scope or completeness, rather than only in the final Boolean judgment\.

Evidence coverage is the main reported failure point \(RQ4\)\.The certificate fields show that the earliest reported error is most often inS^E\\hat\{S\}\_\{E\}, where models mischaracterize the evidence scope or completeness rather than only erring in the final coverage judgment\.

ProceedingsMedicalModelCond\.A\-CN↑\\uparrowB\-UNK↑\\uparrowC\-UNK↑\\uparrowΔB−C\\Delta\_\{B\-C\}A\-CN↑\\uparrowB\-UNK↑\\uparrowQwenNaive35\.093\.673\.719\.924\.7100\.0Def\.0\.0100\.0100\.00\.08\.6100\.0Def\.\+CoT88\.799\.666\.533\.194\.8100\.0Cert\.93\.2100\.091\.78\.371\.598\.5HaikuNaive62\.863\.220\.342\.986\.541\.2Def\.99\.232\.73\.029\.799\.324\.3Def\.\+CoT100\.098\.127\.870\.3100\.0100\.0Cert\.95\.5100\.012\.088\.091\.4100\.0GemmaNaive97\.01\.10\.80\.499\.310\.5Def\.99\.60\.00\.00\.0100\.041\.6Def\.\+CoT100\.093\.63\.490\.298\.1100\.0Cert\.100\.098\.91\.997\.097\.8100\.0Table 8:CROWN\-Real variant\-level correct rates \(%\) for the four core LLM conditions\. Full results are reported in Table[16](https://arxiv.org/html/2608.04591#A4.T16)in Appendix D\.
### Transfer to CROWN\-Real

Table[8](https://arxiv.org/html/2608.04591#Sx6.T8)reports how the models handle the A/B/C variants on two real\-document sources\. A\-CN is the percentage of query\-covering A variants predicted asCertified\-Negative; B\-UNK and C\-UNK are the percentages of B and C variants predicted asUnknown, respectively\. For Proceedings,ΔB−C=B\-UNK−C\-UNK\\Delta\_\{B\-C\}=\\text\{B\-UNK\}\-\\text\{C\-UNK\}; a positive value means that C is more error\-prone than B\. Medical C is an explicit\-partial control and is excluded from the implicit\-partial comparison\.

The results vary across models, prompts, and document sources\. One pattern, however, is consistent\. In Proceedings, B is complete for one queried collection but does not cover the full query scope, corresponding to Scope\-Mismatch in CROWN\-Synth\. C contains only a subset of titles from the queried collections, corresponding to partial coverage\. Across all 15 model–condition cells,ΔB−C\\Delta\_\{B\-C\}is nonnegative, and its bootstrap interval excludes zero in 12 \(Appendix D, Table[20](https://arxiv.org/html/2608.04591#A4.T20)\)\. Because B and C are both labeledUnknown, the lower C\-UNK rates show that models more often predictCertified\-Negativefor partial C than for narrower\-complete B\. This mirrors the ordering in CROWN\-Synth, where complete–partial pairs have lower CS and higher Both\-CN rates than complete–Scope\-Mismatch pairs\.

Partial evidence remains at least as difficult as Scope\-Mismatch evidence \(RQ5\)\.In CROWN\-Synth and the Proceedings component of CROWN\-Real, partial evidence is at least as error\-prone as evidence that is complete only for a narrower scope\. However, accuracy on query\-covering evidence and the size of the partial\-coverage error differ across models, prompting conditions, and document sources\.

## Conclusion

We introduced completeness\-sensitive negative reasoning and CROWN\-QA, combining a controlled paired core with a real\-document contrast\-set evaluation\. LLMs show partial but unstable ability to distinguishCertified\-NegativefromUnknown\. CROWN\-Synth exposes a pronounced asymmetry in the implicit source\-type regime: models usually classify complete framings correctly, but often also label the matched partial framings asCertified\-Negativeinstead ofUnknown\. Prompting redistributes errors between over\- and under\-closure without a consistent remedy, while certificate analysis shows that the earliest reported error most often lies in evidence\-coverage characterization\. On CROWN\-Real Proceedings, partial evidence is at least as error\-prone as narrower\-scope complete evidence across every model and condition, although accuracy and error balance vary\. Taken together, these results isolate a challenge beyond detecting absent support: determining whether the available evidence completely covers the query scope\.

## References

- A\. N\. Angelopoulos, S\. Bates, M\. Jordan, and J\. Malik \(2021\)Uncertainty sets for image classifiers using conformal prediction\.InProceedings of the International Conference on Learning Representations,Cited by:[Related Work](https://arxiv.org/html/2608.04591#Sx2.p4.1)\.
- Association for Computational Linguistics \(2026\)ACL Anthology\.Note:https://aclanthology\.org/Cited by:[CROWN\-Real: Real\-Document Transfer Set](https://arxiv.org/html/2608.04591#Sx4.SSx2.p1.1)\.
- J\. Chen, H\. Lin, X\. Han, and L\. Sun \(2024\)Benchmarking large language models in retrieval\-augmented generation\.InProceedings of the AAAI Conference on Artificial Intelligence,pp\. 17754–17762\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p1.1)\.
- L\. Chen, Z\. Liang, X\. Wang, J\. Liang, Y\. Xiao, F\. Wei, J\. Chen, Z\. Hao, B\. Han, and W\. Wang \(2025\)Teaching large language models to express knowledge boundary from their own signals\.InProceedings of the 3rd Workshop on Towards Knowledgeable Foundation Models,Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p4.1)\.
- J\. R\. Cole, M\. J\.Q\. Zhang, D\. Gillick, J\. M\. Eisenschlos, B\. Dhingra, and J\. Eisenstein \(2023\)Selectively answering ambiguous questions\.InProceedings of the Conference on Empirical Methods in Natural Language Processing,pp\. 530–543\.Cited by:[Related Work](https://arxiv.org/html/2608.04591#Sx2.p4.1)\.
- F\. Darari, W\. Nutt, G\. Pirrò, and S\. Razniewski \(2013\)Completeness statements about rdf data sources and their use for query answering\.InProceedings of the International Semantic Web Conference,pp\. 66–83\.Cited by:[Related Work](https://arxiv.org/html/2608.04591#Sx2.p3.1)\.
- R\. El\-Yaniv and Y\. Wiener \(2010\)On the foundations of noise\-free selective classification\.Journal of Machine Learning Research11,pp\. 1605–1641\.Cited by:[Related Work](https://arxiv.org/html/2608.04591#Sx2.p4.1)\.
- S\. Es, J\. James, L\. Espinosa\-Anke, and S\. Schockaert \(2024\)RAGAS: automated evaluation of retrieval augmented generation\.InProceedings of the Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations,pp\. 150–158\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p1.1)\.
- F\. Fatahi Bayat, L\. Zhang, S\. Munir, and L\. Wang \(2025\)FactBench: a dynamic benchmark for in\-the\-wild language model factuality evaluation\.InProceedings of the Annual Meeting of the Association for Computational Linguistics,pp\. 33090–33110\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p4.1)\.
- H\. Y\. Fu, A\. Shrivastava, J\. Moore, P\. West, C\. Tan, and A\. Holtzman \(2025\)AbsenceBench: language models can’t tell what’s missing\.arXiv preprint arXiv:2506\.11440\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p4.1),[Table 1](https://arxiv.org/html/2608.04591#Sx2.T1.1.1.1.2),[Related Work](https://arxiv.org/html/2608.04591#Sx2.p2.1)\.
- I\. García\-Ferrero, B\. Altuna, J\. Alvez, I\. Gonzalez\-Dios, and G\. Rigau \(2023\)This is not a dataset: a large negation benchmark to challenge large language models\.InProceedings of the Conference on Empirical Methods in Natural Language Processing,pp\. 8596–8615\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p4.1),[Table 1](https://arxiv.org/html/2608.04591#Sx2.T1.2.2.2.2),[Related Work](https://arxiv.org/html/2608.04591#Sx2.p2.1)\.
- Y\. Geifman and R\. El\-Yaniv \(2017\)Selective classification for deep neural networks\.InProceedings of the International Conference on Neural Information Processing Systems,pp\. 4885–4894\.Cited by:[Related Work](https://arxiv.org/html/2608.04591#Sx2.p4.1)\.
- M\. Glockner, X\. Jiang, L\. F\. R\. Ribeiro, I\. Gurevych, and M\. Dreyer \(2025\)NeoQA: evidence\-based question answering with generated news events\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 11842–11926\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p4.1),[Related Work](https://arxiv.org/html/2608.04591#Sx2.p1.1)\.
- H\. Joren, J\. Zhang, C\. Ferng, D\. Juan, A\. Taly, and C\. Rashtchian \(2025\)Sufficient context: a new lens on retrieval augmented generation systems\.InProceedings of the International Conference on Learning Representations,Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p1.1),[Introduction](https://arxiv.org/html/2608.04591#Sx1.p4.1),[Table 1](https://arxiv.org/html/2608.04591#Sx2.T1.3.3.5.1),[Related Work](https://arxiv.org/html/2608.04591#Sx2.p1.1)\.
- S\. Kadavath, T\. Conerly, A\. Askell, T\. Henighan,et al\.\(2022\)Language models \(mostly\) know what they know\.arXiv preprint arXiv:2207\.05221\.Cited by:[Related Work](https://arxiv.org/html/2608.04591#Sx2.p4.1)\.
- P\. Kirichenko, M\. Ibrahim, K\. Chaudhuri, and S\. J\. Bell \(2025\)AbstentionBench: reasoning llms fail on unanswerable questions\.arXiv preprint arXiv:2506\.09038\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p4.1),[Table 1](https://arxiv.org/html/2608.04591#Sx2.T1.3.3.7.1),[Related Work](https://arxiv.org/html/2608.04591#Sx2.p1.1)\.
- P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Kuttler, M\. Lewis, W\. Yih, T\. Rocktaschel, S\. Riedel, and D\. Kiela \(2020\)Retrieval\-augmented generation for knowledge\-intensive nlp tasks\.InProceedings of the International Conference on Neural Information Processing Systems,pp\. 9459–9474\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p1.1),[Related Work](https://arxiv.org/html/2608.04591#Sx2.p1.1)\.
- Z\. Li, J\. Zhang, C\. Yan, K\. Das, S\. Kumar, M\. Kantarcioglu, and B\. A\. Malin \(2024\)Do you know what you are talking about? characterizing query\-knowledge relevance for reliable retrieval augmented generation\.InProceedings of the Conference on Empirical Methods in Natural Language Processing,pp\. 6130–6151\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p3.1),[Related Work](https://arxiv.org/html/2608.04591#Sx2.p1.1)\.
- N\. Madhusudhan, S\. T\. Madhusudhan, V\. Yadav, and M\. Hashemi \(2025\)Do LLMs know when to NOT answer? investigating abstention abilities of large language models\.InProceedings of the International Conference on Computational Linguistics,pp\. 9329–9345\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p4.1),[Related Work](https://arxiv.org/html/2608.04591#Sx2.p4.1)\.
- C\. Niu, Y\. Wu, J\. Zhu, S\. Xu, K\. Shum, R\. Zhong, J\. Song, and T\. Zhang \(2024\)RAGTruth: a hallucination corpus for developing trustworthy retrieval\-augmented language models\.InProceedings of the Annual Meeting of the Association for Computational Linguistics,pp\. 10862–10878\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p1.1),[Related Work](https://arxiv.org/html/2608.04591#Sx2.p1.1)\.
- X\. Peng, P\. K\. Choubey, C\. Xiong, and C\. Wu \(2025\)Unanswerability evaluation for retrieval augmented generation\.InProceedings of the Annual Meeting of the Association for Computational Linguistics,pp\. 8452–8472\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p4.1),[Table 1](https://arxiv.org/html/2608.04591#Sx2.T1.3.3.6.1),[Related Work](https://arxiv.org/html/2608.04591#Sx2.p1.1)\.
- P\. Rajpurkar, R\. Jia, and P\. Liang \(2018\)Know what you don’t know: unanswerable questions for squad\.InProceedings of the Annual Meeting of the Association for Computational Linguistics,pp\. 784–789\.Cited by:[Related Work](https://arxiv.org/html/2608.04591#Sx2.p4.1)\.
- S\. Razniewski, H\. Arnaout, S\. Ghosh, and F\. Suchanek \(2024\)Completeness, recall, and negation in open\-world knowledge bases: a survey\.ACM Computing Surveys56\(6\),pp\. 1–42\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p3.1),[Table 1](https://arxiv.org/html/2608.04591#Sx2.T1.3.3.3.2),[Related Work](https://arxiv.org/html/2608.04591#Sx2.p3.1)\.
- S\. Razniewski, F\. Korn, W\. Nutt, and D\. Srivastava \(2015\)Identifying the extent of completeness of query answers over partially complete databases\.InProceedings of the ACM SIGMOD International Conference on Management of Data,pp\. 561–576\.Cited by:[Related Work](https://arxiv.org/html/2608.04591#Sx2.p3.1)\.
- J\. Saad\-Falcon, O\. Khattab, C\. Potts, and M\. Zaharia \(2024\)ARES: an automated evaluation framework for retrieval\-augmented generation systems\.InProceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 338–354\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p1.1),[Related Work](https://arxiv.org/html/2608.04591#Sx2.p1.1)\.
- M\. Shafiei, H\. Saffari, and N\. S\. Moosavi \(2025\)MultiHoax: a dataset of multi\-hop false\-premise questions\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 10169–10187\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p4.1)\.
- U\.S\. National Library of Medicine \(2026\)DailyMed\.Note:https://dailymed\.nlm\.nih\.gov/dailymed/Cited by:[CROWN\-Real: Real\-Document Transfer Set](https://arxiv.org/html/2608.04591#Sx4.SSx2.p1.1)\.
- X\. Yang, K\. Sun, H\. Xin, Y\. Sun,et al\.\(2024\)CRAG – comprehensive RAG benchmark\.InProceedings of the International Conference on Neural Information Processing Systems,pp\. 10470–10490\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p1.1)\.
- Z\. Yin, Q\. Sun, Q\. Guo, J\. Wu, X\. Qiu, and X\. Huang \(2023\)Do large language models know what they don’t know?\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 8653–8665\.Cited by:[Related Work](https://arxiv.org/html/2608.04591#Sx2.p4.1)\.
- W\. Yu, H\. Zhang, X\. Pan, P\. Cao, K\. Ma, J\. Li, H\. Wang, and D\. Yu \(2024\)Chain\-of\-note: enhancing robustness in retrieval\-augmented language models\.InProceedings of the Conference on Empirical Methods in Natural Language Processing,pp\. 14672–14685\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p4.1)\.
- T\. Zhanget al\.\(2024\)CLAMBER: a benchmark of identifying and clarifying ambiguous information needs in large language models\.InProceedings of the Annual Meeting of the Association for Computational Linguistics,pp\. 10746–10766\.Cited by:[Introduction](https://arxiv.org/html/2608.04591#Sx1.p4.1),[Related Work](https://arxiv.org/html/2608.04591#Sx2.p4.1)\.

## Appendix AAppendix A\. Dataset Examples

### Dataset Format and Worked Examples

CROWN\-Synth and CROWN\-Real are stored as JSONL files, with one example per line\. The examples below show shortened versions of the text given to the models\. Full records and metadata are included in the submitted files\.

##### CROWN\-Synth worked pair\.

The following example is the L3 pairproceedings\_0252\. MLConf and the listed paper titles are synthetic\. Both members use the same question and the same three listed titles\. Only the source\-type phrase in the first sentence changes\.

Question: Did MLConf 2025 publish a paper whose title contains “structured prompt plans”?

Shared fictional paper titles:

- •Reliable Evaluation for Planning Agents
- •Incremental Updates for Knowledge Stores
- •Scoring Citations in Literature Reviews

No listed title contains “structured prompt plans\.”

Complete memberPartial memberCoverage sentenceThe master index for the MLConf 2025 proceedings lists these paper titles\.The highlights digest for the MLConf 2025 proceedings lists these paper titles\.Gold labelCertified\-NegativeUnknownTable 9:Shortened CROWN\-Synth L3 pair\. The question and listed titles are identical; only the source\-type phrase changes\.
##### CROWN\-Real worked contrast set\.

The following medical contrast set uses the same question and target phrase in all three variants\. The target phrase is absent from every displayed context\.

Question: For Atropine label version 2026\-03\-10, does Section 5, Warnings and Precautions, of the Full Prescribing Information contain the exact phrase “Impaired Renal Function”?

Variant AVariant BVariant CEvidence shownFull Section 5, beginning with “5\. WARNINGS AND PRECAUTIONS” and ending at “6\. ADVERSE REACTIONS\.”Subsection 5\.2, “Elevation of Blood Pressure,” ending at the heading for subsection 5\.3\.“HIGHLIGHTS OF PRESCRIBING INFORMATION,” including the statement that the Highlights do not contain all information needed to use the drug\.CoverageComplete for the full queried Section 5Complete only for subsection 5\.2Non\-exhaustive HighlightsGold labelCertified\-NegativeUnknownUnknownTable 10:Shortened CROWN\-Real medical contrast set\. The question and target phrase are fixed, while the displayed document range changes\.

## Appendix BAppendix B\. Prompt Templates

Table[B](https://arxiv.org/html/2608.04591#A2)summarizes the seven LLM conditions, the number of calls required by each condition, and the intervention being tested\. The conditions range from direct\-label prediction to reasoning, abstention\-aware prompting, self\-revision, and structured scope\-and\-coverage elicitation\. We next provide the exact prompt templates and shared components used across conditions\.

ConditionCallsInterventionNaive1Direct label without the task rule\.Definition\-aware1Direct label with the task rule\.Definition\-aware\+CoT1Adds step\-by\-step reasoning\.Definition\-aware\+Abstain1Adds a conservative default toUnknown\.Definition\-aware\+Self\-check2Reviews and may revise an initial label\.Certificate1Elicits\(S^q,S^E,c^\)\(\\hat\{S\}\_\{q\},\\hat\{S\}\_\{E\},\\hat\{c\}\)\.Certificate\+CoT1Adds reasoning to certificate elicitation\.Table 11:Summary of the evaluation conditions and their corresponding interventions\.All LLM conditions receive the same base question and evidence context, and no few\-shot demonstrations are used\. For the API model, the common system message is supplied through the system role\. For open\-weight models, it is placed in the system segment of the corresponding chat template\.

The prompt templates below use bracketed block names such as \[Shared Task Rule\] and \[Input Block\]\. At runtime, each bracketed name is replaced by the complete text of the corresponding common component defined below\. The bracketed names themselves are not sent to the model\. The order shown inside each prompt box is the exact order used at runtime, with one blank line between consecutive blocks\.

### Common Components

#### System Message

The following system message is used for every LLM call\.

You are an evidence\-grounded question\-answering classifier\. Use only the provided evidence and do not use external knowledge\. Follow the requested output format exactly and do not add unrequested text\.

#### Shared Task Rule

The following rule block is used by Definition\-aware and its CoT, Abstain, and Self\-check variants\. It is omitted from Naive, Certificate, and Certificate\+CoT\.

Apply the following rules:The queried fact is not supported by the observed factual content\. Your task is to decide whether this observed absence is licensed as a negative answer\.Use CERTIFIED\_NEGATIVE only when the evidence is complete for the exact scope of the question\.Use UNKNOWN when the evidence is partial, sampled, incomplete, unspecified, ambiguous, or complete only for a different scope\.Evidence that is complete only for a different entity, time period, attribute, or collection is not complete for the question\.

#### Direct\-Label Output Block

The following output block is used when the LLM directly returns one of the two task labels\.

Return exactly one of the following labels and nothing else:CERTIFIED\_NEGATIVEUNKNOWN

#### Input Block

The following input block is appended to every LLM user message\.

Evidence:\{context\}Question:\{question\}

Direct\-label outputs are parsed according to the condition\-specific format given below\. The machine\-readable labels CERTIFIED\_NEGATIVE and UNKNOWN correspond respectively toCertified\-NegativeandUnknown\. Output normalization is specified in Appendix C\.

### Naive Prompt

The Naive condition omits the shared task rule and measures the model’s default treatment of missing evidence\. The following box shows the complete user\-message assembly\.

Classify whether the observed absence licenses a negative answer using only the provided evidence\.\[Direct\-Label Output Block\]\[Input Block\]

Thus, the Naive condition consists of the common system message and one user message containing the three components above\.

### Definition\-Aware Prompt

The Definition\-aware condition provides the task rule but does not request intermediate scope or coverage judgments\.

Classify whether the observed absence licenses a negative answer using only the provided evidence\.\[Shared Task Rule\]\[Direct\-Label Output Block\]\[Input Block\]

The LLM directly generates the final label in one call\.

### Definition\-Aware\+CoT Prompt

The Definition\-aware\+CoT condition adds a general step\-by\-step reasoning instruction to the Definition\-aware condition\. The LLM still generates the final label directly; no structured certificate or fixed decision rule is used\.

Classify whether the observed absence licenses a negative answer using only the provided evidence\.\[Shared Task Rule\]Reason step by step before choosing the final label\.End the response with exactly one of the following lines:FINAL LABEL: CERTIFIED\_NEGATIVEFINAL LABEL: UNKNOWN\[Input Block\]

Only the final line beginning with “FINAL LABEL:” is used for scoring\.

### Definition\-Aware\+Abstain Prompt

The Definition\-aware\+Abstain condition uses the shared task rule and adds a conservative default for uncertain cases\.

Classify whether the observed absence licenses a negative answer using only the provided evidence\.\[Shared Task Rule\]Adopt a conservative policy\. If it is uncertain whether the evidence is complete for the exact scope of the question, return UNKNOWN\.\[Direct\-Label Output Block\]\[Input Block\]

This condition tests whether over\-closure can be reduced by returningUnknownmore readily and whether this increases under\-closure\.

### Definition\-Aware\+Self\-Check Prompt

Definition\-aware\+Self\-check uses two LLM calls\. The first call uses the complete Definition\-aware prompt above\. Its normalized output is inserted as \{initial\_label\} in the second\-call prompt below\.

Review the following initial label and revise it if necessary:\{initial\_label\}\[Shared Task Rule\]Do not assume that the initial label is correct\. Verify whether the evidence is complete for the exact scope of the question\. Return the label implied by the task rule\.\[Direct\-Label Output Block\]\[Input Block\]

Only the second\-call response is used as the final prediction\.

### Certificate Prompt

The Certificate condition uses one structured LLM call\. Unlike the direct\-label conditions, it does not include the Shared Task Rule, because the LLM is not asked to choose between the two labels\. Instead, the model returns the query scope, evidence coverage scope, and coverage judgment\. A fixed decision rule then maps thecomplete\_for\_queryfield to the final prediction\.

Do not output a final answer label\.Using only the provided evidence, return a valid JSON object with exactly the following three keys:\{"query\_scope": "…","evidence\_coverage\_scope": "…","complete\_for\_query": true or false\}Use the fields as follows:"query\_scope" must describe all constraints required by the question, such as the relevant entity, time period, attribute, and collection\."evidence\_coverage\_scope" must describe what the evidence covers and whether that coverage is complete, partial, or unspecified\.Set "complete\_for\_query" to true only when the evidence establishes complete coverage of the entire query scope\. Evidence that is complete only for a different scope is not complete for the question\. Set the field to false when coverage is partial, sampled, unspecified, ambiguous, or mismatched with the query scope\.Return only the JSON object\. Do not include an explanation, markdown formatting, a final label, or additional keys\.\[Input Block\]

### Certificate\+CoT Prompt

The Certificate\+CoT condition uses one structured LLM call\. It asks the model to reason about the query scope, evidence coverage scope, and coverage relation before returning the same three certificate fields as Certificate\. It does not include the Shared Task Rule or the Direct\-Label Output Block\.

Do not output a final answer label\.First reason step by step about the following three points:1\. What is the exact query scope?2\. What scope does the evidence claim to cover?3\. Does the evidence coverage close the entire query scope?Then return a final JSON object\.Use only the provided evidence\. Do not use external knowledge\.The final JSON object must have exactly the following three keys:\{"query\_scope": "…","evidence\_coverage\_scope": "…","complete\_for\_query": true or false\}Use the fields as follows:"query\_scope" must describe all constraints required by the question, such as the relevant entity, time period, attribute, and collection\."evidence\_coverage\_scope" must describe what the evidence covers and whether that coverage is complete, partial, or unspecified\.Set "complete\_for\_query" to true only when the evidence establishes complete coverage of the entire query scope\. Evidence that is complete only for a different scope is not complete for the question\. Set the field to false when coverage is partial, sampled, unspecified, ambiguous, or mismatched with the query scope\.Return the response in exactly this format:REASONING:<brief step\-by\-step reasoning\>FINAL\_JSON:\{"query\_scope": "…","evidence\_coverage\_scope": "…","complete\_for\_query": true or false\}Do not wrap the JSON in markdown code fences\. Do not include a final label, additional keys, or any text after the JSON object\. The value of "complete\_for\_query" must be a JSON boolean, not a string\.\[Input Block\]

Only the JSON object following “FINAL\_JSON:” is parsed for scoring\. Its Boolean field is mapped to the two\-label space using the same fixed mapping as Certificate\. The reasoning text and the two scope fields are retained in the experiment logs for diagnostic analysis\.

## Appendix CAppendix C\. Implementation Details

### Models and Runtime Environment

We evaluate three instruction\-tuned LLMs spanning open\-weight and API access: Qwen3\.5\-9B, Gemma\-4\-12B, and Claude Haiku 4\.5\. For each benchmark, all models are evaluated on the same frozen dataset version and receive no in\-context examples\.

The open\-weight models were run on one NVIDIA A100 GPU in Google Colab\. Claude Haiku 4\.5 was accessed through the Anthropic API\. Definition\-aware\+Self\-check normally used two sequential calls per example; all other conditions normally used one call\. If a response could not be parsed, the same call was repeated once with the same input, prompt, and generation settings\. No further retries were made\. All calls used temperature0, and the open\-weight models used greedy decoding\. The submitted artifact contains the datasets, exact prompts, raw outputs, inference and scoring code, model configuration files, and the Python environment used for the reported experiments\.

### Generation Limit

All calls used a maximum output length of 5120 new tokens\. No input question or evidence context was truncated\.

### Output Parsing and Scoring

All predictions are scored in the two\-label space\{Certified\-Negative,Unknown\}\\\{\\textsc\{Certified\-Negative\},\\textsc\{Unknown\}\\\}\. The parser follows the output format specified for each condition in Appendix B\. For Naive, Definition\-aware, and Definition\-aware\+Abstain, the parser reads the direct label\. For Self\-check, the first response provides the initial label, and the second response is used as the final prediction for scoring\. For Definition\-aware\+CoT, the parser reads the final line beginning withFINAL LABEL:\.

For Certificate, the parser reads the returned JSON object\. For Certificate\+CoT, it reads the JSON object followingFINAL\_JSON:\. The parsed object must containquery\_scope,evidence\_coverage\_scope, andcomplete\_for\_query\. The Boolean field is mapped as follows:

true↦Certified\-Negative,false↦Unknown\.\\texttt\{true\}\\mapsto\\textsc\{Certified\-Negative\},\\qquad\\texttt\{false\}\\mapsto\\textsc\{Unknown\}\.
If the first response could not be parsed, the same call was repeated once with the same input, prompt, and generation settings\. When the second response was parseable, it was used for scoring\. If the second response also could not be parsed, the example was counted as incorrect and remained in the denominator; it was never mapped toUnknown\. The “Off” column of Table[12](https://arxiv.org/html/2608.04591#A4.T12)reports the rate of examples that remained unparseable after this retry\.

For the Certificate condition in Table[7](https://arxiv.org/html/2608.04591#Sx6.T7), each incorrect output is assigned to the first field that fails the rule\-based check\. The script first comparesquery\_scopewith the gold query scope\. If that check passes, it comparesevidence\_coverage\_scopewith the gold evidence scope and coverage status\. If both scope fields pass butcomplete\_for\_queryis incorrect, the error is assigned to the Boolean field\.

## Appendix DAppendix D\. Additional Experimental Results and Analyses

### Full Pair Outcome Decomposition

Table[12](https://arxiv.org/html/2608.04591#A4.T12)extends the pair\-level analysis in the main text to all evaluation conditions\. Across all evaluation conditions, complete–partial pairs produce Both\-CN more often than complete–Scope\-Mismatch pairs\. This indicates that models are more prone to over\-close on partial evidence than on Scope\-Mismatch evidence\. Reversed outcomes and output\-format failures are negligible, so most pair failures arise from assigning the same label to both members\.

Pair: Complete vs\. PartialPair: Complete vs\. Scope\-MismatchModelConditionCS↑\\uparrowBoth\-CN↓\\downarrowBoth\-UNK↓\\downarrowRev\.↓\\downarrowOff↓\\downarrowCS↑\\uparrowBoth\-CN↓\\downarrowBoth\-UNK↓\\downarrowRev\.↓\\downarrowOff↓\\downarrowQwenNaive28\.056\.613\.42\.00\.057\.614\.626\.61\.10\.0Def\.44\.643\.811\.60\.00\.075\.59\.614\.50\.40\.0Def\.\+CoT58\.634\.56\.60\.20\.177\.620\.31\.30\.50\.3Def\.\+Abstain53\.722\.123\.90\.30\.059\.53\.436\.90\.20\.0Def\.\+Self\-check48\.614\.336\.40\.70\.050\.22\.247\.30\.20\.0Cert\.59\.330\.210\.00\.60\.072\.99\.817\.10\.20\.0Cert\.\+CoT65\.232\.31\.10\.21\.178\.614\.45\.00\.71\.4HaikuNaive26\.167\.96\.00\.00\.069\.013\.317\.30\.40\.0Def\.66\.218\.615\.00\.10\.067\.97\.924\.20\.00\.0Def\.\+CoT72\.620\.57\.00\.00\.087\.04\.38\.60\.00\.0Def\.\+Abstain67\.815\.716\.60\.00\.064\.78\.426\.70\.20\.0Def\.\+Self\-check71\.219\.98\.90\.00\.069\.613\.117\.10\.20\.0Cert\.70\.421\.97\.60\.10\.081\.24\.614\.20\.00\.0Cert\.\+CoT64\.229\.95\.80\.00\.087\.65\.07\.20\.20\.0GemmaNaive2\.197\.90\.00\.00\.044\.253\.91\.60\.20\.0Def\.33\.461\.84\.20\.60\.067\.726\.35\.60\.40\.0Def\.\+CoT66\.731\.80\.60\.90\.075\.016\.78\.10\.20\.0Def\.\+Abstain55\.339\.45\.30\.00\.069\.422\.28\.10\.30\.0Def\.\+Self\-check34\.565\.30\.00\.20\.058\.737\.53\.10\.60\.0Cert\.17\.082\.70\.20\.10\.068\.228\.23\.40\.20\.0Cert\.\+CoT36\.461\.91\.50\.20\.080\.218\.31\.40\.20\.0Table 12:Full pair\-level outcome decomposition across models and conditions \(%\)\. CS denotes the correct paired outcome\(Certified\-Negative,Unknown\)\(\\textsc\{Certified\-Negative\},\\textsc\{Unknown\}\)\. Both\-CN and Both\-UNK assign the same label to both members; Rev\. denotes the reversed\(Unknown,Certified\-Negative\)\(\\textsc\{Unknown\},\\textsc\{Certified\-Negative\}\)outcome, and Off denotes output\-format failures\.The full\-condition decomposition therefore reinforces the RQ2 finding that partial evidence is a more persistent source of over\-closure than Scope\-Mismatch evidence\.

### Full Complete–Partial Regime Analysis

Regime\-SpecificCSL3 MemberCorrect RateModelCond\.L1L2L3L4Comp\.\-CN↑\\uparrowPart\.\-UNK↑\\uparrowQwenNaive45\.436\.712\.517\.386\.524\.7Def\.58\.870\.09\.639\.789\.719\.9Def\.\+CoT94\.997\.18\.733\.391\.716\.0Def\.\+Abstain58\.576\.79\.669\.972\.136\.2Def\.\+Self\-check52\.459\.714\.467\.653\.857\.7Cert\.85\.079\.27\.765\.186\.918\.6Cert\.\+CoT96\.8100\.011\.552\.296\.813\.8HaikuNaive38\.030\.49\.326\.696\.213\.1Def\.97\.899\.410\.657\.182\.427\.9Def\.\+CoT100\.099\.415\.175\.692\.922\.1Def\.\+Abstain98\.499\.016\.357\.176\.639\.7Def\.\+Self\-check99\.099\.415\.770\.593\.322\.4Cert\.94\.699\.717\.669\.699\.418\.3Cert\.\+CoT98\.799\.416\.042\.696\.519\.6GemmaNaive3\.80\.00\.04\.5100\.00\.0Def\.49\.842\.80\.040\.7100\.00\.0Def\.\+CoT99\.499\.46\.461\.598\.76\.7Def\.\+Abstain86\.674\.10\.060\.3100\.00\.0Def\.\+Self\-check56\.242\.20\.039\.4100\.00\.0Cert\.22\.416\.90\.028\.8100\.00\.0Cert\.\+CoT43\.158\.10\.643\.699\.70\.6Table 13:Full coverage\-regime results on 1,250 complete–partial pairs \(%\)\. L1–L4 report regime\-specific CS\. Comp\.\-CN and Part\.\-UNK report the separate member\-level correct rates within L3 complete–partial pairs\.Table[13](https://arxiv.org/html/2608.04591#A4.T13)extends Table[5](https://arxiv.org/html/2608.04591#Sx6.T5)to the three additional LLM conditions omitted from the main\-text table, while preserving the same analysis units and reported columns\. Across all seven conditions, L3 remains the lowest or tied\-lowest CS in every model–condition row\. Thus, the L3 bottleneck identified for RQ2 is not limited to the four conditions reported in the main text\.

### Source\-Type Family Breakdown of the L3 Asymmetry

To test whether the L3 gap in Table[5](https://arxiv.org/html/2608.04591#Sx6.T5)is concentrated in a particular source\-type family, Table[14](https://arxiv.org/html/2608.04591#A4.T14)reports member\-level correct rates for each of the four complete and four partial families under the same four conditions\. Each family appears in all five domains, avoiding a fixed family–domain association\. Across all model–condition cells, complete\-family correct rates range from 76\.3 to 100\.0%, whereas partial\-family correct rates range from 0\.0 to 31\.6%\. Thus, the RQ2 asymmetry is observed across the full source\-type inventory rather than being concentrated in a single family\.

Complete\-CN by familyPartial\-UNK by familyModelCond\.C1C2C3C4P1P2P3P4QwenNaive85\.089\.582\.189\.717\.928\.828\.923\.1Def\.81\.288\.296\.293\.620\.520\.019\.719\.2Def\.\+CoT90\.085\.594\.996\.29\.018\.822\.414\.1Cert\.88\.881\.679\.597\.417\.922\.518\.415\.4HaikuNaive95\.097\.494\.997\.411\.516\.213\.211\.5Def\.80\.076\.378\.294\.926\.930\.031\.623\.1Def\.\+CoT91\.288\.292\.3100\.020\.520\.030\.317\.9Cert\.100\.097\.4100\.0100\.020\.520\.021\.111\.5GemmaNaive100\.0100\.0100\.0100\.00\.00\.00\.00\.0Def\.100\.0100\.0100\.0100\.00\.00\.00\.00\.0Def\.\+CoT96\.2100\.098\.7100\.05\.12\.510\.59\.0Cert\.100\.0100\.0100\.0100\.00\.00\.00\.00\.0Table 14:L3 member\-level correct rates by source\-type family on the 312 complete–partial pairs \(%\)\. C1–C4 are “system of record,” “master index,” “canonical register,” and “definitive index”; P1–P4 are “activity feed,” “update bulletin,” “highlights digest,” and “briefing summary\.” Each family appears in all five domains; rates are aggregated over domains and over the family used by the paired member\.
### Full Prompting Transition Analysis

Table[15](https://arxiv.org/html/2608.04591#A4.T15)extends Table[6](https://arxiv.org/html/2608.04591#Sx6.T6)by adding the Abstain and Self\-check transitions omitted from the main\-text table and by reporting all L1–L4 partial groups\. The four partial columns together cover the 1,250 partial members of the complete–partial branch\. Complete pools the 2,500 query\-covering complete members, and SM contains the 1,250 Scope\-Mismatch members\. Each entry is the after\-minus\-before accuracy change within the corresponding item group; positive values indicate accuracy gains\.

ModelTransitionCompleteL1\-Part\.L2\-Part\.L3\-Part\.L4\-Part\.SMQwenDef\.→\\rightarrowDef\.\+CoT\+8\.9\+34\.8\+13\.1\-3\.8\-7\.7\-11\.1Def\.→\\rightarrowDef\.\+Abstain\-17\.4\+18\.8\+12\.1\+16\.3\+38\.5\+6\.4Def\.→\\rightarrowDef\.\+Self\-check\-29\.1\+32\.3\+10\.2\+37\.8\+34\.9\+7\.5Cert\.→\\rightarrowCert\.\+CoT\+9\.9\+11\.8\+11\.2\-4\.8\-29\.8\-5\.7HaikuDef\.→\\rightarrowDef\.\+CoT\+11\.8\+2\.2\+0\.6\-5\.8\-4\.2\+3\.6Def\.→\\rightarrowDef\.\+Abstain\-2\.1\+0\.6\-0\.3\+11\.9\+0\.0\-0\.6Def\.→\\rightarrowDef\.\+Self\-check\+6\.6\+1\.3\+0\.0\-5\.4\-0\.6\-5\.4Cert\.→\\rightarrowCert\.\+CoT\+4\.4\+4\.2\+0\.0\+1\.3\-37\.2\-0\.6GemmaDef\.→\\rightarrowDef\.\+CoT\+0\.5\+49\.5\+56\.5\+6\.7\+6\.1\+9\.8Def\.→\\rightarrowDef\.\+Abstain\-1\.4\+36\.7\+32\.6\+0\.0\+22\.4\+4\.2Def\.→\\rightarrowDef\.\+Self\-check\+3\.4\+6\.4\-0\.6\+0\.0\-18\.3\-11\.4Cert\.→\\rightarrowCert\.\+CoT\+0\.3\+20\.8\+41\.5\+0\.6\+19\.9\+9\.9Table 15:Accuracy change after each within\-family prompting transition by item group \(percentage points\)\. Complete denotes query\-covering complete members; L1\-Part\.–L4\-Part\. denote partial members in the corresponding regimes; SM denotes Scope\-Mismatch members\. Positive values indicate accuracy gains\.The expanded breakdown shows that the overall results reflect very different changes across item groups\. For Qwen, adding CoT repairs complete, L1, and L2 cases but causes regressions on L3, L4, and Scope\-Mismatch; Abstain and Self\-check improve non\-closing cases at a substantial cost on complete evidence\. For Haiku, Definition\-aware\+CoT improves complete cases but degrades L3 and L4, while Certificate\+CoT sharply degrades L4 with comparatively small changes in the other non\-closing groups\. Gemma instead obtains broader gains from CoT, although Self\-check regresses on L4 and Scope\-Mismatch\. Thus, no transition improves all item groups consistently across models, supporting the RQ3 conclusion that prompting redistributes rather than uniformly removes closure errors\.

### Full CROWN\-Real Results

ProceedingsMedicalModelConditionA\-CN↑\\uparrowB\-UNK↑\\uparrowC\-UNK↑\\uparrowΔB−C\\Delta\_\{B\-C\}A\-CN↑\\uparrowB\-UNK↑\\uparrowC\-UNK↑\\uparrowQwenNaive35\.093\.673\.719\.924\.7100\.0100\.0Def\.0\.0100\.0100\.00\.08\.6100\.0100\.0Def\.\+CoT88\.799\.666\.533\.194\.8100\.097\.4Cert\.93\.2100\.091\.78\.371\.598\.5100\.0Cert\.\+CoT97\.4100\.075\.224\.895\.5100\.096\.3HaikuNaive62\.863\.220\.342\.986\.541\.218\.7Def\.99\.232\.73\.029\.799\.324\.319\.9Def\.\+CoT100\.098\.127\.870\.3100\.0100\.085\.0Cert\.95\.5100\.012\.088\.091\.4100\.095\.9Cert\.\+CoT100\.0100\.047\.452\.690\.6100\.0100\.0GemmaNaive97\.01\.10\.80\.499\.310\.599\.3Def\.99\.60\.00\.00\.0100\.041\.698\.9Def\.\+CoT100\.093\.63\.490\.298\.1100\.0100\.0Cert\.100\.098\.91\.997\.097\.8100\.099\.6Cert\.\+CoT100\.098\.98\.390\.696\.3100\.0100\.0Table 16:Full CROWN\-Real variant\-level correct rates \(%\)\.ΔB−C=B\-UNK−C\-UNK\\Delta\_\{B\-C\}=\\text\{B\-UNK\}\-\\text\{C\-UNK\}is reported in percentage points\. Medical C is an explicit\-partial control and is not used in the implicit\-partial transfer comparison\.Table[16](https://arxiv.org/html/2608.04591#A4.T16)extends Table[8](https://arxiv.org/html/2608.04591#Sx6.T8)with Certificate\+CoT and Medical C\. Within Proceedings, B represents narrower\-scope completeness and C represents partial coverage\. Across all three models and all five CROWN\-Real conditions, B\-UNK is never lower than C\-UNK: the difference is positive in 13 model–condition cells and zero in two, with no reversal\. Thus, the partial\-versus\-Scope\-Mismatch ordering reported for RQ5 is not limited to the four conditions shown in the main\-text table\.

Medical C is reported only as an explicit\-partial control and is not used to test implicit\-partial transfer\. For Medical, both non\-closing variants are generally classified correctly under Definition\-aware\+CoT, Certificate, and Certificate\+CoT, whereas A\-CN remains lower for some model–condition combinations\.

### Bootstrap Uncertainty Analysis

Tables[17](https://arxiv.org/html/2608.04591#A4.T17)–[20](https://arxiv.org/html/2608.04591#A4.T20)report percentile 95% confidence intervals based on 10,000 bootstrap samples\. CROWN\-Synth analyses resample matched\-pair identifiers from the relevant subset, whereas the CROWN\-Real analysis resamples A/B/C contrast\-set identifiers and retains all three variants of each selected set\. When prompting conditions are compared, both conditions are evaluated on the same resampled pairs\. For the Proceedings B–C comparison, both variants are evaluated on the same resampled contrast sets\. Branch and regime comparisons resample their corresponding subsets separately\. Confidence intervals including zero do not indicate a clear directional difference\.

##### Prompting contrasts\.

As shown in Table[17](https://arxiv.org/html/2608.04591#A4.T17), Definition\-aware prompting improves Acc and CS for all three models, but through different directional changes: it reduces both OCR and UCR for Qwen, while trading large OCR reductions for higher UCR in Haiku and Gemma\. Adding CoT further improves Acc and CS across models, but its OCR change is inconclusive for Qwen and Haiku and strongly negative for Gemma\. Abstention\-aware prompting consistently lowers OCR while increasing UCR, whereas Self\-check remains model\-dependent\. Certificate improves Acc and CS for Qwen and Haiku but degrades both for Gemma\. Adding CoT to Certificate improves Acc and CS for Qwen and Gemma, with no reliable change in either metric for Haiku\.

ModelContrastΔ\\DeltaAccΔ\\DeltaOCRΔ\\DeltaUCRΔ\\DeltaCSQwenDef\.−\-Naive\+9\.3 \[8\.1, 10\.5\]\-10\.3 \[\-12\.2, \-8\.4\]\-8\.3 \[\-10\.4, \-6\.4\]\+17\.2 \[15\.0, 19\.5\]Def\.\+CoT−\-Def\.\+3\.9 \[2\.9, 4\.9\]\+1\.0 \[\-0\.6, 2\.6\]\-8\.9 \[\-10\.1, \-7\.6\]\+8\.0 \[6\.1, 9\.9\]Def\.\+Abstain−\-Def\.\-1\.7 \[\-2\.8, \-0\.7\]\-13\.9 \[\-15\.3, \-12\.6\]\+17\.4 \[15\.9, 18\.9\]\-3\.4 \[\-5\.5, \-1\.4\]Def\.\+Self\-check−\-Def\.\-5\.5 \[\-6\.6, \-4\.3\]\-18\.2 \[\-19\.7, \-16\.6\]\+29\.1 \[27\.3, 30\.9\]\-10\.6 \[\-12\.9, \-8\.3\]Cert\.−\-Def\.\+2\.9 \[1\.9, 3\.9\]\-6\.6 \[\-8\.0, \-5\.1\]\+0\.7 \[\-0\.7, 2\.1\]\+6\.0 \[4\.1, 7\.9\]Cert\.\+CoT−\-Cert\.\+2\.8 \[1\.9, 3\.8\]\+4\.3 \[2\.8, 5\.8\]\-9\.9 \[\-11\.2, \-8\.6\]\+5\.8 \[4\.0, 7\.6\]HaikuDef\.−\-Naive\+9\.8 \[8\.7, 10\.9\]\-27\.5 \[\-29\.3, \-25\.7\]\+7\.8 \[6\.4, 9\.2\]\+19\.5 \[17\.3, 21\.7\]Def\.\+CoT−\-Def\.\+6\.4 \[5\.6, 7\.2\]\-0\.9 \[\-1\.9, 0\.0\]\-11\.8 \[\-13\.2, \-10\.4\]\+12\.7 \[11\.1, 14\.4\]Def\.\+Abstain−\-Def\.\-0\.4 \[\-1\.0, 0\.2\]\-1\.2 \[\-1\.9, \-0\.5\]\+2\.1 \[1\.1, 3\.1\]\-0\.8 \[\-2\.0, 0\.4\]Def\.\+Self\-check−\-Def\.\+1\.6 \[1\.0, 2\.3\]\+3\.3 \[2\.5, 4\.1\]\-6\.6 \[\-7\.8, \-5\.4\]\+3\.3 \[2\.0, 4\.7\]Cert\.−\-Def\.\+4\.4 \[3\.6, 5\.2\]0\.0 \[\-1\.0, 0\.9\]\-8\.7 \[\-10\.0, \-7\.3\]\+8\.7 \[7\.2, 10\.3\]Cert\.\+CoT−\-Cert\.\+0\.0 \[\-0\.7, 0\.7\]\+4\.3 \[3\.2, 5\.4\]\-4\.4 \[\-5\.3, \-3\.4\]\+0\.1 \[\-1\.2, 1\.5\]GemmaDef\.−\-Naive\+13\.5 \[12\.6, 14\.5\]\-31\.5 \[\-33\.3, \-29\.6\]\+4\.5 \[3\.7, 5\.3\]\+27\.4 \[25\.5, 29\.2\]Def\.\+CoT−\-Def\.\+10\.2 \[9\.2, 11\.1\]\-19\.8 \[\-21\.6, \-17\.9\]\-0\.5 \[\-1\.3, 0\.3\]\+20\.4 \[18\.5, 22\.2\]Def\.\+Abstain−\-Def\.\+6\.1 \[5\.3, 6\.8\]\-13\.6 \[\-15\.0, \-12\.2\]\+1\.4 \[0\.9, 2\.0\]\+11\.8 \[10\.4, 13\.3\]Def\.\+Self\-check−\-Def\.\-1\.9 \[\-2\.8, \-1\.1\]\+7\.3 \[5\.6, 9\.0\]\-3\.4 \[\-4\.2, \-2\.7\]\-3\.9 \[\-5\.6, \-2\.2\]Cert\.−\-Def\.\-3\.8 \[\-4\.6, \-3\.0\]\+11\.0 \[9\.4, 12\.7\]\-3\.5 \[\-4\.3, \-2\.7\]\-7\.9 \[\-9\.5, \-6\.3\]Cert\.\+CoT−\-Cert\.\+7\.8 \[7\.0, 8\.6\]\-15\.3 \[\-16\.9, \-13\.8\]\-0\.3 \[\-0\.9, 0\.3\]\+15\.6 \[14\.0, 17\.3\]Table 17:Pair\-level bootstrap 95% confidence intervals for the prompting contrasts \(percentage points; 10,000 replicates\)\. NegativeΔ\\DeltaOCR andΔ\\DeltaUCR indicate error reduction\. Intervals excluding zero support a directional difference under pair\-level resampling\.
##### Partial versus Scope\-Mismatch\.

Table[18](https://arxiv.org/html/2608.04591#A4.T18)quantifies uncertainty for the RQ2 branch comparison reported in Table[4](https://arxiv.org/html/2608.04591#Sx6.T4)and extended in Table[12](https://arxiv.org/html/2608.04591#A4.T12)\. Across every model and condition, the OCR and Both\-CN differences are positive and their intervals exclude zero, showing consistently greater over\-closure on partial evidence than on Scope\-Mismatch evidence\. The CS difference is also positive and excludes zero in most conditions, but is inconclusive for Qwen Self\-check and for Haiku under Definition\-aware, Abstain, and Self\-check\. Thus, although the pair\-discrimination gap weakens under a few interventions, partial coverage remains the more persistent source of over\-closure\.

ModelConditionOCR\(part\)−\-OCR\(SM\)CS\(SM\)−\-CS\(part\)Both\-CN\(part\)−\-Both\-CN\(SM\)QwenNaive\+42\.9 \[39\.5, 46\.2\]\+29\.6 \[25\.9, 33\.4\]\+42\.0 \[38\.6, 45\.4\]Def\.\+33\.8 \[30\.7, 37\.0\]\+31\.0 \[27\.4, 34\.6\]\+34\.2 \[31\.1, 37\.4\]Def\.\+CoT\+13\.6 \[10\.2, 17\.0\]\+19\.0 \[15\.5, 22\.6\]\+13\.8 \[10\.4, 17\.4\]Def\.\+Abstain\+18\.8 \[16\.3, 21\.4\]\+5\.8 \[1\.9, 9\.8\]\+18\.6 \[16\.2, 21\.1\]Def\.\+Self\-check\+12\.6 \[10\.4, 14\.7\]\+1\.7 \[\-2\.2, 5\.6\]\+12\.1 \[10\.0, 14\.2\]Cert\.\+20\.7 \[17\.7, 23\.8\]\+13\.6 \[10\.0, 17\.3\]\+20\.4 \[17\.4, 23\.5\]Cert\.\+CoT\+17\.9 \[14\.6, 21\.3\]\+13\.4 \[9\.9, 16\.9\]\+18\.6 \[15\.3, 21\.9\]HaikuNaive\+54\.2 \[51\.1, 57\.5\]\+43\.0 \[39\.4, 46\.5\]\+54\.6 \[51\.5, 57\.9\]Def\.\+10\.8 \[8\.2, 13\.4\]\+1\.7 \[\-2\.1, 5\.4\]\+10\.7 \[8\.1, 13\.4\]Def\.\+CoT\+16\.2 \[13\.7, 18\.6\]\+14\.5 \[11\.4, 17\.6\]\+16\.2 \[13\.7, 18\.6\]Def\.\+Abstain\+7\.1 \[4\.6, 9\.7\]\-3\.0 \[\-6\.8, 0\.7\]\+7\.3 \[4\.8, 9\.8\]Def\.\+Self\-check\+6\.6 \[3\.8, 9\.6\]\-1\.6 \[\-5\.2, 2\.0\]\+6\.8 \[3\.9, 9\.8\]Cert\.\+17\.4 \[14\.9, 20\.0\]\+10\.8 \[7\.5, 14\.2\]\+17\.4 \[14\.8, 19\.9\]Cert\.\+CoT\+24\.7 \[21\.9, 27\.6\]\+23\.4 \[20\.2, 26\.6\]\+24\.9 \[22\.1, 27\.8\]GemmaNaive\+43\.8 \[40\.9, 46\.6\]\+42\.2 \[39\.4, 45\.0\]\+44\.0 \[41\.2, 46\.9\]Def\.\+35\.7 \[32\.0, 39\.4\]\+34\.3 \[30\.6, 38\.0\]\+35\.5 \[31\.9, 39\.2\]Def\.\+CoT\+15\.8 \[12\.5, 19\.1\]\+8\.3 \[4\.8, 11\.8\]\+15\.0 \[11\.8, 18\.3\]Def\.\+Abstain\+16\.9 \[13\.4, 20\.5\]\+14\.1 \[10\.3, 17\.8\]\+17\.2 \[13\.7, 20\.8\]Def\.\+Self\-check\+27\.4 \[23\.5, 31\.1\]\+24\.2 \[20\.4, 28\.1\]\+27\.8 \[23\.9, 31\.5\]Cert\.\+54\.4 \[51\.1, 57\.7\]\+51\.2 \[47\.8, 54\.6\]\+54\.6 \[51\.3, 57\.9\]Cert\.\+CoT\+43\.6 \[40\.2, 47\.0\]\+43\.8 \[40\.3, 47\.2\]\+43\.6 \[40\.2, 47\.0\]Table 18:Bootstrap 95% confidence intervals for partial versus Scope\-Mismatch branch differences \(percentage points\)\. Positive OCR and Both\-CN differences indicate more over\-closure on partial evidence; a positive CS difference indicates lower pair discrimination on the partial branch\.
##### Coverage\-regime contrasts\.

Table[19](https://arxiv.org/html/2608.04591#A4.T19)quantifies uncertainty for the regime\-level and L3 member\-level contrasts underlying the RQ2 results in Table[5](https://arxiv.org/html/2608.04591#Sx6.T5)\. L3 has lower CS than L1 in every model–condition cell and lower CS than L2 in every cell except the exact L2–L3 tie for Gemma under Naive prompting; all nonzero L1 and L2 contrasts exclude zero\. The L3 CS point estimate is also lower than L4 in every condition, although the Qwen Naive interval includes zero\. The L3 complete\-CN minus partial\-UNK gap is positive and excludes zero in every condition except Qwen Self\-check\. Thus, the bootstrap analysis supports the RQ2 conclusion that L3 is the most persistent regime\-level bottleneck and that its failure is usually one\-sided, concentrated on implicit\-partial members\.

ModelConditionCS\(L1\)−\-CS\(L3\)CS\(L2\)−\-CS\(L3\)CS\(L4\)−\-CS\(L3\)L3 cCN−\-pUNKQwenNaive\+32\.9 \[26\.2, 39\.6\]\+24\.2 \[17\.8, 30\.6\]\+4\.8 \[\-0\.6, 10\.3\]\+61\.9 \[54\.2, 69\.6\]Def\.\+49\.2 \[42\.8, 55\.6\]\+60\.4 \[54\.3, 66\.4\]\+30\.1 \[23\.7, 36\.5\]\+69\.9 \[62\.8, 76\.6\]Def\.\+CoT\+86\.2 \[82\.1, 90\.1\]\+88\.5 \[84\.6, 92\.0\]\+24\.7 \[18\.6, 30\.8\]\+75\.6 \[69\.2, 81\.7\]Def\.\+Abstain\+48\.9 \[42\.5, 55\.2\]\+67\.1 \[61\.3, 72\.8\]\+60\.3 \[53\.8, 66\.3\]\+35\.9 \[26\.3, 45\.2\]Def\.\+Self\-check\+38\.0 \[31\.3, 44\.7\]\+45\.3 \[38\.6, 52\.0\]\+53\.2 \[46\.5, 59\.6\]\-3\.8 \[\-13\.8, 6\.1\]Cert\.\+77\.3 \[72\.2, 82\.1\]\+71\.5 \[66\.1, 77\.0\]\+57\.4 \[51\.3, 63\.1\]\+68\.3 \[61\.2, 75\.3\]Cert\.\+CoT\+85\.3 \[81\.1, 89\.1\]\+88\.5 \[84\.9, 91\.7\]\+40\.7 \[34\.3, 47\.1\]\+83\.0 \[78\.2, 87\.5\]HaikuNaive\+28\.7 \[22\.3, 34\.8\]\+21\.1 \[15\.0, 27\.1\]\+17\.3 \[11\.5, 23\.1\]\+83\.0 \[77\.9, 87\.8\]Def\.\+87\.2 \[83\.3, 91\.0\]\+88\.8 \[85\.3, 92\.3\]\+46\.5 \[40\.1, 52\.9\]\+54\.5 \[45\.8, 62\.8\]Def\.\+CoT\+84\.9 \[80\.8, 88\.8\]\+84\.3 \[80\.1, 88\.1\]\+60\.6 \[54\.5, 66\.3\]\+70\.8 \[64\.1, 77\.2\]Def\.\+Abstain\+82\.1 \[77\.6, 86\.5\]\+82\.7 \[78\.5, 86\.9\]\+40\.7 \[34\.0, 47\.4\]\+36\.9 \[27\.6, 45\.8\]Def\.\+Self\-check\+83\.3 \[78\.9, 87\.2\]\+83\.7 \[79\.5, 87\.5\]\+54\.8 \[48\.1, 60\.9\]\+70\.8 \[64\.4, 77\.2\]Cert\.\+76\.9 \[71\.8, 81\.7\]\+82\.1 \[77\.9, 86\.2\]\+51\.9 \[45\.2, 58\.3\]\+81\.1 \[76\.6, 85\.6\]Cert\.\+CoT\+82\.7 \[78\.5, 86\.9\]\+83\.3 \[79\.2, 87\.5\]\+26\.6 \[19\.9, 33\.3\]\+76\.9 \[71\.5, 82\.4\]GemmaNaive\+3\.8 \[1\.9, 6\.1\]\+0\.0 \[0\.0, 0\.0\]\+4\.5 \[2\.2, 6\.7\]\+100\.0 \[100\.0, 100\.0\]Def\.\+49\.8 \[44\.4, 55\.3\]\+42\.8 \[37\.4, 48\.2\]\+40\.7 \[35\.3, 46\.2\]\+100\.0 \[100\.0, 100\.0\]Def\.\+CoT\+93\.0 \[90\.1, 95\.5\]\+93\.0 \[90\.1, 95\.8\]\+55\.1 \[49\.0, 60\.9\]\+92\.0 \[88\.8, 94\.9\]Def\.\+Abstain\+86\.6 \[82\.7, 90\.4\]\+74\.1 \[69\.3, 78\.9\]\+60\.3 \[54\.8, 65\.4\]\+100\.0 \[100\.0, 100\.0\]Def\.\+Self\-check\+56\.2 \[50\.8, 61\.7\]\+42\.2 \[36\.7, 47\.6\]\+39\.4 \[34\.0, 44\.9\]\+100\.0 \[100\.0, 100\.0\]Cert\.\+22\.4 \[17\.9, 27\.2\]\+16\.9 \[12\.8, 21\.1\]\+28\.8 \[23\.7, 34\.0\]\+100\.0 \[100\.0, 100\.0\]Cert\.\+CoT\+42\.5 \[37\.1, 47\.9\]\+57\.5 \[51\.8, 62\.9\]\+42\.9 \[37\.5, 48\.4\]\+99\.0 \[97\.8, 100\.0\]Table 19:Bootstrap 95% confidence intervals for coverage\-regime contrasts on the 1,250 complete–partial pairs \(percentage points\)\. The first three columns reportCS​\(L​k\)−CS​\(L​3\)\\mathrm\{CS\}\(Lk\)\-\\mathrm\{CS\}\(L3\); positive values indicate lower CS for L3 than for the compared regime\. The final column reports the L3 complete\-CN minus partial\-UNK correct\-rate gap\.
##### CROWN\-Real transfer contrast\.

For Proceedings, we resample the 266 contrast\-set identifiers with replacement and retain the A, B, and C variants of each selected set\. Table[20](https://arxiv.org/html/2608.04591#A4.T20)reports bootstrap intervals forΔB−C=B​\-​UNK−C​\-​UNK\\Delta\_\{B\-C\}=\\mathrm\{B\\mbox\{\-\}UNK\}\-\\mathrm\{C\\mbox\{\-\}UNK\}\. A positive difference indicates that partial C is more error\-prone than narrower\-complete B\.

ConditionQwenHaikuGemmaNaive\+19\.9 \[13\.9, 25\.9\]\+42\.9 \[36\.1, 49\.2\]\+0\.4 \[\-1\.1, 1\.9\]Def\.\+0\.0 \[0\.0, 0\.0\]\+29\.7 \[23\.7, 35\.7\]\+0\.0 \[0\.0, 0\.0\]Def\.\+CoT\+33\.1 \[27\.4, 38\.7\]\+70\.3 \[64\.7, 75\.9\]\+90\.2 \[86\.5, 93\.6\]Cert\.\+8\.3 \[5\.3, 11\.7\]\+88\.0 \[83\.8, 91\.7\]\+97\.0 \[94\.7, 98\.9\]Cert\.\+CoT\+24\.8 \[19\.9, 30\.1\]\+52\.6 \[46\.6, 58\.6\]\+90\.6 \[86\.8, 94\.0\]Table 20:Contrast\-set bootstrap 95% confidence intervals forΔB−C=B​\-​UNK−C​\-​UNK\\Delta\_\{B\-C\}=\\mathrm\{B\\mbox\{\-\}UNK\}\-\\mathrm\{C\\mbox\{\-\}UNK\}on the 266 Proceedings A/B/C sets \(percentage points; 10,000 replicates\)\. Positive values indicate that partial C is more error\-prone than narrower\-complete B\. The A/B/C variants of each selected set are retained together\. Confidence intervals including zero do not indicate a clear directional difference\.Across the 15 model–condition cells, all point estimates ofΔB−C\\Delta\_\{B\-C\}are nonnegative, and the confidence intervals exclude zero in 12 cells\. Of the remaining three, two are exact ties under Definition\-aware prompting: B\-UNK and C\-UNK are both 100\.0% for Qwen and both 0\.0% for Gemma\. The third is a near\-tie for Gemma under Naive prompting\. Thus, partial C is more error\-prone with a confidence interval excluding zero in 12 cells, while the remaining three show no clear directional difference\. No observed point estimate is negative\.

Similar Articles

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning

arXiv cs.CL

This paper characterizes 'futile reasoning' in large language models, where models produce superficially valid but incorrect reasoning on tasks beyond their capability. They introduce CaRL, a capability-aligned reinforcement learning method that trains LLMs to abstain from futile reasoning while preserving performance.