Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
Summary
This research paper proposes Provenance Density, an evidence-visualization interface to help users distinguish AI-generated text by showing verified claims, mitigating the Fluency Trap where users trust fluent but fabricated content.
View Cached Full Text
Cached at: 09/04/26, 06:05 AM
# Beyond “Made with AI”: Visualizing Provenance Density to Mitigate the Transparency Penalty Source: [https://arxiv.org/html/2609.03460](https://arxiv.org/html/2609.03460) Yifei HuangAffiliation:Institute of Industrial Science, The University of TokyoEmail:[hyf015@gmail\.com](mailto:[email protected])Juyoung LeeAffiliation:Korea Advanced Institute of Science and TechnologyEmail:[ejuyoung@gmail\.com](mailto:[email protected])Thad StarnerAffiliation:Georgia Institute of TechnologyEmail:[thad\.starner@gmail\.com](mailto:[email protected])Jun RekimotoAffiliation:Sony CSL KyotoEmail:[rekimoto@acm\.org](mailto:[email protected]) ###### Abstract As generative AI makes polished prose cheap to produce, users can no longer rely on fluency as a proxy for truth\. We call this failure mode the Fluency Trap: users trust fluent hallucinations while also discounting accurate content once it is disclosed as AI\-generated\. Binary “Made with AI” labels respond with authorship disclosure, but they do not show what supports a claim\. We propose Provenance Density, an evidence\-visualization interface that shows the density of verified claims in a text\. In a user study with 81 participants, an idealized Provenance Density interface produced a large discernment gap between truth and fabrication \(\+4\.15\+4\.15points,d=1\.82d=1\.82\), whereas participants given no signal showed no detectable discrimination\. A technical audit with 200 samples shows that retrieval density alone is insufficient; unexpectedly, the Consistency Veto carries most of the discriminative signal on dynamic queries\. As AI\-generated content becomes indistinguishable from human writing, effective transparency must move from authorship disclosure toward evidence visualization\. ††footnotetext:The user study was approved by the University of Tokyo institutional review process\. Code, data, prompts, and audit materials are available at[https://github\.com/artisticsciencex/ijcai\-ecai\-2026\-provenance\-density](https://github.com/artisticsciencex/ijcai-ecai-2026-provenance-density)\.## 1Introduction Readers often use processing fluency—the subjective ease with which information is processed—as a cue when judging truth[Reber and Schwarz \(1999\)](https://arxiv.org/html/2609.03460#bib.bib26);[Dechêne et al\. \(2010\)](https://arxiv.org/html/2609.03460#bib.bib27);[Reber and Unkelbach \(2010\)](https://arxiv.org/html/2609.03460#bib.bib28)\. Related work on retrieval fluency shows that ease\-based heuristics can be ecologically useful when fluency covaries with properties of the environment[Hertwig et al\. \(2008\)](https://arxiv.org/html/2609.03460#bib.bib25);[Marewski and Schooler \(2011\)](https://arxiv.org/html/2609.03460#bib.bib31)\. We extend this logic to linguistic presentation\. Before generative AI, producing high\-quality, articulate prose typically required substantial education and editorial labor, allowing polish to function as an imperfect signal of competence\.Costly Signaling Theoryexplains this relationship: a signal is trustworthy only when faking it is prohibitively expensive[Spence \(1978\)](https://arxiv.org/html/2609.03460#bib.bib10);[Zahavi \(1975\)](https://arxiv.org/html/2609.03460#bib.bib11);[Gintis et al\. \(2001\)](https://arxiv.org/html/2609.03460#bib.bib30)\. Generative AI disrupts this mechanism by reducing the marginal cost of fluency to near\-zero[Galdin and Silbert \(2025\)](https://arxiv.org/html/2609.03460#bib.bib4)\. Large Language Models \(LLMs\) enable the mass production of professional\-sounding text regardless of the author’s underlying expertise\. This collapse of the “separating equilibrium” creates what we term aFluency Trap: a structural vulnerability where users continue to trust fluent text as if it were costly, even when it is generated cheaply by systems indifferent to truth\. Psychological evidence on theIllusion of Truthsuggests that ease of processing suppresses epistemic vigilance[Hasher et al\. \(1977\)](https://arxiv.org/html/2609.03460#bib.bib6);[Reber and Schwarz \(1999\)](https://arxiv.org/html/2609.03460#bib.bib26);[Dechêne et al\. \(2010\)](https://arxiv.org/html/2609.03460#bib.bib27);[Sperber et al\. \(2010\)](https://arxiv.org/html/2609.03460#bib.bib38)\. This tendency leaves humans susceptible to “hallucinated plausibility”, text that is syntactically perfect but semantically ungrounded[Reber and Unkelbach \(2010\)](https://arxiv.org/html/2609.03460#bib.bib28);[Unkelbach et al\. \(2011\)](https://arxiv.org/html/2609.03460#bib.bib29)\. Current governance responses, specifically binary “Made with AI” disclosures, fail to address this decoupling\. By focusing onidentity\(“Who wrote this?”\) rather thanprovenance\(“What supports this?”\), such labels can shift reader perceptions without supplying evidence about individual claims[Nakano et al\. \(2026\)](https://arxiv.org/html/2609.03460#bib.bib12)\. In experiments with news headlines, AI labels reduced perceived accuracy even when the headlines were true or human\-written[Altay and Gilardi \(2024\)](https://arxiv.org/html/2609.03460#bib.bib33)\. We characterize this accuracy\-independent discounting as a “Transparency Penalty\.” We frame Provenance Density as acognitive affordancefor reading in the LLM era\. Instead of asking users to evaluate veracity from prose alone, the interface shifts attention toward extrinsic evidence: high\-contrast indicators visualize the density of verified claims, offloading part of the verification burden from working memory to the interface[Chirayath et al\. \(2025\)](https://arxiv.org/html/2609.03460#bib.bib35);[Clark and Chalmers \(1998\)](https://arxiv.org/html/2609.03460#bib.bib39)\. We validate this approach through a dual\-stream evaluation\. First, we conduct an automated technical audit \(N=200N=200\) on a composite ofTruthfulQA[Lin et al\. \(2022\)](https://arxiv.org/html/2609.03460#bib.bib37)andFreshQA[Vu et al\. \(2024\)](https://arxiv.org/html/2609.03460#bib.bib24)to test robustness against both adversarial misconceptions and dynamic ambiguity\. Second, we run a within\-subjects user experiment with 81 participants to measure truth discernment\. Our results empirically confirm the Fluency Trap: in the absence of signals, users failed to distinguish high\-fluency hallucinations \(M=6\.28M=6\.28\) from ground truth \(M=5\.78M=5\.78;p=\.43p=\.43\)\. While binary labels acted as a blunt warning, Provenance Density restored truth discernment under correct signaling \(d=1\.82d=1\.82,p<\.001p<\.001\)\. We make three main contributions:Theory:We synthesize the mechanics of the Fluency Trap \(Section[2](https://arxiv.org/html/2609.03460#S2)\), detailing how RLHF\-driven sycophancy structures the decoupling of fluency from veracity\.Design:We proposeProvenance Density\(Section[3](https://arxiv.org/html/2609.03460#S3)\), a formalized metric \(D\(T\)D\(T\)\) and interaction paradigm that imposes a computational verification handicap on generated text\.Evidence:We evaluate the approach through a technical audit \(N=200N=200\) and a within\-subjects user study \(N=81N=81\), jointly examining metric behavior and interface\-supported truth discernment \(Sections[4](https://arxiv.org/html/2609.03460#S4)&[5\.1](https://arxiv.org/html/2609.03460#S5.SS1)\)\. ## 2Related Works We argue that the decoupling of fluency from veracity is not an accidental byproduct of LLM scaling, but a structural inevitability driven by two converging factors: an economic shift from costly signaling to cheap talk, and a technical objective function that prioritizes plausibility over truth\. #### From Hallucination to Indifference\. While early critiques of Large Language Models \(LLMs\) focused on “hallucinations” as sporadic errors, recent scholarship suggests a more structural diagnosis\. The framework of “Machine Bullshit” has been proposed to distinguish these outputs from lying, defined instead by the model’s fundamental indifference to truth value[Liang et al\. \(2025\)](https://arxiv.org/html/2609.03460#bib.bib3)\. Analysis of the “Bullshit Index” reveals that Reinforcement Learning from Human Feedback \(RLHF\) exacerbates this issue, incentivizing models to prioritize rhetorical plausibility and “paltering” \(misleading use of truth\) over factual grounding\. Consequently, the resulting text is optimized to bypass human epistemic vigilance\. This structural indifference renders traditional governance mechanisms, such as binary warning labels, largely ineffective\. Empirical evaluations demonstrate a “failure of inoculation”: while pre\-emptive warnings about AI fallibility successfully reduce global trust in the system, they fail to mitigate reliance on specific misleading articles once the user is engaged with the content[Spearing et al\. \(2025\)](https://arxiv.org/html/2609.03460#bib.bib9)\. This discrepancy, where users theoretically acknowledge AI bias but practically accept AI fluency, underscores the limitations of heuristic warnings and motivates our proposal for granular Provenance Density indicators\. #### The Structural Decoupling of Fluency\. Our analysis is grounded in theHandicap Principle, which posits that reliable signals must impose a cost on the signaler[Zahavi \(1975\)](https://arxiv.org/html/2609.03460#bib.bib11)\. Historically, linguistic polish functioned as this handicap, creating a Separating Equilibrium where articulate text correlated with competence[Spence \(1978\)](https://arxiv.org/html/2609.03460#bib.bib10)\. Generative AI collapses this balance into a Pooling Equilibrium, where expert testimony and fabrication can share the same polished form[Spence \(1978\)](https://arxiv.org/html/2609.03460#bib.bib10);[Galdin and Silbert \(2025\)](https://arxiv.org/html/2609.03460#bib.bib4)\. This shift is sustained by the technical alignment of modern LLMs\. Reinforcement Learning from Human Feedback \(RLHF\) can incentivize sycophancy, where models prioritize user agreement over factual accuracy[Sharma et al\. \(2024\)](https://arxiv.org/html/2609.03460#bib.bib40)\. Evidence shows that models will agree with illogical premises to remain “helpful”[Chen et al\. \(2025\)](https://arxiv.org/html/2609.03460#bib.bib5)and flip arguments to match user views[Kaur \(2025\)](https://arxiv.org/html/2609.03460#bib.bib18)\. This results in “Machine Bullshit”—text optimized for rhetorical persuasion rather than truth[Liang et al\. \(2025\)](https://arxiv.org/html/2609.03460#bib.bib3)\. This decoupling can become self\-amplifying: recursive training on such high\-fluency, low\-entropy data leads toModel Collapse, where the “tails” of human variance are lost to a homogenized mean[Shumailov et al\. \(2024\)](https://arxiv.org/html/2609.03460#bib.bib16)\. #### The Failure of Post\-Hoc Governance\. Current governance relies on the assumption that users, once warned, can critically evaluate AI\-generated text\. However, AI\-generated self\-presentations can be difficult to distinguish from human\-written ones[Jakesch et al\. \(2023\)](https://arxiv.org/html/2609.03460#bib.bib7), and AI literacy does not reliably translate into consistent fact\-checking behavior[Rheu and Cho \(2025\)](https://arxiv.org/html/2609.03460#bib.bib17)\. Binary disclosures \(e\.g\., “Made with AI”\) can exacerbate the issue through Epistemic Stigmatization\. Studies report a “transparency penalty” where disclosure reduces perceived trustworthiness even when content quality is constant[Nakano et al\. \(2026\)](https://arxiv.org/html/2609.03460#bib.bib12);[Cheong et al\. \(2025\)](https://arxiv.org/html/2609.03460#bib.bib13)\. Related advertising experiments found that identical ads received more critical evaluations when labeled AI\-generated[Buder et al\. \(2024\)](https://arxiv.org/html/2609.03460#bib.bib20)\. Such effects can produce broad skepticism without enabling claim\-level verification\. #### Limitations of Automated Detection\. While automated detectors are proposed as an enforcement mechanism, they exhibit substantial bias against linguistically marginalized groups\. Detectors relying on perplexity heuristics frequently misclassify the lower\-perplexity lexical patterns of non\-native speakers as AI\-generated[Liang et al\. \(2023\)](https://arxiv.org/html/2609.03460#bib.bib14), leading to “hermeneutical access injustice”[Kay et al\. \(2024\)](https://arxiv.org/html/2609.03460#bib.bib15)\. The adversarial nature of detection also creates an arms race that generators inevitably win\. We therefore shift the paradigm frompost\-hoc detectionof authorship tointrinsic signalingof Provenance Density\. ## 3Defining Provenance Density To function as a costly signal in the game\-theoretic sense, Provenance Density must be computationally hard to fake\. We define the Provenance Density scoreDDfor a given textTTas a theoretical function of verifiable claims weighted by source reputation, contextual relevance, and internal semantic consistency\. ### 3\.1Proposed Metric:D\(T\)D\(T\) We propose a density metric that integrates external verification \(Retrieval Augmented Generation\)[Lewis et al\. \(2020\)](https://arxiv.org/html/2609.03460#bib.bib36)with internal uncertainty quantification\. Drawing on the semantic\-consistency motivation of Farquhar et al\.[Farquhar et al\. \(2024\)](https://arxiv.org/html/2609.03460#bib.bib19), we introduce a penalty termPintP\_\{int\}based on an NLI cross\-sample consistency heuristic that down\-weights generations whose stochastic samples disagree\. Unlike semantic entropy, our heuristic does not estimate entropy over probability\-weighted semantic clusters\. LetCCbe the set of distinct factual claims inTT, andCv⊆CC\_\{v\}\\subseteq Cbe the subset of claims successfully grounded in external evidence\. LetScS\_\{c\}be the set of independent root domains supporting a verified claimcc\. We define the Provenance Density scoreD\(T\)∈\[0,1\)D\(T\)\\in\[0,1\)as: D\(T\)=\(1−Pint\)⋅tanh\(1β∑c∈C\(∑s∈Scw\(s,c\)\)λ\)D\(T\)=\(1\-P\_\{int\}\)\\cdot\\tanh\\left\(\\frac\{1\}\{\\beta\}\\sum\_\{c\\in C\}\\left\(\\sum\_\{s\\in S\_\{c\}\}w\(s,c\)\\right\)^\{\\lambda\}\\right\)\(1\) Equivalently, we first raise each claim’s source sum to the powerλ\\lambda, then sum across claims\. Our design choices target specific theoretical properties revealed during technical validation: #### NLI Consistency Heuristic \(“Consistency Veto,”1−Pint1\-P\_\{int\}\): For each query, we drawK=5K=5stochastic samples and computeρconsistent\\rho\_\{\\mathrm\{consistent\}\}, the fraction of unordered sample pairs for which the NLI contradiction probability remains below0\.50\.5in both directions\. We definePint=1−ρconsistentP\_\{int\}=1\-\\rho\_\{\\mathrm\{consistent\}\}, so internally stable generations havePint≈0P\_\{int\}\\approx 0and mutually inconsistent generations approachPint→1P\_\{int\}\\to 1\. As pairwise inconsistency rises, the gate reduces the score regardless of external evidence\. This heuristic is inspired by semantic entropy but neither forms semantic\-equivalence clusters nor weights them by generation probabilities\. It can suppress inconsistent confabulations, but it cannot detect false answers that recur consistently across samples\. #### Cubic Contextual Weight \(w\(s,c\)w\(s,c\)\): Unlike traditional citation metrics that rely solely on domain authority, we definew\(s,c\)w\(s,c\)as the product of source reputation and semantic relevance: w\(s,c\)=Reputation\(s\)⋅\(MatchRatio\)3w\(s,c\)=\\text\{Reputation\}\(s\)\\cdot\(\\text\{MatchRatio\}\)^\{3\}We impose a cubic exponent to aggressively suppress weak conceptual matches\. Our preliminary tests indicated that quadratic weighting \(x2x^\{2\}\) remained too permissible of “keyword stuffing,” where irrelevant sources share surface\-level vocabulary \(MatchRatio≈0\.5\\approx 0\.5\)\. The cubic term \(0\.53=0\.1250\.5^\{3\}=0\.125\) acts as a noise filter, ensuring that only sources with high semantic alignment contribute meaningfully to the Provenance Density score\. #### Saturation Scalar \(β\\beta\): We introduceβ\\betaas a density temperature parameter \(set toβ=5\.0\\beta=5\.0in our reference implementation\)\. This term scales the input to the hyperbolic tangent function, preventing the score from saturating to1\.01\.0based on a single citation\. It forces the model to provide accumulated, multi\-source evidence to achieve a perfect density signal\. #### Concentration Exponent \(λ\\lambda\): We setλ=1\.2\>1\\lambda=1\.2\>1, applied to the per\-claim source sum before aggregation\. Withλ\>1\\lambda\>1, the term is superlinear in per\-claim evidence: a claim supported by several aligned sources contributes more than the same total evidence spread thinly across several claims\. This concentrates score on well\-corroborated claims rather than rewarding shallow citation across many claims\. We note that citation stuffing within a single claim is suppressed upstream by the cubic relevance penalty inw\(s,c\)w\(s,c\), not byλ\\lambda\. This formulation imposes a significant “handicap” on the generator: achieving a highD\(T\)D\(T\)requires the model to perform costly verification and achieve strict semantic alignment between claims and sources\. ### 3\.2Implementation ofD\(T\)D\(T\) Equation 1 is instantiated by four concrete steps: claim segmentation, evidence retrieval, source\-relevance scoring, and aggregation; key implementation parameters are summarized in Table[1](https://arxiv.org/html/2609.03460#S3.T1)\. We segmentTTinto atomic factual claims with a deterministicgpt\-4o\-minipass atτ=0\\tau=0\. Claims shorter than five tokens are discarded as scaffolding\. For each remaining claimcc, we issue one search query: the claim text itself when it contains at least eight tokens, otherwise the original question concatenated withcc\. Retrieved evidence is scored by a coarse but explicit relevance prior\. For each claimcc, letKcK\_\{c\}be the set of rare keywords extracted by matching capitalized tokens \(regular expression\\b\[A\-Z\]\[a\-zA\-Z0\-9\-\]\+\\b\) and removing a sentence\-leading stoplist\. For each resultsswith URLu\(s\)u\(s\)and snippetσ\(s\)\\sigma\(s\), we compute MatchRatio\(σ,Kc\)=\|\{k∈Kc:k↓∈σ↓\}\|max\(\|Kc\|,1\),\\mathrm\{MatchRatio\}\(\\sigma,K\_\{c\}\)=\\frac\{\|\\\{k\\in K\_\{c\}:k\_\{\\downarrow\}\\in\\sigma\_\{\\downarrow\}\\\}\|\}\{\\max\(\|K\_\{c\}\|,1\)\},where↓\\downarrowdenotes case\-folded substring matching\. The source contribution is then w\(s,c\)=Reputation\(u\(s\)\)⋅MatchRatio\(σ,Kc\)3\.w\(s,c\)=\\mathrm\{Reputation\}\(u\(s\)\)\\cdot\\mathrm\{MatchRatio\}\(\\sigma,K\_\{c\}\)^\{3\}\.Reputationis a three\-level domain prior:1\.01\.0for high\-trust domains \(e\.g\.,\.gov,\.edu,wikipedia\.org,nih\.gov,reuters\.com,apnews\.com,nature\.com\),0\.10\.1for low\-trust domains \(e\.g\.,reddit\.com,quora\.com,medium\.com,twitter\.com\), and0\.50\.5otherwise\. The cubic exponent is intentionally aggressive: it suppresses weak keyword overlap that would otherwise permit “keyword stuffing” from superficially related pages\. ParameterValueRoleβ\\beta5\.0Saturation scalar in Eq\. 1λ\\lambda1\.2Concentration exponent on per\-claim sumSearch depth3Top\-kkorganic results per claimConsistency samplesKK5Stochastic samples forPintP\_\{int\}Generator temperature1\.0 / 0\.0Audit sampling / segmentationNLI threshold0\.5Contradiction cutoff for consistencyHigh\-trust prior1\.0Gov/edu/Wikipedia/major news/scienceLow\-trust prior0\.1Reddit/Quora/Medium/Twitter/XDefault prior0\.5All other domainsTable 1:Key implementation choices for the technical audit\.For reproducibility, the audit usesgpt\-4o\-minifor generation and deterministic segmentation, a DeBERTa\-based NLI checker to estimatePintP\_\{int\}from pairwise agreement acrossK=5K=5samples, and Serper as the retrieval backend for evidence collection\. The released code and data include the exact prompts, model calls, judge labels, and scripts used in the audit\. ### 3\.3Operationalization: The Oracle Protocol Equation 1 defines the target signal for a deployed system, but a user study introduces a separate methodological question: does the*visual signal itself*improve truth discernment when its underlying value is correct? To separate that interaction question from retriever noise, the empirical study in Section[5](https://arxiv.org/html/2609.03460#S5)uses a Wizard\-of\-Oz \(Oracle\) protocol\. In this protocol, participants do not see the live output of the auditing pipeline\. Instead, the interface displays idealized endpoint values of the signal: grounded summaries are paired with high\-density indicators and fabricated summaries with null indicators\. This lets the user study estimate the*interaction effect*of PDI under known\-correct signaling, while the technical audit in Section[4](https://arxiv.org/html/2609.03460#S4)independently evaluates whether the real pipeline can approximate that ideal in practice\. ### 3\.4Design Rationale: Countering Pseudo\-Profound Fluency The visualization ofD\(T\)D\(T\)is designed to counter “pseudo\-profound bullshit”—syntactically persuasive but semantically vacuous text[Pennycook et al\. \(2015\)](https://arxiv.org/html/2609.03460#bib.bib8)\. Rather than asking users to infer truth from fluency, PDI presents evidence density as a high\-contrast cue \(Figure[2](https://arxiv.org/html/2609.03460#S5.F2)\) that can interrupt fastSystem 1judgments and invite more deliberateSystem 2evaluation[Kahneman \(2011\)](https://arxiv.org/html/2609.03460#bib.bib2)\. ## 4Technical Validation Before evaluating the interaction paradigm with users, we audited the proposed metricD\(T\)D\(T\)to characterize how its retrieval and consistency components behave across static and dynamic queries\. #### Experimental Setup\. To evaluate the metric across both established knowledge and emerging information, we constructed a composite dataset \(N=200N=200\): Static Knowledge \(N=150N=150\), randomly sampled from theTruthfulQAgeneration task, and Dynamic Knowledge \(N=50N=50\), a curatedFreshQAsubset focused on post\-2024 events \(e\.g\., elections, stock prices\) designed to probe the “Cold Start” regime where model training weights are outdated\. Per\-row hallucination labels were produced bygpt\-4o\(temperature00\) using the dataset’s official reference answers when available\. For the static subset, the judge was given TruthfulQA’s official correct / incorrect answer lists; for the dynamic subset, judgement was closed\-book\. Audit latency measured during this pipeline ranged from 17\.0s \(Dynamic\) to 26\.4s \(Static\), with exact timing logs included in the released audit materials\. Static items were slower because TruthfulQA’s adversarial prompts elicited longer generations fromgpt\-4o\-mini\(mean 98 vs\. 54 words\), which produced more atomic claims and therefore more retrieval and scoring calls\. ### 4\.1Results: Ecological Sensitivity and Hallucination Detection Figure 1:Ecological Sensitivity of the Metric\. Distribution of Provenance Density scoresD\(T\)D\(T\)for Static Knowledge \(TruthfulQA, green\) and Dynamic/Fresh Knowledge \(FreshQA, red\)\. Static items cluster above the high\-trust threshold \(M=0\.79M=0\.79\), whereas FreshQA items have a lower mean \(M=0\.64M=0\.64\) and broader spread, allowing the metric to signal when information is newer or less settled\. The dashed high\-trust threshold denotesD\(T\)≥0\.7D\(T\)\\geq 0\.7\.Ecological Sensitivity \(Static vs\. Dynamic\)\.As shown in Figure[1](https://arxiv.org/html/2609.03460#S4.F1), the metric reproduces the established / emerging knowledge separation reported in prior work\. Static TruthfulQA items cluster above the high\-trust threshold \(M=0\.79M=0\.79\), while dynamic FreshQA items yield a lower mean \(M=0\.64M=0\.64\) and broader spread\. This confirms thatD\(T\)D\(T\)tracks epistemic maturity as well as truth, warning users when information is novel and consensus remains unsettled\. Table 2:Hallucination\-detection performance for three detectors derived from Eq\. 1\. Evidence density alone is near chance overall and anti\-discriminative on FreshQA, while the Consistency VetoPintP\_\{int\}is the strongest discriminator, especially on dynamic queries\. FreshQA estimates should be interpreted cautiously because the split contains only three hallucinated items\.#### Hallucination Detection\. To quantify the metric’s discriminative power, we labelled every audited generation with agpt\-4ojudge\. Across all 200 generations, the judge markednH=41n\_\{H\}=41as hallucinations \(nHstatic=38n\_\{H\}^\{\\text\{static\}\}=38,nHdynamic=3n\_\{H\}^\{\\text\{dynamic\}\}=3\)\. Table[2](https://arxiv.org/html/2609.03460#S4.T2)reports ROC\-AUC and average precision for three candidate detectors derived from Eq\. 1: the fullD\(T\)D\(T\)score, evidence density alone, and the Consistency VetoPintP\_\{int\}\. Three findings stand out\. First,retrieval density is not the workhorse\. Evidence density alone is near chance overall \(AUC=0\.47=0\.47\) and anti\-discriminative on FreshQA \(AUC=0\.28=0\.28\), indicating that high\-density retrieval can still support false claims when a misconception is popular or a topic is in transition\. Second,the Consistency VetoPintP\_\{int\}is the principal discriminator on dynamic queries and contributes most of the signal in the combined score\. It achieves the strongest overall AUC \(0\.600\.60\) and reaches AUC=0\.92=0\.92on dynamic queries, precisely where the model is most likely to be uncertain about post\-cutoff claims\. However, because the FreshQA split contains only three hallucinated items, these dynamic\-query AUC estimates should be interpreted as suggestive rather than definitive\. Third,the combinedD\(T\)D\(T\)score inherits some of the Veto’s signal through the\(1−Pint\)\(1\-P\_\{int\}\)gate\. It reaches AUC=0\.72=0\.72on FreshQA while preserving sensitivity to epistemic maturity on static knowledge, although it remains weaker than the veto alone because density is noisy on adversarial misconceptions\. In this sense,PintP\_\{int\}is the stronger standalone hallucination detector, whereasD\(T\)D\(T\)is retained as a composite signal because it adds ecological sensitivity to epistemic maturity that the veto alone does not express\. TheIndonesia\-capital transitioncase illustrates the role of the Consistency Veto: retrieval can return high\-density evidence for competing time\-sensitive answers, while the consistency gate down\-weights internally unstable generations\. ## 5Empirical Evaluation: Interaction Paradigm Having validated the computational feasibility of theD\(T\)D\(T\)metric \(Section[4](https://arxiv.org/html/2609.03460#S4)\), we turn to the critical Human\-Centered AI question: Does visualizing this “costly signal” actually enable users to overcome the Fluency Trap? We conducted a within\-subjects controlled experiment to measure Trust Calibration, the alignment between a user’s confidence and the factual veracity of the content—under different interface conditions\. Table 3:3×\\times3 Latin Square Design\. Participants were counterbalanced across topics and conditions so that PDI was evaluated on both grounded \(True\) and hallucinated \(Hall\.\) content, reducing topic\-specific bias through partial counterbalancing\.Figure 2:Experimental stimuli \(Group 1; schematic\)\. Simplified mock\-ups of the three interaction paradigms tested in the user study: \(a\) Control, shown without an external signal; \(b\) Binary Disclosure, shown with a static “AI\-Generated Content” banner; and \(c\) Provenance Density \(PDI\), shown with a density sidebar indicating a high proportion of verified claims \(D\(T\)≈1\.0D\(T\)\\\!\\approx\\\!1\.0\)\. Body text is abbreviated for layout clarity; full\-fidelity stimulus screens are included in the supplementary release\.#### Experimental Design\. We utilized a3×33\\times 3Latin Square within\-subjects design \(N=81N=81\) \(Table[3](https://arxiv.org/html/2609.03460#S5.T3)\)\. The independent variable was the Interaction Paradigm with three levels: Control \(The Status Quo\): Standard text with no disclosure, representing the current “Pooling Equilibrium” where fluency is the only visible signal\. Binary Disclosure: A static “Made with AI” banner\. This represents the current industry standard for transparency\. Provenance Density \(PDI\): Our proposed interface \(Figure[2](https://arxiv.org/html/2609.03460#S5.F2)\), visualizing the density of verification traces \(e\.g\., “5/5 Claims Verified”\)\. #### Stimuli and Oracle Assignment\. The key design goal was to separatealgorithmic errorfrominteraction failure\. Accordingly, the user study did not expose participants to live retrieval outputs\. Instead, in the PDI condition we used the Oracle protocol introduced in Section[3\.3](https://arxiv.org/html/2609.03460#S3.SS3): grounded summaries were shown with high\-density indicators \(D\(T\)≈1\.0D\(T\)\\approx 1\.0\), and hallucinated summaries were shown with null indicators \(D\(T\)≈0\.0D\(T\)\\approx 0\.0\)\. This makes the user study a test of whether the visualization can override the fluency heuristic when the signal is correct, rather than a test of search quality\. We generated six academic summaries\. For the hallucinated items, we wrote high\-fluency fabrications \(e\.g\., a fictional “Hua Dynasty” or “Theanine\-B” molecule\) matched to the grounded summaries in length, tone, and apparent scholarly style\. These adversarial stimuli intentionally preserve surface plausibility so that any improvement in ratings can be attributed to the interface rather than to obvious textual defects\. #### Layout Control\. To avoid a layout confound, all conditions used the same fixed\-width text column \(approximately 520 px\)\. In the Control and Binary conditions, the sidebar region was preserved as an invisible container so that line breaks, whitespace, and reading flow remained constant across conditions\. #### Procedure and Participants\. Participants \(N=81N=81\) were recruited via Prolific \(Fluent English, Approval Rate\>95%\>95\\%\)\. To ensure we measured the “layperson’s reliance on fluency,” we excluded participants with self\-reported expert domain knowledge in the specific topics \(History, Chemistry, Physics\)\. Each participant rated the Perceived Veracity of claims on a scale of 0 \(Completely Fabricated\) to 10 \(Completely Factual\)\. Because each participant contributed three ratings, all inferential analyses used a linear mixed model with a random intercept for participant\. ### 5\.1Results: Restoring the Separating Equilibrium A linear mixed model fit withlme4in R and a random intercept for participant \(rating∼\\simInterface×\\timesVeracity \+ \(1∣\\midparticipant\)\) revealed a significant Interface×\\timesVeracity interaction \(likelihood\-ratio testχ2\(2\)=31\.32\\chi^\{2\}\(2\)=31\.32,p<\.001p<\.001\)\. This substantial effect confirms that the choice of interaction paradigm fundamentally alters the user’s epistemic stance\. #### Confirming the Fluency Trap \(Control\): In the absence of signals, participants showed no detectable discrimination\. High\-Fluency Hallucinations \(M=6\.28M=6\.28\) were rated statistically indistinguishable from Ground Truth \(M=5\.78M=5\.78;p=\.43p=\.43,d=−0\.21d=\-0\.21\)\. This pattern is consistent with a collapse of the separating equilibrium: without external scaffolding, participants’ ratings treated fluency as statistically indistinguishable from veracity\. Figure 3:User\-study results \(N=243N=243ratings: 81 participants×\\times3 ratings each\)\.\(a\)Mean rating±\\pm95% CI shows the Interface×\\timesVeracity interaction \(χ2\(2\)=31\.32\\chi^\{2\}\(2\)=31\.32,p<\.001p<\.001\)\.\(b\)Discernment gapΔ=MTrue−MHall\\Delta=M\_\{\\mathrm\{True\}\}\-M\_\{\\mathrm\{Hall\}\}with Cohen’sddannotated; PDI produces the largest separation \(d=1\.82d=1\.82\)\. #### Stigmatization vs\. Calibration \(Binary vs\. PDI\): While both interventions reduced trust in hallucinations, they did so via distinct cognitive mechanisms \(Figure[3](https://arxiv.org/html/2609.03460#S5.F3)\):Binary Disclosure \(Stigmatization\):The “Made with AI” label functioned as a crude penalty\. It reduced trust in hallucinations \(M=5\.04M=5\.04\) but failed to meaningfully restore confidence in the truth \(M=6\.70M=6\.70\)\.Provenance Density \(Calibration\):In contrast, PDI produced a large separation between truth and fabrication ratings\. It penalized ungrounded content nearly twice as severely as binary labels \(M=3\.93M=3\.93\), while simultaneously elevating trust in verified content to the highest observed level \(M=8\.08M=8\.08\)\. Post\-hoc Tukey HSD comparisons across the Interface×\\timesVeracity estimated marginal means confirm that PDI significantly restores confidence in verified content compared to standard Binary Disclosure \(p=\.039p=\.039\), effectively reversing the Transparency Penalty\. Table 4:Quantifying Discernment\. PDI creates the largest separation \(d=1\.82d=1\.82\) between truth and fabrication\.Taken together, these results show that PDI improved discernment by raising confidence in grounded content while reducing confidence in hallucinations, whereas Binary Disclosure produced a smaller separation\. ## 6Discussion Our quantitative results confirm that Provenance Density indicators \(PDI\) significantly outperform both Control and Binary Disclosure conditions in restoring truth discernment\. By triangulating these statistical findings with the technical audit \(N=200N=200\) and qualitative participant feedback \(N=81N=81\), we infer the cognitive mechanisms driving these effects\. #### Mechanism of the Fluency Trap: Heuristic Substitution\. The failure of participants to distinguish between truth and hallucination in the Control condition \(p=\.43p=\.43\) is consistent with a related form of heuristic substitution: in the absence of provenance signals, participants may have substitutedveracity\(is this true?\) withfluency\(does this sound professional?\)\. This pattern parallels research on receptivity to superficially impressive but semantically vacuous statements[Pennycook et al\. \(2015\)](https://arxiv.org/html/2609.03460#bib.bib8), although our stimuli presented meaningful but false factual claims rather than vacuous prose\. Multiple participants explicitly described this strategy\. P25 noted, “They sound believable and well written… It didn’t seem made up to me,” while P81 judged accuracy based on “how the text was worded and flowed\.” These comments support the interpretation that when the cost of generating professional\-sounding text approaches zero, style ceases to be a reliable proxy for substance\. #### The “Warning” Effect of Binary Labels\. Prior work shows context\-dependent effects of AI labels\. In experiments with news headlines, AI labels reduced perceived accuracy even for true or human\-generated headlines[Altay and Gilardi \(2024\)](https://arxiv.org/html/2609.03460#bib.bib33); labels also reduced belief in misleading AI\-generated image posts[Wittenberg et al\. \(2025\)](https://arxiv.org/html/2609.03460#bib.bib32), whereas a health\-content experiment found no significant overall label effect and only nonsignificant reductions for accurate content[Li and Yang \(2024\)](https://arxiv.org/html/2609.03460#bib.bib34)\. Our analysis suggests that Binary Disclosure operated through a blunt mechanism ofepistemic stigmatization: users interpreted the label not as a transparency aid, but as a risk marker\. P23 described the interaction vividly: “The AI banner seemed like a warning vs being informative\.” In our study, this description supports the interpretation that the binary label operated as a risk cue rather than a verification aid\. This context\-specific penalty also aligns with concerns about the normative fairness of mandatory disclosure[Hosseini et al\. \(2025\)](https://arxiv.org/html/2609.03460#bib.bib21), as it punishes theuseof the tool rather than theaccuracyof the content[Cheong et al\. \(2025\)](https://arxiv.org/html/2609.03460#bib.bib13)\. #### Calibrated Reliance and Cognitive Offloading\. In the PDI condition, participant strategies shifted from intrinsic text evaluation to extrinsic evidence evaluation, with users treating the Provenance Density score as an evidence cue\. A critical question in HAI is whether such “cognitive offloading” functions as extended cognition[Chirayath et al\. \(2025\)](https://arxiv.org/html/2609.03460#bib.bib35);[Clark and Chalmers \(1998\)](https://arxiv.org/html/2609.03460#bib.bib39)\. Our technical validation \(Section[4](https://arxiv.org/html/2609.03460#S4)\) suggests cautious support, but with important limits\. The metric’sConsistency Veto—demonstrated by the suppression of scores in ambiguous queries such as the Indonesia capital transition—provides a partial safety rail for this offloading, not a guarantee\. Users like P6 \(“One of the passages also had a measurement of claims verified… I was more inclined to believe that”\) are not blindly trusting the AI; they are trusting averified attributeof the information\. However, because high\-density misconceptions can still evade the veto, the interface should be understood as improving truth discernment rather than eliminating the “False Assurance” pitfall common in retrieval systems\. #### Ecological Alignment: The metric’s sensitivity to information age \(Ecological Sensitivity\) aligns user trust with the stability of knowledge\. As shown in our technical results, established facts yielded higher density signals \(M=0\.79M=0\.79\) than emerging dynamic topics \(M=0\.64M=0\.64\)\. PDI therefore functions as a “Time\-Aware” signal, implicitly guiding users to be more skeptical of breaking news \(FreshQA\) than settled science \(TruthfulQA\), mirroring the actual epistemic reliability of those domains\. ## 7Limitations and Future Work #### From Oracle to Deployment\. The Oracle protocol in Section[3\.3](https://arxiv.org/html/2609.03460#S3.SS3)shows that PDI has interaction value when the signal is correct, but deployment depends on how closely the live pipeline approaches that ideal\. Retrieval can reinforce popular misconceptions, while the cross\-sample consistency gate cannot detect false answers that recur consistently\. The user\-study result should therefore be interpreted as an upper bound on interface efficacy under correct signaling; a key next step is to study near\-miss cases, especially false\-positive high\-density signals\. A second design caveat is that the3×33\\times 3Latin\-square assignment did not fully cross interface, veracity, and topic: the PDI–Hallucinated cell was measured on Matcha, whereas the Control–Hallucinated cell averaged over Silk and Matcha\. Thus, the observed PDI penalty \(M=3\.93M=3\.93\) should be interpreted as strong evidence for the interface under these stimuli, rather than as a fully topic\-independent estimate; a fully crossed replication is a clear next step\. #### Ecological Conservatism \(The Cold Start Problem\)\. Provenance Density is inherently conservative: it privileges established consensus \(M=0\.79M=0\.79\) over emerging novelty \(M=0\.64M=0\.64\)\. Legitimate new information, such as breaking news or novel scientific hypotheses, may lack the citation network required to achieve high density \(Sc=∅S\_\{c\}=\\emptyset\), risking classification of “the new” as “the unverified\.” Future iterations should integrate temporal metadata to distinguish low provenance due to fabrication from low provenance due to novelty\. #### Latency Constraints\. The “Consistency Veto” is computationally costly: with audit\-measured inference latencies ranging from 17\.0s \(Dynamic\) to 26\.4s \(Static\),D\(T\)D\(T\)is unsuitable for synchronous, turn\-by\-turn chat\. As noted in Section[4](https://arxiv.org/html/2609.03460#S4), the higher Static latency reflects longer TruthfulQA generations rather than a swapped label\. Future work should explore asynchronous interface patterns, such as “Check in Background” or “Deep Check,” that integrate verification latency without blocking interaction\. Because latency was not experimentally manipulated in our study, future experiments should also test whether this friction functions as a costly signal that improves discernment, or merely as an engineering constraint\. #### Adversarial Gaming and Automation Bias\. PDI also introduces second\-order risks\. The cubic MatchRatio term in Section[3\.1](https://arxiv.org/html/2609.03460#S3.SS1)is designed to reduce citation\-laundering by suppressing weak surface overlap, but the technical audit also surfaces vulnerabilities including domain spoofing, empty\-keyword exploitation, SEO\-style inflation, and mimetic misconception inheritance from open\-web consensus\. Hardening the reference implementation will require stricter domain parsing, conservative handling of vague claims, caps on per\-domain contribution, and checks for duplicated evidence\. The risk ofautomation bias[Cummings \(2017\)](https://arxiv.org/html/2609.03460#bib.bib1)remains: if users transition from “reading the text” to “scanning for the green bar,” future designs should add frictional elements that periodically force manual verification\. ## 8Conclusion This paper argues that AI\-generated misinformation is a problem ofsignalingas well asgeneration\. As fluency becomes cheap, style no longer reliably signals substance: without external scaffolding, users in our study could not distinguish grounded truth from high\-fluency hallucination\. Binary “Made with AI” labels address this crisis through identity disclosure, but our findings show that they act as crude warnings, producing a Transparency Penalty in which accurate AI\-generated information is discounted alongside fabrication\. Provenance Density offers a more calibrated alternative by visualizing verification rather than authorship\. In the user study, PDI created a large positive discernment gap \(\+4\.15\), restoring users’ ability to separate truth from fabrication under correct signaling\. The technical audit qualifies this promise by showing that high\-density evidence can still support misconceptions\. AI transparency should therefore shift from detectingwhowrote a text toward visualizingwhatsupports it, with interfaces that make verification visible, intuitive, and contestable\. ## Ethical Statement This research was conducted in accordance with the ethical standards of the University of Tokyo institutional review process\. We recruited 81 participants for the user study via Prolific, all of whom provided informed consent before engaging with the study materials\. Prolific submission IDs were collected solely for payment processing, stored separately from response data, and are not retained in the released datasets; no other personally identifying information was collected\. All analyses and shared materials use pseudonymous participant codes \(e\.g\.,P\_G1\_01\); raw survey exports are excluded from the open\-science release in accordance with the consent form, while the verbatim consent protocol and the de\-identified participant ratings are included\. Participants were compensated for their time in accordance with fair labor standards\. Beyond procedural compliance, we acknowledge the broader sociotechnical implications of proposing Provenance Density as a credibility standard\. While PDI effectively mitigates hallucinations in digitized high\-resource domains, it introduces a risk ofSystemic Authority Bias[Milgram \(1963\)](https://arxiv.org/html/2609.03460#bib.bib22);[Jackson et al\. \(2012\)](https://arxiv.org/html/2609.03460#bib.bib23)\. By algorithmically privileging information with a traceable digital lineage, citation\-based metrics may inadvertently disenfranchise non\-digitized knowledge systems—including indigenous epistemologies, oral histories, and low\-resource languages—that lack dense citation networks\. There is a specific risk that users may conflate “unverified” \(lack of digital evidence\) with “untrue” \(falsehood\)\. Consequently, we advocate that PDI be deployed strictly as a verification tool for digitized consensus, rather than a universal arbiter of truth, to prevent the exacerbation of existing epistemic exclusions\. ## Acknowledgments This work was supported in part by JSPS KAKENHI Grant Numbers 25KK0001 \(Fund for the Promotion of Joint International Research \(Fostering Joint International Research\)\) and 25K21241 \(Grant\-in\-Aid for Early\-Career Scientists\), and by JST Moonshot R&D Grant JPMJMS2012\. ## References - Altay and Gilardi \(2024\)S\. Altay and F\. GilardiPeople are skeptical of headlines labeled as ai\-generated, even if true or human\-made, because they assume full ai automation\.PNAS Nexus3\(10\),pp\. pgae403\.External Links:ISSN 2752\-6542,[Document](https://dx.doi.org/10.1093/pnasnexus/pgae403),[Link](https://doi.org/10.1093/pnasnexus/pgae403),https://academic\.oup\.com/pnasnexus/article\-pdf/3/10/pgae403/59961427/pgae403\.pdfCited by:[§1](https://arxiv.org/html/2609.03460#S1.p3.1),[§6](https://arxiv.org/html/2609.03460#S6.SS0.SSS0.Px2.p1.1)\. - Buderet al\.\(2024\)F\. Buder, N\. Hesel, and H\. DietrichBeyond the buzz: creating marketing value with generative ai\.NIM Marketing Intelligence Review16\(1\),pp\. 50–55\.Cited by:[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px3.p2.1)\. - Chenet al\.\(2025\)S\. Chen, M\. Gao, K\. Sasse, T\. Hartvigsen, B\. Anthony, L\. Fan, H\. Aerts, J\. Gallifant, and D\. S\. BittermanWhen helpfulness backfires: llms and the risk of false medical information due to sycophantic behavior\.npj Digital Medicine8\(1\),pp\. 605\.Cited by:[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px2.p2.1)\. - Cheonget al\.\(2025\)I\. Cheong, A\. Guo, M\. Lee, Z\. Liao, K\. Kadoma, D\. Go, J\. C\. Chang, P\. Henderson, M\. Naaman, and A\. X\. ZhangPenalizing transparency? how ai disclosure and author demographics shape human and ai judgments about writing\.arXiv preprint arXiv:2507\.01418\.Cited by:[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px3.p2.1),[§6](https://arxiv.org/html/2609.03460#S6.SS0.SSS0.Px2.p2.1)\. - Chirayathet al\.\(2025\)G\. Chirayath, K\. Premamalini, and J\. JosephCognitive offloading or cognitive overload? how ai alters the mental architecture of coping\.Frontiers in Psychology16,pp\. 1699320\.External Links:[Document](https://dx.doi.org/10.3389/fpsyg.2025.1699320)Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p4.1),[§6](https://arxiv.org/html/2609.03460#S6.SS0.SSS0.Px3.p1.1)\. - Clark and Chalmers \(1998\)A\. Clark and D\. ChalmersThe extended mind\.analysis58\(1\),pp\. 7–19\.Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p4.1),[§6](https://arxiv.org/html/2609.03460#S6.SS0.SSS0.Px3.p1.1)\. - Cummings \(2017\)M\. L\. CummingsAutomation bias in intelligent time critical decision support systems\.InDecision making in aviation,pp\. 289–294\.Cited by:[§7](https://arxiv.org/html/2609.03460#S7.SS0.SSS0.Px4.p1.1)\. - Dechêneet al\.\(2010\)A\. Dechêne, C\. Stahl, J\. Hansen, and M\. WänkeThe truth about the truth: a meta\-analytic review of the truth effect\.Personality and Social Psychology Review14\(2\),pp\. 238–257\.Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p1.1),[§1](https://arxiv.org/html/2609.03460#S1.p2.1)\. - Farquharet al\.\(2024\)S\. Farquhar, J\. Kossen, L\. Kuhn, and Y\. GalDetecting hallucinations in large language models using semantic entropy\.Nature630\(8017\),pp\. 625–630\.Cited by:[§3\.1](https://arxiv.org/html/2609.03460#S3.SS1.p1.1)\. - Galdin and Silbert \(2025\)A\. Galdin and J\. SilbertMaking talk cheap: generative ai and labor market signaling\.arXiv preprint arXiv:2511\.08785\.Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p2.1),[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px2.p1.1)\. - Gintiset al\.\(2001\)H\. Gintis, E\. A\. Smith, and S\. BowlesCostly signaling and cooperation\.Journal of theoretical biology213\(1\),pp\. 103–119\.Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p1.1)\. - Hasheret al\.\(1977\)L\. Hasher, D\. Goldstein, and T\. ToppinoFrequency and the conference of referential validity\.Journal of verbal learning and verbal behavior16\(1\),pp\. 107–112\.Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p2.1)\. - Hertwiget al\.\(2008\)R\. Hertwig, S\. M\. Herzog, L\. J\. Schooler, and T\. ReimerFluency heuristic: a model of how the mind exploits a by\-product of information retrieval\.\.Journal of Experimental Psychology: Learning, memory, and cognition34\(5\),pp\. 1191\.Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p1.1)\. - Hosseiniet al\.\(2025\)M\. Hosseini, B\. Gordijn, G\. E\. Kaebnick, and K\. HolmesDisclosing generative ai use for writing assistance should be voluntary\.Research Ethics21\(4\),pp\. 728–735\.External Links:[Document](https://dx.doi.org/10.1177/17470161251345499)Cited by:[§6](https://arxiv.org/html/2609.03460#S6.SS0.SSS0.Px2.p2.1)\. - Jacksonet al\.\(2012\)J\. Jackson, B\. Bradford, M\. Hough, A\. Myhill, P\. Quinton, and T\. R\. TylerWhy do people comply with the law? legitimacy and the influence of legal institutions\.British journal of criminology52\(6\),pp\. 1051–1071\.Cited by:[Ethical Statement](https://arxiv.org/html/2609.03460#Sx1.p2.1)\. - Jakeschet al\.\(2023\)M\. Jakesch, J\. T\. Hancock, and M\. NaamanHuman heuristics for ai\-generated language are flawed\.Proceedings of the National Academy of Sciences120\(11\),pp\. e2208839120\.Cited by:[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px3.p1.1)\. - Kahneman \(2011\)D\. KahnemanThinking, fast and slow\.macmillan\.Cited by:[§3\.4](https://arxiv.org/html/2609.03460#S3.SS4.p1.1)\. - Kaur \(2025\)A\. KaurEchoes of agreement: argument driven sycophancy in large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 22803–22812\.Cited by:[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px2.p2.1)\. - Kayet al\.\(2024\)J\. Kay, A\. Kasirzadeh, and S\. MohamedEpistemic injustice in generative ai\.InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society,Vol\.7,pp\. 684–697\.Cited by:[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px4.p1.1)\. - Lewiset al\.\(2020\)P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel,et al\.Retrieval\-augmented generation for knowledge\-intensive nlp tasks\.Advances in neural information processing systems33,pp\. 9459–9474\.Cited by:[§3\.1](https://arxiv.org/html/2609.03460#S3.SS1.p1.1)\. - Li and Yang \(2024\)F\. Li and Y\. YangImpact of artificial intelligence–generated content labels on perceived accuracy, message credibility, and sharing intentions for misinformation: web\-based, randomized, controlled experiment\.JMIR Formative Research8,pp\. e60024\.Cited by:[§6](https://arxiv.org/html/2609.03460#S6.SS0.SSS0.Px2.p1.1)\. - Lianget al\.\(2025\)K\. Liang, H\. Hu, X\. Zhao, D\. Song, T\. L\. Griffiths, and J\. F\. FisacMachine bullshit: characterizing the emergent disregard for truth in large language models\.arXiv preprint arXiv:2507\.07484\.Cited by:[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px2.p2.1)\. - Lianget al\.\(2023\)W\. Liang, M\. Yuksekgonul, Y\. Mao, E\. Wu, and J\. ZouGPT detectors are biased against non\-native english writers\.Patterns4\(7\)\.Cited by:[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px4.p1.1)\. - Linet al\.\(2022\)S\. Lin, J\. Hilton, and O\. EvansTruthfulQA: measuring how models mimic human falsehoods\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 3214–3252\.Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p5.1)\. - Marewski and Schooler \(2011\)J\. N\. Marewski and L\. J\. SchoolerCognitive niches: an ecological model of strategy selection\.\.Psychological review118\(3\),pp\. 393\.Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p1.1)\. - Milgram \(1963\)S\. MilgramBehavioral study of obedience\.\.The Journal of abnormal and social psychology67\(4\),pp\. 371\.Cited by:[Ethical Statement](https://arxiv.org/html/2609.03460#Sx1.p2.1)\. - Nakanoet al\.\(2026\)H\. Nakano, J\. Takezawa, F\. Matulic, C\. Yang, and K\. YataniUnderstanding reader perception shifts upon disclosure of ai authorship\.InProceedings of the 31st International Conference on Intelligent User Interfaces,pp\. 2131–2146\.Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p3.1),[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px3.p2.1)\. - Pennycooket al\.\(2015\)G\. Pennycook, J\. A\. Cheyne, N\. Barr, D\. J\. Koehler, and J\. A\. FugelsangOn the reception and detection of pseudo\-profound bullshit\.Judgment and Decision making10\(6\),pp\. 549–563\.Cited by:[§3\.4](https://arxiv.org/html/2609.03460#S3.SS4.p1.1),[§6](https://arxiv.org/html/2609.03460#S6.SS0.SSS0.Px1.p1.1)\. - Reber and Schwarz \(1999\)R\. Reber and N\. SchwarzEffects of perceptual fluency on judgments of truth\.Consciousness and cognition8\(3\),pp\. 338–342\.Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p1.1),[§1](https://arxiv.org/html/2609.03460#S1.p2.1)\. - Reber and Unkelbach \(2010\)R\. Reber and C\. UnkelbachThe epistemic status of processing fluency as source for judgments of truth\.Review of philosophy and psychology1\(4\),pp\. 563–581\.Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p1.1),[§1](https://arxiv.org/html/2609.03460#S1.p2.1)\. - Rheu and Cho \(2025\)M\. Rheu and J\. ChoThe trap of ai literacy: the paradoxical relationships between college students’ use of llms, ai literacy, and fact\-checking behavior\.InProceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems,pp\. 1–7\.Cited by:[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px3.p1.1)\. - Sharmaet al\.\(2024\)M\. Sharma, M\. Tong, T\. Korbak, D\. Duvenaud, A\. Askell, S\. R\. Bowman, E\. DURMUS, Z\. Hatfield\-Dodds, S\. R\. Johnston, S\. M\. Kravec, T\. Maxwell, S\. McCandlish, K\. Ndousse, O\. Rausch, N\. Schiefer, D\. Yan, M\. Zhang, and E\. PerezTowards understanding sycophancy in language models\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=tvhaxkMKAn)Cited by:[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px2.p2.1)\. - Shumailovet al\.\(2024\)I\. Shumailov, Z\. Shumaylov, Y\. Zhao, N\. Papernot, R\. Anderson, and Y\. GalAI models collapse when trained on recursively generated data\.Nature631\(8022\),pp\. 755–759\.Cited by:[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px2.p2.1)\. - Spearinget al\.\(2025\)E\. R\. Spearing, C\. I\. Gile, A\. L\. Fogwill, T\. Prike, B\. Swire\-Thompson, S\. Lewandowsky, and U\. K\. EckerCountering ai\-generated misinformation with pre\-emptive source discreditation and debunking\.Royal Society Open Science12\(6\),pp\. 242148\.External Links:[Document](https://dx.doi.org/10.1098/rsos.242148)Cited by:[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px1.p2.1)\. - Spence \(1978\)M\. SpenceJob market signaling\.InUncertainty in economics,pp\. 281–306\.Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p1.1),[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px2.p1.1)\. - Sperberet al\.\(2010\)D\. Sperber, F\. Clément, C\. Heintz, O\. Mascaro, H\. Mercier, G\. Origgi, and D\. WilsonEpistemic vigilance\.Mind & language25\(4\),pp\. 359–393\.Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p2.1)\. - Unkelbachet al\.\(2011\)C\. Unkelbach, M\. Bayer, H\. Alves, A\. Koch, and C\. StahlFluency and positivity as possible causes of the truth effect\.Consciousness and cognition20\(3\),pp\. 594–602\.Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p2.1)\. - Vuet al\.\(2024\)T\. Vu, M\. Iyyer, X\. Wang, N\. Constant, J\. Wei, J\. Wei, C\. Tar, Y\. Sung, D\. Zhou, Q\. Le, and T\. LuongFreshLLMs: refreshing large language models with search engine augmentation\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 13697–13720\.External Links:[Link](https://aclanthology.org/2024.findings-acl.813/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.813)Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p5.1)\. - Wittenberget al\.\(2025\)C\. Wittenberg, Z\. Epstein, G\. Péloquin\-Skulski, A\. J\. Berinsky, and D\. G\. RandLabeling ai\-generated media online\.PNAS nexus4\(6\),pp\. pgaf170\.Cited by:[§6](https://arxiv.org/html/2609.03460#S6.SS0.SSS0.Px2.p1.1)\. - Zahavi \(1975\)A\. ZahaviMate selection—a selection for a handicap\.Journal of theoretical Biology53\(1\),pp\. 205–214\.Cited by:[§1](https://arxiv.org/html/2609.03460#S1.p1.1),[§2](https://arxiv.org/html/2609.03460#S2.SS0.SSS0.Px2.p1.1)\.
Similar Articles
Provenance: A survival toolkit for an AI dominant information landscape
The article discusses the growing threat of AI-generated deception in the information landscape and proposes provenance—an ecosystem-level adoption of content authentication—as a remedy, highlighting risks like AI catfishing, fabricated scientific data, and coordinated disinformation campaigns.
ProvenAI: Provenance-Native Traces of Evidence in Generated Answers
ProvenAI introduces a framework for decomposing transparency in multi-hop question answering into three independently measurable layers: answer correctness, citation fidelity, and per-document influence, revealing a citation-influence gap where cited sources may have weak influence while uncited sources significantly shape the output.
Advancing content provenance for a safer, more transparent AI ecosystem
OpenAI announces new initiatives for content provenance, including C2PA conformance, integration of Google DeepMind's SynthID watermarking for images, and a preview of a verification tool to help users identify AI-generated content.
@OpenAI: We’re adding new ways for people to identify AI-generated images and understand where they came from. In addition to C2…
OpenAI announces new content provenance features including C2PA Content Credentials, SynthID watermarking from Google DeepMind, and a public verification tool to identify AI-generated images from its products, aiming to enhance transparency and trust.
Every AI Visibility Tool Is Lying to You
This article critically examines the accuracy of AI visibility tools that claim to measure brand presence in generative AI responses, arguing that they provide false precision due to nondeterminism, personalization, and scraping biases. It calls for transparency in methodology and warns against treating opaque dashboards as stable truth.