Tag
This paper defines a hierarchy of faithfulness criteria for knowledge base completion models and evaluates current embedding models, finding they are not logically faithful.
This paper uses activation patching to causally measure if reasoning steps in chain-of-thought are load-bearing, finding that behavioral tests overestimate faithfulness and larger models like Qwen3-4B maintain better faithfulness across reasoning depths.
LoomSum is a training-free framework that improves faithful summarization of long text-table documents by explicitly linking quantitative facts with narrative evidence, and introduces a new metric TGF for evaluating faithfulness.
FACE-Eval evaluation reveals that chain-of-thought monitoring is less reliable when preference cues come from tool outputs or implicit artifacts, with models showing lower verbalized commitment and higher unverbalized adoption across diverse open-weight models.
This paper evaluates adaptive computation in transformers using SEWN, a two-stream model with a learned gate for token routing, demonstrating that SEWN-sparse achieves significant throughput improvements over BERT-base and DistilBERT with modest accuracy trade-offs while providing interpretable token-importance signals.
This paper demonstrates that Chain-of-Thought reasoning in large language models is not always faithful, with models producing plausible but incorrect reasoning chains in response to naturally worded prompts, raising concerns for AI safety and transparency.
Introduces FaithformBench, a benchmark for assessing the faithfulness of mathematical chain-of-thought autoformalisation systems by measuring validity and invalidity preservation on perturbed steps. Applied to eight AF systems, it reveals widespread sycophancy where invalid inputs are silently corrected.
This paper presents a grounded and decomposed framework for evaluating relation-level hallucinations in abstractive summarization, introducing a normalized Relation Hallucination Index (RHI) with linguistically informed relation extraction enhancements.
This paper introduces MirageBench, a benchmark showing that LLMs with persistent memory fabricate user profiles through over-inference 35-49% of the time, and reveals that model self-reported confidence is inversely correlated with actual over-inference, making self-monitoring unreliable for comparing models.
This paper from Stony Brook University identifies 'Averaging Bias' in human faithfulness annotations for text summarization: global human labels correlate better with the average of per-sentence LLM judgments than with a strict conjunctive rule, meaning humans often label summaries as faithful even when they contain local factual errors.
This paper introduces LAWFUL, a framework for verifying whether neural networks learn and causally use physical laws over continuous variables, addressing gaps in coverage-aware causal consistency and domain-of-validity testing. It demonstrates the approach on a Mocap2Radar transformer and the Doppler frequency law.
This paper investigates how well activation steering for improving chain-of-thought faithfulness generalizes across cue types, datasets, and steering vector construction methods across several Gemma and Qwen models. The authors find that when steering is effective, it generalizes broadly, and the evaluation setting matters more than the training setting or vector construction method.
This paper introduces CoT-Mediate, a behavioral framework to test whether chain-of-thought reasoning in medical vision-language models actually drives predictions or merely decorates them. Auditing LLaVA-Med and MedGemma on VQA-RAD, it finds that how reasoning is injected (prefix-forcing vs re-prompting) and the attributed source (self vs expert) significantly affect model faithfulness and sycophancy.
This paper proposes a k-order relaxation of the faithfulness assumption for learning graphical Markov blankets, and introduces a proof-of-concept algorithm (kOMB) that can recover Markov blankets even under violations of faithfulness, such as parity-type relationships.
This paper introduces LedgerMind, a provenance-constrained multimodal agentic reasoning framework that uses a Structured Evidence Ledger to ensure grounded, faithful reasoning in visual question answering, addressing failure patterns like hallucination and over-reasoning.
This paper introduces an evaluation protocol with four new metrics and a benchmark dataset to assess context attribution methods for LLMs, showing they fail when context overlaps with training data.
This paper studies behavioral detection of unfaithful chain-of-thought reasoning in LLMs, finding that answer correctness structures detection performance: on incorrect answers, where most unfaithfulness occurs, behavioral signals are at chance, while on correct answers they offer modest separation.
This paper presents the first systematic study of faithfulness in document-grounded podcast generation, introducing a turn-level LLM-as-a-judge evaluation framework and a model-agnostic catch-n-repair method that improves faithfulness across domains.
ReFact proposes an adaptive fact-restatement citation framework that trains LLMs to decide when reasoning steps need contextual grounding, improving faithfulness and compactness in chain-of-thought reasoning while reducing token consumption.
Proposes CASE, a framework combining training-time causal alignment and inference-time structural enforcement to improve faithfulness of chain-of-thought reasoning in large language models, achieving a 37% average improvement in CoT faithfulness across benchmarks.