Tag
This paper introduces TKFQA, a counterfactual benchmark of 10,130 QA pairs over tables, texts, and knowledge graphs for evaluating LLM factuality consistency and order-robust reasoning, and proposes ORLF, a training framework that improves reasoning-chain accuracy and reduces input-order sensitivity.
Proposes three attention-guided strategies for layer selection in contrastive decoding for large language models, improving factuality on TruthfulQA over the DoLa baseline.
This paper proposes a novel online reinforcement learning method to improve factuality in reasoning LLMs by designing a reward function that balances factual precision, detail, and relevance, achieving a 23.1 percentage point reduction in hallucination rate on six benchmarks.
This paper investigates whether neural network probes that predict correctness of language model outputs actually capture objective correctness or the model's own self-judgement, using conflict cases where the two disagree. The authors find that transferable directions predominantly preserve self-judgement polarity, challenging the interpretation of correctness readouts.
Arena.ai has added Factuality to model rankings, supporting weighting of human preference and factuality, and showing changes in model rankings.
This paper proposes a semantic framework to describe AI systems, distinguishing justified claims from misleading outputs, and defines common failures such as hallucination and unsupported assertions.
This paper systematically examines cross-lingual disparities in LLM-based healthcare question answering across five languages, finding significant gaps in factual alignment and proposing the MultiWikiHealthCare dataset.
Introduces MARGO, a reinforcement learning framework that uses mixed-mode advantage regularization to mitigate thinking-induced hallucinations in large reasoning models by comparing thinking and non-thinking trajectories.
This paper proposes D2R-RAG, a model-agnostic and resource-aware framework that diagnoses and repairs factual errors in RAG systems under latency and VRAM constraints, achieving better accuracy-efficiency trade-offs on FEVER and HotpotQA.
ConflictScore is a new metric that quantifies how well language models acknowledge conflicting evidence in their grounding documents, decomposing responses into atomic claims and measuring conflict balance. The paper also introduces ConflictBench, a benchmark covering diverse conflict forms, and shows the metric can improve truthfulness on TruthfulQA.
Report claims that GPT-5.5 Instant shows significant improvements in factual accuracy, particularly in high-stakes fields like medicine, law, and finance.
MoshiRAG combines a compact full-duplex speech language model with asynchronous retrieval-augmented generation to improve factuality while maintaining real-time interactivity. The approach leverages natural temporal gaps in conversation to retrieve external knowledge without disrupting the natural flow of dialogue.
Google DeepMind and Kaggle have launched the FACTS Benchmark Suite, a comprehensive set of evaluations including parametric, search, multimodal, and grounding benchmarks to systematically measure the factuality of large language models.
DeepMind introduces FACTS Grounding, a comprehensive benchmark with 1,719 examples for evaluating how accurately large language models ground their responses in source material and avoid hallucinations. The benchmark includes a public dataset and an online Kaggle leaderboard tracking LLM performance on factual accuracy and grounding tasks.
OpenAI introduces SimpleQA, a new factuality benchmark dataset with 4,326 short fact-seeking questions designed to evaluate frontier language models on their ability to provide accurate answers without hallucination. The dataset achieves high quality through dual independent annotation, rigorous criteria, and achieves only ~3% estimated error rate, with GPT-4o scoring less than 40%.