factuality

Tag

Cards List
#factuality

Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge

arXiv cs.AI · 2026-08-11 Cached

This paper introduces TKFQA, a counterfactual benchmark of 10,130 QA pairs over tables, texts, and knowledge graphs for evaluating LLM factuality consistency and order-robust reasoning, and proposes ORLF, a training framework that improves reasoning-chain accuracy and reduces input-order sensitivity.

0 favorites 0 likes
#factuality

Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models

arXiv cs.CL · 2026-07-28 Cached

Proposes three attention-guided strategies for layer selection in contrastive decoding for large language models, improving factuality on TruthfulQA over the DoLa baseline.

0 favorites 0 likes
#factuality

Learning to Reason for Factuality

arXiv cs.CL · 2026-07-27 Cached

This paper proposes a novel online reinforcement learning method to improve factuality in reasoning LLMs by designing a reward function that balances factual precision, detail, and relevance, achieving a 23.1 percentage point reduction in hallucination rate on six benchmarks.

0 favorites 0 likes
#factuality

Diagnosing Correctness Probes under Self-Judgement Confounding

arXiv cs.CL · 2026-07-21 Cached

This paper investigates whether neural network probes that predict correctness of language model outputs actually capture objective correctness or the model's own self-judgement, using conflict cases where the two disagree. The authors find that transferable directions predominantly preserve self-judgement polarity, challenging the interpretation of correctness readouts.

0 favorites 0 likes
#factuality

@mylifcc: http://Arena.ai just officially added 'Factuality' to model rankings. The leaderboard now supports weighted viewing of 'Human Preference + Factuality' (default 25% factuality weight). They have annotated over 2 million claims from real conversations (Text Ar…

X AI KOLs Timeline · 2026-07-15 Cached

Arena.ai has added Factuality to model rankings, supporting weighting of human preference and factuality, and showing changes in model rankings.

0 favorites 0 likes
#factuality

Ceci n'est pas une pipe: AI systems as semantic abstractions

arXiv cs.AI · 2026-07-13 Cached

This paper proposes a semantic framework to describe AI systems, distinguishing justified claims from misleading outputs, and defines common failures such as hallucination and unsupported assertions.

0 favorites 0 likes
#factuality

Zoom In Disparities in Healthcare LLM Q&A

arXiv cs.CL · 2026-07-09 Cached

This paper systematically examines cross-lingual disparities in LLM-based healthcare question answering across five languages, finding significant gaps in factual alignment and proposing the MultiWikiHealthCare dataset.

0 favorites 0 likes
#factuality

Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization

arXiv cs.CL · 2026-07-08 Cached

Introduces MARGO, a reinforcement learning framework that uses mixed-mode advantage regularization to mitigate thinking-induced hallucinations in large reasoning models by comparing thinking and non-thinking trajectories.

0 favorites 0 likes
#factuality

Diagnosing and Repairing Factual Errors in RAG under Budget Constraints

arXiv cs.AI · 2026-06-30 Cached

This paper proposes D2R-RAG, a model-agnostic and resource-aware framework that diagnoses and repairs factual errors in RAG systems under latency and VRAM constraints, achieving better accuracy-efficiency trade-offs on FEVER and HotpotQA.

0 favorites 0 likes
#factuality

ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence

arXiv cs.CL · 2026-06-26 Cached

ConflictScore is a new metric that quantifies how well language models acknowledge conflicting evidence in their grounding documents, decomposing responses into atomic claims and measuring conflict balance. The paper also introduces ConflictBench, a benchmark covering diverse conflict forms, and shows the metric can improve truthfulness on TruthfulQA.

0 favorites 0 likes
#factuality

@FinanceYF5: 3/ Improved Accuracy: GPT-5.5 Instant shows significant improvements in factual accuracy, particularly in fields with high accuracy requirements such as medicine, law, and finance.

X AI KOLs Following · 2026-05-10 Cached

Report claims that GPT-5.5 Instant shows significant improvements in factual accuracy, particularly in high-stakes fields like medicine, law, and finance.

0 favorites 0 likes
#factuality

MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models

arXiv cs.CL · 2026-04-20 Cached

MoshiRAG combines a compact full-duplex speech language model with asynchronous retrieval-augmented generation to improve factuality while maintaining real-time interactivity. The approach leverages natural temporal gaps in conversation to retrieve external knowledge without disrupting the natural flow of dialogue.

0 favorites 0 likes
#factuality

FACTS Benchmark Suite: Systematically evaluating the factuality of large language models

Google DeepMind Blog · 2025-12-09 Cached

Google DeepMind and Kaggle have launched the FACTS Benchmark Suite, a comprehensive set of evaluations including parametric, search, multimodal, and grounding benchmarks to systematically measure the factuality of large language models.

0 favorites 0 likes
#factuality

FACTS Grounding: A new benchmark for evaluating the factuality of large language models

Google DeepMind Blog · 2024-12-17 Cached

DeepMind introduces FACTS Grounding, a comprehensive benchmark with 1,719 examples for evaluating how accurately large language models ground their responses in source material and avoid hallucinations. The benchmark includes a public dataset and an online Kaggle leaderboard tracking LLM performance on factual accuracy and grounding tasks.

0 favorites 0 likes
#factuality

Introducing SimpleQA

OpenAI Blog · 2024-10-30 Cached

OpenAI introduces SimpleQA, a new factuality benchmark dataset with 4,326 short fact-seeking questions designed to evaluate frontier language models on their ability to provide accurate answers without hallucination. The dataset achieves high quality through dual independent annotation, rigorous criteria, and achieves only ~3% estimated error rate, with GPT-4o scoring less than 40%.

0 favorites 0 likes
← Back to home

Submit Feedback