truthfulness

Tag

Cards List
#truthfulness

@josephdecker: The winner of my 16-model eval fabricated 5 times in the audit. Second place, a tenth of a point back: zero fabrication…

X AI KOLs Timeline · 2026-07-24 Cached

Joseph Decker evaluates 16 AI models on truthfulness for his product Condensr, discovers that the leaderboard winner fabricated content five times in an audit, and instead ships the second-place model which had zero fabrications. The post details the evaluation process, a bug in the LLM judge that penalized accurate summaries due to truncated transcripts, and the importance of custom evals over generic benchmarks.

0 favorites 0 likes
#truthfulness

Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models

Hugging Face Daily Papers · 2026-07-24 Cached

Introduces ReDe, a framework that denoises reasoning traces by filtering irrelevant and repetitive steps to improve hallucination detection in large reasoning models, achieving up to 87.32 AUROC on TruthfulQA.

0 favorites 0 likes
#truthfulness

GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs

arXiv cs.CL · 2026-07-13 Cached

GrAInS is a contrastive gradient-based method that uses Integrated Gradients to identify influential tokens and construct steering vectors for inference-time steering of LLMs and VLMs, improving truthfulness and reducing hallucinations without degrading fluency.

0 favorites 0 likes
#truthfulness

@jchudnov: Pass@k and self-consistency work great for math and code; sample more and verify. So we asked: can the same trick scale…

X AI KOLs Following · 2026-07-05 Cached

A new paper shows that scaling inference compute via methods like self-consistency improves LLM accuracy in math and code but fails to improve truthfulness in domains without external verifiers, as model errors are too correlated.

0 favorites 0 likes
#truthfulness

ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence

arXiv cs.CL · 2026-06-26 Cached

ConflictScore is a new metric that quantifies how well language models acknowledge conflicting evidence in their grounding documents, decomposing responses into atomic claims and measuring conflict balance. The paper also introduces ConflictBench, a benchmark covering diverse conflict forms, and shows the metric can improve truthfulness on TruthfulQA.

0 favorites 0 likes
#truthfulness

@elonmusk: Grok is maximally truthful

X AI KOLs Following · 2026-06-11 Cached

Elon Musk claims Grok is maximally truthful, referencing a claim that Fable 5 (hypothetical) lies 96% of the time.

0 favorites 0 likes
#truthfulness

TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness Assessment

arXiv cs.AI · 2026-06-03 Cached

TriEval is a new pipeline for evaluating LLMs across bias, toxicity, and truthfulness simultaneously, designed to be resource-efficient and run on standard laptops. It has been tested on Llama 3 8B, Mistral 7B, Gemma 2 9B, and Claude Haiku, and is released as open source.

0 favorites 0 likes
#truthfulness

Hallucination Is Linearly Decodable from Mid-Layer Hidden States in Quantized LLMs

arXiv cs.LG · 2026-06-03 Cached

This paper investigates whether open-source quantized LLMs encode a linearly separable truthfulness signal in their hidden states. Across three 7B-8B instruction-tuned models, a linear probe on a single mid-network layer achieves 0.904-1.000 AUROC on hallucination detection benchmarks, outperforming sampling-based methods.

0 favorites 0 likes
#truthfulness

Claude made me realize most AI models optimize for confidence, not truth

Reddit r/artificial · 2026-05-22

A reflection on how many AI models prioritize sounding confident over being truthful, using Claude as an example of a model that seems more focused on internal consistency and logical honesty.

0 favorites 0 likes
#truthfulness

Lying Is Just a Phase: The Hidden Alignment Transition in Language Model Scaling

arXiv cs.LG · 2026-05-20

This paper identifies a phase transition in language model scaling where below a critical parameter count, reasoning and truthfulness are anticorrelated, but above it they cooperate. It provides diagnostics and interventions for improving alignment across model families.

0 favorites 0 likes
#truthfulness

FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language Models

arXiv cs.CL · 2026-04-20 Cached

FineSteer is a novel inference-time steering framework that decomposes steering into conditional steering and fine-grained vector synthesis stages, using Subspace-guided Conditional Steering (SCS) and Mixture-of-Steering-Experts (MoSE) mechanisms to improve safety and truthfulness while preserving model utility. Experiments show 7.6% improvement over state-of-the-art methods on TruthfulQA with minimal utility loss.

0 favorites 0 likes
#truthfulness

TruthfulQA: Measuring how models mimic human falsehoods

OpenAI Blog · 2021-09-08 Cached

TruthfulQA is a benchmark of 817 questions across 38 categories designed to measure whether language models generate truthful answers. The study found that the best model achieved only 58% truthfulness compared to 94% for humans, and larger models were generally less truthful—suggesting scaling alone is insufficient for improving truthfulness.

0 favorites 0 likes
← Back to home

Submit Feedback