hallucination

Tag

Cards List
#hallucination

AI is confidently wrong way more than people give it credit for, change my mind

Reddit r/ArtificialInteligence ↗ · 2026-08-12

A user shares concerns about AI models presenting thin or ambiguous data with the same confidence as well-supported findings, citing a case where a complaint appearing only twice in 200 comments was ranked as a top concern. The piece questions whether this is a fixable prompting issue or a fundamental limitation requiring manual verification.

0 favorites 0 likes
#hallucination

Can Gemma and Qwen models catch hallucinations by looking at their own logprobs?

Reddit r/LocalLLaMA ↗ · 2026-08-11

The author shares experiments using a custom WebUI to let Gemma and Qwen models inspect their own logprobs to detect hallucinations. Initial observations suggest that first-recall token probabilities can indicate uncertainty, though both models struggle to read their own logprobs.

0 favorites 0 likes
#hallucination

A Grounded and Decomposed Framework for Relation-Level Hallucination Evaluation in Abstractive Summarization

arXiv cs.CL ↗ · 2026-08-11 Cached

This paper presents a grounded and decomposed framework for evaluating relation-level hallucinations in abstractive summarization, introducing a normalized Relation Hallucination Index (RHI) with linguistically informed relation extraction enhancements.

0 favorites 0 likes
#hallucination

Unified Hallucination Fuzzing for Multimodal Large Language Models

arXiv cs.CL ↗ · 2026-08-11 Cached

This paper presents UniHall, a fine-grained hallucination benchmark with a unified taxonomy, and Self-Adaptive Multimodal Fuzzing (SAMF), a self-evolving stress-testing framework for multimodal LLMs. Experiments show SOTA models degrade significantly under fuzzing and reveal a helpfulness-hallucination trade-off.

0 favorites 0 likes
#hallucination

REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment

arXiv cs.AI ↗ · 2026-08-11 Cached

This paper introduces REIN, an alignment framework that reduces hallucination in large reasoning models by training them to explicitly reflect before answering and to abstain when knowledge is insufficient. Experiments show consistent gains in selective accuracy and hallucination reduction across benchmarks.

0 favorites 0 likes
#hallucination

Claude Voice Mode Did Something Concerning

Reddit r/ArtificialInteligence ↗ · 2026-08-10

A user reports that Claude's voice mode produced a suspicious tool-call result containing an apparent prompt-injection message claiming to be from Anthropic's security team, asking for access to sensitive files.

0 favorites 0 likes
#hallucination

Don't classify, hallucinate!

Hacker News Top ↗ · 2026-08-10 Cached

The article describes a technique for classifying e-commerce queries with LLMs by having the model hallucinate hypothetical classifications, then mapping them to the real taxonomy using embeddings, which is cheaper and simpler than constrained structured outputs.

0 favorites 0 likes
#hallucination

I tested 32 models at extraction, the results are surprising

Reddit r/AI_Agents ↗ · 2026-08-06

A developer benchmarks 32 local models on fact extraction for agent memory, showing that F1 hides a critical failure mode: models with similar scores differ greatly in how often they invent facts on inputs that should output nothing. The article argues agent memory evaluation must include empty-output and retraction cases.

0 favorites 0 likes
#hallucination

Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks

arXiv cs.CL ↗ · 2026-08-06 Cached

This paper proposes a framework to elicit intrinsic hallucinations in LLMs using semantically equivalent adversarial perturbations, showing that state-of-the-art models degrade significantly in contextual faithfulness even with meaning-preserving query variations.

0 favorites 0 likes
#hallucination

Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary

arXiv cs.CL ↗ · 2026-08-06 Cached

This paper introduces ACT-Eval, a tool-augmented evaluation framework for LLM chess commentary, and releases a benchmark of 325 position-move pairs. It finds that factual hallucinations remain pervasive in LLM chess commentary, and tool augmentation improves factual correctness but not expert-level strategic coverage.

0 favorites 0 likes
#hallucination

FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact

arXiv cs.CL ↗ · 2026-08-05 Cached

Introduces factwash, an open-source write-time gate that deterministically catches AI rewrites that turn hearsay into fact by stripping attribution, hedging, or temporal context. The paper analyzes when simple checks suffice versus when an LLM witness is needed, and evaluates the approach on large annotated corpora.

0 favorites 0 likes
#hallucination

@BenjaminDEKR: Google's built-in AI Overview answerbot just told me it is made by OpenAI.

X AI KOLs Timeline ↗ · 2026-08-04 Cached

A user shares that Google's AI Overview answerbot claimed it was made by OpenAI, highlighting a hallucination or misattribution in the AI system.

0 favorites 0 likes
#hallucination

GeoArbiter: Verifiability-Guided Grounding for Remote-Sensing Multimodal LLMs

arXiv cs.LG ↗ · 2026-08-04 Cached

GeoArbiter proposes a training-free pipeline that selectively injects image-unverifiable geographic facts into remote-sensing multimodal LLMs to reduce knowledge hallucinations while preserving retrieval accuracy gains.

0 favorites 0 likes
#hallucination

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning

arXiv cs.LG ↗ · 2026-08-04 Cached

This paper proves that using error-penalized scoring rules with abstention as a discrete action can kill both the reward gradient and the KL anchor, causing models to collapse toward refusing everything. It proposes a structural repair — training a mandatory confidence report — and validates the mechanism with simulations and language model experiments.

0 favorites 0 likes
#hallucination

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning

arXiv cs.CL ↗ · 2026-08-03 Cached

This paper characterizes 'futile reasoning' in large language models, where models produce superficially valid but incorrect reasoning on tasks beyond their capability. They introduce CaRL, a capability-aligned reinforcement learning method that trains LLMs to abstain from futile reasoning while preserving performance.

0 favorites 0 likes
#hallucination

The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?

arXiv cs.CL ↗ · 2026-08-03 Cached

This paper introduces the 'Agentic Formalism Trap' and an Evaluative Dissonance Index, showing how LLM-as-a-Judge systems can be misled by structural formalism and consensus mimicry rather than semantic truth, based on 22,500 trajectories across multiple domains.

0 favorites 0 likes
#hallucination

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

Hugging Face Daily Papers ↗ · 2026-08-03 Cached

Introduces AD-MCQ and DEFT-RLVR, a method for verifiable reasoning in autonomous driving VLMs that defers future trajectory exposure to post-decision verification, improving reasoning faithfulness while reducing hallucinations.

0 favorites 0 likes
#hallucination

My personal AI benchmark: "Generate an SVG of a frog with a Habsburg jaw."

Hacker News Top ↗ · 2026-08-02 Cached

A personal AI benchmark asking models to generate an SVG of a frog with a Habsburg jaw reveals that the model adds elaborate editorializing, inventing royal status and medical commentary beyond the literal request.

0 favorites 0 likes
#hallucination

The EU AI Act makes failure to disclose AI-generated content (especially if it's hallucinated) illegal and costly.

Reddit r/artificial ↗ · 2026-08-02

Article 50 of the EU AI Act takes effect, requiring disclosure of AI-generated text on public-interest matters, with fines for non-compliance. The article highlights cases where consulting firms like PwC and Deloitte used hallucinated AI content and now face legal consequences.

0 favorites 0 likes
#hallucination

Exploring the "Dario and Amanda" Prompt

Reddit r/singularity ↗ · 2026-07-31 Cached

An exploration of a strange prompt that causes Claude Opus 5 to hallucinate and reproduce content resembling leaked private chats between Anthropic users and employees, raising questions about training data and AI behavior.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback