Show HN: 2 Weeks of Hallucinate – The Photo Gallery
Summary
A photo gallery showcasing two weeks of AI-generated hallucinatory images, hosted on hallucinate.site.
View Cached Full Text
Cached at: 06/13/26, 02:38 PM
Similar Articles
AI is confidently wrong way more than people give it credit for, change my mind
A user shares concerns about AI models presenting thin or ambiguous data with the same confidence as well-supported findings, citing a case where a complaint appearing only twice in 200 comments was ranked as a top concern. The piece questions whether this is a fixable prompting issue or a fundamental limitation requiring manual verification.
Can Gemma and Qwen models catch hallucinations by looking at their own logprobs?
The author shares experiments using a custom WebUI to let Gemma and Qwen models inspect their own logprobs to detect hallucinations. Initial observations suggest that first-recall token probabilities can indicate uncertainty, though both models struggle to read their own logprobs.
A Grounded and Decomposed Framework for Relation-Level Hallucination Evaluation in Abstractive Summarization
This paper presents a grounded and decomposed framework for evaluating relation-level hallucinations in abstractive summarization, introducing a normalized Relation Hallucination Index (RHI) with linguistically informed relation extraction enhancements.
Unified Hallucination Fuzzing for Multimodal Large Language Models
This paper presents UniHall, a fine-grained hallucination benchmark with a unified taxonomy, and Self-Adaptive Multimodal Fuzzing (SAMF), a self-evolving stress-testing framework for multimodal LLMs. Experiments show SOTA models degrade significantly under fuzzing and reveal a helpfulness-hallucination trade-off.
REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment
This paper introduces REIN, an alignment framework that reduces hallucination in large reasoning models by training them to explicitly reflect before answering and to abstain when knowledge is insufficient. Experiments show consistent gains in selective accuracy and hallucination reduction across benchmarks.