hallucination-benchmark

Tag

Cards List
#hallucination-benchmark

One of the most interesting benchmarks and its implications for alignment

Reddit r/ArtificialInteligence ↗ · 2026-09-10

The article discusses how public benchmarking of hallucination in LLMs has driven rapid improvements and explores implications for alignment, emphasizing the need for benchmarks on transparency and honesty.

0 favorites 0 likes
#hallucination-benchmark

ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning

Hugging Face Daily Papers ↗ · 2026-06-12 Cached

ClinHallu is a benchmark for diagnosing and mitigating hallucinations in medical multimodal large language models by decomposing reasoning into visual recognition, knowledge recall, and reasoning integration stages, using trace-supervised fine-tuning to reduce errors.

0 favorites 0 likes
#hallucination-benchmark

ERRORQUAKE: Heavy-Tailed Error Severity Distributions in Open-Weight Large Language Models

arXiv cs.LG ↗ · 2026-06-05 Cached

The paper introduces Errorquake-10k, a benchmark for evaluating error severity in open-weight LLMs, showing that models with matched accuracy can have vastly different error severity distributions, and argues that severity should be reported alongside accuracy.

0 favorites 0 likes
← Back to home

Submit Feedback