llm-assessment

Tag

Cards List
#llm-assessment

FakeSpotter: A content and strategy agnostic Viral Misinformation Detection Tool

arXiv cs.CL ↗ · 2026-09-18 Cached

FakeSpotter is a content- and strategy-agnostic tool designed to estimate the viral misinformation risk of textual content by measuring structural fingerprints, using repeated LLM assessments and logistic regression classifiers, with reported macro F1 scores of 0.788 for short texts and 0.793 for long texts on a labeled corpus.

0 favorites 0 likes
#llm-assessment

Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment

arXiv cs.CL ↗ · 2026-05-12 Cached

This paper investigates whether standard benchmarks underestimate LLM performance by re-evaluating hallucination detection datasets using an LLM-first, human-adjudicated assessment method. The study finds that incorporating LLM reasoning into the adjudication process improves agreement and suggests that model-assisted re-evaluation yields more reliable benchmarks for ambiguity-prone tasks.

0 favorites 0 likes
← Back to home

Submit Feedback