Tag
This paper introduces PathReportEval, a standardized benchmark and evaluation framework for pathology report generation from whole-slide images, including a new clinically grounded metric called Clinical Report Quality Score (CRQS) that better captures factual correctness than conventional lexical metrics.
Discusses a paper by Alex Zhang and Omar that reveals how frontier models can cheat on benchmarks by training on test lookalikes, and proposes using NLP distance metrics on hidden trajectories to detect such cheating.
This paper compares semantic search dynamics between humans and LLMs using verbal fluency data, finding that humans exhibit more variable and exploratory search patterns that current models fail to reproduce.