nlp-metrics

Tag

Cards List
#nlp-metrics

PathReportEval: A Systematic Benchmark for Pathology Report Generation

arXiv cs.CL · 4d ago Cached

This paper introduces PathReportEval, a standardized benchmark and evaluation framework for pathology report generation from whole-slide images, including a new clinically grounded metric called Clinical Report Quality Score (CRQS) that better captures factual correctness than conventional lexical metrics.

0 favorites 0 likes
#nlp-metrics

@swyx: very notable trajectory comparison writeup here buried in the RLM paper from @a1zhang and @lateinteraction. an open sec…

X AI KOLs Following · 5d ago Cached

Discusses a paper by Alex Zhang and Omar that reveals how frontier models can cheat on benchmarks by training on test lookalikes, and proposes using NLP distance metrics on hidden trajectories to detect such cheating.

0 favorites 0 likes
#nlp-metrics

Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing

arXiv cs.CL · 2026-07-15 Cached

This paper compares semantic search dynamics between humans and LLMs using verbal fluency data, finding that humans exhibit more variable and exploratory search patterns that current models fail to reproduce.

0 favorites 0 likes
← Back to home

Submit Feedback