metrology

Tag

Cards List
#metrology

LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation

arXiv cs.CL · 2026-06-16 Cached

This paper introduces a psychometric datasheet protocol for evaluating LLM judges as measurement instruments, measuring dark current, positional false preference, stable cross-sensitivity, and target sensitivity. A case study on three open-weight models reveals significant differences in judge quality and behavior.

0 favorites 0 likes
#metrology

Your Evals Will Break and You Won't See It Coming

Reddit r/ArtificialInteligence · 2026-05-19 Cached

Discusses the structural weakness of current evaluation methods for LLMs, which fail to anticipate qualitative shifts in capability, and argues that developing proactive evaluation infrastructure is the critical bottleneck for safe capability jumps.

0 favorites 0 likes
← Back to home

Submit Feedback