Tag
CruxBench is a new benchmark that evaluates LLMs on their ability to discover valuable information — decomposing forecasting questions into informative subquestions ("cruxes") graded by Value of Information. Evaluations on 293 forecasting questions show VOI strongly correlates with model capability (r=0.90), yet frontier models still barely beat a random-timing baseline.
This paper introduces an experimental protocol to measure open-ended LLM conformity, showing that wrong peer input degrades revision quality and that evaluators are not neutral when shown peer context, highlighting the need for anchor calibration.
Introduces DualEval, a framework that jointly calibrates model ability and item difficulty/sharpness to unify static benchmark and arena-style evaluation, enabling more reliable rankings and downstream applications like benchmark compression and anomaly detection.
RubricsTree proposes a scalable, expert-aligned evaluation framework for personal health agents using over 100 atomic Boolean rubrics, achieving up to 66% relative gains on HealthBench across Gemini, GPT, and Qwen model families.