open-ended-evaluation

Tag

Cards List
#open-ended-evaluation

CruxBench: A Benchmark of Information Discovery

arXiv cs.CL ↗ · 3d ago Cached

CruxBench is a new benchmark that evaluates LLMs on their ability to discover valuable information — decomposing forecasting questions into informative subquestions ("cruxes") graded by Value of Information. Evaluations on 293 forecasting questions show VOI strongly correlates with model capability (r=0.90), yet frontier models still barely beat a random-timing baseline.

0 favorites 0 likes
#open-ended-evaluation

The Evaluator Is Part of the Experiment: Measuring Open-Ended LLM Conformity

arXiv cs.CL ↗ · 2026-08-06 Cached

This paper introduces an experimental protocol to measure open-ended LLM conformity, showing that wrong peer input degrades revision quality and that evaluators are not neutral when shown peer context, highlighting the need for anchor calibration.

0 favorites 0 likes
#open-ended-evaluation

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation

arXiv cs.LG ↗ · 2026-06-26 Cached

Introduces DualEval, a framework that jointly calibrates model ability and item difficulty/sharpness to unify static benchmark and arena-style evaluation, enabling more reliable rankings and downstream applications like benchmark compression and anomaly detection.

0 favorites 0 likes
#open-ended-evaluation

RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills

arXiv cs.CL ↗ · 2026-06-17 Cached

RubricsTree proposes a scalable, expert-aligned evaluation framework for personal health agents using over 100 atomic Boolean rubrics, achieving up to 66% relative gains on HealthBench across Gemini, GPT, and Qwen model families.

0 favorites 0 likes
← Back to home

Submit Feedback