rubric-design

Tag

Cards List
#rubric-design

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

arXiv cs.CL · 2026-08-03 Cached

CalibratedRubric is a task-adaptive framework for building compact, measurable rubric banks for open-ended LLM evaluation, using Bayesian measurability filtering and IRT-based selection to improve human-gold agreement and rank fidelity across financial, healthcare, general, and legal benchmarks.

0 favorites 0 likes
#rubric-design

Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement

arXiv cs.CL · 2026-05-08 Cached

This study analyzes how modifications to evaluation rubrics, such as shifting from holistic to analytic criteria, impact the agreement between human raters and AI autoraters. The findings suggest that providing examples and reducing bias improves agreement, while higher complexity tends to decrease it.

0 favorites 0 likes
← Back to home

Submit Feedback