Tag
This study examines how AI raters (LLMs) score clinical AI outputs under different protocols in complex type 2 diabetes pharmacotherapy, finding that rubric-anchored scoring provides greater discriminative power than rubric-free scoring.
MIRA is a data selection framework for the mid-training stage of LLM development that adaptively constructs quality rubrics per data source, using a teacher model to propose dimensions and distilling into lightweight scorers. It achieves superior performance using only half the tokens compared to full-corpus training.