标签
This paper introduces an experimental protocol to measure open-ended LLM conformity, showing that wrong peer input degrades revision quality and that evaluators are not neutral when shown peer context, highlighting the need for anchor calibration.
介绍了DualEval框架,该框架联合校准模型能力与项目难度/锐度,以统一静态基准和竞技场式评估,从而实现更可靠的排名以及基准压缩和异常检测等下游应用。
RubricsTree 提出了一种可扩展且与专家对齐的个人健康智能体评估框架,使用超过100个原子布尔规则,在Gemini、GPT和Qwen模型系列的HealthBench上实现了高达66%的相对提升。