Tag
Introduces DecoEvo, a score-decoupled co-evolution method for LLM optimization in text space that jointly improves solver and rubric-generator skills without gold rubrics, achieving 2.8–5.0% relative gains over baselines across five benchmarks.
This paper presents the first systematic meta-evaluation of LLM-generated rubrics for reproducing experiments from research papers. It reformulates rubrics into a checklist format and evaluates generation settings both intrinsically (semantic similarity) and extrinsically (score alignment), finding that augmented settings improve downstream evaluation alignment but generated rubrics are often overly fine-grained and biased toward high scores.
Introduces FinResearchBench II, a benchmark for evaluating deep research agents' financial report quality using a scalable pipeline that automatically generates and filters rubrics via LLM consensus, enabling system differentiation without human experts.
This paper proposes a training-free method to automatically generate fine-grained evaluation rubrics for LLM-as-a-judge without human annotation, and further introduces an iterative fine-tuning strategy for a rubric generator that outperforms larger proprietary models.
SCOPE is a self-play framework for open-ended tasks that co-evolves a Challenger and Solver policy, achieving up to +10.4 points on benchmarks without external supervision.