rubric-generation

Tag

Cards List
#rubric-generation

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

Hugging Face Daily Papers · 2026-07-28 Cached

Introduces DecoEvo, a score-decoupled co-evolution method for LLM optimization in text space that jointly improves solver and rubric-generator skills without gold rubrics, achieving 2.8–5.0% relative gains over baselines across five benchmarks.

0 favorites 0 likes
#rubric-generation

Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction

arXiv cs.CL · 2026-07-15 Cached

This paper presents the first systematic meta-evaluation of LLM-generated rubrics for reproducing experiments from research papers. It reformulates rubrics into a checklist format and evaluates generation settings both intrinsically (semantic similarity) and extrinsically (score alignment), finding that augmented settings improve downstream evaluation alignment but generated rubrics are often overly fine-grained and biased toward high scores.

0 favorites 0 likes
#rubric-generation

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality

arXiv cs.CL · 2026-07-15 Cached

Introduces FinResearchBench II, a benchmark for evaluating deep research agents' financial report quality using a scalable pipeline that automatically generates and filters rubrics via LLM consensus, enabling system differentiation without human experts.

0 favorites 0 likes
#rubric-generation

Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge

arXiv cs.CL · 2026-06-01 Cached

This paper proposes a training-free method to automatically generate fine-grained evaluation rubrics for LLM-as-a-judge without human annotation, and further introduces an iterative fine-tuning strategy for a rubric generator that outperforms larger proprietary models.

0 favorites 0 likes
#rubric-generation

SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks

Hugging Face Daily Papers · 2026-05-29 Cached

SCOPE is a self-play framework for open-ended tasks that co-evolves a Challenger and Solver policy, achieving up to +10.4 points on benchmarks without external supervision.

0 favorites 0 likes
← Back to home

Submit Feedback