Tag
This paper introduces RM-EVAL, a reward model trained on human preference data for reference-free meta-evaluation of grammatical error correction, and shows how it can improve GEC systems via reward-guided text generation.
This paper investigates agentic data cleaning without a clean reference, proposing an evidence-grounded framework and evaluating trade-offs across multiple configurations.
Introduces Q-CARE, a query-agnostic and reference-free framework for fine-grained RAG evaluation that decomposes queries into sub-queries and answers into atomic claims, achieving higher correlation with human judgment than existing metrics like RAGEval and RAGChecker.
This paper shows that LLM judges tend to over-credit incorrect answers when no reference answer is provided, and adding a reference can flip verdicts by up to 85%, aligning more with human judgments. The authors propose calibration steps for using LLM judges in reference-free settings.
This paper identifies a structural flaw in reference-free LLM judges used in self-play training, showing they score plausibility rather than correctness, leading to reward hacking where policies learn to produce plausible-but-wrong answers. The authors propose a hidden-anchor audit and a de-anchored reward to mitigate this issue.
Introduces PoQ-Judge, a multi-architecture evaluation framework with reference-free judge models (TextCNN, MiniLM, DeBERTa) for cost-aware Proof-of-Quality in decentralized LLM inference, achieving high correlation with ground-truth proxies while eliminating the need for reference answers.
Granuscore is a reference-free measure of granularity for text analysis and question answering. It uses hierarchical embedding spaces to capture fine-grained vs. coarse language and demonstrates consistent differences in model behavior across QA benchmarks.
This paper applies Group Relative Policy Optimization (GRPO) to encoder-decoder Seq2Seq models for machine translation fine-tuning, using reference-free rewards (LaBSE and COMET-Kiwi) that require no parallel data, and achieves consistent improvements across 13 languages.