Tag
The paper introduces SWARM, a multilingual human-annotated dataset for detecting Russian propaganda in search engine results across nine languages. It benchmarks various models, finding that content-level analysis with LLMs outperforms traditional methods for propaganda detection.
This paper introduces a factuality-specific annotation policy using failure-space analysis to improve human-anchored factuality evaluation for LLMs under limited annotation budgets, achieving significant efficiency gains on benchmark systems.
The paper presents UPHELD, a benchmark with extensive human annotations for evaluating conversational LLMs, and a Mixture-of-Judges framework that enhances evaluation accuracy by 30%.
This paper from Stony Brook University identifies 'Averaging Bias' in human faithfulness annotations for text summarization: global human labels correlate better with the average of per-sentence LLM judgments than with a strict conjunctive rule, meaning humans often label summaries as faithful even when they contain local factual errors.
PrivacyAlign introduces a human-annotated dataset and training framework for aligning LLM agents to respect contextual privacy norms, showing that frontier models still leak sensitive information and that human-grounded evaluation improves alignment.
This paper introduces Metric Match, a method for selecting a subset of samples for human annotation to estimate LLM judge reliability more efficiently, reducing annotation costs by 32.5% and achieving a win-rate of 0.838 against random selection.
This paper introduces PRECISE, an extension of Prediction-Powered Inference that combines a small set of human labels with a large set of LLM judgments to produce unbiased and variance-reduced estimates of ranking evaluation metrics like Precision@K. The method is validated on the ESCI benchmark and in a production A/B test, where it correctly identified the best system variant using only 100 human labels, confirmed by a +407 bps sales improvement.
This paper empirically examines when to interrupt autonomous AI agents during software execution, finding that affective-state thresholds saturate quickly, LLM judges achieve low F1 scores (0.17–0.40) at high cost, and human annotators themselves show near-chance agreement on intervention timing, making the construct unreliable as an optimization target.
This paper presents a large-scale audit of human annotation reporting in NLP from 2018-2025, showing inconsistent documentation of critical details but improvements over time, and provides a framework and recommendations for better reporting.