human-annotation

Tag

Cards List
#human-annotation

SWARM: A Multilingual Human-Annotated Dataset for Russian Propaganda Detection in Search Engine Results

arXiv cs.CL · 2026-09-14 Cached

The paper introduces SWARM, a multilingual human-annotated dataset for detecting Russian propaganda in search engine results across nine languages. It benchmarks various models, finding that content-level analysis with LLMs outperforms traditional methods for propaganda detection.

0 favorites 0 likes
#human-annotation

Human-Anchored Factuality Evaluation with Strategic Annotation

arXiv cs.CL · 2026-09-02 Cached

This paper introduces a factuality-specific annotation policy using failure-space analysis to improve human-anchored factuality evaluation for LLMs under limited annotation budgets, achieving significant efficiency gains on benchmark systems.

0 favorites 0 likes
#human-annotation

@dair_ai: // Your LLM judge disagrees with the experts // LLM Judges can be tricky to build. Here is an interesting showcasing wh…

X AI KOLs Timeline · 2026-08-30 Cached

The paper presents UPHELD, a benchmark with extensive human annotations for evaluating conversational LLMs, and a Mixture-of-Judges framework that enhances evaluation accuracy by 30%.

0 favorites 0 likes
#human-annotation

Averaging Bias: Human Faithfulness Annotations are not Locally Faithful

arXiv cs.CL · 2026-08-04 Cached

This paper from Stony Brook University identifies 'Averaging Bias' in human faithfulness annotations for text summarization: global human labels correlate better with the average of per-sentence LLM judgments than with a strict conjunctive rule, meaning humans often label summaries as faithful even when they contain local factual errors.

0 favorites 0 likes
#human-annotation

PrivacyAlign: Contextual Privacy Alignment for LLM Agents

Hugging Face Daily Papers · 2026-06-19 Cached

PrivacyAlign introduces a human-annotated dataset and training framework for aligning LLM agents to respect contextual privacy norms, showing that frontier models still leak sensitive information and that human-grounded evaluation improves alignment.

0 favorites 0 likes
#human-annotation

Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability

arXiv cs.AI · 2026-06-16 Cached

This paper introduces Metric Match, a method for selecting a subset of samples for human annotation to estimate LLM judge reliability more efficiently, reducing annotation costs by 32.5% and achieving a win-rate of 0.838 against random selection.

0 favorites 0 likes
#human-annotation

Statistically Reliable LLM-Based Ranking Evaluation via Prediction-Powered Inference

arXiv cs.LG · 2026-06-05 Cached

This paper introduces PRECISE, an extension of Prediction-Powered Inference that combines a small set of human labels with a large set of LLM judgments to produce unbiased and variance-reduced estimates of ranking evaluation metrics like Precision@K. The method is validated on the ESCI benchmark and in a production A/B test, where it correctly identified the best system variant using only 100 human labels, confirmed by a +407 bps sales improvement.

0 favorites 0 likes
#human-annotation

The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents

arXiv cs.AI · 2026-06-04 Cached

This paper empirically examines when to interrupt autonomous AI agents during software execution, finding that affective-state thresholds saturate quickly, LLM judges achieve low F1 scores (0.17–0.40) at high cost, and human annotators themselves show near-chance agreement on intervention timing, making the construct unreliable as an optimization target.

0 favorites 0 likes
#human-annotation

Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025

Hugging Face Daily Papers · 2026-06-01 Cached

This paper presents a large-scale audit of human annotation reporting in NLP from 2018-2025, showing inconsistent documentation of critical details but improvements over time, and provides a framework and recommendations for better reporting.

0 favorites 0 likes
← Back to home

Submit Feedback