Tag
The paper proposes information satisfaction as a reader-centered axis for summarization evaluation, showing that current metrics fail to capture user-specific informational needs and agree poorly with human judgment.
This paper introduces EuroExec, a human-expert benchmark for evaluating frontier LLMs on open-ended European executive decision tasks. It finds that the strongest model solves only 56.9% of tasks, falling well short of expert-written reference answers, highlighting gaps in real-world open-ended problem-solving.
This paper formalizes multi-model human evaluation as a best-arm identification problem in a multi-armed bandit setup, adaptively allocating annotation effort to focus on competitive models and improve ranking discrimination.
FriendBench is a new benchmark for evaluating whether humans and multimodal LLMs can infer if two people are familiar or strangers from a 20-second video clip of an ice-breaker conversation. Results show the best models match human accuracy but differ in bias, and only humans benefit from richer visual behavior.
This paper presents a controlled comparison of span-guided and unguided text detoxification via human evaluation, finding a trade-off: span-guided rewriting is preferred when preserving original stance is important, while unguided rewriting is favored for more complete mitigation, especially on milder toxicity. The study also assesses automatic evaluators as diagnostics rather than substitutes for human judgment.
Introduces Contrastive Error Span Annotation (cESA), a protocol for human evaluation of multiple translations simultaneously, reducing annotation time and noise compared to standard methods.
A large-scale human evaluation of 10 frontier LLMs across 62 questions finds that all models exhibit response drift, with most converging to a 78-81% deviation ceiling, while two achieve lower deviation. Drift varies by domain and question, and automated metrics explain little of human judgments, highlighting the need for human evaluation.
Real World VoiceEQ is a new benchmark for evaluating the human quality of voice AI, based on over a million human ratings, assessing models across speech recognition, synthesis, and understanding in real-world conditions.
A study comparing human and AI translations of literary works shows that while machine translations are deemed 'fine', readers still prefer human translations for their immersiveness and clarity. Automatic metrics fail to capture reader preferences.
The author praises GLM-5.2, an MIT open-weights model, for its exceptional real-world performance in human evaluation benchmarks, claiming it rivals the best closed-source models like those from Claude.
This exploratory study evaluates whether augmenting AI agents with a medical research skill package improves the quality of transcriptomic research analysis outputs compared to native AI, using a multi-model human evaluation in an NSCLC biomarker task. Results show a directional but statistically non-significant improvement, highlighting the need for larger, more robust evaluations.
This paper introduces RQ-Bench, a benchmark to evaluate LLMs' ability to assess the novelty of scientific research questions. It finds that LLM judges consistently rate generated questions as more novel than human experts do, raising concerns about the reliability of using LLMs for scientific novelty evaluation.
This paper investigates the effectiveness of LLM personalization by putting real humans back into the evaluation loop, revealing systematic gaps between human judgments and LLM outputs at every stage of the personalization pipeline, and highlighting the limitations of synthetic data and LLM judges.
This paper proposes a two-stage sampling design where LLM evaluations are used to augment, rather than replace, human ratings, and provides guidance on determining sample sizes for human and LLM reviews using a doubly robust estimator from missing data literature.
A human review of TranslateGemma-12b's translations revealed that 71% of segments rated clean by automated metrics actually contained errors, highlighting significant gaps in metric-only evaluation for multilingual translation quality.
OpenAI researchers found that optimizing language models purely for correct answers reduces human interpretability, and propose 'prover-verifier games' where a prover generates solutions and a verifier checks them, improving legibility for both humans and AI systems.