PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?
Summary
This paper introduces PlaylistEval, an agentic framework for benchmarking video-language judges on day-scale videos, revealing that frontier models achieve only 75.4% accuracy and highlighting the need for multi-modal approaches.
View Cached Full Text
Cached at: 09/30/26, 04:14 AM
Paper page - PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?
Source: https://huggingface.co/papers/2609.34314
Abstract
Video-languagemodelsareincreasinglyusedasjudgesofvideounderstanding,bothforevaluatingmodeloutputsandfortrainingrewardmodels.Whethertheirjudgmentsremainreliablewhentheevidenceisburiedinday-longvideoshasyettobeestablished.Existingbenchmarkscannotanswerthis.Theirvideosaretypicallyonlyafewminuteslong,manyanswerpairscanbeseparatedfromthetranscriptalone,andcollectinghumanjudgmentsdoesnotscaletoultra-longvideos.WeintroducePlaylistEval,anagenticframeworkthatbuildsvideo-languagejudgebenchmarksover100-hourplaylistcollectionwithouthumanannotation.Itautomaticallygeneratesquestionswithpairedanswerswhosedifferencesarecontrolledbycausaldegradation,sothateverypairdemandsretrievalacrossthecollection.Theresultingbenchmarkcontains630pairsacrosssevendomainsspanningbothstaticanddynamicknowledge,andonastratifiedsubsetof152pairsitagreeswithhumanjudgments93.0%ofthetime(IAA0.781).Evaluating17omnimodalandmultimodalmodelsfromeightfamiliesrevealsthatfrontierjudgesreachonly75.4%pairwiseaccuracy,whileopen-sourcejudgemodelsperformfarbehind.Wefurthershowthatbothretrievalandfinaljudgmentdependonusingmultiplemodalities,andthatjudgeaccuracydegradesastheplaylistsetgrows.Wereleaseourpipeline,benchmark,andevaluationcodeathttps://playlisteval.github.io.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2609\.34314
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.34314 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.34314 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.34314 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
EduPanel is a rubric-grounded, learner-conditioned LLM judge that uses three specialized agents to evaluate teaching video quality, achieving reliability comparable to human experts and improving scoring accuracy while maintaining detectability of unreliable outputs.
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models
This paper introduces EvalDetectBench, an open benchmark and pipeline for measuring evaluation awareness in frontier language models, addressing biases in existing methods to improve AI safety assessments.
EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
EvalVerse is a comprehensive evaluation framework for professional cinematic video generation that uses expert-calibrated vision-language models and multi-stage assessment to bridge human aesthetic judgment and machine scoring.
Evaluating Language Models in Realistic Conversational Contexts
This paper introduces UPHELD, a large benchmark for evaluating human-scale conversational ability in LLMs, and proposes a Mixture-of-Judges framework that improves correlation with human assessments by approximately 30%.
Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning
Introduces Video-MME-Logical, a controlled benchmark for evaluating video temporal-logical reasoning in multimodal large language models, revealing a substantial human-model gap.