PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?

Hugging Face Daily Papers Papers

Summary

This paper introduces PlaylistEval, an agentic framework for benchmarking video-language judges on day-scale videos, revealing that frontier models achieve only 75.4% accuracy and highlighting the need for multi-modal approaches.

Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Existing benchmarks cannot answer this. Their videos are typically only a few minutes long, many answer pairs can be separated from the transcript alone, and collecting human judgments does not scale to ultra-long videos. We introduce PlaylistEval, an agentic framework that builds video-language judge benchmarks over 100-hour playlist collection without human annotation. It automatically generates questions with paired answers whose differences are controlled by causal degradation, so that every pair demands retrieval across the collection. The resulting benchmark contains 630 pairs across seven domains spanning both static and dynamic knowledge, and on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781). Evaluating 17 omnimodal and multimodal models from eight families reveals that frontier judges reach only 75.4% pairwise accuracy, while open-source judge models perform far behind. We further show that both retrieval and final judgment depend on using multiple modalities, and that judge accuracy degrades as the playlist set grows. We release our pipeline, benchmark, and evaluation code at https://playlisteval.github.io.
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:14 AM

Paper page - PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?

Source: https://huggingface.co/papers/2609.34314

Abstract

Video-languagemodelsareincreasinglyusedasjudgesofvideounderstanding,bothforevaluatingmodeloutputsandfortrainingrewardmodels.Whethertheirjudgmentsremainreliablewhentheevidenceisburiedinday-longvideoshasyettobeestablished.Existingbenchmarkscannotanswerthis.Theirvideosaretypicallyonlyafewminuteslong,manyanswerpairscanbeseparatedfromthetranscriptalone,andcollectinghumanjudgmentsdoesnotscaletoultra-longvideos.WeintroducePlaylistEval,anagenticframeworkthatbuildsvideo-languagejudgebenchmarksover100-hourplaylistcollectionwithouthumanannotation.Itautomaticallygeneratesquestionswithpairedanswerswhosedifferencesarecontrolledbycausaldegradation,sothateverypairdemandsretrievalacrossthecollection.Theresultingbenchmarkcontains630pairsacrosssevendomainsspanningbothstaticanddynamicknowledge,andonastratifiedsubsetof152pairsitagreeswithhumanjudgments93.0%ofthetime(IAA0.781).Evaluating17omnimodalandmultimodalmodelsfromeightfamiliesrevealsthatfrontierjudgesreachonly75.4%pairwiseaccuracy,whileopen-sourcejudgemodelsperformfarbehind.Wefurthershowthatbothretrievalandfinaljudgmentdependonusingmultiplemodalities,andthatjudgeaccuracydegradesastheplaylistsetgrows.Wereleaseourpipeline,benchmark,andevaluationcodeathttps://playlisteval.github.io.

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2609\.34314

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.34314 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.34314 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.34314 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles