Content Depth Matters in Short-Video Recommendation: Rethinking the Attention Economy

arXiv cs.AI Papers

Summary

This paper introduces Content Depth Score (CDS) and SCOPE-Bench to evaluate short-video recommender systems for cognitive depth, revealing that existing systems prioritize shallow content over deeper engagement.

arXiv:2608.13990v1 Announce Type: new Abstract: Driven by the attention economy, short-video Recommender Systems (RSs) are primarily optimized to maximize user engagement by promoting videos that capture attention within seconds. These systems inherently favor shallow-content videos that are effective at attracting immediate attention. However, growing evidence suggests that prolonged exposure to such content may negatively affect users' cognitive engagement and mental well-being, raising concerns about the long-term societal impact of the short-video platform. To tackle this challenge, this paper introduces a new metric, the \textbf{Content Depth Score (CDS)}, to quantify the content depth of short videos. CDS measures the extent to which a video is expected to stimulate higher-order cognitive processes, using a seven-level scale grounded in established theories of cognitive psychology and learning. As an initial step toward this vision, we present \textbf{SCOPE-Bench}, the first benchmark for content-depth evaluation in short-video recommendation. Built upon a large-scale open-source short-video dataset, SCOPE-Bench provides CDS annotations for 150K videos, enabling systematic evaluation of RSs from a cognitive-content perspective. Leveraging SCOPE-Bench, we evaluate 13 representative RSs and reveal a consistent preference for shallow-content videos. Moreover, we find that these algorithms recommending cognitively deep content are only marginally better than random selection, highlighting a previously overlooked limitation of existing recommendation objectives. Our code and datasets are available at https://liweidengdavid.github.io/SCOPE-Bench/.
Original Article
View Cached Full Text

Cached at: 08/17/26, 10:00 AM

# Content Depth Matters in Short-Video Recommendation: Rethinking the Attention Economy
Source: [https://arxiv.org/html/2608.13990](https://arxiv.org/html/2608.13990)
Jing JiangZhiwei LiThanks:Corresponding author: Zhiwei Li \(zhw\.li@outlook\.com\)\.Yang WangGuodong Long

###### Abstract

Driven by the attention economy, short\-video Recommender Systems \(RSs\) are primarily optimized to maximize user engagement by promoting videos that capture attention within seconds\. These systems inherently favor shallow\-content videos that are effective at attracting immediate attention\. However, growing evidence suggests that prolonged exposure to such content may negatively affect users’ cognitive engagement and mental well\-being, raising concerns about the long\-term societal impact of the short\-video platform\. To tackle this challenge, this paper introduces a new metric, theContent Depth Score \(CDS\), to quantify the content depth of short videos\. CDS measures the extent to which a video is expected to stimulate higher\-order cognitive processes, using a seven\-level scale grounded in established theories of cognitive psychology and learning\. As an initial step toward this vision, we presentSCOPE\-Bench, the first benchmark for content\-depth evaluation in short\-video recommendation\. Built upon a large\-scale open\-source short\-video dataset, SCOPE\-Bench provides CDS annotations for 150K videos, enabling systematic evaluation of RSs from a cognitive\-content perspective\. Leveraging SCOPE\-Bench, we evaluate 13 representative RSs and reveal a consistent preference for shallow\-content videos\. Moreover, we find that these algorithms recommending cognitively deep content are only marginally better than random selection, highlighting a previously overlooked limitation of existing recommendation objectives\. Our code and datasets are available athttps://liweidengdavid\.github\.io/SCOPE\-Bench/\.

## 1Introduction

In the attention economy, short\-video feeds on platforms such as YouTube\([Ørmen and Gregersen 2023](https://arxiv.org/html/2608.13990#bib.bib35)\), TikTok\([TikTok Team 2021](https://arxiv.org/html/2608.13990#bib.bib43)\), and Kuaishou\([Kuaishou Technology 2025](https://arxiv.org/html/2608.13990#bib.bib26)\)have become a central channel for everyday content consumption\. Adult users now spend more than one hour per day on these platforms\([Kemp 2025](https://arxiv.org/html/2608.13990#bib.bib23)\)\. Their Recommender Systems \(RSs\)\([Zou and Sun 2025](https://arxiv.org/html/2608.13990#bib.bib62);[Li et al\. 2026](https://arxiv.org/html/2608.13990#bib.bib27)\)optimize engagement signals, such as watch time and clicks, to turn user attention into revenue\([García\-Rapp 2017](https://arxiv.org/html/2608.13990#bib.bib7)\), and therefore favor videos capable of rapidly capturing users’ attention\. Over time, continuous exposure to such content may weaken users’ sustained attention\([Mahakud and Thapliyal 2026](https://arxiv.org/html/2608.13990#bib.bib31)\)and affect their long\-term well\-being\([Tang et al\. 2026](https://arxiv.org/html/2608.13990#bib.bib42)\)\.

Figure 1:Evaluating 13 Recommender Systems \(RSs\) on both Engagement metric \(X axis\) and our Content\-Depth metric \(Y axis\)\. The regions are split into three levels corresponding to content\-depth metric: Low \(blue\), Medium \(green\), and High \(red\)\. The dashed line marks the content depth of a random selection baseline\. Existing RSs perform competitively on the engagement metric while their content\-depth metric is low and near to the random selection\.Optimizing for engagement rewards videos that retain user attention, which raises a more fundamental question:Do RSs favor attention\-grabbing content over videos with greater depth?Content depth reflects how deeply a video develops its information, providing users with richer opportunities for understanding, reasoning, and reflection\([Chi and Wylie 2014](https://arxiv.org/html/2608.13990#bib.bib6)\)\. We examine this question by comparing the engagement of 13 representative RSs with the content depth of the videos they recommend\. As Figure[1](https://arxiv.org/html/2608.13990#S1.F1)shows, the content\-depth performance of every RS remains close to that of random recommendation\. Therefore, stronger engagement does not translate into greater recommended content depth\.

The situation has already drawn awareness and responses beyond research\. For example, Australia\([Australian Government 2025](https://arxiv.org/html/2608.13990#bib.bib3)\)and the UK\([UK Government 2026](https://arxiv.org/html/2608.13990#bib.bib44)\)have imposed under\-16 restrictions on short\-video platforms, and platforms themselves cap adolescents’ daily usage\([Keenan 2023](https://arxiv.org/html/2608.13990#bib.bib22)\)\. This convergence of regulators signals that the harm to young users is now widely acknowledged\. Yet these interventions act on access rather than content\. One important reason is the absence of any measure for the content itself\.

Inspired by the need to promote healthier short video and build a sustainable short\-video ecosystem, we propose a new metric to make content depth measurable, enabling it to serve as an explicit objective in RS evaluation and optimization\. This paper proposes theContentDepthScore \(CDS\), a novel metric to measure how deeply a short video presents and delivers information to humans\. CDS scores a video on a seven\-level rubric grounded in theories of cognition and learning\([Wason and Evans 1974](https://arxiv.org/html/2608.13990#bib.bib49);[Anderson and Krathwohl 2001](https://arxiv.org/html/2608.13990#bib.bib2);[Biggs and Collis 2014](https://arxiv.org/html/2608.13990#bib.bib5)\), ranging from low\-level emotional stimulation to higher\-order cognitive processes111Specifically, in this paper, cognitive processes refer to the mental operations through which individuals acquire, process, store, and use information\([Sternberg, Sternberg, and Mio 2006](https://arxiv.org/html/2608.13990#bib.bib41)\)\. CDS measures it as the opportunities a video’s content provides for these operations\.\. To validate CDS at both the item and recommendation\-list levels, we constructSCOPE\-Bench, aShort\-videoCOntent dePthEvaluationBenchmark, by extending the existing open\-source ShortVideo dataset\([Shang et al\. 2025](https://arxiv.org/html/2608.13990#bib.bib40)\)\. Specifically, we annotate its 150K videos with CDS labels, covering about 1M user\-item interactions from 10K users\.

Our main contributions are summarized as follows:

- •To the best of our knowledge, we are the first to formulate and systematically investigate the lack of attention to content depth in engagement\-optimized RSs\.
- •We propose CDS, the first quantitative metric for video content depth, based on a seven\-level rubric grounded in theories of cognition and learning\. We further extended CDS to a list\-level metric for recommended lists\.
- •We establish a comprehensive evaluation framework for video content depth and construct SCOPE\-Bench, a benchmark of 150K videos with human\-aligned CDS annotations and about 1M interactions from 10K users\.
- •Experiments on SCOPE\-Bench confirm that the CDS evaluation protocol closely aligns with human judgments\. Using CDS to evaluate 13 representative RSs, we find that their recommended lists achieve content\-depth performance close to that of random recommendation, revealing that engagement and content depth are decoupled\.

## 2Related Works

### 2\.1Content Quality Assessment

Content quality assessment can be broadly categorized into*metric\-based assessment*for explicitly quantifiable properties and*judgment\-based assessment*for open\-ended or interpretive properties that are difficult to capture with conventional metrics\. The former mainly coversvideo quality assessment, including perceptual fidelity\([Wang et al\. 2004](https://arxiv.org/html/2608.13990#bib.bib47)\), temporal consistency\([Wang, Lu, and Bovik 2004](https://arxiv.org/html/2608.13990#bib.bib48)\), generative realism\([Heusel et al\. 2017](https://arxiv.org/html/2608.13990#bib.bib17)\), and cross\-modal alignment\([Radford et al\. 2021](https://arxiv.org/html/2608.13990#bib.bib38)\), as well astext quality assessment, covering linguistic\([Napoles, Sakaguchi, and Tetreault 2017](https://arxiv.org/html/2608.13990#bib.bib33)\)and semantic properties\([Barzilay and Lapata 2008](https://arxiv.org/html/2608.13990#bib.bib4)\), trustworthiness\([Lin, Hilton, and Evans 2022](https://arxiv.org/html/2608.13990#bib.bib28)\), and safety\-related dimensions\([Gehman et al\. 2020](https://arxiv.org/html/2608.13990#bib.bib9)\)\. For open\-ended and interpretive properties, recent studies have increasingly adoptedLLM\-as\-a\-Judgeas a flexible judgment\-based evaluation paradigm\. Existing approaches range from scalar scoring\([Liu et al\. 2023](https://arxiv.org/html/2608.13990#bib.bib29)\)and pairwise comparison\([Zheng et al\. 2023](https://arxiv.org/html/2608.13990#bib.bib58)\)to rubric\-based evaluation\([Kim et al\. 2024](https://arxiv.org/html/2608.13990#bib.bib24);[Wang et al\. 2024](https://arxiv.org/html/2608.13990#bib.bib46)\)and fine\-grained assessment protocols\([Ye et al\. 2023](https://arxiv.org/html/2608.13990#bib.bib53)\)\. While effective, these methods provide only a limited conceptualization of content depth\. In the paper, we systematically define content depth and propose the first quantitative metric, termed the CDS, to measure the content depth\.

### 2\.2Evaluation of Recommender Systems

Existing evaluation of RSs can be broadly categorized into*system\-centric evaluation*and*user\-centric evaluation*\. System\-centric evaluation coversaccuracy\-oriented evaluationandbeyond\-accuracy evaluation\([Zangerle and Bauer 2022](https://arxiv.org/html/2608.13990#bib.bib56)\)\. The former assesses whether relevant items are accurately retrieved and ranked using metrics such as Precision\([Herlocker et al\. 2004](https://arxiv.org/html/2608.13990#bib.bib16)\), Recall\([Allen et al\. 1955](https://arxiv.org/html/2608.13990#bib.bib1)\), and NDCG\([Järvelin and Kekäläinen 2002](https://arxiv.org/html/2608.13990#bib.bib20)\), whereas the latter considers complementary recommendation qualities, including diversity\([Ziegler et al\. 2005](https://arxiv.org/html/2608.13990#bib.bib61)\), novelty\([Vargas and Castells 2011](https://arxiv.org/html/2608.13990#bib.bib45)\), and catalog coverage\([Ge, Delgado\-Battenfeld, and Jannach 2010](https://arxiv.org/html/2608.13990#bib.bib8)\)\. In contrast,user\-centric evaluationexamines users’ subjective perceptions and experiences, including choice difficulty\([Knijnenburg et al\. 2012](https://arxiv.org/html/2608.13990#bib.bib25)\), usefulness\([Pu, Chen, and Hu 2011](https://arxiv.org/html/2608.13990#bib.bib36)\), and satisfaction\([Hijikata, Kai, and Nishida 2012](https://arxiv.org/html/2608.13990#bib.bib18)\)\. Although prior studies have attempted to evaluate RSs beyond conventional engagement metrics, little research has quantitatively examined the content depth of recommended videos\. To address this gap, we introduce the List\-wise CDS \(LCDS\), the first list\-level metric for quantifying the content depth of recommended lists\.

## 3Content Depth Score Metric

CDSLevelRubricLabelOperationalCriterionDualProcessTheoryBloom’sTaxonomySOLOTaxonomyLevel 0AffectMainly evokes affect, humor, spectacle, or atmosphere\.System 1––Level 1PointPresents an isolated opinion, or label without explaining why or how\.System 2RememberPrestructuralLevel 2ConceptDefines or illustrates a single idea with simple background or explanation\.System 2UnderstandUnistructuralLevel 3ProcedureShows how a method can be used in concrete cases, typically through multiple steps or illustrative examples\.System 2ApplyMultistructuralLevel 4MechanismExplains mechanisms, variables, constraints, conditions, causal links, or system relationships\.System 2AnalyzeRelationalLevel 5JudgmentWeighs evidence or competing explanations, including limitations, uncertainty, or counterexamples\.System 2EvaluateRelational /Extended AbstractLevel 6ModelBuilds a generalizable model, framework, principle, or decision rule that can transfer across contexts\.System 2CreateExtended AbstractTable 1:Seven\-level scoring rubric for CDS and its approximate theoretical anchors\. Each level is defined by an operational criterion for consistent annotation\. A dash indicates that no direct theoretical correspondence is assigned\. The rubric captures a progression from affective responses to increasingly complex reasoning and transferable model construction\.### 3\.1Content Depth

To measure content depth, we begin by defining it precisely\.

###### Definition 1\(Content Depth\)\.

Content depth of a video is the degree to which it develops a topic from isolated information into structured understanding by explaining concepts, demonstrating procedures, analyzing mechanisms, forming evaluative judgments and generalizable insights\.

Based on Definition[1](https://arxiv.org/html/2608.13990#Thmdefinition1), we propose CDS, the first metric to quantify the content depth of short videos\. A video receives a higher CDS when it presents richer semantic information, clearer explanations, and higher\-order reasoning structures\. Content depth is cognitively meaningful because deeper content offers viewers richer opportunities for higher\-order cognitive processes\([Anderson and Krathwohl 2001](https://arxiv.org/html/2608.13990#bib.bib2)\), which support cognitive development and intellectual growth\.

### 3\.2Seven\-level Scoring Rubric

Theoretical Foundation\.We operationalize CDS as a seven\-level rubric grounded in Dual Process Theory\([Wason and Evans 1974](https://arxiv.org/html/2608.13990#bib.bib49)\), the Revised Bloom’s Taxonomy222For simplicity, we refer to it as “Bloom’s Taxonomy” hereafter\.\([Anderson and Krathwohl 2001](https://arxiv.org/html/2608.13990#bib.bib2)\), and the SOLO Taxonomy\([Biggs and Collis 2014](https://arxiv.org/html/2608.13990#bib.bib5)\)\. Specifically, Dual Process Theory distinguishes between immediate affective and deliberate processing\. Bloom’s Taxonomy organizes cognitive processes into six levels of increasing complexity, whereas the SOLO Taxonomy classifies the degree of understanding demonstrated by learners into five levels\. Based on these complementary perspectives, we use System 1 in Dual Process Theory to describe the lowest level and the six cognitive levels of Bloom’s Taxonomy to define the remaining levels\. We further use the SOLO Taxonomy as complementary guidance for distinguishing the complexity of understanding across levels\. The resulting rubric ranges from immediate affective responses to increasingly complex cognitive processes\.

Rubric Structure\.Table[1](https://arxiv.org/html/2608.13990#S3.T1)presents the full rubric and its approximate theoretical anchors, grouping the seven levels into three groups\. Low\-CDS group \(Level 0\) represents content that mainly elicits affective or entertaining responses, with limited explicit knowledge value\. Medium\-CDS group \(Levels 1–3\) captures progressively more structured knowledge transmission, moving from isolated information to conceptual explanation and practical application\. High\-CDS group \(Levels 4–6\) reflects higher\-order cognitive processes, ranging from analytical reasoning to evaluative judgment and the development of transferable models or frameworks\.

Detailed indicators and examples, together with the theoretical foundations and construction of the rubric, are provided in Appendices A and B, respectively\.

## 4Benchmark for CDS Evaluation

![Refer to caption](https://arxiv.org/html/2608.13990v1/Overview_a.png)Figure 2:Overview of the scalable CDS annotation workflow, which applies the rubric grounded in theories of cognition and learning\. A gold\-standard subset is used to select the LLM evaluator best aligned with human judgments for large\-scale labeling\.We introduce SCOPE\-Bench, the first benchmark that supports content\-depth assessment from two different perspectives\. At theitem level\(Figure[2](https://arxiv.org/html/2608.13990#S4.F2)\), SCOPE\-Bench assesses the content depth of individual short videos from the user perspective\. At therecommendation\-list level\(Figure[3](https://arxiv.org/html/2608.13990#S4.F3)\), it evaluates the content depth of recommendation lists from the platform perspective, enabling platforms to assess RSs beyond conventional engagement\-oriented performance\.

Dataset\#Users\#Items\#Inter\.CaptionImageASRSparsityFull10,000153,5611,007,746100\.00%88\.56%88\.56%99\.93%Sampled6,65431,496128,105100\.00%99\.26%99\.26%99\.94%
Note:The modality coverage is not complete because captions, visual features, and Automatic Speech Recognition \(ASR\) transcripts in the publicly released data are not uniformly available for all videos\.

Table 2:Statistics and modality coverage of the dataset\.### 4\.1Video Dataset

We build SCOPE\-Bench upon ShortVideo333https://github\.com/tsinghua\-fib\-lab/ShortVideo˙dataset\([Shang et al\. 2025](https://arxiv.org/html/2608.13990#bib.bib40)\), a publicly available short\-video dataset containing user\-item interactions and multimodal information\. ShortVideo records approximately 1M chronologically ordered interactions generated by 10K real\-world users over a one\-week period\. The dataset provides two versions,FullandSampled, which cover the same data collection period\. TheSampledversion contains a subset of the users included inFull, although the sampling strategy used to construct this subset is not documented by the\([Shang et al\. 2025](https://arxiv.org/html/2608.13990#bib.bib40)\)\. We perform necessary preprocessing to improve data consistency and retain the content signals required for CDS assessment, as detailed in Appendix C\. The statistics of the resultingFullandSampledversions are summarized in Table[2](https://arxiv.org/html/2608.13990#S4.T2)\.

### 4\.2Evaluation Protocol

Our protocol adopts a content\-based setting, where the video content is the primary object of CDS assessment\. Instead of directly processing raw visual and audio streams, we use three textual signals as input, which capture complementary aspects of the video content: the captionccsummarizes the main theme, the category labelkkprovides topic\-level context, and the ASR transcriptaacaptures the spoken content\. Given a videox=\(c,k,a\)x=\(c,k,a\)and the system prompt𝒫\\mathcal\{P\}, an LLM evaluatorℳ\\mathcal\{M\}produces a structured outputy=\(s,ℓ,r,e,q\)y=\(s,\\ell,r,e,q\), comprising the CDSss, its level nameℓ\\ell, a reasonrr, supporting evidenceee, and a confidenceqq\. The scoressis ordinal, taking values in\{0,1,…,6,∅\}\\\{0,1,\\dots,6,\\varnothing\\\}, where∅\\varnothingmarks insufficient information for a reliable assessment\. Because this setting relies on text, videos whose depth resides mainly in visual or auditory content, or that lack a usable transcript, receive∅\\varnothing\.

### 4\.3Annotation

Gold\-Standard Subset Construction\.To establish reliable reference labels for CDS assessment, we construct a gold\-standard subset \(gold set\) through human annotation\. Specifically, we conducted a human evaluation with five human evaluators from computer science and psychology\. Following the Delphi method\([Hong et al\. 2019](https://arxiv.org/html/2608.13990#bib.bib19)\), the annotation process consists of three steps\. In Step ①, the evaluators independently annotated the sampled videos according to scoring rubric based on video content\. In Step ②, the human evaluators conducted a second round of annotation for videos showing substantial disagreement across evaluators\. In Step ③, we aggregated the annotations using the median to reduce the influence of outlier judgments and improve the reliability of the resulting gold\-standard labels\. The gold set serves two purposes: validating interpretability and practical usability of our rubric and providing a human reference for selecting the LLM evaluator used in automatic annotation\. Details about human evaluators, the annotation procedure, and the size of the gold set are provided in Appendix E\.

Automatic Scalable Annotation\.To enable scalable annotation while maintaining alignment with human judgments, we first evaluate a set of candidate LLMs on the gold set\. We compare their CDS predictions with human annotations using multiple agreement and error metrics, and select the model with the strongest overall alignment as the default evaluator\. The evaluation results are reported in Section[5\.2](https://arxiv.org/html/2608.13990#S5.SS2)\. We then apply the selected evaluator to the short\-video dataset to generate CDS annotations\. By augmenting the original dataset with annotations, we construct SCOPE\-Bench dataset, which supports two complementary evaluation settings at both item and list levels for recommendation\.

![Refer to caption](https://arxiv.org/html/2608.13990v1/Overview_b.png)Figure 3:Overview of the content\-depth\-aware RS evaluation framework for jointly assessing engagement and content depth\.
### 4\.4List\-Level Content\-Depth Evaluation

As illustrated in Figure[3](https://arxiv.org/html/2608.13990#S4.F3), the list\-level content\-depth evaluation workflow evaluates the content depth of top\-KKrecommendation lists\. Specifically, we introduce*List\-wiseContentDepthScore \(LCDS\)*, a complementary metric that aggregates the CDS of the top\-KKlists\. Letsi∈\{0,…,6\}s\_\{i\}\\in\\\{0,\\ldots,6\\\}444We assign a score of 0 to items withsi=∅s\_\{i\}=\\varnothing\. Such cases are primarily caused by insufficient information from ASR, so the resulting values should be interpreted as lower\-bound estimation\.denote the CDS of the video at rankii\. LCDS@KKis defined as:

LCDS𝜶,β,𝒘⁡@​K\\displaystyle\\operatorname\{LCDS\}\_\{\\boldsymbol\{\\alpha\},\\beta,\\boldsymbol\{w\}\}@K=\(∑i=1Kwi​\[g𝜶​\(si\)\]β∑i=1Kwi\)1/β,\\displaystyle=\\left\(\\frac\{\\sum\_\{i=1\}^\{K\}w\_\{i\}\\left\[g\_\{\\boldsymbol\{\\alpha\}\}\(s\_\{i\}\)\\right\]^\{\\beta\}\}\{\\sum\_\{i=1\}^\{K\}w\_\{i\}\}\\right\)^\{1/\\beta\},\(1\)where​g𝜶​\(s\)\\displaystyle\\text\{where\}\\;g\_\{\\boldsymbol\{\\alpha\}\}\(s\)=\{0,s=0,∑ℓ=1sαℓ∑r=16αr,s∈\{1,…,6\}\.\\displaystyle=\\begin\{cases\}0,&s=0,\\\\ \\displaystyle\\frac\{\\sum\_\{\\ell=1\}^\{s\}\\alpha\_\{\\ell\}\}\{\\sum\_\{r=1\}^\{6\}\\alpha\_\{r\}\},&s\\in\\\{1,\\ldots,6\\\}\.\\end\{cases\}\(2\)In Eq\. \([2](https://arxiv.org/html/2608.13990#S4.E2)\),αℓ\>0\\alpha\_\{\\ell\}\>0denotes the marginal value of moving from Levelℓ−1\\ell\-1to Levelℓ\\ell, andg𝜶​\(s\)∈\[0,1\]g\_\{\\boldsymbol\{\\alpha\}\}\(s\)\\in\[0,1\]maps an ordinal CDS label to a normalized gain\.𝒘\\boldsymbol\{w\}model position\-dependent exposure and satisfy bothwi≥0w\_\{i\}\\geq 0and∑i=1Kwi\>0\\sum\_\{i=1\}^\{K\}w\_\{i\}\>0\.

Letw¯i=wi/∑j=1Kwj\\bar\{w\}\_\{i\}=\{w\_\{i\}\}/\{\\sum\_\{j=1\}^\{K\}w\_\{j\}\}denote the normalized rankiiexposure weight, andℐ\+:=\{i∈\{1,…,K\}∣wi\>0\}\\mathcal\{I\}\_\{\+\}:=\\left\\\{i\\in\\\{1,\\ldots,K\\\}\\mid w\_\{i\}\>0\\right\\\}\. The aggregation behavior of LCDS can be characterized by:

LCDSβ​@​K\\displaystyle\\text\{LCDS\}\_\{\\beta\}@K=\{∏i∈ℐ\+g𝜶​\(si\)w¯i,β→0\+,∑i=1Kw¯i​g𝜶​\(si\),β=1,maxi∈ℐ\+⁡g𝜶​\(si\),β→∞\.\\displaystyle=\\begin\{cases\}\\displaystyle\\prod\_\{i\\in\\mathcal\{I\}\_\{\+\}\}g\_\{\\boldsymbol\{\\alpha\}\}\(s\_\{i\}\)^\{\\bar\{w\}\_\{i\}\},&\\beta\\rightarrow 0^\{\+\},\\\\\[5\.69054pt\] \\displaystyle\\sum\_\{i=1\}^\{K\}\\bar\{w\}\_\{i\}g\_\{\\boldsymbol\{\\alpha\}\}\(s\_\{i\}\),&\\beta=1,\\\\\[5\.69054pt\] \\displaystyle\\max\_\{i\\in\\mathcal\{I\}\_\{\+\}\}g\_\{\\boldsymbol\{\\alpha\}\}\(s\_\{i\}\),&\\beta\\rightarrow\\infty\.\\end\{cases\}\(3\)Hence, larger values ofβ\\betaproduce a more peak\-oriented evaluation, whereas smaller values make LCDS more sensitive to low\-CDS positions and therefore favor lists that sustain CDS\. Following the grouping of the original CDS rubric, we further define three corresponding interpretive levels for LCDS\. The thresholds are defined directly on the transformed scale as Low LCDS:\[0,τL\)\[0,\\tau\_\{\\mathrm\{L\}\}\), Medium LCDS:\[τL,τH\)\[\\tau\_\{\\mathrm\{L\}\},\\tau\_\{\\mathrm\{H\}\}\), and High LCDS:\[τH,1\]\[\\tau\_\{\\mathrm\{H\}\},1\], whereτL=g𝜶​\(1\)\\tau\_\{\\mathrm\{L\}\}=g\_\{\\boldsymbol\{\\alpha\}\}\(1\),τH=g𝜶​\(4\)\\tau\_\{\\mathrm\{H\}\}=g\_\{\\boldsymbol\{\\alpha\}\}\(4\)\.

For the default setting, we assume equal marginal values, i\.e\.,𝜶=𝟏\\boldsymbol\{\\alpha\}=\\boldsymbol\{1\}, and use arithmetic aggregation withβ=1\\beta=1, yieldingg⁡\(si\)=si/6g\(s\_\{i\}\)=s\_\{i\}/6\. Uniform rank weights \(wi=1w\_\{i\}=1\) give Top\-KKAverage LCDS \(A\-LCDS@KK\) definition as follows:

A​\-​LCDS⁡@​K=1K​∑i=1Kg⁡\(si\),\\operatorname\{A\\text\{\-\}LCDS\}@K=\\frac\{1\}\{K\}\\sum\_\{i=1\}^\{K\}g\(s\_\{i\}\),\(4\)which measures the average content depth of the recommendation list\. Moreover, to account for greater exposure at higher ranks, we setwi=1/log2⁡\(i\+1\)w\_\{i\}=1/\\log\_\{2\}\(i\+1\)followed by NDCG\([Järvelin and Kekäläinen 2002](https://arxiv.org/html/2608.13990#bib.bib20)\), and define Top\-KKExposure\-weighted LCDS \(E\-LCDS@KK\) as follows:

E​\-​LCDS⁡@​K=∑i=1Kg⁡\(si\)log2⁡\(i\+1\)∑j=1K1log2⁡\(j\+1\)\.\\operatorname\{E\\text\{\-\}LCDS\}@K=\\frac\{\\sum\_\{i=1\}^\{K\}\\frac\{g\(s\_\{i\}\)\}\{\\log\_\{2\}\(i\+1\)\}\}\{\\sum\_\{j=1\}^\{K\}\\frac\{1\}\{\\log\_\{2\}\(j\+1\)\}\}\.\(5\)Both metrics lie in\[0,1\]\[0,1\], with larger values indicating greater content depth of top\-KKrecommendation lists\.

## 5Experiment

\(a\)Agreement under score\-difference tolerance\.\(b\)Bootstrap distribution of Ordinal Krippendorff’sα\\alpha\.
Figure 4:Human annotation reliability analysis, showing high agreement in their CDS annotations\.### 5\.1Human Annotation Reliability

We assess the reliability of the independent human annotations collected during the construction of the gold set\. Specifically, we analyze the Step ① annotations before conflict resolution, thereby evaluating whether different human evaluators can consistently apply the proposed CDS rubric without consensus\-based adjustment, following prior evaluation practices\([Wang et al\. 2024](https://arxiv.org/html/2608.13990#bib.bib46);[Han et al\. 2025](https://arxiv.org/html/2608.13990#bib.bib12)\)\. As shown in Figure[4](https://arxiv.org/html/2608.13990#S5.F4), the human evaluators exhibit a high level of agreement in their CDS annotations, which remains strong even after accounting for chance agreement\. Overall, these results indicate that the proposed scoring rubric can be reliably applied by different human evaluators\.

ModelSpearman↑\\uparrowKendall↑\\uparrowPearson↑\\uparrowExact↑\\uparrowMAE↓\\downarrowGLM\-5\.10\.66580\.64100\.766371\.67%0\.3337GPT\-5\.50\.64140\.61080\.731267\.35%0\.4418Gemini3\.1\-Pro0\.72700\.70200\.768376\.71%0\.3097Kimi\-K2\.60\.70030\.67140\.761071\.43%0\.3505MiMo\-V2\.5\-Pro0\.63560\.60270\.694467\.35%0\.4538Qwen3\.7\-Max0\.71770\.69600\.793478\.03%0\.2605Table 3:Comparison of LLM evaluators against human CDS judgments\.↑\\uparrowand↓\\downarrowindicate that higher and lower values are better, respectively, andbolddenotes the best result\.
### 5\.2LLMs\-Human Agreement

To identify the most suitable LLMs for our evaluation protocol, we compare six leading models from the leaderboard555https://artificialanalysis\.ai/leaderboards/models: three open\-weight models,Kimi\-K2\.6\([Moonshot AI 2026](https://arxiv.org/html/2608.13990#bib.bib32)\),MiMo\-V2\.5\-Pro\([Xiaomi MiMo Team 2026](https://arxiv.org/html/2608.13990#bib.bib51)\), andGLM\-5\.1\([Z\.ai 2026](https://arxiv.org/html/2608.13990#bib.bib55)\); and three proprietary models,Gemini 3\.1\-Pro\([Google DeepMind 2026](https://arxiv.org/html/2608.13990#bib.bib10)\),GPT\-5\.5\([OpenAI 2026](https://arxiv.org/html/2608.13990#bib.bib34)\), andQwen3\.7\-Max\([Qwen Team 2026](https://arxiv.org/html/2608.13990#bib.bib37)\)\. Following prior work\([Liu et al\. 2023](https://arxiv.org/html/2608.13990#bib.bib29);[Kim et al\. 2024](https://arxiv.org/html/2608.13990#bib.bib24);[Ye et al\. 2023](https://arxiv.org/html/2608.13990#bib.bib53)\), we evaluate all models on the same gold set and compare their scoressis\_\{i\}with human annotations using Spearman’sρ\\rho, Kendall’sτ\\tau, Pearson correlation, exact\-match accuracy, and MAE\. As shown in Table[3](https://arxiv.org/html/2608.13990#S5.T3),Qwen3\.7\-Maxachieves the strongest human alignment in three of the five metrics and is therefore adopted as the default evaluator\. Most LLM models also correlate well with the human scores, suggesting that our protocol enables LLMs to reproduce judgments broadly shared by human evaluators\. Additional agreement results are reported in Appendix F\.

### 5\.3Analysis of CDS

Figure 5:Distribution of CDS across all videos\. Most videos fall into the low\-CDS group, with few in the high\-CDS group\.CDS Distribution\.Figure[5](https://arxiv.org/html/2608.13990#S5.F5)shows that the majority of videos fall into the low\-CDS or NaN group666NaN cases mainly result from missing raw videos or ASR transcripts that are too short or noisy for reliable CDS assessment\.\. Overall, the distribution suggests that a large ratio of videos in short\-video platforms imposes limited content depth\. This observation is consistent with the attention\-economy nature of platforms, where entertaining and attention\-grabbing content is prevalent\.

RankLowMediumHigh1DanceHealthFinance2ComedyLawMilitary3BeautyScienceHistory4MusicHistoryLaw5Short DramasFinanceScience6Casual VideosReal EstateInformation7AnimeDigital ProductsHealthTable 4:Top seven categories in each CDS group, ranked by proportions\. Low\-CDS categories differ from the others, while medium\- and high\-CDS categories largely overlap\.Category Distribution\.Table[4](https://arxiv.org/html/2608.13990#S5.T4)presents the top categories across three CDS group\. We observe that medium\- and high\-CDS videos are more frequently associated with knowledge\-intensive categories, such as Finance, History, and Law\. In contrast, low\-CDS videos are more commonly associated with entertainment\-oriented categories, such as Dance, Comedy, and Beauty\. These category\-level patterns are consistent with our design intuition of CDS: videos involving domain knowledge tend to receive higher CDS, whereas videos primarily designed for entertainment tend to receive lower CDS\. See Appendix G for distribution details\.

![Refer to caption](https://arxiv.org/html/2608.13990v1/cpd_theory_signal_heatmap_aaai.png)Figure 6:Lexicon\-based enrichment across CDS levels\. Each cell reports the relative enrichment of signal words at CDS levels, with positive and negative values indicating values above and below the overall average, respectively\. The smallest enrichment value in each row is shown inunderlined\.CDS Word Frequency\.Figure[6](https://arxiv.org/html/2608.13990#S5.F6)shows the enrichment patterns of different CDS levels based on our predefined theory\-oriented lexicon777We built the lexicon by extracting the top 500 words in gold subset and the top 200 words at each CDS level, grouping them into theoretical categories, and expanding each category\.\. Entertainment reaction and plot/dialogue show a decreasing trend as the CDS level rises, whereas most other categories exhibit the opposite trend\. These findings are consistent with our expectations: low\-CDS videos are more likely to focus on entertainment\-oriented reactions and plot\-level descriptions, whereas high\-CDS videos are more likely to contain signals related to explanation, evidence evaluation, generalization, transfer, and conceptualized expression\. We also provide details representative words for each CDS level in Appendix H\.

VariableAssociationExplained varianceCaption lengthSpearmanρ=0\.0828∗\\rho=0\.0828^\{\*\}R2=0\.57%R^\{2\}=0\.57\\%Log\-ASR lengthaSpearmanρ=0\.2182∗\\rho=0\.2182^\{\*\}R2=6\.09%R^\{2\}=6\.09\\%CategoryCramérV=0\.2469∗V=0\.2469^\{\*\}η2=25\.34%\\eta^\{2\}=25\.34\\%
- aASR length\|a\|\|a\|exhibits a long\-tailed distribution\. Therefore, we use log\-transformed formlog⁡\(1\+\|a\|\)\\log\(1\+\|a\|\)to estimate the relationship\.
- •∗denotes statistical significance atp<0\.001p<0\.001\.

Table 5:Associations between valid CDS \(s≠∅s\\neq\\varnothing\) and surface\-level input attributes\. CDS is more strongly associated with ASR length and category than with caption length\.Surface\-Level Correlates of CDS\.Table[5](https://arxiv.org/html/2608.13990#S5.T5)summarizes the relationships between valid CDS \(s≠∅s\\neq\\varnothing\) and three surface\-level input attributes: Caption length, ASR length, and category\. The results show that caption length has only a weak association with CDS, whereas ASR length and category exhibit stronger associations\. This pattern is consistent with high\-CDS content requiring sufficient textual space to express explanations, procedures, and reasoning structures\. Category differences may similarly arise from their inherent content orientation, with entertainment\-oriented and knowledge\-oriented categories tending to receive lower and higher CDS scores, respectively, as shown in Table[4](https://arxiv.org/html/2608.13990#S5.T4)\.

### 5\.4Evaluation Protocol Robustness

In Section[5\.3](https://arxiv.org/html/2608.13990#S5.SS3), we observe that CDS exhibit certain correlations with video category and ASR length\. These observations naturally raise two robustness concerns:Whether the evaluation protocol has category prior bias and verbosity bias\. To examine these issues, we conduct paired counterfactual robustness tests\([Zheng et al\. 2023](https://arxiv.org/html/2608.13990#bib.bib58)\), where only one input field is perturbed at a time\. For each videoii, we compare the original CDS scoresis\_\{i\}and ASR length\|ai\|\|a\_\{i\}\|with their counterfactual valuessi′s\_\{i\}^\{\\prime\}and\|ai′\|\|a\_\{i\}^\{\\prime\}\|\. We define the absolute score change and relative ASR length change asΔ​si=\|si′−si\|\\Delta s\_\{i\}=\|s\_\{i\}^\{\\prime\}\-s\_\{i\}\|andδ​\|ai\|=\(\|ai′\|−\|ai\|\)/\|ai\|\\delta\|a\_\{i\}\|=\{\(\|a\_\{i\}^\{\\prime\}\|\-\|a\_\{i\}\|\)\}/\{\|a\_\{i\}\|\}, respectively\.

PerturbationNNP\(Δ​s=1\\Delta s=1\)P\(Δ​s=2\\Delta s=2\)P\(Δ​s\>2\\Delta s\>2\)Low→\\rightarrowMedium500\.00%0\.00%0\.00%Low→\\rightarrowHigh500\.00%0\.00%0\.00%Medium→\\rightarrowLow606\.67%3\.33%0\.00%High→\\rightarrowLow4613\.04%13\.04%0\.00%Table 6:Category counterfactual perturbation results measured by CDS shift magnitude\.Source→Target\\text\{Source\}\\rightarrow\\text\{Target\}indicates perturbing videos from the Source CDS group to Target\-associated categories\. Category perturbations leave low\-CDS videos unchanged and cause only limited one\- or two\-level shifts for medium\- and high\-CDS videos\.Robustness to Category Prior Bias\.We examine whether the evaluation protocol relies on category\-level shortcuts by overemphasizing category labels while underutilizing ASR transcripts, which provide more direct evidence of a video’s semantic content\. As shown in Table[6](https://arxiv.org/html/2608.13990#S5.T6), changing low\-CDS videos to either medium\- or high\-CDS\-associated categories does not change their CDS\. For the reverse direction, mapping high\- or medium\-CDS videos to low\-CDS\-associated categories may introduce slight downward shifts, with more pronounced changes observed for high\-CDS videos\. This is consistent with our conservative scoring principle: when the available evidence does not fully support a higher CDS level, the evaluator tends to assign a lower score\.

PerturbationGroupNNδ​\|a\|\\delta\|a\|P\(Δ​s=1\\Delta s=1\)P\(Δ​s=2\\Delta s=2\)P\(Δ​s\>2\\Delta s\>2\)ASR length↑\\uparrowLow50\+52\.05%2\.00%0\.00%0\.00%Medium60\+51\.39%5\.00%0\.00%0\.00%ASR length↓\\downarrowMedium60\-24\.49%8\.33%0\.00%0\.00%High46\-27\.35%23\.91%15\.22%0\.00%Table 7:ASR counterfactual perturbation results measured by CDS shifts\.Δ​si\\Delta s\_\{i\}andδ​\|ai\|\\delta\|a\_\{i\}\|denote the absolute CDS change and relative ASR length change, respectively\. CDS remains largely stable under lengthening and changes modestly under shortening, with no shifts exceeding two levels\.Robustness to Verbosity Bias\.We examine whether the evaluation protocol favors longer ASR transcripts without additional meaningful information\. Table[7](https://arxiv.org/html/2608.13990#S5.T7)shows that meaningless length expansion has minimal impact, whereas transcript shortening affects high\-CDS videos more than medium\-CDS videos, with no CDS change exceeding two levels\. This sensitivity is consistent with our conservative scoring principle, as transcript compression may weaken the reasoning evidence required for high CDS levels, including argument logic and evidence connections\.

Overall, these results suggest that our evaluation protocol primarily responds to meaningful informational and reasoning structures in the videos, rather than superficial cues such as category labels or ASR transcript length\. Details and stability experiment are provided in Appendix I\.

## 6SCOPE\-Bench Leaderboard

MethodShortVideoSampledShortVideoFullR@20↑\\uparrowA@20↑\\uparrowE@20↑\\uparrowR@20↑\\uparrowA@20↑\\uparrowE@20↑\\uparrowID\-based RecommendationBPR3\.306\.906\.852\.227\.067\.10NCF3\.016\.496\.422\.097\.467\.53LightGCN3\.547\.367\.332\.376\.897\.08Multimodal RecommendationVBPR2\.736\.225\.701\.807\.447\.95GRCN2\.837\.927\.851\.646\.726\.73LATTICE2\.707\.137\.042\.256\.977\.10BM32\.886\.726\.662\.416\.706\.80FREEDOM3\.307\.137\.302\.516\.816\.84MGCN3\.827\.387\.582\.656\.896\.99LGMRec3\.406\.716\.632\.427\.247\.52DiffMM3\.247\.247\.352\.397\.107\.18REARM3\.767\.717\.762\.426\.816\.93FITMM3\.258\.128\.352\.417\.597\.88Random0\.086\.906\.900\.026\.596\.59Table 8:Engagement and content\-depth performance of baselines on two datasets\. All metrics are scaled by×100\\times 100, where R denotes Recall, A/E denote A\-LCDS/E\-LCDS, and the best results arebolded\. Baselines are competitive on engagement but remain in the low\-LCDS range and close to random recommendation on content\-depth metric\.We evaluate 13 baselines on SCOPE\-Bench, including three ID\-based methods, BPR\([Rendle et al\. 2009](https://arxiv.org/html/2608.13990#bib.bib39)\), NCF\([He et al\. 2017](https://arxiv.org/html/2608.13990#bib.bib15)\), and LightGCN\([He et al\. 2020](https://arxiv.org/html/2608.13990#bib.bib14)\), and ten multimodal methods, VBPR\([He and McAuley 2016](https://arxiv.org/html/2608.13990#bib.bib13)\), GRCN\([Wei et al\. 2020](https://arxiv.org/html/2608.13990#bib.bib50)\), LATTICE\([Zhang et al\. 2021](https://arxiv.org/html/2608.13990#bib.bib57)\), BM3\([Zhou et al\. 2023](https://arxiv.org/html/2608.13990#bib.bib60)\), FREEDOM\([Zhou and Shen 2023](https://arxiv.org/html/2608.13990#bib.bib59)\), MGCN\([Yu et al\. 2023](https://arxiv.org/html/2608.13990#bib.bib54)\), LGMRec\([Guo et al\. 2024](https://arxiv.org/html/2608.13990#bib.bib11)\), DiffMM\([Jiang et al\. 2024](https://arxiv.org/html/2608.13990#bib.bib21)\), REARM\([Ma et al\. 2025](https://arxiv.org/html/2608.13990#bib.bib30)\), and FITMM\([Yang et al\. 2025](https://arxiv.org/html/2608.13990#bib.bib52)\)\. The Random baseline reports the expected performance of uniformly sampling items from each candidate set\. We use an 8:1:1 training\-validation\-test split\([Shang et al\. 2025](https://arxiv.org/html/2608.13990#bib.bib40)\)\. As shown in Table[8](https://arxiv.org/html/2608.13990#S6.T8), existing methods achieve competitive performance on engagement metrics, whereas their A\-LCDS and E\-LCDS scores consistently remain within the Low\-LCDS range, i\.e,\[0,1/6\)\[0,1/6\)\. Moreover, most methods perform close to the Random baseline\. These results indicate that stronger engagement performance does not necessarily translate into the recommendation of content with greater depth\. Appendix C presents the experimental setup, results under an alternative treatment ofsi=∅s\_\{i\}=\\varnothing, training trajectories of engagement and content\-depth metrics, and CDS\-aware optimization\. These results further demonstrate that engagement and content depth are currently decoupled\.

## 7Conclusion

In this paper, we propose a new metric, termedCDS, to measure the content depth of short videos\. CDS provides a principled basis for evaluating and optimizing content depth in existing RSs\. To comprehensively evaluate this dimension, we construct SCOPE\-Bench, the first benchmark that supports both item\- and list\-level content\-depth evaluation\. Empirical results demonstrate the interpretability, practical utility, and robustness of the proposed evaluation framework\. Our experiments further reveal that existing RSs tend to favor low\-depth videos, and remain limited in recommending videos with high CDS\. Hereby, CDS and SCOPE\-Bench establish a new evaluation axis for developing short\-video RSs that jointly consider content depth and user engagement\.

## References

- Allen et al\. \(1955\)Allen, K\.; Berry, M\. M\.; Luehrs Jr, F\. U\.; and Perry, J\. W\. 1955\.Machine literature searching VIII\. Operational criteria for designing information retrieval systems\.*American documentation \(pre\-1986\)*, 6\(2\): 93\.
- Anderson and Krathwohl \(2001\)Anderson, L\. W\.; and Krathwohl, D\. R\. 2001\.*A taxonomy for learning, teaching, and assessing: A revision of Bloom’s taxonomy of educational objectives: complete edition*\.Addison Wesley Longman, Inc\.
- Australian Government \(2025\)Australian Government\. 2025\.Social media minimum age\.https://www\.infrastructure\.gov\.au/media\-communications/internet/online\-safety/social\-media\-minimum\-age\.
- Barzilay and Lapata \(2008\)Barzilay, R\.; and Lapata, M\. 2008\.Modeling local coherence: An entity\-based approach\.*Computational Linguistics*, 34\(1\): 1–34\.
- Biggs and Collis \(2014\)Biggs, J\. B\.; and Collis, K\. F\. 2014\.*Evaluating the quality of learning: The SOLO taxonomy \(Structure of the Observed Learning Outcome\)*\.Academic press\.
- Chi and Wylie \(2014\)Chi, M\. T\.; and Wylie, R\. 2014\.The ICAP framework: Linking cognitive engagement to active learning outcomes\.*Educational psychologist*, 49\(4\): 219–243\.
- García\-Rapp \(2017\)García\-Rapp, F\. 2017\.Popularity markers on YouTube’s attention economy: the case of Bubzbeauty\.*Celebrity studies*, 8\(2\): 228–245\.
- Ge, Delgado\-Battenfeld, and Jannach \(2010\)Ge, M\.; Delgado\-Battenfeld, C\.; and Jannach, D\. 2010\.Beyond accuracy: Evaluating recommender systems by coverage and serendipity\.In*Proceedings of the Fourth ACM Conference on Recommender Systems*, 257–260\.
- Gehman et al\. \(2020\)Gehman, S\.; Gururangan, S\.; Sap, M\.; Choi, Y\.; and Smith, N\. A\. 2020\.Realtoxicityprompts: Evaluating neural toxic degeneration in language models\.In*Findings of the association for computational linguistics: EMNLP 2020*, 3356–3369\.
- Google DeepMind \(2026\)Google DeepMind\. 2026\.Gemini 3\.1 Pro\.https://deepmind\.google/models/model\-cards/gemini\-3\-1\-pro/\.
- Guo et al\. \(2024\)Guo, Z\.; Li, J\.; Li, G\.; Wang, C\.; Shi, S\.; and Ruan, B\. 2024\.Lgmrec: Local and global graph learning for multimodal recommendation\.In*AAAI*, 8454–8462\.
- Han et al\. \(2025\)Han, H\.; Li, S\.; Chen, J\.; Yuan, Y\.; Wu, Y\.; Deng, Y\.; Leong, C\. T\.; Du, H\.; Fu, J\.; Li, Y\.; et al\. 2025\.Video\-bench: Human\-aligned video generation benchmark\.In*Proceedings of the Computer Vision and Pattern Recognition Conference*, 18858–18868\.
- He and McAuley \(2016\)He, R\.; and McAuley, J\. 2016\.VBPR: visual bayesian personalized ranking from implicit feedback\.In*AAAI*, volume 30\.
- He et al\. \(2020\)He, X\.; Deng, K\.; Wang, X\.; Li, Y\.; Zhang, Y\.; and Wang, M\. 2020\.Lightgcn: Simplifying and powering graph convolution network for recommendation\.In*Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval*, 639–648\.
- He et al\. \(2017\)He, X\.; Liao, L\.; Zhang, H\.; Nie, L\.; Hu, X\.; and Chua, T\.\-S\. 2017\.Neural collaborative filtering\.In*Proceedings of the 26th international conference on world wide web*, 173–182\.
- Herlocker et al\. \(2004\)Herlocker, J\. L\.; Konstan, J\. A\.; Terveen, L\. G\.; and Riedl, J\. T\. 2004\.Evaluating collaborative filtering recommender systems\.*ACM Transactions on Information Systems*, 22\(1\): 5–53\.
- Heusel et al\. \(2017\)Heusel, M\.; Ramsauer, H\.; Unterthiner, T\.; Nessler, B\.; and Hochreiter, S\. 2017\.Gans trained by a two time\-scale update rule converge to a local nash equilibrium\.*Advances in neural information processing systems*, 30\.
- Hijikata, Kai, and Nishida \(2012\)Hijikata, Y\.; Kai, Y\.; and Nishida, S\. 2012\.The relation between user intervention and user satisfaction for information recommendation\.In*Proceedings of the 27th Annual ACM Symposium on Applied Computing*, 2002–2007\.
- Hong et al\. \(2019\)Hong, Q\. N\.; Pluye, P\.; Fàbregues, S\.; Bartlett, G\.; Boardman, F\.; Cargo, M\.; Dagenais, P\.; Gagnon, M\.\-P\.; Griffiths, F\.; Nicolau, B\.; et al\. 2019\.Improving the content validity of the mixed methods appraisal tool: a modified e\-Delphi study\.*Journal of clinical epidemiology*, 111: 49–59\.
- Järvelin and Kekäläinen \(2002\)Järvelin, K\.; and Kekäläinen, J\. 2002\.Cumulated gain\-based evaluation of IR techniques\.*ACM Transactions on Information Systems \(TOIS\)*, 20\(4\): 422–446\.
- Jiang et al\. \(2024\)Jiang, Y\.; Xia, L\.; Wei, W\.; Luo, D\.; Lin, K\.; and Huang, C\. 2024\.DiffMM: Multi\-Modal Diffusion Model for Recommendation\.*arXiv preprint arXiv:2406\.11781*\.
- Keenan \(2023\)Keenan, C\. 2023\.New Features for Teens and Families on TikTok\.https://newsroom\.tiktok\.com/new\-features\-for\-teens\-and\-families\-on\-tiktok\-au?lang=en\-AU\.
- Kemp \(2025\)Kemp, S\. 2025\.Digital 2025 July Global Statshot Report\.https://datareportal\.com/reports/digital\-2025\-july\-global\-statshot\.
- Kim et al\. \(2024\)Kim, S\.; Shin, J\.; Jang, J\.; Longpre, S\.; Lee, H\.; Yun, S\.; Shin, R\.; Kim, S\.; Thorne, J\.; Seo, M\.; et al\. 2024\.Prometheus: Inducing fine\-grained evaluation capability in language models\.In*International Conference on Learning Representations*, volume 2024, 29927–29962\.
- Knijnenburg et al\. \(2012\)Knijnenburg, B\. P\.; Willemsen, M\. C\.; Gantner, Z\.; Soncu, H\.; and Newell, C\. 2012\.Explaining the user experience of recommender systems\.*User Modeling and User\-Adapted Interaction*, 22\(4\): 441–504\.
- Kuaishou Technology \(2025\)Kuaishou Technology\. 2025\.Kuaishou Technology Announces Third Quarter 2025 Unaudited Financial Results\.https://ir\.kuaishou\.com/news\-releases/news\-release\-details/kuaishou\-technology\-announces\-third\-quarter\-2025\-unaudited/\.
- Li et al\. \(2026\)Li, Z\.; Long, G\.; Jiang, J\.; Zhang, C\.; and Yang, Q\. 2026\.Federated Vision\-Language\-Recommendation with Personalized Fusion\.In*AAAI*, volume 40, 23337–23345\.
- Lin, Hilton, and Evans \(2022\)Lin, S\.; Hilton, J\.; and Evans, O\. 2022\.Truthfulqa: Measuring how models mimic human falsehoods\.In*Proceedings of the 60th annual meeting of the association for computational linguistics \(volume 1: long papers\)*, 3214–3252\.
- Liu et al\. \(2023\)Liu, Y\.; Iter, D\.; Xu, Y\.; Wang, S\.; Xu, R\.; and Zhu, C\. 2023\.G\-eval: NLG evaluation using gpt\-4 with better human alignment\.In*Proceedings of the 2023 conference on empirical methods in natural language processing*, 2511–2522\.
- Ma et al\. \(2025\)Ma, S\.; Zeng, Y\.; Wu, S\.; and Xu, G\. 2025\.Refining Contrastive Learning and Homography Relations for Multi\-Modal Recommendation\.In*Proceedings of the 33rd ACM International Conference on Multimedia*, 6316–6324\.
- Mahakud and Thapliyal \(2026\)Mahakud, G\. C\.; and Thapliyal, A\. 2026\.Assessing the Impact of Different Social Media Video Content Formats on Sustained Attention and Working Memory\.*Annals of Neurosciences*, 09727531261424994\.
- Moonshot AI \(2026\)Moonshot AI\. 2026\.Meet Kimi K2\.6: Advancing Open\-Source Coding\.https://forum\.moonshot\.ai/t/meet\-kimi\-k2\-6\-advancing\-open\-source\-coding/369\.
- Napoles, Sakaguchi, and Tetreault \(2017\)Napoles, C\.; Sakaguchi, K\.; and Tetreault, J\. 2017\.JFLEG: A fluency corpus and benchmark for grammatical error correction\.In*Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers*, 229–234\.
- OpenAI \(2026\)OpenAI\. 2026\.GPT\-5\.5 System Card\.https://openai\.com/index/gpt\-5\-5\-system\-card/\.
- Ørmen and Gregersen \(2023\)Ørmen, J\.; and Gregersen, A\. 2023\.Towards the engagement economy: interconnected processes of commodification on YouTube\.*Media, Culture & Society*, 45\(2\): 225–245\.
- Pu, Chen, and Hu \(2011\)Pu, P\.; Chen, L\.; and Hu, R\. 2011\.A user\-centric evaluation framework for recommender systems\.In*Proceedings of the Fifth ACM Conference on Recommender Systems*, 157–164\.
- Qwen Team \(2026\)Qwen Team\. 2026\.Qwen3\.7: The Agent Frontier\.https://qwen\.ai/blog?id=qwen3\.7\.
- Radford et al\. \(2021\)Radford, A\.; Kim, J\. W\.; Hallacy, C\.; Ramesh, A\.; Goh, G\.; Agarwal, S\.; Sastry, G\.; Askell, A\.; Mishkin, P\.; Clark, J\.; et al\. 2021\.Learning transferable visual models from natural language supervision\.In*International conference on machine learning*, 8748–8763\. PmLR\.
- Rendle et al\. \(2009\)Rendle, S\.; Freudenthaler, C\.; Gantner, Z\.; and Schmidt\-Thieme, L\. 2009\.BPR: Bayesian Personalized Ranking from Implicit Feedback\.In*Proceedings of the Twenty\-Fifth Conference on Uncertainty in Artificial Intelligence*, 452–461\. AUAI Press\.
- Shang et al\. \(2025\)Shang, Y\.; Gao, C\.; Li, N\.; and Li, Y\. 2025\.A Large\-scale Dataset with Behavior, Attributes, and Content of Mobile Short\-video Platform\.In*Companion Proceedings of the ACM on Web Conference 2025*, 793–796\. New York, NY, USA: Association for Computing Machinery\.
- Sternberg, Sternberg, and Mio \(2006\)Sternberg, R\. J\.; Sternberg, K\.; and Mio, J\. 2006\.*Cognitive psychology*\.Thomson/Wadsworth Belmont, CA\.
- Tang et al\. \(2026\)Tang, D\.; Zhang, X\.; Gou, P\.; Feng, J\.; Hu, R\.; and Sum, K\.\-w\. R\. 2026\.Association between short\-form video use and mental health: systematic review and Meta\-analysis\.*Journal of Medical Internet Research*, 28: e82503\.
- TikTok Team \(2021\)TikTok Team\. 2021\.Thanks a Billion\!https://newsroom\.tiktok\.com/1\-billion\-people\-on\-tiktok?lang=en\.
- UK Government \(2026\)UK Government\. 2026\.Fact sheet: New rules to protect children online\.https://www\.gov\.uk/government/publications/fact\-sheet\-new\-rules\-to\-protect\-children\-online\.
- Vargas and Castells \(2011\)Vargas, S\.; and Castells, P\. 2011\.Rank and relevance in novelty and diversity metrics for recommender systems\.In*Proceedings of the Fifth ACM Conference on Recommender Systems*, 109–116\.
- Wang et al\. \(2024\)Wang, Y\.; Yu, Z\.; Yao, W\.; Zeng, Z\.; Yang, L\.; Wang, C\.; Chen, H\.; Jiang, C\.; Xie, R\.; Wang, J\.; et al\. 2024\.Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization\.In*International Conference on Learning Representations*, volume 2024, 43573–43593\.
- Wang et al\. \(2004\)Wang, Z\.; Bovik, A\. C\.; Sheikh, H\. R\.; and Simoncelli, E\. P\. 2004\.Image quality assessment: from error visibility to structural similarity\.*TIP*, 13\(4\): 600–612\.
- Wang, Lu, and Bovik \(2004\)Wang, Z\.; Lu, L\.; and Bovik, A\. C\. 2004\.Video Quality Assessment Based on Structural Distortion Measurement\.*Signal Processing: Image Communication*, 19\(2\): 121–132\.
- Wason and Evans \(1974\)Wason, P\. C\.; and Evans, J\. S\. B\. 1974\.Dual processes in reasoning?*Cognition*, 3\(2\): 141–154\.
- Wei et al\. \(2020\)Wei, Y\.; Wang, X\.; Nie, L\.; He, X\.; and Chua, T\.\-S\. 2020\.Graph\-refined convolutional network for multimedia recommendation with implicit feedback\.In*Proceedings of the 28th ACM international conference on multimedia*, 3541–3549\.
- Xiaomi MiMo Team \(2026\)Xiaomi MiMo Team\. 2026\.Xiaomi MiMo\-V2\.5\-Pro\.https://mimo\.xiaomi\.com/mimo\-v2\-5\-pro/\.
- Yang et al\. \(2025\)Yang, W\.; Zhong, R\.; Chen, Y\.; Li, S\.; Ping, H\.; Lu, C\.; and Jiang, P\. 2025\.FITMM: Adaptive Frequency\-Aware Multimodal Recommendation via Information\-Theoretic Representation Learning\.In*Proceedings of the 33rd ACM International Conference on Multimedia*, 6193–6202\.
- Ye et al\. \(2023\)Ye, S\.; Kim, D\.; Kim, S\.; Hwang, H\.; Kim, S\.; Jo, Y\.; Thorne, J\.; Kim, J\.; and Seo, M\. 2023\.Flask: Fine\-grained language model evaluation based on alignment skill sets\.*arXiv preprint arXiv:2307\.10928*\.
- Yu et al\. \(2023\)Yu, P\.; Tan, Z\.; Lu, G\.; and Bao, B\.\-K\. 2023\.Multi\-view graph convolutional network for multimedia recommendation\.In*Proceedings of the 31st ACM international conference on multimedia*, 6576–6585\.
- Z\.ai \(2026\)Z\.ai\. 2026\.GLM\-5\.1: Towards Long\-Horizon Tasks\.https://z\.ai/blog/glm\-5\.1\.
- Zangerle and Bauer \(2022\)Zangerle, E\.; and Bauer, C\. 2022\.Evaluating recommender systems: survey and framework\.*ACM computing surveys*, 55\(8\): 1–38\.
- Zhang et al\. \(2021\)Zhang, J\.; Zhu, Y\.; Liu, Q\.; Wu, S\.; Wang, S\.; and Wang, L\. 2021\.Mining latent structures for multimedia recommendation\.In*Proceedings of the 29th ACM international conference on multimedia*, 3872–3880\.
- Zheng et al\. \(2023\)Zheng, L\.; Chiang, W\.\-L\.; Sheng, Y\.; Zhuang, S\.; Wu, Z\.; Zhuang, Y\.; Lin, Z\.; Li, Z\.; Li, D\.; Xing, E\.; et al\. 2023\.Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.*Advances in neural information processing systems*, 36: 46595–46623\.
- Zhou and Shen \(2023\)Zhou, X\.; and Shen, Z\. 2023\.A tale of two graphs: Freezing and denoising graph structures for multimodal recommendation\.In*Proceedings of the 31st ACM international conference on multimedia*, 935–943\.
- Zhou et al\. \(2023\)Zhou, X\.; Zhou, H\.; Liu, Y\.; Zeng, Z\.; Miao, C\.; Wang, P\.; You, Y\.; and Jiang, F\. 2023\.Bootstrap latent representations for multi\-modal recommendation\.In*Proceedings of the ACM web conference 2023*, 845–854\.
- Ziegler et al\. \(2005\)Ziegler, C\.\-N\.; McNee, S\. M\.; Konstan, J\. A\.; and Lausen, G\. 2005\.Improving recommendation lists through topic diversification\.In*Proceedings of the 14th International Conference on World Wide Web*, 22–32\.
- Zou and Sun \(2025\)Zou, K\.; and Sun, A\. 2025\.A Survey of Real\-World Recommender Systems: Challenges, Constraints, and Industrial Perspectives\.*arXiv preprint arXiv:2509\.06002*\.

Similar Articles

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling

Hugging Face Daily Papers

This paper introduces DeScore, a video reward model that decouples reasoning and scoring processes to improve training efficiency and generalization. It addresses the limitations of existing discriminative and generative reward models by using a 'think-then-score' paradigm with multimodal large language models.

SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence

Hugging Face Daily Papers

Introduces SVI-Bench, a large-scale benchmark for strategic video intelligence using team sports, designed to evaluate models on dynamic scene understanding, causal reasoning, strategic simulation, and agentic synthesis. The benchmark reveals a capability cliff where models perform well on perceptual tasks but sharply degrade on higher-level strategic reasoning.