VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing
Summary
Introduces VlogReward, a reward model for evaluating vlog editing plans across six dimensions, along with a large-scale dataset and benchmark, achieving state-of-the-art results against GPT-5 and Gemini-3-Pro.
View Cached Full Text
Cached at: 07/28/26, 06:26 AM
# VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing
Source: [https://arxiv.org/html/2607.22632](https://arxiv.org/html/2607.22632)
Wen ZhongSijie ZhuXin GuFan ChenJunxian DuanJie CaoLongyin WenZhenfang Chen
###### Abstract
The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans\. However, vlog assessment is highly subjective and remains challenging due to a lack of standardized criteria, dataset and benchmark, and effective reward models\. To address these challenges, we define a comprehensive vlog evaluation framework guided by professional vlog creators and product managers, establishing a taxonomy of six key dimensions,i\.e\.,Creativity,Consistency,Concept Design,Cinematography,Narration, andPacing\. Subsequently, we curate a large\-scale dataset of 100k vlog edits and a dedicated benchmark,VRMBench, to evaluate the vlog rewarding capabilities of Multimodal Large Language Models \(MLLMs\)\. Finally, we presentVlogReward, a robust vlog reward model that can provide both fine\-grained multi\-dimensional scores and actionable feedback for iterative refinement\. Technically, we enhance the Group Relative Policy Optimization \(GRPO\) framework by introducing an adjustable inter\-group comparison reward, which mitigates the “direction blindness” issue of standard GRPO and enables the model to better distinguish varied\-quality edits\. VlogReward achieves state\-of\-the\-art results that significantly outperform existing MLLMs, including GPT\-5 and Gemini\-3\-Pro\. We hope that our study can help vlog creators and foster automated vlog evaluation and refinement systems\.
Machine Learning, ICML
Figure 1:Different from regular video understanding tasks whose processes are often simple and separated, vlog editing invovles more complicated and interwoven processes\. To conduct a comprehensive assessment for each step and element of vlog editing, the evaluation criteria would be multi\-layered and highly subjective, showcasing unique challenges for vlog reward modeling\.## 1Introduction
Nowadays, the rapid proliferation of personal media creation has positioned video blogs \(i\.e\., vlogs\) as a dominant and highly personalized medium for online storytelling\(Ladhariet al\.,[2020](https://arxiv.org/html/2607.22632#bib.bib48); Duanet al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib47)\)\. Concurrently, the comprehensive capabilities of Multimodal Large Language Models \(MLLMs\) have developed significantly\(Zhanget al\.,[2025a](https://arxiv.org/html/2607.22632#bib.bib24); Honget al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib22); Comaniciet al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib29)\), offering new possibilities for complex video\-language tasks\. This synergy raises a new question:can we leverage the advanced multimodal reasoning capabilities of MLLMs to evaluate vlog drafts and distill high\-quality editing plans?This research would help vlog creators improve their artworks, while paving the way for automated vlog edit refinement systems to facilitate commercial production\.
Unlike traditional video understanding tasks that typically deal with single, continuous, and well\-structured videos\(Tanget al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib35); Madanet al\.,[2024](https://arxiv.org/html/2607.22632#bib.bib36)\), vlog editing is more complicated, as shown in Figure[1](https://arxiv.org/html/2607.22632#S0.F1)\. It starts with a vast collection of raw, unorganized, and often redundant user\-provided video assets\. To transform these discrete footages into a compelling vlog, creators must navigate a complex decision\-making space: designing the storyline and narrative style, selecting the most expressive segments from cluttered assets, precisely cropping timestamps, determining sequence order, composing scripts, and orchestrating multi\-modal editing elements such as subtitles, voiceovers, transitions, music, emojis, font style, special effects, etc\.
Consequently, evaluating a vlog’s quality is inherently multi\-layered and more subjective, presenting unique challenges that transcend standard image/video evaluation\(Xuet al\.,[2023](https://arxiv.org/html/2607.22632#bib.bib26); Fuet al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib25)\)\. Beyond objective factors like subtitle\-visual correspondence and logic of shot order, the essence of a montage vlog lies in highly subjective dimensions: the concept design of the theme, overall narrative creativity, the cinematographic aesthetic of shot selection, the storyline pacing, etc\. These factors are difficult to quantify but essential for distinguishing varied\-quality vlog edits\.
Therefore, despite the compelling potential of MLLM\-based evaluation\(Wanget al\.,[2025a](https://arxiv.org/html/2607.22632#bib.bib32),[e](https://arxiv.org/html/2607.22632#bib.bib33); Wuet al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib39)\), the automatic assessment of vlog edits remains an arduous task\.First, there is a critical shortage of standardized vlog evaluation criteria for reference\.Second, this field suffers from a severe deficit of dedicated benchmarks and annotated vlog data\.Lastly, these collectively result in a lack of effective models for such vlog reward modeling\.
To systemically overcome these challenges, wefirstengage professional vlog designers and artistic experts to consolidate the evaluation metrics and define a standardized taxonomy of six key dimensions, establishing a multidimensional and granular assessment framework\.Second, we curate a large\-scale training dataset with 100k vlog edits, and a dedicated benchmark to evaluate the vlog rewarding capabilities of MLLMs\.Lastly, we develop a robust Vlog Reward Model \(VRM\) that can provide multi\-dimensional scores and feedback for any given raw footage assets and corresponding editing plan\.
Following recent progress in reward modeling\(Guoet al\.,[2025b](https://arxiv.org/html/2607.22632#bib.bib46); Wanget al\.,[2025d](https://arxiv.org/html/2607.22632#bib.bib41)\), we first perform Supervised Fine\-Tuning \(SFT\) and then utilize reinforcement learning \(RL\) to further enhance our VRM’s vlog reward capability\. However, vanilla Group Relative Policy Optimization \(GRPO\)\(Shaoet al\.,[2024](https://arxiv.org/html/2607.22632#bib.bib3)\)is inherently designed to calculate intra\-group relative advantages, thus lacking cross\-sample reference signals for different edits\. We observe that this would lead to “direction blindness” issue and introduceinter\-group comparison reward\. By involving reward signals of comparing inter\-group rollouts, it enables the model to better distinguish vlog edits of varying quality and more effectively identify the optimal editing plan\.
Our contributions can be summarized as follows:
- •Novel Task Formulation and Systematic Evaluation Framework\.We propose a new task of automated vlog editing evaluation, and establish a comprehensive taxonomy across six key dimensions in collaboration with professional vlog designers and product managers\.
- •Large\-Scale Dataset and Benchmark Construction\.We curate a high\-quality training dataset of 100k vlog edits, and introduceVRMBench, a dedicated benchmark including both subjective, aesthetic evaluation cases and objective, factual\-error ones\. VRMBench comprises 400 unique groups, each containing a set of raw video assets paired with four distinct editing plans of descending quality, annotated with scores and textual feedback\. These involve a two\-stage laborious data collection with3\.6k person\-hourverification\.
- •Multi\-Dimensional Vlog Reward Model\.We developVlogRewardthat can discriminate varied\-quality vlog edits with multi\-dimensional scores and feedback\. With our inter\-group comparison reward design, VlogReward achieves superior comparison accuracy \(69\.4%→\\rightarrow73\.5%\), Best\-of\-N accuracy \(65\.0%→\\rightarrow67\.6%\) and score accuracy \(50\.3%→\\rightarrow51\.3%\)\. Moreover, adjusting the intensity of comparison reward enables a flexible trade\-off between scoring and discriminating capability\.
Figure 2:Data collection pipeline\. Subjective Dataset: We first collect raw assetsv\{v\}and use MLLMs to synthesize diverse editing plans\(x1,x2,…,xm\)\(\{x\}\_\{1\},\{x\}\_\{2\},\\dots,\{x\}\_\{m\}\)differing from subjective domains,e\.g\., aesthetics and creativity\. Then we reward them to obtain multi\-dimensional scoresssand feedbackff\. The entire process involves human verification to check for hallucinations, ensure alignment between scores and feedback, and filter out low\-quality data\. Objective Dataset: We additionally collect 8k assets and annotate the editing plans with the same pipeline\. Subsequently, we progressively introduce tampering such as altering cut timestamps or subtitles to create objective factual errors, resulting in 4 plans of decreasing quality for each footage\. Both the subjective and objective datasets are split separately\. The resulting splits are then merged to form the final training dataset and VRMBench benchmark\.
## 2Related Work
Vlog Editing and Evaluation\.Due to the complexity of vlog editing, various studies mainly focus on its decomposed sub‑tasks such as highlight detection\(Ganet al\.,[2023](https://arxiv.org/html/2607.22632#bib.bib65)\), logical ordering of shots\(Pardoet al\.,[2024](https://arxiv.org/html/2607.22632#bib.bib62)\), video content summarization\(Xieet al\.,[2022](https://arxiv.org/html/2607.22632#bib.bib66)\), and next‑shot prediction\(Huet al\.,[2023](https://arxiv.org/html/2607.22632#bib.bib67)\)\. With the rapid progress of MLLMs, several recent researches have begun to explore agentic end‑to‑end frameworks to automate the entire editing pipeline in a single pass\(Yanget al\.,[2024](https://arxiv.org/html/2607.22632#bib.bib63); Sandoval\-Castanedaet al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib68); Pardoet al\.,[2024](https://arxiv.org/html/2607.22632#bib.bib62); Wanget al\.,[2025c](https://arxiv.org/html/2607.22632#bib.bib64)\)\. However, existing work widely evaluates vlog editing with IoU‑type overlap metrics or adopts custom, labor‑intensive human rating, which limits reproducibility and promotion\. Moreover, vlog editing is an inherently creative and flexible process, and there should be no absolute ground truth label for cut timestamps\. This necessitates a standardized, comprehensive, and general evaluation framework, and a corresponding vlog reward model to foster automatic evaluation\.
Multimodal Reward Modeling\.Multimodal reward modeling plays a pivotal role for aligning MLLMs with human preferences\(Zhanget al\.,[2025e](https://arxiv.org/html/2607.22632#bib.bib30); Sunet al\.,[2024](https://arxiv.org/html/2607.22632#bib.bib10); Zhanget al\.,[2025d](https://arxiv.org/html/2607.22632#bib.bib13); Duanet al\.,[2026](https://arxiv.org/html/2607.22632#bib.bib77)\)\. Early approaches use fine\-tuned CLIP models to produce preference scores\(Kirstainet al\.,[2023](https://arxiv.org/html/2607.22632#bib.bib11); Wuet al\.,[2023](https://arxiv.org/html/2607.22632#bib.bib12)\)\. Subsequent works append a linear regression head to the model’s backbone to output a scalar reward value\(Liuet al\.,[2025a](https://arxiv.org/html/2607.22632#bib.bib43); Zanget al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib42)\)\. With the rapid development of MLLMs, generative reward modeling \(GRM\) represents a paradigm shift from scalar regressive reward models \(RRM\), enabling more flexible and interpretative output\. Recently, several GRM studies have attempted to apply some reasoning strategies, such as Test\-Time scaling and RL approaches\(Liuet al\.,[2025b](https://arxiv.org/html/2607.22632#bib.bib78); Zhanget al\.,[2026](https://arxiv.org/html/2607.22632#bib.bib79); Liuet al\.,[2025c](https://arxiv.org/html/2607.22632#bib.bib57)\)to further enhance the reward capability\(Chenet al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib44); Xueet al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib45)\)\. However, most existing multimodal reward modeling frameworks directly utilize a pairwise comparison paradigm to identify relative preferences\(Zhanget al\.,[2025c](https://arxiv.org/html/2607.22632#bib.bib31),[f](https://arxiv.org/html/2607.22632#bib.bib38)\)\. While pairwise ranking can efficiently capture comparative signals, it lacks the capacity for fine\-grained, multi\-dimensional scoring, and straightforward diagnostic feedback\. Therefore, we propose a GRM framework beyond binary preferences, providing comprehensive scores and feedback across six distinct dimensions while reflecting relative preferences\.
GRPO Variants\.Several variants of GRPO have emerged to address its stability and scalability challenges\. GSPO\(Zhenget al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib75)\)shifts to sequence\-level optimization to stabilize Mixture\-of\-Experts models, and DAPO\(Yuet al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib72)\)introduces decoupled clipping for large\-scale RL\. BNPO\(Xiaoet al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib71)\)introduces adaptive reward normalization using a dynamic Beta distribution\. GMPO\(Zhaoet al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib70)\)replaces the arithmetic mean with a geometric mean to suppress token\-level outliers\. For decentralized environments, GEPO\(Zhanget al\.,[2025b](https://arxiv.org/html/2607.22632#bib.bib73)\)utilizes group expectation weighting to mitigate high KL divergence caused by network latency\. Similar to our inter\-group reward idea, some work utilizes inter\-group rollouts to calculate advantages\. For example, DyKnow\-RAG\(Xuet al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib69)\)integrates GRPO to retrieval\-augmented generation with inter\-group advantage\. GRPOformer\(Guoet al\.,[2025a](https://arxiv.org/html/2607.22632#bib.bib76)\)adapts GRPO for efficient hyperparameter optimization using inter\-group relative advantages\.
Figure 3:Data example\. To facilitate multimodal processing, raw video assets\{v1,v2,…,vn\}\\\{\{v\}\_\{1\},\{v\}\_\{2\},\\dots,\{v\}\_\{n\}\\\}are concatenated into a single videov\{v\}, with each clip separated by a 3\-second black screen\. The editing planx\{x\}is a JSON encapsulating the cut timestamps, playback speed, voiceover, shot order, title and subtitle\. We restrict editable variables to these core structural attributes to accommodate MLLMs’ frame\-sampling mechanism and minimize confounds\. Stylistic elements such as voice timbre, transitions, visual effects, and customized typography \(e\.g\., font color, size, and style\) are standardized\. In this scheme, “shot\_timestamp” and “final\_timestamp” correspond to the time range in the stitched videov\{v\}and final edited vlog, respectively\. A duration mismatch between them triggers a playback speed adjustment\. The accompanying feedback provides scoring rationale, where evident, flagsnotable issuesandimprovement suggestions\.
## 3Method
### 3\.1Task Design
Focusing on Montage vlogs, our core objective is to develop a robust VRM to evaluate the quality of editing plans\. Given a set of raw video assetsv=\{v1,v2,…,vn\}\{v\}=\\\{\{v\}\_\{1\},\{v\}\_\{2\},\\dots,\{v\}\_\{n\}\\\}, the VRMℳ\\mathcal\{M\}should provide a comprehensive assessment of any candidate editing planx\{x\}\.
Unlike scalar reward models used in simpler tasks, our VRM is designed as a multi\-dimensional generative reward model\. For a given pair\(v,x\)\(\{v\},\{x\}\), the model generates a set of numerical scoress=\{s1,s2,…,sk\}s=\\\{s\_\{1\},s\_\{2\},\\dots,s\_\{k\}\\\}representingkkdistinct dimensions of quality, and an interpretable textual feedback sequencef=\{f1,f2,…,fk\}f=\\\{f\_\{1\},f\_\{2\},\\dots,f\_\{k\}\\\}\. The feedbackffserves as a diagnostic critique, explicitly identifying factual issues and providing refinement instructions within the editing planx\{x\}, thereby enhancing the transparency and explainability of the evaluation process,i\.e\.,ℳ\(v,x\)→\(s,f\)\\mathcal\{M\}\(\{v\},\{x\}\)\\rightarrow\(s,f\)\. We use this mechanism as it can be utilized in two ways\. First, the numerical scoresscan promote to select the optimal editing plan through Test\-Time Scaling \(TTS\) or serve as reward signals to enhance the editing model’s performance through RL\. Second, the textual feedbackffcan return to the policy model for iterative refinement of editing plans\.
With the guidance of professional vlog creators and product managers, we consolidate the vlog evaluation metrics into six key dimensions, and assign integer ratings from 1 to 5,i\.e\.,k=6k=6,sk∈\{1,2,3,4,5\}s\_\{k\}\\in\\\{1,2,3,4,5\\\}\. The six evaluation dimensions areCreativity\(narrative inventiveness and originality\),Consistency\(text\-visual alignment\),Concept Design\(script element appropriateness\),Cinematography\(shot selection and framing\),Narration\(storyline and narration quality\), andPacing\(temporal rhythm and sequence logic\)\. Details of the evaluation criteria are provided in Appendix[E](https://arxiv.org/html/2607.22632#A5)\.
### 3\.2Dataset
The dataset collection follows a two\-stage pipeline\. In the 1st stage, we collect raw footagev\{v\}and corresponding varied\-quality editing plans\{x1,x2,…,xm\}\\\{\{x\}\_\{1\},\{x\}\_\{2\},\\dots,\{x\}\_\{m\}\\\}\. In the 2nd stage, we reward each pair of\(v,x\),x∈\{x1,x2,…,xm\}\(\{v\},\{x\}\),\{x\}\\in\\\{\{x\}\_\{1\},\{x\}\_\{2\},\\dots,\{x\}\_\{m\}\\\}to obtain the score and feedback\(s,f\)\(s,f\)\. Edits of better overall quality should receive higher total score sums∑i=1ksi\\sum\_\{i=1\}^\{k\}s\_\{i\}, maintaining a consistent quality\-to\-score mapping\. To enhance data diversity, we curate two sets of edits that differ along subjective \(SB\) or objective \(OB\) dimensions\. Details are expounded in the following and displayed in Figure[2](https://arxiv.org/html/2607.22632#S1.F2)\.
Figure 4:Vlog types\. Our dataset contains various vlog types including daily life, recreation, hobbies, pets, sports, emotion, travel, and interview\.Figure 5:Data Statistics\. Our dataset contains various vlog types, including daily life, recreation, hobbies, pets, sports, emotion, travel, and interview\. On VRMBench, the footage durations are appropriately distributed, ranging from 10 seconds to 5 minutes\. Both VRMBench\-SB and VRMBench\-OB contain balanced data across all score values, demonstrating the diversity of our dataset\.#### 3\.2\.1Raw Assets and Editing Plans\(v,x\)\(\{v\},\{x\}\)
We collect extensive raw footage from public platforms spanning diverse categories\. Leveraging advanced MLLMs, we identify narrative templates for each group of assets and synthesize the corresponding editing plans\. After a rigorous human\-in\-the\-loop verification phase to eliminate low\-quality samples, we finally obtainEdits\-Datasetcomprising raw video assets with standard editing plans\. Figure[4](https://arxiv.org/html/2607.22632#S3.F4)summarizes the various vlog types our dataset contains\.
#### 3\.2\.2Scores and Textual Feedback\(s,f\)\(s,f\)
Given the high subjectivity of vlog evaluation, the annotated scores even vary with the well\-defined rubric with the same annotator\. To mitigate this issue, we employ an AI\-assisted annotation pipeline with human verification\. We first design MLLM pipeline to evaluate each\(v,x\)\(\{v\},\{x\}\)from Edits\-Dataset to obtainVRM\-Dataset\-SB\. Then we select samples with 4 edits, where the score difference between any two plans≥\\geq2\. These data are manually audited to ensure that there are no hallucinations and the feedbackffis well\-aligned with the scoress\. Finally, we divided them into two parts:VRM\-Dataset\-SB\-RL\(20k\) andVRM\-Dataset\-SB\-Test\(8k\)\. We exclude VRM\-Dataset\-SB\-Test from VRM\-Dataset\-SB, and obtainVRM\-Dataset\-SFT\(100k\)\.
To introduce objective criteria, we additionally collect 8k video assets with standard editing plans\. We then randomly select a segment to tamper with shot timestamps or subtitles to create factual errors, ensuring each modification progressively degrades the plan’s quality\. After automatic scoring, we filter out samples where the human\-judged quality ranking contradicted the scores\. This results inVRM\-Dataset\-OBconsisting of 5k assets, each paired with 4 editing plans of decreasing quality\. We randomly sample 200 assets to obtainVRM\-Dataset\-OB\-Test, remaining others asVRM\-Dataset\-OB\-RL\. Finally, VRM\-Dataset\-SB\-RL and VRM\-Dataset\-OB\-RL are merged intoVRM\-Dataset\-RL\(40k\), and VRM\-Dataset\-SB\-Test and VRM\-Dataset\-OB\-Test are merged intoVRMBench\(1\.6k\)\.
### 3\.3Benchmark
As described in Section[3\.2](https://arxiv.org/html/2607.22632#S3.SS2), we use VRMBench to test and benchmark the vlog rewarding capability of mainstream MLLMs\. VRMBench comprises 400 unique data groups, each consisting of a set of raw footage \(3\-10 clips\) paired with 4 editing plans of descending quality\. Each plan is annotated with scores and textual feedback across six dimensions, where the cumulative score strictly follows the decreasing order of quality\. Each group features distinct raw assets with total durations ranging from 10 seconds to 5 minutes\. The benchmark is categorized into two subsets: VRMBench\-SB and VRMBench\-OB\. VRMBench\-SB contains 200 groups where quality variance is driven by subjective attributes \(e\.g\., creativity and aesthetic appeal\), while VRMBench\-OB includes 200 groups focusing on objective discrepancies \(e\.g\., factual errors in subtitle descriptions\)\. More statistics are displayed in Figure[5](https://arxiv.org/html/2607.22632#S3.F5)\.
### 3\.4GRPO with Inter\-Group Comparison Reward
We use the pipeline of SFT \+ GRPO to train our VRM\. However, GRPO inherently computes rewards and relative advantages within intra\-group rollouts and lacks inter\-group comparative information\. Despite improving the accuracy of absolute scores, it suffers from “direction blindness” issue \(explained in Figure[6](https://arxiv.org/html/2607.22632#S3.F6)\)\. One potential remedy is to calculate inter\-group advantages in a batch\(Guoet al\.,[2025a](https://arxiv.org/html/2607.22632#bib.bib76)\)\. However, the high variance between different samples may lead to training instability, where a positive\-advantage rollout for a single sample may be improperly penalized with negative advantage relative to inter\-group rollouts\. To address this, we introduce inter\-group comparison reward while ensuring relative advantages are still computed within intra\-group rollouts of the single sample\. Given a sample of raw assets and an editing planq=\(v,x\)q=\(\{v\},\{x\}\), the old VRM policy modelπold\\pi\_\{\\mathrm\{old\}\}generates a group of rollouts\{o1,o2,…,oG\}\\\{o\_\{1\},o\_\{2\},\\dots,o\_\{G\}\\\}\. We optimize the current policy modelπθ\\pi\_\{\\theta\}by maximizing the objective of standard GRPO, as described in Appendix[B](https://arxiv.org/html/2607.22632#A2)\. Our reward designs are elaborated in the following\.
Format Reward\.We do not use the common thinking mode that encloses the thought process with<think\>\.\.\.</think\>and the answer with<answer\>\.\.\.</answer\>, as it yeilds sub\-optimal performance compared to only using the answer content\. Details can be found in Appendix[C](https://arxiv.org/html/2607.22632#A3)\. The VRM is required to provide integer scores from 1 to 5 and feedback for all six dimensions in its output answer\.
rformat=\{0,if output matches format,1,if output doesn’t match format\.r\_\{\\mathrm\{format\}\}=\\begin\{cases\}0,&\\text\{if output matches format\},\\\\ 1,&\\text\{if output doesn't match format\}\.\\end\{cases\}\(1\)
Score Reward \(SR\)\.The score reward is formulated as the proportion of the six evaluation dimensions that match the ground\-truth labels\. Specifically, lets=\{s1,s2,…,sk\}s=\\\{s\_\{1\},s\_\{2\},\\dots,s\_\{k\}\\\}be the predicted scores and𝐬^=\{s^1,s^2,…,s^k\}\\mathbf\{\\hat\{s\}\}=\\\{\\hat\{s\}\_\{1\},\\hat\{s\}\_\{2\},\\dots,\\hat\{s\}\_\{k\}\\\}be the corresponding ground truths acrossk=6k=6dimensions\. We assign a nominal rewardα\\alpha\(i\.e\., 0\.1 in our setting\) to rollouts that maintain the correct format but fail to match any ground\-truth scores\. This serves to distinguish such instances from malformed outputs, which are penalized with a reward of 0\. The score reward is defined as
rscore=1k∑i=1k𝟙\(si=si^\)×\(1−α\)\+α\.r\_\{score\}=\\frac\{1\}\{k\}\\sum\_\{i=1\}^\{k\}\\mathds\{1\}\(s\_\{i\}=\\hat\{s\_\{i\}\}\)\\times\(1\-\\alpha\)\+\\alpha\.\(2\)
Figure 6:Illustration of the role of comparison reward \(CR\) beyond score reward \(SR\)\. For ease of analysis, we consider a simplified scenario with 2 rollouts and only score in a single dimension from 0 to 100\. We assume the vlog edit sample 1 is better than sample 2, with ground truth \(GT\) score 60 and 70, respectively\. For both samples, the predicted score of rollout 1 is closer to GT, resulting in higher SR and positive relative advantage\. Consequently, the policy update direction favors rollout 1 and approaches GT\. However, this would lead to lower predicted scores of sample 1 compared to sample 2,i\.e\.,o11<o12o^\{1\}\_\{1\}<o^\{2\}\_\{1\}for their scores\. In contrast, CR adds a direction signal of score difference between sample 1 and 2 that SR alone lacks\. The limitation of SR rises to this “direction blindness” issue, failing to recognize the preference direction\. With SR \+ CR,o21o^\{1\}\_\{2\}exhibits higher CR and provides an opposing policy update direction to restrain excessive updates of SR\. By appropriately adjusting the intensity of CR, we can maintain score accuracy while ensuring score differences align with sample quality preferences\.Comparison Reward \(CR\)\.To further enhance the VRM’s ability to differentiate varied\-quality editing plans, we design a comparison reward to introduce comparative information among inter\-group rollouts\. In VRM\-Dataset\-RL, each raw assetsv\{v\}is paired withm=4m=4editing plans\{x1,x2,x3,x4\}\\\{\{x\}^\{1\},\{x\}^\{2\},\{x\}^\{3\},\{x\}^\{4\}\\\}of descending quality\. Note the sum of predicted scores for\(v,xj\)\(\{v\},\{x\}^\{j\}\)asSj=∑i=1ksijS^\{j\}=\\sum\_\{i=1\}^\{k\}s^\{j\}\_\{i\}and the ground\-truth sum asS^j=∑i=1ks^ij\\hat\{S\}^\{j\}=\\sum\_\{i=1\}^\{k\}\\hat\{s\}\_\{i\}^\{j\}, it satisfiesS^1\>S^2\>S^3\>S^4\\hat\{S\}^\{1\}\>\\hat\{S\}^\{2\}\>\\hat\{S\}^\{3\}\>\\hat\{S\}^\{4\}\. The core idea is to encourage VRM to assign higher predicted scores to better samples\(v,xj\)\(\{v\},\{x\}^\{j\}\)than worse samples\{\(v,xl\)∣j\>l\}\\\{\(\{v\},\{x\}^\{l\}\)\\mid j\>l\\\},i\.e\.,Sj\>SlS^\{j\}\>S^\{l\}\. The comparison reward for theiith rolloutoijo\_\{i\}^\{j\}of\(v,xj\)\(\{v\},\{x\}^\{j\}\)is defined as
rcomp\(oij,l\)=\{1,ifsgn\[\(Sij−∑i′=1G′Si′lG′\)\(S^j−S^l\)\]=1,0,otherwise\.\\\!r\_\{\\mathrm\{comp\}\}\(o\_\{i\}^\{j\},l\)=\\begin\{cases\}1,\\text\{if\}\\,\\,\\text\{sgn\}\\big\[\(S\_\{i\}^\{j\}\-\\frac\{\\sum\_\{i^\{\\prime\}=1\}^\{G^\{\\prime\}\}S\_\{i^\{\\prime\}\}^\{l\}\}\{G^\{\\prime\}\}\)\(\\hat\{S\}^\{j\}\-\\hat\{S\}^\{l\}\)\\big\]=1,\\\\ 0,\\,\\text\{otherwise\}\.\\end\{cases\}\(3\)rcomp\(oij\)=1m−1∑l=1l≠jmrcomp\(oij,l\),r\_\{\\mathrm\{comp\}\}\(o\_\{i\}^\{j\}\)=\\frac\{1\}\{m\-1\}\\sum\_\{\\begin\{subarray\}\{c\}l=1\\\\ l\\neq j\\end\{subarray\}\}^\{m\}r\_\{comp\}\(o^\{j\}\_\{i\},l\),\(4\)whereSijS^\{j\}\_\{i\}is the sum of predicted scores ofoijo\_\{i\}^\{j\},m=4m=4is the number of candidate plans, andG′G^\{\\prime\}is the number of rollouts in the correct format\.
Total Reward\.The total reward is defined as
rtotal=rformat×\[\(1−λ\)⋅rscore\+λ⋅rcomp\],r\_\{\\mathrm\{total\}\}=r\_\{\\mathrm\{format\}\}\\times\\big\[\(1\-\\lambda\)\\cdot r\_\{\\mathrm\{score\}\}\+\\lambda\\cdot r\_\{\\mathrm\{comp\}\}\\big\],\(5\)whereλ\\lambdais the hyper\-parameter to adjust the intensity of the comparative reward signal\.
Table 1:Comprehensive evaluations of different MLLMs and baselines on VRMBench\. SB refers to VRMBench\-SB, OB represents VRMBench\-OB, and AVG indicates the average performance between SB and OB\. Compared with other SOTA MLLMs and different VRM counterparts, our model VlogReward achieves the highest average performance on all metrics\.ModelsScore AccuracyComparison AccuracyBest\-of\-N AccuracySBOBAVGSBOBAVGSBOBAVGClosed\-Source MLLMsGPT\-4o24\.432\.428\.438\.151\.844\.924\.342\.133\.2GPT\-529\.743\.236\.453\.772\.563\.137\.465\.751\.5o4\-mini24\.338\.231\.346\.164\.655\.326\.547\.737\.1Gemini\-3\-Pro35\.042\.938\.660\.583\.371\.935\.570\.152\.8Open\-Source MLLMsLLaVA\-OneVision\-7B24\.026\.725\.440\.944\.842\.818\.921\.920\.4LLaVA\-NeXT\-Video\-7B25\.531\.528\.539\.844\.742\.224\.428\.026\.2InternVL3\.5\-14B23\.229\.526\.330\.236\.833\.520\.335\.127\.7Kimi\-VL\-A3B\-Thinking\-250622\.930\.426\.741\.644\.843\.222\.723\.623\.1MiMo\-VL\-7B\-RL22\.531\.426\.944\.453\.949\.228\.539\.133\.8Qwen2\.5\-VL\-7B22\.432\.927\.734\.842\.838\.820\.531\.926\.2Qwen2\.5\-VL\-72B26\.138\.132\.136\.855\.346\.024\.844\.534\.6Qwen3\-VL\-8B22\.831\.627\.239\.354\.847\.024\.947\.035\.9Regressive Vlog Reward ModelsRRM\-MSE23\.625\.424\.542\.345\.543\.931\.641\.036\.3RRM\-BT\-\-\-58\.483\.170\.836\.569\.553\.0Generative Vlog Reward ModelsGRM\-SFT37\.441\.439\.454\.765\.059\.840\.344\.142\.2GRM\-GRPO44\.056\.550\.354\.884\.069\.443\.786\.265\.0GRM\-GRPO\-Inter\-Advantage43\.654\.649\.138\.873\.356\.035\.775\.655\.7VlogReward \(Ours\)45\.956\.751\.360\.486\.673\.548\.386\.867\.6
## 4Experiments
### 4\.1Settings
Dataset and Benchmark\.As described in Section[3\.2](https://arxiv.org/html/2607.22632#S3.SS2), VRM\-Dataset\-SFT \(100k\) is employed as SFT training data for cold start initialization\. After that, we use the refined dataset VRM\-Dataset\-RL \(40k\) to perform RL\. VRMBench is used to test and benchmark other MLLMs\.
Implementation Details\.We use Qwen2\.5\-VL\-7B\-Instruct\(Baiet al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib20)\)as our base model\. For the trade\-off between performance and efficiency, we limit the fps to 1 and the maximum video frames to 64\. Each frame is processed at a maximum resolution of64×28×2864\\times 28\\times 28pixels\.
For all evaluations, we follow the decoding configuration used in the official Qwen2\.5\-VL demo, with top\_p = 0\.001 and temperature = 0\.01\. For SFT training, the learning rate is set to 2e\-6, and the total batchsize is 128\. After 2 epochs of SFT training on VRM\-Dataset\-SFT, we conduct RL on VRM\-Dataset\-RL for 1 epoch, with learning rate 1e\-6 and batchsize 64\. The rollout temperature is set to 1, and the rollout numberGGis 8\. The hype\-parameterλ\\lambdais set to 0\.25\.
Baselines\.The evaluations are compared with recent state\-of\-the\-art MLLMs, including GPT\-4o\-2024\-05\-13\(Hurstet al\.,[2024](https://arxiv.org/html/2607.22632#bib.bib51)\), GPT\-5\-2025\-08\-07\(Singhet al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib52)\), o4\-mini\-2025\-04\-16\(Open AI,[2025](https://arxiv.org/html/2607.22632#bib.bib53)\), Gemini\-3\-Pro\(Google DeepMind,[2025](https://arxiv.org/html/2607.22632#bib.bib28)\), LLaVA\-OneVision\(Liet al\.,[2024](https://arxiv.org/html/2607.22632#bib.bib4)\), LLaVA\-Next\-Video\(Zhanget al\.,[2024](https://arxiv.org/html/2607.22632#bib.bib15)\), InternVL3\.5\(Wanget al\.,[2025b](https://arxiv.org/html/2607.22632#bib.bib17)\), Kimi\-VL\-Thinking\(Teamet al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib19)\), MiMo\-VL\-RL\(Xiaomi,[2025](https://arxiv.org/html/2607.22632#bib.bib18)\), Qwen2\.5\-VL\(Baiet al\.,[2025](https://arxiv.org/html/2607.22632#bib.bib20)\)and Qwen3\-VL\(Team,[2025](https://arxiv.org/html/2607.22632#bib.bib21)\)\. Inference parameters are consistent as described above across all models, except in two cases: 1\) GPT\-5 and o4\-mini use temperature=1 and top\_p=0\.7 for official mandatory specifications, and 2\) LLaVA\-One\-Vision and LLaVA\-Next\-Video use a fixed 16 frames because of their context length constraint\. What’s more, We also compare different paradigms of our VRM:
- •RRM\-MSE: Replace the LLM head with a trainable linear head to predict scores with mean square error loss\.
- •RRM\-BT: Replace the LLM head with a trainable linear head to predict pairwise preferences with Bradley\-Terry loss\(Bradley and Terry,[1952](https://arxiv.org/html/2607.22632#bib.bib34)\)\.
- •GRM\-SFT:The base MLLM is cold\-start initialized with VRM\-Dataset\-SFT, resulting in this model\.
- •GRM\-GRPO:This represents the baseline approach that applies GRPO with only SR on GRM\-SFT\.
- •GRM\-GRPO\-Inter\-Advantage:This calculates the relative advantage in inter\-group rollouts with the same settings of GRM\-GRPO\.
- •VlogReward:This is our final vlog reward model, which performs GRPO with both SR and CR\.
### 4\.2Benchmark Metrics
- •Score Accuracy:The average accuracy of predicted scores on the benchmark𝒟\\mathcal\{D\},i\.e\.,𝔼\(v,x\)∼𝒟1k∑i=1k𝟙\(si=s^i\)\\mathbb\{E\}\_\{\(\{v\},\{x\}\)\\sim\\mathcal\{D\}\}\\,\\frac\{1\}\{k\}\\sum\_\{i=1\}^\{k\}\\mathds\{1\}\(s\_\{i\}=\\hat\{s\}\_\{i\}\)\.
- •Comparison Accuracy:The proportion of assigning higher total score to superior editing plans over inferior ones,i\.e\.,𝔼\(v,xj,xl\)∼𝒟,j\>l1\(Sj\>Sl\)\\mathbb\{E\}\_\{\(\{v\},\{x\}\_\{j\},\{x\}\_\{l\}\)\\sim\\mathcal\{D\},\\,j\>l\}\\,\\mathds\{1\}\(S^\{j\}\>S^\{l\}\)\.
- •Best\-of\-N Accuracy:The proportion of assigning the highest total score to the best editing plan,i\.e\.,𝔼\(v,x1,x2,…,xm\)∼𝒟1\[\(argmaxjSj\)=1;1≤j≤m\]\\mathbb\{E\}\_\{\(\{v\},\{x\}\_\{1\},\{x\}\_\{2\},\\dots,\{x\}\_\{m\}\)\\sim\\mathcal\{D\}\}\\,\\mathds\{1\}\[\(\\operatorname\*\{argmax\}\\limits\_\{j\}\\,S^\{j\}\)=1;1\\leq j\\leq m\]\. In cases of tied score sums, we randomly select one of the tied plans as the predicted best candidate\. This repeats 5 times to report average Best\-of\-N Accuracy\.
Figure 7:Comparison of reward scores of Qwen2\.5\-VL\-7B and VlogReward for different editing plans\. Plan 1 suffers from chronological disorder \(day→\\rightarrownight→\\rightarrowday\), while Plan 2 maintains a coherent temporal flow\. Qwen assigns identical scores to both plans, failing to distinguish the edit quality\. Conversely, VlogReward accurately identifies the sequence logic flaws in Plan 1 \(assigningpacingscore to 1\), and the score difference reflects the correct preference\.
### 4\.3Main Results
Table[1](https://arxiv.org/html/2607.22632#S3.T1)summarizes the detailed performance of our VRM against various state\-of\-the\-art \(SOTA\) closed\-source and open\-source MLLMs\. Our analysis yields several key insights into the current state of multimodal vlog evaluation\.
General Vlog Editing Evaluation Limitations of Existing MLLMs\.The experiment results reveal that the general performance for vlog editing evaluation of existing MLLMs remains deficient\. Even when disregarding potential scoring bias, most models exhibit a limited capacity to distinguish varied\-quality editing plans, as evidenced by the relatively low Comparison Accuracy and Best\-of\-N Accuracy across the board\. For instance, all tested open\-source models fail to surpass 50% in average Comparison Accuracy, suggesting that current MLLMs struggle to identify the better editing plan among candidates,i\.e\., assigning equal or even lower scores for better vlog edits\. These findings underscore a significant gap in the comprehension capabilities of current MLLMs for vlog rewarding tasks, indicating substantial room for future improvement\.
Performance Disparity on Objective vs\. Subjective Domains\.A consistent trend across all evaluated models is that performance on VRMBench\-OB is generally superior to that on VRMBench\-SB\. Existing MLLMs demonstrate higher accuracy in identifying factual errors, but struggle more with subjective criteria about aesthetic feelings\. For example, GPT\-5 achieves a 72\.5% Comparison Accuracy on objective dimensions, while dropping to 53\.7% when evaluating nuanced, aesthetic preferences\. This disparity highlights the inherent difficulty in modeling complex, human\-centric quality metrics compared to explicit factual inconsistencies\.
Closed\-Source Performance Dominance\.Closed\-source MLLMs generally outperform open\-source counterparts across most metrics\. Among the closed\-source MLLMs, Gemini\-3\-Pro emerges as the strongest baselines\. It demonstrates superior Comparison Accuracy of 71\.9% and Best\-of\-N Accuracy of 52\.8%\. In contrast, open\-source MLLMs show a noticeable performance lag\.
Superiority of Our Proposed VlogReward and CR\.VlogReward achieves state\-of\-the\-art results across almost all evaluated metrics, significantly outperforming both open\-source and closed\-source models, and other VRM counterparts\. Figure[7](https://arxiv.org/html/2607.22632#S4.F7)shows an example of predicted scores of Qwen2\.5\-VL\-7B and VlogReward\. Qwen assigns identical scores to two clearly differentiated editing plans of varying quality, while VlogReward provides targeted scores reflecting editing issues\. These demonstrate the prominent ability of our model and our comparison reward design to provide multi\-demenisonal scores and identify the best editing plan\.
Table 2:Detailed performances across differentλ\\lambdavalues\.MetricDatasetλ\\lambda00\.250\.50\.751ScoreAccuracySB44\.045\.944\.143\.421\.7OB56\.556\.756\.154\.218\.9AVG50\.351\.350\.148\.820\.3ComparisonAccuracySB54\.860\.461\.265\.965\.6OB84\.086\.688\.890\.090\.2AVG69\.473\.575\.078\.077\.9Best\-of\-NAccuracySB43\.748\.348\.549\.449\.3OB86\.286\.893\.493\.293\.5AVG65\.067\.671\.071\.371\.4Table 3:Sensitivity analysis of the Comparison Reward \(CR\) under different configurations of group sizemmand rollout numberGG\.𝒎m𝑮GRewardScoreComparisonBoN38SR49\.164\.864\.338SR\+CR50\.268\.867\.044SR51\.067\.965\.344SR\+CR50\.272\.066\.8
### 4\.4Analysis and Discussions
Effect of Different Values ofλ\\lambda\.We investigate the influence of the hyper\-parameterλ\\lambdain the total reward formulation \(Equation[5](https://arxiv.org/html/2607.22632#S3.E5)\), which balances the granular score rewardrscorer\_\{score\}and the inter\-group comparison rewardrcompr\_\{comp\}\. As summarized in Table[2](https://arxiv.org/html/2607.22632#S4.T2), the model’s performance exhibits a non\-monotonic trend on Score Accuracy asλ\\lambdaincreases from 0 to 1\. Specifically, Score Accuracy initially improves and then gradually declines, reaching its peak performance of 51\.3% atλ=0\.25\\lambda=0\.25\. Both Comparison Accuracy and Best\-of\-N Accuracy ascend with the enhancement of the comparison signal, achieving their respective optima of 78% and 71\.4% atλ=0\.75\\lambda=0\.75andλ=1\\lambda=1, respectively\. These indicate that the appropriate introduction of comparison reward can simultaneously enhance both scoring accuracy and the capability to distinguish between superior and inferior samples \(e\.g\.,λ=0\.25\\lambda=0\.25\)\. The hyper\-parameterλ\\lambdaserves as an effective mechanism for balancing performance trade\-offs across different evaluation metrics\. What’s more, Figure[8](https://arxiv.org/html/2607.22632#S4.F8)shows the Best\-of\-N Accuracy curve of RL training\. The performance forλ=0\.75,1\\lambda=0\.75,1shows continuous improvement\. However, it exhibits an initial increase followed by a decline forλ=0\\lambda=0, which is most likely attributable to the “direction blindness” issue inherent to the score reward\.
Figure 8:Left:Best\-of\-N Accuracy curve for RL with differentλ\\lambdaon random 100 samples from VRMBench\.Right:Average score on all dimensions and samples vs refinement round\.Impact of Group Sizemm\.We evaluate the performance when the group size is reduced tom=3m=3while keepingG=8G=8\. As shown in Table[3](https://arxiv.org/html/2607.22632#S4.T3), although the reduction in group size slightly limits the comparison space, the integration of CR \(SR\+CR\) still yields substantial improvements over the baseline SR\. Specifically, the Comparison accuracy and BoN accuracy increase from 64\.8% and 64\.3% to 68\.8% and 67\.0%, respectively, demonstrating that CR remains highly effective even with fewer candidate plans\.
Sentivity on Rollout NumberGG\.We restrict the rollout number toG=4G=4while maintainingm=4m=4\. Under this resource\-constrained setting, the baseline SR achieves a score of 51\.0, but its comparison and BoN accuracies are limited to 67\.9% and 65\.3%\. When applying our CR, we observe a significant boost in comparison accuracy \(72\.0%\) and BoN accuracy \(66\.8%\)\. We note a minor decrease in the absolute reward score \(from 51\.0 to 50\.2\), which we attribute to the fact that fewer rollouts \(G=4G=4\) lead to a less stable estimation of the mean baseline in GRPO, thereby slightly perturbing the absolute score scale\. Nevertheless, the consistent gains in comparison and BoN selection highlight the robust preference\-distinguishing capability of CR\.
Utilization for TTS Vlog Editing Inference\.Our VlogReward can also be used to enchance the vlog editing performance through Test\-Time Scaling\. To prove this, we use Edits\-Dataset to train a Vlog Editing Model \(VEM\) to generate vlog editing plans for given raw footage\. For evaluation, we randomly sample 100 assets from the VRMBench\. The VEM is used to generateN=8N=8candidate editing plans per asset using a sampling temperature of 0\.5\. We then employ VlogReward to score these candidates and select the plan with the highest total score\. We compare this “Best\-of\-8” selection against a baseline generated via greedy decoding\. Our user study indicates that the TTS approach with VlogReward scoring was preferred in 62% of cases, demonstrating a significant improvement in editing quality over standard greedy decoding\.
Iterative Refinement with Textual Feedback\.To evaluate the quality of the textual feedback provided by VlogReward, we conducte an experiment using the same 100 samples\. Specifically, we compare the feedback capabilities of VlogReward against the base model\. In each iteration, the VRM provides diagnostic critiques, which the VEM then uses to refine the previous editing plan\. As illustrated in Figure[8](https://arxiv.org/html/2607.22632#S4.F8), the average scores of the editing plans refined by VlogReward show a consistent upward trend and achieve higher overall quality\. In contrast, the refinement process guided solely by the base model’s feedback results in unstable score trajectories and lower performance\. These results demonstrate that our VRM provides more accurate and actionable feedback, effectively driving the iterative improvement\.
### 4\.5Transferability to General Preference Learning
To evaluate the generalization capability of our proposed CR beyond the domain of vlog editing, we apply it to a broader reward preference learning task within the Reinforcement Learning from Human Feedback framework\.
Experimental Setup\.We adopt Qwen2\.5\-7B\-Instruct as our backbone model\. For the training and evaluation, we utilize theUltraFeedbackdataset\(Cuiet al\.,[2024](https://arxiv.org/html/2607.22632#bib.bib80)\), a large\-scale, high\-quality benchmark designed for training robust reward models\. Each sample comprises a prompt, four model\-generated responses, and corresponding fine\-grained human/GPT\-4 feedback scored across four dimensions: instruction\-following, truthfulness, honesty, and helpfulness\. In our setup, we allocate 40k samples for SFT, 10k samples for GRPO, and reserve 1k samples for evaluation\.
Results and Discussion\.Table[4](https://arxiv.org/html/2607.22632#S4.T4)summarizes the evaluation reults on raw reward score accuracy, pairwise comparison accuracy, and Best\-of\-N \(BoN\) selection accuracy\. Upon integrating SR\+CR, the performance significantly improves across all metrics, achieving 47\.0 in score, 56\.4% in comparison accuracy, and 53\.6% in BoN accuracy\. These results demonstrate that our CR mechanism is not limited to vlog editing but generalizes effectively to general preference alignment tasks, successfully mitigating direction blindness in broader policy optimization scenarios\.
Table 4:Generalization results on theUltraFeedbackdataset\.MethodRewardScoreComparisonBoNSFT—43\.345\.944\.0GRPOSR46\.154\.751\.2GRPOSR\+CR47\.056\.453\.6
## 5Conclusion
In this paper, we establish a comprehensive six\-dimensional taxonomy for vlog editing evaluation, curate a large\-scale training dataset and VRMBench benchmark, and develop a robust Vlog Reward Model that provides both fine\-grained scores and diagnostic feedback\. Our proposed inter\-group comparison reward effectively alleviates the ”direction blindness” issue of standard GRPO, enabling superior performance in distinguishing varied\-quality edits\.
## Acknowledgement
We thank Jin Liu for his insights and technical assistance, and the strong support and efforts of all creators, managers, and annotators\. The work is supported by New Generation Artificial Intelligence\-National Science and Technology Major Project \(2025ZD0123505\), National Natural Science Foundation of China \(Grant Nos\. 62576338, 62550062, 62425606, 32341009, 62506362\), and Beijing Natural Science Foundation \(Grant Nos\. L252145, L257008\)\.
## Impact Statement
This research utilizes personal video blog content, which inherently involves sensitive data including individuals’ likenesses, personal environments, activities, and associated metadata\. We acknowledge the potential privacy risks, such as unauthorized identity exposure or inference of personal habits, especially if such data were to be mishandled or released without safeguards\. To mitigate these risks and ensure ethical compliance, our data collection and usage strictly adhere to the following protocols:
- •Data Compliance and Ethical Adherence:All personal vlog content incorporated into our dataset is collected from public platform under permission of the original creators, in accordance with relevant copyright and license\.
- •Data Anonymization:During preprocessing, we implement measures to anonymize sensitive elements where feasible, focusing the model’s evaluation on editing structure and narrative quality rather than personal identity\.
- •Controlled Usage Scope:The dataset is used solely for training and benchmarking vlog reward models in this research\. It should not be publicly distributed in its raw form to prevent potential misuse\.
Research Implications\.Our contributions support the research community by providing a standardized evaluation criteria and a benchmark for vlog editing assessment\. For creators, VlogReward offers an assistive tool to generate detailed, interpretable feedback on editing drafts, potentially lowering barriers to high\-quality production and enhancing creative workflows\.
However, automated evaluation systems may inadvertently promote stylistic homogenization or embed biases present in the training data\. To mitigate this, we advocate for VlogReward’s use as a supportive tool within a human\-in\-the\-loop creative process, not as an autonomous arbiter of quality\. The multi\-dimensional, feedback\-rich output is designed to augment, not replace human judgment\. Future work must focus on expanding dataset diversity across cultures and genres, implementing debiasing techniques, and ensuring transparency in model limitations\.
JSON Editing Plans vs Rendered Vlogs\.We initially compared the 2 approaches in detail, and chose JSON for two primary reasons:
- •Efficiency in Production Scenarios:In real\-world applications like ”one\-click” vlog creation in editing software, the internal automated evaluation algorithm must quickly select the best among multiple AI\-generated plans\. Waiting to render each plan into a final video would cause unacceptable delays\. Since all real\-world editing actions can be represented as structured data, using JSON would be more immediate and efficient\.
- •Capturing Fine\-Grained Editing Details:Current multimodal models struggle to accurately perceive dynamic effects in edited vlogs due to their discrete frame\-sampling nature\. In contrast, a JSON file explicitly documents every editing operation, allowing the model to understand all actions more reliably\.
## References
- S\. Bai, K\. Chen, X\. Liu, J\. Wang, W\. Ge, S\. Song, K\. Dang, P\. Wang, S\. Wang, J\. Tang,et al\.\(2025\)Qwen2\. 5\-vl technical report\.arXiv preprint arXiv:2502\.13923\.Cited by:[§4\.1](https://arxiv.org/html/2607.22632#S4.SS1.p2.1),[§4\.1](https://arxiv.org/html/2607.22632#S4.SS1.p4.1)\.
- R\. A\. Bradley and M\. E\. Terry \(1952\)Rank analysis of incomplete block designs: i\. the method of paired comparisons\.Biometrika\.Cited by:[2nd item](https://arxiv.org/html/2607.22632#S4.I1.i2.p1.1)\.
- X\. Chen, G\. Li, Z\. Wang, B\. Jin, C\. Qian, Y\. Wang, H\. Wang, Y\. Zhang, D\. Zhang, T\. Zhang,et al\.\(2025\)Rm\-r1: reward modeling as reasoning\.arXiv preprint arXiv:2505\.02387\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p2.1)\.
- G\. Comanici, E\. Bieber, M\. Schaekermann, I\. Pasupat, N\. Sachdeva, I\. Dhillon, M\. Blistein, O\. Ram, D\. Zhang, E\. Rosen,et al\.\(2025\)Gemini 2\.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.arXiv preprint arXiv:2507\.06261\.Cited by:[§1](https://arxiv.org/html/2607.22632#S1.p1.1)\.
- G\. Cui, L\. Yuan, N\. Ding, G\. Yao, B\. He, W\. Zhu, Y\. Ni, G\. Xie, R\. Xie, Y\. Lin,et al\.\(2024\)ULTRAFEEDBACK: boosting language models with scaled ai feedback\.InProceedings of the 41st International Conference on Machine Learning,Cited by:[§4\.5](https://arxiv.org/html/2607.22632#S4.SS5.p2.1)\.
- J\. Duan, S\. Liu, Y\. Hao, H\. Huang, and R\. He \(2026\)Dual frequency\-guided spatiotemporal feature learning for face forgery detection\.IEEE Trans\. Biom\. Behav\. Identity Sci\.8\(2\),pp\. 179–191\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p2.1)\.
- X\. Duan, C\. Lou, and H\. K\. Kim \(2025\)Vlog my life: how vloggers’ self\-disclosure and co\-creation with consumers influence relationship\-building and advertising outcomes\.International Journal of Advertising\.Cited by:[§1](https://arxiv.org/html/2607.22632#S1.p1.1)\.
- C\. Fu, Y\. Dai, Y\. Luo, L\. Li, S\. Ren, R\. Zhang, Z\. Wang, C\. Zhou, Y\. Shen, M\. Zhang,et al\.\(2025\)Video\-mme: the first\-ever comprehensive evaluation benchmark of multi\-modal llms in video analysis\.InProceedings of the Computer Vision and Pattern Recognition Conference,pp\. 24108–24118\.Cited by:[§1](https://arxiv.org/html/2607.22632#S1.p3.1)\.
- B\. Gan, X\. Shu, R\. Qiao, H\. Wu, K\. Chen, H\. Li, and B\. Ren \(2023\)Collaborative noisy label cleaner: learning scene\-aware trailers for multi\-modal highlight detection in movies\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p1.1)\.
- Google DeepMind \(2025\)o4\-mini\.[https://deepmind\.google/models/gemini/pro](https://deepmind.google/models/gemini/pro)\.Cited by:[§4\.1](https://arxiv.org/html/2607.22632#S4.SS1.p4.1)\.
- H\. Guo, J\. Pan, and W\. Zhai \(2025a\)GRPOformer: advancing hyperparameter optimization via group relative policy optimization\.arXiv preprint arXiv:2509\.17105\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p3.1),[§3\.4](https://arxiv.org/html/2607.22632#S3.SS4.p1.4)\.
- J\. Guo, Z\. Chi, L\. Dong, Q\. Dong, X\. Wu, S\. Huang, and F\. Wei \(2025b\)Reward reasoning model\.arXiv preprint arXiv:2505\.14674\.Cited by:[§1](https://arxiv.org/html/2607.22632#S1.p6.1)\.
- W\. Hong, W\. Yu, X\. Gu, G\. Wang, G\. Gan, H\. Tang, J\. Cheng, J\. Qi, J\. Ji, L\. Pan,et al\.\(2025\)GLM\-4\.1 v\-thinking: towards versatile multimodal reasoning with scalable reinforcement learning\.arXiv preprint arXiv:2507\.01006\.Cited by:[§1](https://arxiv.org/html/2607.22632#S1.p1.1)\.
- P\. Hu, N\. Xiao, F\. Li, Y\. Chen, and R\. Huang \(2023\)A reinforcement learning\-based automatic video editing method using pre\-trained vision\-language model\.InProceedings of the 31st ACM International Conference on Multimedia,Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p1.1)\.
- A\. Hurst, A\. Lerer, A\. P\. Goucher, A\. Perelman, A\. Ramesh, A\. Clark, A\. Ostrow, A\. Welihinda, A\. Hayes, A\. Radford,et al\.\(2024\)GPT\-4o system card\.arXiv preprint arXiv:2410\.21276\.Cited by:[§4\.1](https://arxiv.org/html/2607.22632#S4.SS1.p4.1)\.
- Y\. Kirstain, A\. Polyak, U\. Singer, S\. Matiana, J\. Penna, and O\. Levy \(2023\)Pick\-a\-pic: an open dataset of user preferences for text\-to\-image generation\.Advances in neural information processing systems\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p2.1)\.
- R\. Ladhari, E\. Massa, and H\. Skandrani \(2020\)YouTube vloggers’ popularity and influence: the roles of homophily, emotional attachment, and expertise\.Journal of Retailing and Consumer Services\.Cited by:[§1](https://arxiv.org/html/2607.22632#S1.p1.1)\.
- P\. Langley \(2000\)Crafting papers on machine learning\.InProceedings of the 17th International Conference on Machine Learning \(ICML 2000\),P\. Langley \(Ed\.\),Stanford, CA,pp\. 1207–1216\.Cited by:[Appendix E](https://arxiv.org/html/2607.22632#A5.p11.1)\.
- B\. Li, Y\. Zhang, D\. Guo, R\. Zhang, F\. Li, H\. Zhang, K\. Zhang, P\. Zhang, Y\. Li, Z\. Liu,et al\.\(2024\)Llava\-onevision: easy visual task transfer\.arXiv preprint arXiv:2408\.03326\.Cited by:[§4\.1](https://arxiv.org/html/2607.22632#S4.SS1.p4.1)\.
- J\. Liu, G\. Liu, J\. Liang, Z\. Yuan, X\. Liu, M\. Zheng, X\. Wu, Q\. Wang, M\. Xia, X\. Wang,et al\.\(2025a\)Improving video generation with human feedback\.arXiv preprint arXiv:2501\.13918\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p2.1)\.
- Y\. Liu, J\. Cao, Z\. Li, R\. He, and T\. Tan \(2025b\)Breaking mental set to improve reasoning through diverse multi\-agent debate\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p2.1)\.
- Y\. Liu, Z\. Li, Z\. Fang, N\. Xu, R\. He, and T\. Tan \(2025c\)Rethinking the role of prompting strategies in llm test\-time scaling: a perspective of probability theory\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 27962–27994\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p2.1)\.
- N\. Madan, A\. Møgelmose, R\. Modi, Y\. S\. Rawat, and T\. B\. Moeslund \(2024\)Foundation models for video understanding: a survey\.arXiv preprint arXiv:2405\.03770\.Cited by:[§1](https://arxiv.org/html/2607.22632#S1.p2.1)\.
- Open AI \(2025\)o4\-mini\.[https://platform\.openai\.com/docs/models/o4\-mini](https://platform.openai.com/docs/models/o4-mini)\.Cited by:[§4\.1](https://arxiv.org/html/2607.22632#S4.SS1.p4.1)\.
- A\. Pardo, J\. Wang, B\. Ghanem, J\. Sivic, B\. Russell, and F\. C\. Heilbron \(2024\)Generative timelines for instructed visual assembly\.arXiv preprint arXiv:2411\.12293\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p1.1)\.
- M\. Sandoval\-Castaneda, B\. Russell, J\. Sivic, G\. Shakhnarovich, and F\. Caba Heilbron \(2025\)EditDuet: a multi\-agent system for video non\-linear editing\.InProceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers,Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p1.1)\.
- J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. Klimov \(2017\)Proximal policy optimization algorithms\.arXiv preprint arXiv:1707\.06347\.Cited by:[Appendix B](https://arxiv.org/html/2607.22632#A2.p2.6)\.
- Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. Li, Y\. Wu,et al\.\(2024\)Deepseekmath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[§1](https://arxiv.org/html/2607.22632#S1.p6.1)\.
- A\. Singh, A\. Fry, A\. Perelman, A\. Tart, A\. Ganesh, A\. El\-Kishky, A\. McLaughlin, A\. Low, A\. Ostrow, A\. Ananthram,et al\.\(2025\)OpenAI GPT\-5 System Card\.arXiv preprint arXiv:2601\.03267\.Cited by:[§4\.1](https://arxiv.org/html/2607.22632#S4.SS1.p4.1)\.
- Z\. Sun, S\. Shen, S\. Cao, H\. Liu, C\. Li, Y\. Shen, C\. Gan, L\. Gui, Y\. Wang, Y\. Yang,et al\.\(2024\)Aligning large multimodal models with factually augmented rlhf\.InFindings of the Association for Computational Linguistics: ACL 2024,Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p2.1)\.
- Y\. Tang, J\. Bi, S\. Xu, L\. Song, S\. Liang, T\. Wang, D\. Zhang, J\. An, J\. Lin, R\. Zhu,et al\.\(2025\)Video understanding with large language models: a survey\.IEEE Transactions on Circuits and Systems for Video Technology\.Cited by:[§1](https://arxiv.org/html/2607.22632#S1.p2.1)\.
- K\. Team, A\. Du, B\. Yin, B\. Xing, B\. Qu, B\. Wang, C\. Chen, C\. Zhang, C\. Du, C\. Wei, C\. Wang, D\. Zhang, D\. Du, D\. Wang, E\. Yuan, E\. Lu, F\. Li, F\. Sung, G\. Wei, G\. Lai, H\. Zhu, H\. Ding, H\. Hu, H\. Yang, H\. Zhang, H\. Wu, H\. Yao, H\. Lu, H\. Wang, H\. Gao, H\. Zheng, J\. Li, J\. Su, J\. Wang, J\. Deng, J\. Qiu, J\. Xie, J\. Wang, J\. Liu, J\. Yan, K\. Ouyang, L\. Chen, L\. Sui, L\. Yu, M\. Dong, M\. Dong, N\. Xu, P\. Cheng, Q\. Gu, R\. Zhou, S\. Liu, S\. Cao, T\. Yu, T\. Song, T\. Bai, W\. Song, W\. He, W\. Huang, W\. Xu, X\. Yuan, X\. Yao, X\. Wu, X\. Zu, X\. Zhou, X\. Wang, Y\. Charles, Y\. Zhong, Y\. Li, Y\. Hu, Y\. Chen, Y\. Wang, Y\. Liu, Y\. Miao, Y\. Qin, Y\. Chen, Y\. Bao, Y\. Wang, Y\. Kang, Y\. Liu, Y\. Du, Y\. Wu, Y\. Wang, Y\. Yan, Z\. Zhou, Z\. Li, Z\. Jiang, Z\. Zhang, Z\. Yang, Z\. Huang, Z\. Huang, Z\. Zhao, and Z\. Chen \(2025\)Kimi\-VL technical report\.External Links:2504\.07491,[Link](https://arxiv.org/abs/2504.07491)Cited by:[§4\.1](https://arxiv.org/html/2607.22632#S4.SS1.p4.1)\.
- Q\. Team \(2025\)Qwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§4\.1](https://arxiv.org/html/2607.22632#S4.SS1.p4.1)\.
- W\. Wang, Z\. Gao, L\. Chen, Z\. Chen, J\. Zhu, X\. Zhao, Y\. Liu, Y\. Cao, S\. Ye, X\. Zhu,et al\.\(2025a\)Visualprm: an effective process reward model for multimodal reasoning\.arXiv preprint arXiv:2503\.10291\.Cited by:[§1](https://arxiv.org/html/2607.22632#S1.p4.1)\.
- W\. Wang, Z\. Gao, L\. Gu, H\. Pu, L\. Cui, X\. Wei, Z\. Liu, L\. Jing, S\. Ye, J\. Shao,et al\.\(2025b\)InternVL3\.5: advancing open\-source multimodal models in versatility, reasoning, and efficiency\.arXiv preprint arXiv:2508\.18265\.Cited by:[§4\.1](https://arxiv.org/html/2607.22632#S4.SS1.p4.1)\.
- X\. Wang, X\. Li, Y\. Wei, X\. Song, Y\. Song, X\. Xia, F\. Zeng, Z\. Chen, L\. Liu, G\. Xu,et al\.\(2025c\)From long videos to engaging clips: a human\-inspired video editing framework with multimodal narrative understanding\.arXiv preprint arXiv:2507\.02790\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p1.1)\.
- Y\. Wang, Z\. Li, Y\. Zang, C\. Wang, Q\. Lu, C\. Jin, and J\. Wang \(2025d\)Unified multimodal chain\-of\-thought reward model through reinforcement fine\-tuning\.arXiv preprint arXiv:2505\.03318\.Cited by:[§1](https://arxiv.org/html/2607.22632#S1.p6.1)\.
- Y\. Wang, Y\. Zang, H\. Li, C\. Jin, and J\. Wang \(2025e\)Unified reward model for multimodal understanding and generation\.arXiv preprint arXiv:2503\.05236\.Cited by:[§1](https://arxiv.org/html/2607.22632#S1.p4.1)\.
- J\. Wu, Y\. Gao, Z\. Ye, M\. Li, L\. Li, H\. Guo, J\. Liu, Z\. Xue, X\. Hou, W\. Liu,et al\.\(2025\)Rewarddance: reward scaling in visual generation\.arXiv preprint arXiv:2509\.08826\.Cited by:[§1](https://arxiv.org/html/2607.22632#S1.p4.1)\.
- X\. Wu, K\. Sun, F\. Zhu, R\. Zhao, and H\. Li \(2023\)Human preference score: better aligning text\-to\-image models with human preference\.InProceedings of the IEEE/CVF International Conference on Computer Vision,Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p2.1)\.
- C\. Xiao, M\. Zhang, and Y\. Cao \(2025\)BNPO: beta normalization policy optimization\.arXiv preprint arXiv:2506\.02864\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p3.1)\.
- L\. Xiaomi \(2025\)MiMo\-vl technical report\.External Links:2506\.03569,[Link](https://arxiv.org/abs/2506.03569)Cited by:[§4\.1](https://arxiv.org/html/2607.22632#S4.SS1.p4.1)\.
- J\. Xie, X\. Chen, T\. Zhang, Y\. Zhang, S\. Lu, P\. Cesar, and Y\. Yang \(2022\)Multimodal\-based and aesthetic\-guided narrative video summarization\.IEEE Transactions on Multimedia\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p1.1)\.
- J\. Xu, X\. Liu, Y\. Wu, Y\. Tong, Q\. Li, M\. Ding, J\. Tang, and Y\. Dong \(2023\)Imagereward: learning and evaluating human preferences for text\-to\-image generation\.Advances in Neural Information Processing Systems36,pp\. 15903–15935\.Cited by:[§1](https://arxiv.org/html/2607.22632#S1.p3.1)\.
- T\. Xu, S\. Yao, C\. Dong, Y\. Jin, Z\. Huang, D\. Ou, and H\. Tang \(2025\)DyKnow\-rag: dynamic knowledge utilization reinforcement framework for noisy retrieval\-augmented generation in e\-commerce search relevance\.arXiv preprint arXiv:2510\.11122\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p3.1)\.
- Z\. Xue, J\. Wu, Y\. Gao, F\. Kong, L\. Zhu, M\. Chen, Z\. Liu, W\. Liu, Q\. Guo, W\. Huang,et al\.\(2025\)DanceGRPO: unleashing grpo on visual generation\.arXiv preprint arXiv:2505\.07818\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p2.1)\.
- L\. Yang, Z\. Chen, X\. Li, P\. Jia, L\. Long, and J\. Yang \(2024\)Agent\-based video trimming\.arXiv preprint arXiv:2412\.09513\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p1.1)\.
- Q\. Yu, Z\. Zhang, R\. Zhu, Y\. Yuan, X\. Zuo, Y\. Yue, W\. Dai, T\. Fan, G\. Liu, L\. Liu,et al\.\(2025\)Dapo: an open\-source llm reinforcement learning system at scale\.arXiv preprint arXiv:2503\.14476\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p3.1)\.
- Y\. Zang, X\. Dong, P\. Zhang, Y\. Cao, Z\. Liu, S\. Ding, S\. Wu, Y\. Ma, H\. Duan, W\. Zhang,et al\.\(2025\)Internlm\-xcomposer2\. 5\-reward: a simple yet effective multi\-modal reward model\.arXiv preprint arXiv:2501\.12368\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p2.1)\.
- B\. Zhang, H\. Li, T\. Zhang, J\. Li, C\. Yan, X\. Liu, J\. Cai, and Y\. Hao \(2026\)Improving the reasoning of multi\-image grounding in mllms via reinforcement learning\.InICASSP 2026\-2026 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p2.1)\.
- B\. Zhang, K\. Li, Z\. Cheng, Z\. Hu, Y\. Yuan, G\. Chen, S\. Leng, Y\. Jiang, H\. Zhang, X\. Li,et al\.\(2025a\)Videollama 3: frontier multimodal foundation models for image and video understanding\.arXiv preprint arXiv:2501\.13106\.Cited by:[§1](https://arxiv.org/html/2607.22632#S1.p1.1)\.
- H\. Zhang, R\. Zheng, Z\. Yi, Z\. Zhang, H\. Peng, H\. Wang, Z\. Yuan, C\. Ke, S\. Chen, J\. Yang,et al\.\(2025b\)GEPO: group expectation policy optimization for stable heterogeneous reinforcement learning\.arXiv preprint arXiv:2508\.17850\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p3.1)\.
- Y\. Zhang, X\. Lu, X\. Hu, C\. Fu, B\. Wen, T\. Zhang, C\. Liu, K\. Jiang, K\. Chen, K\. Tang,et al\.\(2025c\)R1\-reward: training multimodal reward model through stable reinforcement learning\.arXiv preprint arXiv:2505\.02835\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p2.1)\.
- Y\. Zhang, H\. Yang, H\. Zhang, Y\. Shi, Z\. Chen, H\. Tian, C\. Fu, H\. Wang, K\. Wu, B\. Cui,et al\.\(2025d\)Basereward: a strong baseline for multimodal reward model\.arXiv preprint arXiv:2509\.16127\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p2.1)\.
- Y\. Zhang, T\. Yu, H\. Tian, C\. Fu, P\. Li, J\. Zeng, W\. Xie, Y\. Shi, H\. Zhang, J\. Wu,et al\.\(2025e\)MM\-RLHF: the next step forward in multimodal llm alignment\.arXiv preprint arXiv:2502\.10391\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p2.1)\.
- Y\. Zhang, B\. Li, h\. Liu, Y\. j\. Lee, L\. Gui, D\. Fu, J\. Feng, Z\. Liu, and C\. Li \(2024\)LLaVA\-next: a strong zero\-shot video understanding model\.External Links:[Link](https://llava-vl.github.io/blog/2024-04-30-llava-next-video/)Cited by:[§4\.1](https://arxiv.org/html/2607.22632#S4.SS1.p4.1)\.
- Z\. Zhang, X\. Huang, J\. Xu, Z\. Luo, X\. Wang, J\. Wei, and X\. Chen \(2025f\)VideoRewardBench: comprehensive evaluation of multimodal reward models for video understanding\.arXiv preprint arXiv:2509\.00484\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p2.1)\.
- Y\. Zhao, Y\. Liu, J\. Liu, J\. Chen, X\. Wu, Y\. Hao, T\. Lv, S\. Huang, L\. Cui, Q\. Ye,et al\.\(2025\)Geometric\-mean policy optimization\.arXiv preprint arXiv:2507\.20673\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p3.1)\.
- C\. Zheng, S\. Liu, M\. Li, X\. Chen, B\. Yu, C\. Gao, K\. Dang, Y\. Liu, R\. Men, A\. Yang,et al\.\(2025\)Group sequence policy optimization\.arXiv preprint arXiv:2507\.18071\.Cited by:[§2](https://arxiv.org/html/2607.22632#S2.p3.1)\.
## Appendix ALimitations
##### Scope of Vlog Types and Generalization of Editing Attributes\.
Our work primarily focuses on montage\-style vlogs, a dominant but specific genre\. The evaluation framework, dataset, and model may not fully capture the unique narrative structures and editing conventions of other popular formats, such as sit\-down monologues, detailed tutorials, or highly stylized cinematic pieces\. Future work should expand the taxonomy and benchmark to encompass a broader spectrum of vlog genres\. Furthermore, the editable attributes in our editing plan representation are standardized to core structural elements \(e\.g\., cut timestamps, shot order, subtitles\) to accommodate MLLM processing\. This design intentionally excludes fine\-grained stylistic controls \(e\.g\., specific transition effects, color grading, complex overlay animations, music\) that professional editors frequently manipulate\. Extending the model to understand and evaluate these richer attributes is crucial for real\-world applications\.
##### Inherent Subjectivity and the “Ground Truth” Challenge\.
Evaluating creative work is fundamentally subjective\. Our six\-dimensional taxonomy, though developed with experts, cannot fully codify human artistic judgment\. The “ground truth” scores in VRMBench, while carefully curated, represent a consensus rather than an absolute standard\. The model’s performance is therefore measured against this specific consensus, which may not align with all individual viewers’ or creators’ perceptions\. The task itself has no single absolutely correct answer\.
##### Visual Processing Constraints\.
To ensure training efficiency, our model processes videos at a low frame rate \(1 fps\) and a limited resolution\. This coarse visual representation may cause the model to miss subtle cinematic details—such as precise motion, fleeting expressions, or nuanced lighting, that are critical for a professional evaluation of cinematography and pacing\. This represents a trade\-off between computational feasibility and granular visual understanding\. Nevertheless, all our experiments were conducted under identical parameter constraints, thereby ensuring fairness in the evaluation and validating the effectiveness of our proposed inter\-group comparison reward\.
## Appendix BMore Detailed Implementation Settings
In this section, we provide more details about our implementations\.
For GRPO training, we optimize the current policy modelπθ\\pi\_\{\\theta\}by maximizing the following objective
J\(θ\)=𝔼q,\{oi\}i=1G∼πθold\(O∣q\)\[1G∑i=1Gmin\(πθ\(oi∣q\)πθold\(oi∣q\)Ai,clip\(πθ\(oi∣q\)πθold\(oi∣q\),1−ε,1\+ε\)Ai\)−βDKL\(πθ∥πref\)\],J\(\\theta\)=\\mathbb\{E\}\_\{q,\\\{o\_\{i\}\\\}\_\{i=1\}^\{G\}\\sim\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(O\\mid q\)\}\\Biggl\[\\frac\{1\}\{G\}\\sum\_\{i=1\}^\{G\}\\min\\\!\\Bigl\(\\frac\{\\pi\_\{\\theta\}\(o\_\{i\}\\mid q\)\}\{\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(o\_\{i\}\\mid q\)\}\\,A\_\{i\},\\,\\mathrm\{clip\}\\\!\\Bigl\(\\frac\{\\pi\_\{\\theta\}\(o\_\{i\}\\mid q\)\}\{\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(o\_\{i\}\\mid q\)\},1\-\\varepsilon,\\,1\+\\varepsilon\\Bigr\)A\_\{i\}\\Bigr\)\-\\,\\beta\\,D\_\{\\mathrm\{KL\}\}\\bigl\(\\pi\_\{\\theta\}\\,\\big\\\|\\,\\pi\_\{\\mathrm\{ref\}\}\\bigr\)\\Biggr\],\(6\)whereε\\varepsilonandβ\\betaare the clipping hyper\-parameter and coefficient controlling the Kullback–Leibler \(KL\) penalty\(Schulmanet al\.,[2017](https://arxiv.org/html/2607.22632#bib.bib49)\), respectively\.Ai=ri−mean\(\{r1,r2,…,rG\}\)std\(\{r1,r2,…,rG\}\)A\_\{i\}=\\frac\{r\_\{i\}\-\\operatorname\{mean\}\(\\\{r\_\{1\},r\_\{2\},\\dots,r\_\{G\}\\\}\)\}\{\\operatorname\{std\}\(\\\{r\_\{1\},r\_\{2\},\\dots,r\_\{G\}\\\}\)\}is the computed relative advantage using the intra\-group rewards\{r1,r2,⋯,rG\}\\\{r\_\{1\},r\_\{2\},\\cdots,r\_\{G\}\\\}andDKL\(πθ∥πref\)=πref\(oi∣q\)πθ\(oi∣q\)−log\(πref\(oi∣q\)πθ\(oi∣q\)\)−1D\_\{KL\}\(\\pi\_\{\\theta\}\\,\\\|\\,\\pi\_\{\\mathrm\{ref\}\}\)\\\!=\\\!\\frac\{\\pi\_\{\\mathrm\{ref\}\}\(o\_\{i\}\\mid q\)\}\{\\pi\_\{\\theta\}\(o\_\{i\}\\mid q\)\}\\\!\-\\\!\\log\\\!\\Bigl\(\\frac\{\\pi\_\{\\mathrm\{ref\}\}\(o\_\{i\}\\mid q\)\}\{\\pi\_\{\\theta\}\(o\_\{i\}\\mid q\)\}\\Bigr\)\\\!\-\\\!1is the KL divergence\.
For RRM\-MSE, we replace the LLM head with a six‑dimensional linear output head, whereas RRM‑BT requires only a single‑dimensional linear head\. For a fair comparison, we train RRM‑MSE on VRM‑Dataset‑SFT and VRM‑Dataset‑OB, and use the rounded score value for test\. For RRM‑BT, we construct training pairs from each group in VRM‑Dataset‑RL, yielding six pairwise comparisons per raw‑footage group\. We further explore different learning rates for both RRM‑MSE and RRM‑BT, whose results are summarized in Table[5](https://arxiv.org/html/2607.22632#A2.T5), respectively\. We report the relatively best achieved performance in Table[1](https://arxiv.org/html/2607.22632#S3.T1)\. The detailed configuration settings of SFT, GRPO, RRM\-MSE and RRM\-BT are summarized in Table[6](https://arxiv.org/html/2607.22632#A2.T6)\.
Table 5:Performances of RRM\-MSE and RRM\-BT using different learing rates\.ModelsLearning RateScore AccuracyComparison AccuracyBest\-of\-N AccuracySBOBAVGSBOBAVGSBOBAVGRRM\-MSE1e\-425\.423\.324\.435\.526\.531\.025\.824\.625\.21e\-324\.825\.925\.438\.132\.235\.131\.128\.930\.01e\-223\.625\.424\.542\.345\.543\.931\.641\.036\.31e\-123\.429\.826\.643\.044\.843\.928\.037\.532\.82e\-118\.518\.118\.341\.946\.344\.124\.938\.231\.65e\-14\.52\.33\.444\.655\.850\.225\.245\.235\.213\.43\.43\.447\.539\.843\.727\.923\.225\.6RRM\-BT1e\-4\-\-\-44\.270\.157\.124\.549\.437\.01e\-3\-\-\-53\.079\.866\.429\.063\.846\.41e\-2\-\-\-56\.782\.669\.635\.072\.453\.71e\-1\-\-\-58\.483\.170\.836\.569\.553\.02e\-1\-\-\-58\.781\.870\.235\.770\.052\.95e\-1\-\-\-56\.378\.967\.740\.361\.550\.91\-\-\-58\.682\.870\.733\.068\.150\.6Table 6:Detailed configuration settings of SFT, GRPO, RRM\-MSE and RRM\-BT\.ConfigurationSFTGRPORRM\-MSERRM\-BTfreeze\_visual\_encoderTruetune\_mergerFalsefreeze\_llmFalseFalseTrueTruelearning\_rate2e\-61e\-61e\-21e\-1kl\_loss\_coef \(β\\beta\)\-1e\-2\-\-rollout\_number\-8\-\-rollout\_temperature\-1\-\-inference\_temperature0\.010\.01\-\-inference\_top\_p0\.0010\.001\-\-optimizerAdamWAdamW\_betas\(0\.9, 0\.999\)weight\_decay01e\-200warmup\_ratio0\.0300\.050\.05lr\_schedulercosinegroup\_size\-8\-\-batch\_size12864128128number\_of\_epochs2122max\_model\_length163841638461446144max\_total\_pixels64×\\times28×\\times28fps1min\_frames16max\_frames64sample\_num100k40k120k60kDatasetVRM\-Dataset\-SFTVRM\-Dataset\-RLVRM\-Dataset\-SFTVRM\-Dataset\-RLVRM\-Dataset\-OB
## Appendix CExperiments Results Using Thinking Mode
We don’t use R1\-like output mode with<think\>\.\.\.</think\>and<answer\>\.\.\.</answer\>, as we find it does not improve the performance, yet significantly increasing the overhead\. Similar to R1\-series work, we set the system prompt as follows when testing the performance in this thinking mode\.
A conversation between User and Assistant\. The user asks a question, and the assistant solves it\. The assistant first thinks about the reasoning process in the mind and then provides the user with the answer\. The reasoning process and answer are enclosed within <think\> </think\> and <answer\> </answer\> tags, respectively, i\.e\., <think\> reasoning process here </think\><answer\> answer here </answer\>\.
Experiment results are reported in Table[7](https://arxiv.org/html/2607.22632#A3.T7)\. We can see that GRM\-GRPO and VlogReward using thinking mode are comparable to or even worse in overall performance than those without it\. This suggests that the thinking mode is not always suitable for all tasks\. It may be because that vlog evaluation primarily relies on immediate perceptual judgment rather than extended logical reasoning, making the verbose thinking process potentially redundant or distracting\.
Table 7:Accuracy performances of GRM\-GRPO and VlogReward with and without employing the thinking mode\.ModelsModeScore AccuracyComparison AccuracyBest\-of\-N AccuracySBOBAVGSBOBAVGSBOBAVGGRM\-GRPOno thinking44\.056\.550\.354\.884\.069\.443\.786\.265\.0GRM\-GRPOthinking43\.855\.749\.756\.981\.869\.443\.082\.963\.0VlogRewardno thinking45\.956\.751\.360\.486\.673\.548\.386\.867\.6VlogRewardthinking44\.457\.250\.861\.688\.274\.943\.788\.165\.9
## Appendix DData Examples
Figure 9:Examples of edited vlogs in the subjective dataset\. Here we show some segments of the final vlogs edited with corresponding editing plans for the convenience of review\. The vlog script topics are various such as reflections on job and life, urban travel, and game\-style narratives\.Figure 10:Examples of edited vlogs in the objective dataset\. The vlog at the bottom is obtained by tampering with the timestamps of the vlog above, introducing temporal and descriptive errors\.
## Appendix EEvaluation Criteria Details
Under the guidance of professional vlog creators and product managers, we have established a comprehensive taxonomy of six dimensions to evaluate vlog editing plans\. Each dimension is assessed on a 1\-5 integer scale: 5 for outstanding, 4 for excellent, 3 for acceptable, 2 for poor, and 1 for failure\. The specific evaluation emphasis for the six dimensions are summarized as follows\.
- •Creativity & Originality:Evaluates narrative inventiveness, unique perspectives, and the presence of generic clichés\.
- •Script & Visual Plan Consistency:Measures how well the selected visual clips match the specific title, subtitles and voiceover\.
- •Concept Design:Assesses the appropriateness of script elements, theme logic, and the cohesive portrayal of characters and settings\.
- •Cinematography & Clip Selection:Evaluates the aesthetic quality of shot selection and framing, as well as the precision in picking the most expressive segments from raw assets\.
- •Narration & Voice\-Over:Focuses on the quality of the storyline, scriptwriting depth, and the emotional delivery of the narration\.
- •Pacing & Rhythm:Evaluates the temporal rhythm, sequence logic, and the overall chronological flow of the edit\.
When training reward models, the input prompt does not detail the scoring criteria for each dimension\. In contrast, when evaluating other baselines that are not trained on VRMDataset, the input prompt explicitly defines the following scoring rubrics in detail\.
Creativity & OriginalityScore 5 \(Outstanding\)Visionary and masterfully structured\.The narrative is deeply engaging, exceptionally creative, and feels fresh and memorable\. It represents a perfect fusion of idea and execution, offering unique insights\.Score 4 \(Excellent\)Well\-crafted and compelling\.The story has a clear, logical structure and a coherent, engaging plot\. The language is precise and effectively serves the narrative, even if it doesn’t break new creative ground\.Score 3 \(Acceptable\)Functional but lacks impact\.The narrative feels predictable or unengaging, often due to a*flat plot*or*loose structure*\. The story is understandable but fails to leave a lasting impression\.Score 2 \(Poor\)Illogical and confusing\.Key narrative elements are missing or poorly explained, leaving the audience with more questions than answers\. The primary issue is a*lack of clear logic*\.Score 1 \(Failure\)Chaotic and incoherent\.The story completely lacks a discernible plot or structure\. The content is irrelevant to the stated theme\. The primary issue is a*chaotic or non\-existent plot*\.
Script & Visual Plan ConsistencyScore 5 \(Outstanding\)Perfect synergy\.The visuals not only match but elevate the script’s message and emotional intent\. Every frame feels intentional and deeply connected to the narrative, creating a synergistic masterpiece\. The plan is*perfectly consistent*\.Score 4 \(Excellent\)Consistently effective\.The visual plan strongly supports the script\. Nearly all visual choices are well\-justified and enhance the narrative with no significant contradictions\.Score 3 \(Acceptable\)Partially disconnected\.A noticeable gap exists where the visuals don’t fully align with the script’s core message or tone\. For example, the script describes excitement, but the shots are static and low\-energy\. The main storyline, however, remains comprehensible\.Score 2 \(Poor\)Frequently contradictory\.The visual plan often undermines or fails to support the script’s intent, creating confusion and weakening the story’s impact\.Score 1 \(Failure\)Severe conflict\.The visuals are so disconnected from the script that they fundamentally sabotage the intended message\. The primary issue is a*conflicting message*\.
Concept & Asset DesignScore 5 \(Outstanding\)Impeccably immersive world\-building\.The set design, props, and character roles are meticulously planned and seamlessly integrated to create a believable and authentic environment that enhances the story\. Strengths:*flawless set design*,*purposeful props*,*authentic character roles*\.Score 4 \(Excellent\)Cohesive and well\-chosen\.The set design, props, and other assets create a consistent and believable world that supports the narrative\. All key assets are present and used logically\.Score 3 \(Acceptable\)Inconsistent\.Some conceptual elements, character roles, or props feel out of place or conflict with the storyline\. For example, a character meant to be a medieval knight is shown using a smartphone\. Key issues:*illogical role settings*or*inconsistent scene/prop usage*\.Score 2 \(Poor\)Confusing and poorly chosen\.The conceptual assets are mismatched, creating a nonsensical visual experience that actively detracts from the story’s credibility\.Score 1 \(Failure\)Critical assets are missing or illogical\.Key character roles or props are either absent entirely or so inappropriate that they cause a total breakdown in narrative logic\. Key issues:*missing key roles/props*\.
Cinematography & Clip SelectionScore 5 \(Outstanding\)Visionary and precise\.The cinematography is visually stunning \(composition, movement, lens choice\)\. The clip selection is*perfectly timed*to capture the absolute peak of action, emotion, or narrative importance\. Every selected segment is the most impactful choice from the raw footage\. Strengths:*masterful cinematography*,*precise and impactful clip selection*\.Score 4 \(Excellent\)Professional and effective\.The cinematography is well\-executed, and the clip selection is logical and accurate\. The chosen segments effectively convey the intended action and narrative without significant omissions\.Score 3 \(Acceptable\)Functional but suboptimal\.The cinematography is uninspired \(e\.g\.,*simple shot composition*,*lack of dynamism*\)\. The clip selection is relevant but often*misses the key moments*, trimming a shot too early or letting it run too long, weakening its impact\.Score 2 \(Poor\)Technically weak selection\.Shots are frequently out of focus or poorly framed\. The chosen clips often obscure essential information or focus on irrelevant action, failing to capture the core of the moment\.Score 1 \(Failure\)Chaotic and illogical\.The camera work is incoherent, and the chosen clips are either completely irrelevant to the script or trimmed so poorly that they make no sense, resulting in a baffling visual experience\.
Narration & Voice\-OverScore 5 \(Outstanding\)Captivating and flawless\.The writing is exceptional, and the delivery is emotionally resonant with a distinctive style\. The narration adds profound thematic depth and elevates the entire production\. Strengths:*brilliant writing*,*rich emotional delivery*,*deepened thematic meaning*\.Score 4 \(Excellent\)Clear and professional\.The narration effectively tells the story with fluent language and a consistent tone that perfectly matches the content\. The audio is clean and well\-mixed\.Score 3 \(Acceptable\)Basic but functional\.The content may be superficial, the language bland, or the delivery flat and unemotional\. Key issues:*empty content*,*clumsy wording*, or*narration that doesn’t fit the mood*\.Score 2 \(Poor\)Detracts from the experience\.The narration is filled with clichés, awkward phrasing, or is poorly matched with the visuals, actively weakening the final product\.Score 1 \(Failure\)Irrelevant or nonsensical\.The narration is completely disconnected from the story, contains significant factual errors, or is otherwise incomprehensible\.
Pacing & RhythmScore 5 \(Outstanding\)Masterful and dynamic\.The editing has a flawless, instinctive rhythm\. The pacing seamlessly blends shots, sound, and story beats to create a compelling and captivating viewing experience\. The use of all assets is*perfectly integrated*\.Score 4 \(Excellent\)Smooth and effective\.The editing is well\-paced, and the flow logically guides the viewer’s attention and supports the narrative’s emotional arc\. Transitions are purposeful and support the story\.Score 3 \(Acceptable\)Inconsistent\.The editing pace feels uneven, with some sections dragging or feeling rushed\. This occasionally disrupts the narrative flow, but the overall story remains understandable\. Transitions might be abrupt or cliché\.Score 2 \(Poor\)Jarring and amateurish\.Cuts are poorly timed, transitions are awkward, and the overall rhythm feels chaotic, frequently pulling the viewer out of the experience\.Score 1 \(Failure\)Critically flawed\.The editing has a*chaotic flow*that completely disrupts the story’s logic\. Illogical cuts or asset usage make the narrative incomprehensible\.
Vlog Rewarding Prompt## Your Role You are an expert video editor and creative director\. Your task is to provide a rigorous and constructive evaluation of a vlog editing plan\. You will be given a video file containing all raw clips and a detailed editing script\. You must score the plan from 1 to 5 across six core dimensions and provide specific, actionable feedback for each score\. Your feedback must clearly explain what works and what does not, referencing specific timestamps from the script or raw video file\. ## Input 1. 1\.A complete video file:A single video file containing a sequence of all raw clips, stitched together with a 3\-second black screen separating each clip\. 2. 2\.A detailed vlog editing plan:A list containing the selected segments, detailing the cut timestamps, script/narration, shot order, visual effects, and other key editing decisions\. If the duration of a segment’s shot\_timestamp doesn’t match its assigned final\_timestamp, this means adjusting its playback speed \(speed up or slow down\) so its length can match the required final\_timestamp in the edit\. ## Scoring Dimensions and Rubric ### 1\. Creativity & Originality ### 2\. Script & Visual Plan Consistency ### 3\. Concept & Asset Design ### 4\. Cinematography & Clip Selection ### 5\. Narration & Voice\-Over ### 6\. Pacing & Rhythm ## Your Final Task For each of the 6 dimensions, provide a score from 1 to 5 and deliver specific, constructive feedback\. Justify every score with clear examples\. For instance, instead of saying, ‘‘The pacing is off,’’ say, ‘‘In the sequence from 0:45 to 1:10, the editing lingers on wide shots for too long, causing the energy to drop\. The script implies rising tension, which could be better achieved with quicker cuts between the close\-ups in the raw footage\.’’ ## Output Format After reasoning, provide a*plain text*evaluation with scores and feedback for each aspect\. ### Output Example ``` 1. Creativity & Originality Score: 2 Feedback: ... 2. Script & Visual Plan Consistency Score: 1 Feedback: ... 3. Concept & Asset Design Score: 4 Feedback: ... 4. Cinematography & Clip Selection Score: 3 Feedback: ... 5. Narration & Voice-Over Score: 3 Feedback: ... 6. Pacing & Rhythm Score: 1 Feedback: ... ``` ## Vlog Editing Plan \{vlog\_editing\_plan\}Similar Articles
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
This paper introduces a Multi-dimensional Evaluation-Verification Reward (EVR) for reinforcement learning fine-tuning of multi-reference image editing models, improving visual consistency and harmony.
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
This paper introduces a paradigm where Vision-Language Models (VLMs) act as test-time teachers to guide Video Generation Models (VGMs) via differentiable rewards and LoRA optimization, achieving a 16.7-point average improvement on video reasoning benchmarks.
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
VEFX-Bench introduces a large-scale human-annotated video editing dataset (5,049 examples) with multi-dimensional quality labels and a specialized reward model for standardized evaluation of video editing systems. The paper addresses the lack of comprehensive benchmarks in AI-assisted video creation by providing VEFX-Dataset, VEFX-Reward, and a 300-video-prompt benchmark that reveals gaps in current editing models.
Video Models Can Reason with Verifiable Rewards
VideoRLVR optimizes video diffusion models for verifiable reasoning tasks using reinforcement learning with rule-based rewards, achieving better performance than supervised methods in constraint-satisfying video generation.
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
Introduces Edit-Compass and EditReward-Compass, a unified benchmark suite for evaluating image editing models and reward models, with 2,388 annotated instances and 2,251 preference pairs for realistic RL scenarios.