VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes

Hugging Face Daily Papers Papers

Summary

VidaForge is an open research infrastructure that models video data recipes as executable workflows to study their impact on video foundation model pretraining, releasing a dataset of 3.14 million scene-level clips for research.

Video foundation models increasingly rely on large-scale pretraining data, yet the end-to-end data pipelines behind them remain largely closed and difficult to inspect or reuse. Researchers seeking to understand how video data recipes affect model pretraining often need to build substantial infrastructure before testing even a focused hypothesis. We present VIDAFORGE, an open research infrastructure that represents a video data recipe as an executable five-stage workflow from raw videos to training datasets. A decision in this workflow can be varied to construct alternative datasets while preserving how every sample was produced. To demon strate this research workflow, we compare data recipes with different coverage and quality in early from-scratch pretraining of Wan 2.1 and V-JEPA 2.1. Across both learning objectives, the broader-coverage recipe achieves the highest downstream benchmark scores, while loss-based evaluation favors different recipes. This study demonstrates how VidaForge connects data-recipe choices to downstream model performance. We further release VIDAFORGE-3M, containing 3.14 million scene level clips totaling 6,475 hours, with fine-grained annotations and curation signals for video data-recipe research.
Original Article
View Cached Full Text

Cached at: 09/09/26, 04:30 AM

Paper page - VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes

Source: https://huggingface.co/papers/2609.06652 Published on Sep 6

·

Submitted byhttps://huggingface.co/ManTle

Yan Maon Sep 9

Abstract

VIDAFORGE is an open infrastructure that models video data recipes as executable workflows to study how data choices affect pretraining of video foundation models.

Video foundation modelsincreasingly rely on large-scalepretrainingdata, yet the end-to-end data pipelines behind them remain largely closed and difficult to inspect or reuse. Researchers seeking to understand how videodata recipesaffect modelpretrainingoften need to build substantial infrastructure before testing even a focused hypothesis. We presentVIDAFORGE, an open research infrastructure that represents a video data recipe as an executable five-stage workflow from raw videos to training datasets. A decision in this workflow can be varied to construct alternative datasets while preserving how every sample was produced. To demon strate this research workflow, we comparedata recipeswith different coverage and quality in early from-scratchpretrainingofWan 2.1andV-JEPA 2.1. Across both learning objectives, the broader-coverage recipe achieves the highestdownstream benchmark scores, whileloss-based evaluationfavors different recipes. This study demonstrates howVidaForgeconnects data-recipe choices to downstream model performance. We further releaseVIDAFORGE-3M, containing 3.14 million scene level clips totaling 6,475 hours, with fine-grained annotations and curation signals for video data-recipe research.

View arXiv pageView PDFProject pageGitHub88Add to collection

Get this paper in your agent:

hf papers read 2609\.06652

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.06652 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.06652 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.06652 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Laion Big Video Dataset

Hacker News Top

LAION-BVD is a massive open video dataset for multimodal pre-training, containing 80 million videos and achieving competitive performance on standard benchmarks, released to support open research in AI.

ID-V2V: Identity-Preserving Video Restylization

Hugging Face Daily Papers

ID-V2V is a video-to-video generative framework for identity-preserving video restylization. It treats identity preservation as a video relighting problem and uses edited keyframes for style propagation, achieving high-quality results without paired training data.

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

Hugging Face Daily Papers

VEFX-Bench introduces a large-scale human-annotated video editing dataset (5,049 examples) with multi-dimensional quality labels and a specialized reward model for standardized evaluation of video editing systems. The paper addresses the lack of comprehensive benchmarks in AI-assisted video creation by providing VEFX-Dataset, VEFX-Reward, and a 300-video-prompt benchmark that reveals gaps in current editing models.