Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
Summary
This paper introduces Moving Alphabet, a procedural testbed for controlled experiments on how data distribution and caption quality affect text-to-video models, revealing key insights for data curation.
View Cached Full Text
Cached at: 07/24/26, 05:06 AM
Paper page - Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
Source: https://huggingface.co/papers/2607.18789
Abstract
Text-to-videogenerationhasadvancedsignificantlyoverthepastfiveyearsthroughscalingofmodelsize,data,andcompute.Unlikemodelarchitecture,trainingdataisoftenunderexplored.Real-worlddatacurationiscomplexandnon-trivial,involvingclipselectionfromrawvideosandcaptioningtocreatevideo-textpairsforlearningtext-to-videomappings.Westudyhowdatadistributionandcaptionqualityimpacttext-to-videomodels.Toenablecontrolledexperiments,weintroduceMovingAlphabet,aproceduraltestbedthatrendersletterswithvaryingfonts,colors,sizes,andpositions,movingindifferentdirectionsandspeedsagainstablackbackground.Thisdesignallowsprecisecontroloverdatadistributionandcaptionqualitybycorruptingground-truthmetadata.Ourexperimentsyieldthreefindings:a)adiverseandbalanceddistributionofvideocontentanddurationiscriticalforgeneralization;b)captionqualitysignificantlyaffectsbothmodelperformanceandtrainingefficiency,suggestingthattext-to-videomodelsareboundedbyvideounderstandingcapabilities;andc)classifier-freeguidanceandfine-tuningonhigh-qualitydataprovidepartialrecoveryfrommodelstrainedoncorruptedcaptions,butcannotfullycompensateforpoorpre-trainingdata.Webelievetheseinsightscaninformthedevelopmentoflarge-scaletext-to-videomodels,andweadvocateforgreaterattentiontothescienceofpre-trainingdata.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2607\.18789
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.18789 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.18789 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.18789 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Getting video models to learn better, faster
The article explores the evolution of data filtering techniques for generative video pre-training, emphasizing methods like computer vision, LLMs, and reinforcement learning to improve model performance through high-quality data.
Video Generation Models are General-Purpose Vision Learners
This paper proposes that large-scale text-to-video generation can serve as a powerful pre-training paradigm for computer vision, introducing GenCeption which achieves state-of-the-art performance across diverse vision tasks with high data efficiency and emergent generalization to unseen domains.
OSCBench: Benchmarking Object State Change in Text-to-Video Generation
OSCBench is a new benchmark designed to evaluate text-to-video generation models' ability to accurately represent object state changes (transformations caused by actions like peeling or slicing). The paper reveals that current T2V models struggle with temporally consistent state changes, especially in novel and compositional scenarios, identifying this as a key bottleneck in video generation.
Kathleen Writes: Autoregressive Generation and Data Scaling Without Attention
This paper from the Kathleen series shows that an attention-free, byte-level model with ~0.5M parameters can beat a parameter-matched transformer on WikiText-103 language modeling and generation, introduces a non-parametric 'Form Distance' metric for evaluating text realism, and demonstrates that retrieval-augmented decoding from the model's own training corpus improves generation quality.
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
This paper introduces PadCaptioner, a 3B parameter model for omni-modal dense video captioning that uses parallelized autoregressive decoding to achieve high efficiency and quality, outperforming 7B counterparts. A latent planning mechanism enables lossless parallel generation by exploiting weak local dependencies among events.