Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

Hugging Face Daily Papers Papers

Summary

This paper introduces Moving Alphabet, a procedural testbed for controlled experiments on how data distribution and caption quality affect text-to-video models, revealing key insights for data curation.

Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and non-trivial, involving clip selection from raw videos and captioning to create video-text pairs for learning text-to-video mappings. We study how data distribution and caption quality impact text-to-video models. To enable controlled experiments, we introduce Moving Alphabet, a procedural testbed that renders letters with varying fonts, colors, sizes, and positions, moving in different directions and speeds against a black background. This design allows precise control over data distribution and caption quality by corrupting ground-truth metadata. Our experiments yield three findings: a) a diverse and balanced distribution of video content and duration is critical for generalization; b) caption quality significantly affects both model performance and training efficiency, suggesting that text-to-video models are bounded by video understanding capabilities; and c) classifier-free guidance and fine-tuning on high-quality data provide partial recovery from models trained on corrupted captions, but cannot fully compensate for poor pre-training data. We believe these insights can inform the development of large-scale text-to-video models, and we advocate for greater attention to the science of pre-training data.
Original Article
View Cached Full Text

Cached at: 07/24/26, 05:06 AM

Paper page - Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

Source: https://huggingface.co/papers/2607.18789

Abstract

Text-to-videogenerationhasadvancedsignificantlyoverthepastfiveyearsthroughscalingofmodelsize,data,andcompute.Unlikemodelarchitecture,trainingdataisoftenunderexplored.Real-worlddatacurationiscomplexandnon-trivial,involvingclipselectionfromrawvideosandcaptioningtocreatevideo-textpairsforlearningtext-to-videomappings.Westudyhowdatadistributionandcaptionqualityimpacttext-to-videomodels.Toenablecontrolledexperiments,weintroduceMovingAlphabet,aproceduraltestbedthatrendersletterswithvaryingfonts,colors,sizes,andpositions,movingindifferentdirectionsandspeedsagainstablackbackground.Thisdesignallowsprecisecontroloverdatadistributionandcaptionqualitybycorruptingground-truthmetadata.Ourexperimentsyieldthreefindings:a)adiverseandbalanceddistributionofvideocontentanddurationiscriticalforgeneralization;b)captionqualitysignificantlyaffectsbothmodelperformanceandtrainingefficiency,suggestingthattext-to-videomodelsareboundedbyvideounderstandingcapabilities;andc)classifier-freeguidanceandfine-tuningonhigh-qualitydataprovidepartialrecoveryfrommodelstrainedoncorruptedcaptions,butcannotfullycompensateforpoorpre-trainingdata.Webelievetheseinsightscaninformthedevelopmentoflarge-scaletext-to-videomodels,andweadvocateforgreaterattentiontothescienceofpre-trainingdata.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2607\.18789

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2607.18789 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2607.18789 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2607.18789 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Getting video models to learn better, faster

Hacker News Top

The article explores the evolution of data filtering techniques for generative video pre-training, emphasizing methods like computer vision, LLMs, and reinforcement learning to improve model performance through high-quality data.

Video Generation Models are General-Purpose Vision Learners

Hugging Face Daily Papers

This paper proposes that large-scale text-to-video generation can serve as a powerful pre-training paradigm for computer vision, introducing GenCeption which achieves state-of-the-art performance across diverse vision tasks with high data efficiency and emergent generalization to unseen domains.

OSCBench: Benchmarking Object State Change in Text-to-Video Generation

arXiv cs.CL

OSCBench is a new benchmark designed to evaluate text-to-video generation models' ability to accurately represent object state changes (transformations caused by actions like peeling or slicing). The paper reveals that current T2V models struggle with temporally consistent state changes, especially in novel and compositional scenarios, identifying this as a key bottleneck in video generation.

Kathleen Writes: Autoregressive Generation and Data Scaling Without Attention

arXiv cs.CL

This paper from the Kathleen series shows that an attention-free, byte-level model with ~0.5M parameters can beat a parameter-matched transformer on WikiText-103 language modeling and generation, introduces a non-parametric 'Form Distance' metric for evaluating text realism, and demonstrates that retrieval-augmented decoding from the model's own training corpus improves generation quality.

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning

Hugging Face Daily Papers

This paper introduces PadCaptioner, a 3B parameter model for omni-modal dense video captioning that uses parallelized autoregressive decoding to achieve high efficiency and quality, outperforming 7B counterparts. A latent planning mechanism enables lossless parallel generation by exploiting weak local dependencies among events.