Motif-Video 2B: Technical Report
Summary
Motif-Video 2B is a 2B parameter text-to-video generation model that achieves 83.76% on VBench, surpassing Wan2.1 14B while using 7x fewer parameters and trained on fewer than 10M clips with less than 100,000 H200 GPU hours. The model uses a specialized architecture with shared cross-attention and a three-part backbone to separate prompt alignment, temporal consistency, and detail refinement.
View Cached Full Text
Cached at: 04/21/26, 07:21 AM
Paper page - Motif-Video 2B: Technical Report
Source: https://huggingface.co/papers/2604.16503 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Motif-Video 2B achieves high text-to-video generation quality using a specialized architecture with shared cross-attention and three-part backbone, along with efficient training methods, while requiring significantly fewer parameters and training data than larger models.
Training strong video generation models usually requires massive datasets, large parameter counts, and substantial compute. In this work, we ask whether strong text-to-video quality is possible at a much smaller budget: fewer than 10M clips and less than 100,000 H200 GPU hours. Our core claim is that part of the answer lies in how model capacity is organized, not only in how much of it is used. In video generation, prompt alignment, temporal consistency, and fine-detail recovery can interfere with one another when they are handled through the same pathway. Motif-Video 2B addresses this by separating these roles architecturally, rather than relying on scale alone. The model combines two key ideas. First,Shared Cross-Attentionstrengthens text control whenvideo token sequencesbecome long. Second, athree-part backboneseparates early fusion, joint representation learning, and detail refinement. To make this design effective under a limited compute budget, we pair it with an efficient training recipe based ondynamic token routingand early-phasefeature alignmentto afrozen pretrained video encoder. Our analysis shows that later blocks develop clearercross-frame attentionstructure than standard single-stream baselines. OnVBench, Motif-Video~2B reaches 83.76\%, surpassing Wan2.1 14B while using 7times fewer parameters and substantially less training data. These results suggest that careful architectural specialization, combined with an efficiency-oriented training recipe, can narrow or exceed the quality gap typically associated with much larger video models.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2604\.16503
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper1
#### Motif-Technologies/Motif-Video-2B Text-to-Video• Updatedabout 2 hours ago • 1.02k • 72
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2604.16503 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2604.16503 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Motif 3: Technical Report
Motif 3 is a 314B-parameter Mixture-of-Experts language model with 13.2B active parameters per token, featuring Grouped Differential Latent Attention and trained on 12.5T tokens, demonstrating competitive performance across reasoning, coding, and long-context tasks.
Motif-Technologies/Motif-3-Beta
Motif Technologies releases an intermediate beta checkpoint of Motif-3, a large-scale Mixture-of-Experts language model with ~314B total parameters (~13B active), 256K context length, and custom architectures like Grouped Differential Latent Attention (GDLA), openly available for non-commercial research.
LongCat-Video Technical Report
LongCat-Video is a 13.6B parameter video generation model based on Diffusion Transformer, supporting text-to-video, image-to-video, and video-continuation tasks with efficient long video generation using coarse-to-fine and block sparse attention.
@_TobiasLee: Seed 2.1 from Bytedance achieved impressive results on two of our benchmarks. Claw-Eval (Multimodal, https://claw-eval.…
ByteDance's Seed 2.1 model achieved strong results on multimodal agentic (Claw-Eval) and long video understanding (Video-MME) benchmarks, though a gap remains between perception and agentic capabilities.
Kwai Keye-VL-2.0 Technical Report
This technical report presents Kwai Keye-VL-2.0, an open-source Mixture-of-Experts multimodal foundation model designed for long-video understanding and agentic intelligence, leveraging DeepSeek Sparse Attention and cross-modal distillation to achieve state-of-the-art performance among similar-scale models.