Video Generators as General-Purpose Vision Models (8 minute read)

TLDR AI Papers

Summary

GenCeption repurposes pre-trained video generative models into a single unified feed-forward vision model that achieves state-of-the-art performance across multiple tasks with exceptional data efficiency, marking a shift toward general-purpose visual intelligence.

DeepMind researchers introduced GenCeption, which repurposes a pretrained video generation model as a unified vision system controlled through text instructions.
Original Article
View Cached Full Text

Cached at: 07/14/26, 10:55 PM

# Video Generation Models are General-Purpose Vision Learners Source: [https://genception.github.io/](https://genception.github.io/) 1Google DeepMind2University of Toronto3University College London4University of Oxford5MIT6Lund University \*Work done while at Google DeepMind\. ECCV 2026 ## Overview Your browser does not support the video tag\. ## TL;DR Pushing computer vision from the task\-specific era toward general\-purpose visual intelligence — just as NLP evolved from task\-specific models into general\-purpose language intelligence\. GenCeptionrepurposes a pre\-trained video generative model into a single unified, general\-purpose, feed\-forward vision model that solves a wide range of vision tasks with SOTA performance — all steered by text instructions, with exceptional learning efficiency and intriguing emergent behaviors\. ### Visual Generative Pretraining GenCeptionleverages video generation models as representation pretraining, and conducts multi\-task post\-training in a unified architecture\. ### Unified & Feed\-Forward GenCeptionhandles both dense and sparse vision tasks, transforming the multi\-step generative backbone into a single\-step feed\-forward model\. ### SOTA & Emergent GenCeptionis a SOTA unified model on various vision tasks, with exceptional learning efficiency and intriguing emerging behaviors\. ## Methodology Your browser does not support the video tag\. ![Specialized vision models use separate per-task heads, backbones, or losses, whereas GenCeption's unified vision model uses a single head and backbone across all tasks, steered by a task prompt.](https://genception.github.io/assets/paradigm_shift.png) **Methodology\.**GenCeptionleverages a video generative diffusion model as a*pre\-training*base to capture rich spatio\-temporal world priors and native vision\-language alignment at scale\. During multi\-task*post\-training*, the model is adapted to a feed\-forward model fine\-tuned on predominantly synthetic data to handle diverse perception tasks\.GenCeptionshows strong performance on multiple vision tasks with intriguing*Emerging Behaviors*, enabling seamless sim\-to\-real transfer and generalization to out\-of\-distribution object categories\. **Paradigm Shift\.**This highlights a paradigm shift from specialized, task\-specific computer vision models toward fully unified, generalist vision models — just as NLP evolved from task\-specific models into general\-purpose language intelligence\. ## Results ![Radar chart: GenCeption (specialist and generalist) matches or outperforms specialized state-of-the-art models across depth, surface normal, camera pose, segmentation and keypoint tasks.](https://genception.github.io/assets/results_sota.png) ![Data-efficiency plot: GenCeption reaches accuracy comparable to models trained on far more data, using 7x to 500x fewer training frames.](https://genception.github.io/assets/results_efficiency.png) **SOTA performance:**GenCeptionis competitive with or outperforms state\-of\-the\-art models dedicated to individual tasks \(DepthAnything3, D4RT, VGGT\-Ω, SAM3, Sapiens, DAVID, Genmo, Lotus\-2\)\. Our*specialist*denotes a model trained on each task individually, whereas the*generalist*represents a single model trained jointly across multiple tasks\. **Data efficiency:**under matched fine\-tuning data, the video\-generative backbone beats alternative pre\-training paradigms \(V\-JEPA, VideoMAE V2\), shows preliminary scaling properties, where it improves with larger models and more data, and reaches performance comparable to D4RT and VGGT\-Ω with 7× to 500× less training data\. ## Architecture Your browser does not support the video tag\. Architecture overview ofGenCeption, a simple yet powerful architecture adapted from text\-to\-video diffusion models\. Given an input video and a text prompt specifying the desired output, our unified model, trained majorly on synthetic data, is capable of performing a wide range of dense and sparse perception tasks, with a single forward\-pass of the model\. The dense vision tasks are unified in the RGB ambient space where supervision can be applied in latent space efficiently, and the sparse vision tasks are realized by adding learnable tokens as additional inputs to the diffusion transformer \(DiT\)\. ## Video Any\-Task Capabilities GenCeptionis able to seamlessly switch between different vision tasks, just like how humans do in real life\. Your browser does not support the video tag\. ## Native Vision\-Language Alignment Given a natural\-language expression,GenCeptionaccurately segments the referred object — reasoning about colors, spatial relationships and motion — and generalizes to unseen objects \(e\.g\., a rocket\) by leveraging knowledge from text\-to\-video pre\-training\. Your browser does not support the video tag\. ## Understanding the Complex Visual World As an example of pushing the boundaries of understanding the complex visual world,GenCeptionpredicts 4D human keypoints under challenging conditions \(complex motion, occlusion, ego\-centric, multi\-view, etc\.\) — robust 4D human understanding that underpins real\-world applications across robotics and AR/VR\. Your browser does not support the video tag\. ## Grounded 4D Reconstruction From a single video,GenCeptionpredicts per\-pixel geometry and camera pose that lift the scene into a 4D point cloud — enabling free\-viewpoint fly\-throughs and language grounding of objects from a text instruction\. Your browser does not support the video tag\. ## Emergent Behaviors Although fine\-tuned predominantly on synthetic human videos,GenCeptiontransfers seamlessly from simulation to real footage and to out\-of\-distribution categories — evidence of a universal "world model" inside generative video backbones\. Trained purely on synthetic videos, the model transfers**zero\-shot to real\-world footage**— seamless sim\-to\-real transfer\. Trained on synthetic single\-object videos, the model generalizes zero\-shot to real videos with**multiple instances**\. Trained only on humans, it generalizes to**unseen categories**— animals and robots\. ## Citation ``` @inproceedings{wang2026genception, title = {Video Generation Models are General-Purpose Vision Learners}, author = {Wang, Letian and Zhang, Chuhan and Kabra, Rishabh and Uijlings, Jasper and Waslander, Steven and Zisserman, Andrew and Carreira, Joao and He, Kaiming and Andriluka, Misha and Bazavan, Eduard Gabriel and Zanfir, Andrei and Sminchisescu, Cristian}, booktitle = {European Conference on Computer Vision (ECCV)}, year = {2026} } ``` ## Acknowledgements The following acknowledgements are not exhaustive\. We are also grateful to many others whose support, feedback, and discussions contributed to this work but who are not individually acknowledged here\. We are deeply grateful to Jeremiah Harmsen, Joëlle Barral, Chris Dyer, Rahul Sukthankar, Junlin Zhang, Miki Rubinstein, Thabo Beeler, Di Qiu, Jesús Pérez, Alberto García García, Sergio Orts Escolano, Erroll Wood, Iker J\. de los Mozos, and Emily Conn for their invaluable support throughout this project\. We also thank the following individuals \(listed alphabetically by last name\) for insightful research discussions, technical feedback, and valuable conversations on implementation details: Thiemo Alldieck, Yanrui Bin, Ang Cao, Minghao Chen, Carlton Chu, Dima Damen, Shlomi Fruchter, Valentin Gabeur, Robert Geirhos, Tengda Han, Sicong Jiang, Zeren Jiang, Linyi Jin, Ruofan Liang, Ziwei Liao, Haotong Lin, Xingchao Liu, Yibo Liu, Shangbang Long, Viorica Patraucean, Cordelia Schmid, Hao Shao, Jianyuan Wang, Luyu Wang, Thaddäus Wiedemer, Ziyi Wu, Yuxi Xiao, Junyu Xie, Wei Yu, Shangzhan Zhang, and Shuhong Zheng\. Project page for “Video Generation Models are General\-Purpose Vision Learners\.”

Similar Articles

Video Generation Models are General-Purpose Vision Learners

Hugging Face Daily Papers

This paper proposes that large-scale text-to-video generation can serve as a powerful pre-training paradigm for computer vision, introducing GenCeption which achieves state-of-the-art performance across diverse vision tasks with high data efficiency and emergent generalization to unseen domains.

Video generation models as world simulators

OpenAI Blog

OpenAI's technical report on Sora describes a video generation model that unifies diverse visual data through visual patches, enabling large-scale training of generative models capable of producing high-definition videos up to one minute long across variable durations, aspect ratios, and resolutions.

Vision as Unified Multimodal Generation

Hugging Face Daily Papers

This paper presents SenseNova-Vision, a unified multimodal model that formulates computer vision tasks as generation problems, achieving performance comparable to specialized systems across diverse vision tasks. It introduces a large-scale instruction-response corpus and publicly releases the model and datasets.