Kirin: Animal Motion Generation from In-the-Wild Video
Summary
Kirin reconstructs 3D animal motion from in-the-wild videos to create the AiM3D dataset and generate text- and image-conditioned motion for animating 3D meshes.
View Cached Full Text
Cached at: 09/03/26, 03:50 AM
Paper page - Kirin: Animal Motion Generation from In-the-Wild Video
Source: https://huggingface.co/papers/2609.01823
Abstract
Kirin reconstructs 3D animal motion from video to build a large-scale dataset and generate text- and image-conditioned motion for animating 3D meshes.
Understanding animal motion is fundamental to modeling animal behavior and biomechanics, yet progress in this area lags far behind human motion research due to the scarcity of high-quality motion data. While human motion can be captured in controlled environments, it is impractical for most animal species, resulting in small, domain-limited datasets that restrict downstream applications such as animation. To address this challenge, we introduce Kirin, a framework that reconstructs motion from video, learnsmotion priorsat scale, and generates realistic motion that can be directly applied to animated assets. Using large collections of in-the-wild animal videos, we reconstruct 3D motion sequences and pair them with captions to create AiM3D, the first large-scale dataset offering aligned video-text-motion tuples forquadruped animals. Building on this dataset, we develop avisual-guided motion generationmodel that conditions on both text and image to guide the generation of realistic motion across diverse animal species. Finally, by leveraging an off-the-shelf image-to-3D model, we automatically rig and animate 3D meshes using generated motion, producing ready-to-render animated animals. Together, our dataset and framework establish a new foundation for large-scale, text and image conditioned animal motion generation and animation. Project page: https://kirin-ani.github.io/.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2609\.01823
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.01823 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.01823 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.01823 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
MoZoo:Unleashing Video Diffusion power in animal fur and muscle simulation
MoZoo is a generative diffusion model that synthesizes high-fidelity animal videos from coarse meshes, using novel attention mechanisms and a synthetic-to-real data pipeline.
ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
ARDY introduces a streaming generation framework for real-time, high-fidelity 3D human motion generation controlled by text and kinematic constraints, using a hybrid representation and two-stage autoregressive transformer denoiser.
SAM 3D Animal: Promptable Animal 3D Reconstruction from Images in the Wild
SAM 3D Animal introduces a promptable framework for multi-animal 3D reconstruction from single images in the wild, built on the SMAL+ model, achieving state-of-the-art results on multiple datasets.
Japan accelerates video generation with new series of anime generation models.
Japan's AIdea Labs released AnimeGen, a free AI model for anime-style video generation, capable of text-to-video and image-to-video, with commercial use allowed.
PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
PhyMotion proposes a physics-grounded reward system that evaluates kinematic plausibility, contact consistency, and dynamic feasibility of human motion in generated videos, achieving stronger correlation with human judgment and improving motion realism in RL-based post-training.