Tag
Wan-Animate-2 is a new end-to-end character animation framework that consumes driving videos directly in a redesigned Diffusion Transformer, achieving high-fidelity motion generation and identity preservation. It also introduces a lightweight variant for real-time streaming animation, with open-source weights released.
EffectLearner is a semantic-reasoning-enhanced framework for real-world video object removal, combining a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser and introducing the EffectWorld dataset to handle complex object-induced effects.
Xiaomi released XR-1, a robot foundation model trained on over 100K hours of real-world manipulation trajectories. Built on Qwen3-VL and a Diffusion Transformer, it enables out-of-the-box mobile manipulation in unseen environments.
Transformer Transformer is a unified model that generates complete robot embodiments optimized for a given manipulation demonstration, using a diffusion transformer trained on RoboTokens and Dynamics Self-Guidance.
WorldDiT is a unified diffusion transformer architecture that couples action generation with visual world modeling, achieving strong performance on LIBERO simulation suites without relying on large pretrained vision-language models.
NVIDIA Research introduces SANA-Video 2.0, a hybrid video diffusion transformer that generates high-quality 720p video on a single GPU, achieving up to 120× speedup over Wan 2.2-14B via hybrid linear-softmax attention and block attention residuals.
SANA-Video 2.0 introduces a hybrid linear-softmax attention mechanism for video diffusion transformers, achieving high-quality video generation up to 720p on a single GPU with significantly reduced latency compared to full-softmax models, while maintaining competitive VBench scores.
Microsoft releases Mage-Flow-Edit-Turbo, a compact 4B-scale generative model for efficient text-to-image generation and instruction-based image editing, achieving state-of-the-art competitive quality through co-designed tokenizer and backbone.
Microsoft releases Mage-Flow, a compact 4B-parameter foundation model for efficient native-resolution text-to-image generation and instruction-based image editing, achieving competitive quality against much larger models.
Mage-Flow is a compact 4B-parameter generative stack for efficient text-to-image generation and instruction-based image editing, featuring a co-designed lightweight tokenizer (Mage-VAE) and a native-resolution multimodal diffusion transformer trained with rectified flow matching. It achieves competitive performance while enabling high-resolution generation at 0.59s on a single A100 GPU.
Proposes M2GDT, a novel MKGC framework that uses an MLLM-guided diffusion transformer with relation-adaptive mixture-of-experts to align and denoise multimodal features, outperforming baselines on three benchmark datasets.
This paper introduces VideoRAE, a representation autoencoder that leverages frozen video foundation models to create compact, reconstruction-capable, and generation-friendly video latents. It achieves state-of-the-art results on UCF-101 with faster convergence than competing autoencoders.
FreyaTTS is a compact, tokenizer-free Turkish-first text-to-speech model based on a non-autoregressive conditional flow-matching Diffusion Transformer, achieving state-of-the-art performance with a fraction of the parameters of larger systems and released under Apache-2.0.
Introduces digital teleoperation using action-conditioned world models to generate diverse training data for robotics, decoupling data collection from physical hardware. The system achieves real-time generation and enables zero-shot Sim2Real transfer.
PointDiT presents a minimalist pixel-space diffusion transformer using a plain ViT architecture for monocular geometry estimation, outperforming complex latent-based models while maintaining simplicity and robustness in ambiguous regions.
NeuWorld is a new interactive video generation system that uses compact neural implicit scene representations and a transformer VAE with diffusion transformer for trajectory-conditioned rendering, achieving long-horizon consistency.
MirrorPPR introduces an exemplar-based portrait retouching framework using Diffusion Transformer with LoRA adaptation and self-augmented training data, achieving superior quality and identity preservation.
EO-WM proposes a video diffusion transformer for probabilistic Earth observation forecasting that incorporates physically informed conditioning to capture weather-driven uncertainties, achieving improved prediction of vegetation indices under extreme weather.
MaineCoon is a 22B real-time text-to-audio-video model that achieves up to 47.5 FPS on a single H100 GPU, enabling low-cost, long-duration streaming with synchronized speech and visuals for live AI characters.
MeshFlow introduces an equivariant optimal-transport flow matching model for direct triangle mesh generation, achieving state-of-the-art quality while providing approximately 18x inference speedup over autoregressive methods.