Tag
Meta AI introduces Muse Realtime Avatar, a state-of-the-art embodiment technology that turns voice into synchronized, expressive avatars in real time using a streaming Diffusion Transformer.
Mira-Scene introduces a compositional 3D scene reconstruction framework using pixel-aligned canonical coordinate maps for accurate object layouts, achieving significant improvements in layout accuracy over existing methods.
Marigold V2 transforms an image-editing Diffusion Transformer into a single-step model for sharp dense prediction, achieving best zero-shot results in depth estimation and related tasks with Apache 2.0 license.
LynnReal-Omni is a unified multimodal video diffusion framework that integrates agentic visual controls with high-fidelity generation and real-time acceleration for stable, controllable video creation.
A detailed write-up on training a 210M text-to-image diffusion transformer from scratch on a single GPU, sharing key measurements on attention sinks, loss as a health signal, and timestep shifting benefits.
StepAudio 3 Music introduces a large-scale, long-form music generation model with explicit musical planning via ABC-CoT, achieving high scores in audio quality and similarity metrics compared to other systems.
UniMate is a unified diffusion transformer model that generates articulated motion for diverse skeletons from text and rigged 3D assets without per-skeleton retraining, using topology-aware attention and a large curated dataset.
OutageDiT is a generative foundation model for power outage forecasting that uses a Diffusion Transformer architecture to generate seven-day outage trajectories, improving forecast accuracy and enabling zero-shot transfer to new regions.
LLaDA-Image presents a unified framework that combines a 6B diffusion transformer with a frozen vision-language module for generating photorealistic images with precise editing, achieving state-of-the-art results among open-source models through efficient training and fast inference.
World Labs introduces Atlas, a multimodal world model that can generate, reconstruct, and simulate 3D environments with camera control and spatial consistency.
Wan-Animate-2 is a new end-to-end character animation framework that consumes driving videos directly in a redesigned Diffusion Transformer, achieving high-fidelity motion generation and identity preservation. It also introduces a lightweight variant for real-time streaming animation, with open-source weights released.
AtlasVLA introduces a dual-memory architecture with persistent world-ego state modeling to overcome perception and task-progress forgetting in vision-language-action models, achieving state-of-the-art long-horizon manipulation from a single wrist camera.
EffectLearner is a semantic-reasoning-enhanced framework for real-world video object removal, combining a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser and introducing the EffectWorld dataset to handle complex object-induced effects.
Xiaomi released XR-1, a robot foundation model trained on over 100K hours of real-world manipulation trajectories. Built on Qwen3-VL and a Diffusion Transformer, it enables out-of-the-box mobile manipulation in unseen environments.
Transformer Transformer is a unified model that generates complete robot embodiments optimized for a given manipulation demonstration, using a diffusion transformer trained on RoboTokens and Dynamics Self-Guidance.
WorldDiT is a unified diffusion transformer architecture that couples action generation with visual world modeling, achieving strong performance on LIBERO simulation suites without relying on large pretrained vision-language models.
NVIDIA Research introduces SANA-Video 2.0, a hybrid video diffusion transformer that generates high-quality 720p video on a single GPU, achieving up to 120× speedup over Wan 2.2-14B via hybrid linear-softmax attention and block attention residuals.
SANA-Video 2.0 introduces a hybrid linear-softmax attention mechanism for video diffusion transformers, achieving high-quality video generation up to 720p on a single GPU with significantly reduced latency compared to full-softmax models, while maintaining competitive VBench scores.
Microsoft releases Mage-Flow-Edit-Turbo, a compact 4B-scale generative model for efficient text-to-image generation and instruction-based image editing, achieving state-of-the-art competitive quality through co-designed tokenizer and backbone.
Microsoft releases Mage-Flow, a compact 4B-parameter foundation model for efficient native-resolution text-to-image generation and instruction-based image editing, achieving competitive quality against much larger models.