diffusion-transformer

Tag

Cards List
#diffusion-transformer

Wan-Animate-2: Pushing the Application Boundaries of Character Animation Models

Reddit r/LocalLLaMA · 3d ago

Wan-Animate-2 is a new end-to-end character animation framework that consumes driving videos directly in a redesigned Diffusion Transformer, achieving high-fidelity motion generation and identity preservation. It also introduces a lightweight variant for real-time streaming animation, with open-source weights released.

0 favorites 0 likes
#diffusion-transformer

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

Hugging Face Daily Papers · 5d ago Cached

EffectLearner is a semantic-reasoning-enhanced framework for real-world video object removal, combining a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser and introducing the EffectWorld dataset to handle complex object-induced effects.

0 favorites 0 likes
#diffusion-transformer

Xiaomi-Robotics-1: New robotics model released

Reddit r/LocalLLaMA · 5d ago

Xiaomi released XR-1, a robot foundation model trained on over 100K hours of real-world manipulation trajectories. Built on Qwen3-VL and a Diffusion Transformer, it enables out-of-the-box mobile manipulation in unseen environments.

0 favorites 0 likes
#diffusion-transformer

Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-Design

Hacker News Top · 2026-07-29 Cached

Transformer Transformer is a unified model that generates complete robot embodiments optimized for a given manipulation demonstration, using a diffusion transformer trained on RoboTokens and Dynamics Self-Guidance.

0 favorites 0 likes
#diffusion-transformer

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling

Hugging Face Daily Papers · 2026-07-27 Cached

WorldDiT is a unified diffusion transformer architecture that couples action generation with visual world modeling, achieving strong performance on LIBERO simulation suites without relying on large pretrained vision-language models.

0 favorites 0 likes
#diffusion-transformer

Nvidia's New Long-Form Video Generation (12 minute read)

TLDR AI · 2026-07-27 Cached

NVIDIA Research introduces SANA-Video 2.0, a hybrid video diffusion transformer that generates high-quality 720p video on a single GPU, achieving up to 120× speedup over Wan 2.2-14B via hybrid linear-softmax attention and block attention residuals.

0 favorites 0 likes
#diffusion-transformer

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

Hugging Face Daily Papers · 2026-07-23 Cached

SANA-Video 2.0 introduces a hybrid linear-softmax attention mechanism for video diffusion transformers, achieving high-quality video generation up to 720p on a single GPU with significantly reduced latency compared to full-softmax models, while maintaining competitive VBench scores.

0 favorites 0 likes
#diffusion-transformer

microsoft/Mage-Flow-Edit-Turbo

Hugging Face Models Trending · 2026-07-21 Cached

Microsoft releases Mage-Flow-Edit-Turbo, a compact 4B-scale generative model for efficient text-to-image generation and instruction-based image editing, achieving state-of-the-art competitive quality through co-designed tokenizer and backbone.

0 favorites 0 likes
#diffusion-transformer

microsoft/Mage-Flow

Hugging Face Models Trending · 2026-07-21 Cached

Microsoft releases Mage-Flow, a compact 4B-parameter foundation model for efficient native-resolution text-to-image generation and instruction-based image editing, achieving competitive quality against much larger models.

0 favorites 0 likes
#diffusion-transformer

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

Hugging Face Daily Papers · 2026-07-21 Cached

Mage-Flow is a compact 4B-parameter generative stack for efficient text-to-image generation and instruction-based image editing, featuring a co-designed lightweight tokenizer (Mage-VAE) and a native-resolution multimodal diffusion transformer trained with rectified flow matching. It achieves competitive performance while enabling high-resolution generation at 0.59s on a single A100 GPU.

0 favorites 0 likes
#diffusion-transformer

MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion

arXiv cs.AI · 2026-07-20 Cached

Proposes M2GDT, a novel MKGC framework that uses an MLLM-guided diffusion transformer with relation-adaptive mixture-of-experts to align and denoise multimodal features, outperforming baselines on three benchmark datasets.

0 favorites 0 likes
#diffusion-transformer

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

Hugging Face Daily Papers · 2026-07-15 Cached

This paper introduces VideoRAE, a representation autoencoder that leverages frozen video foundation models to create compact, reconstruction-capable, and generation-friendly video latents. It achieves state-of-the-art results on UCF-101 with faster convergence than competing autoencoders.

0 favorites 0 likes
#diffusion-transformer

FreyaTTS Technical Report

arXiv cs.CL · 2026-07-13 Cached

FreyaTTS is a compact, tokenizer-free Turkish-first text-to-speech model based on a non-autoregressive conditional flow-matching Diffusion Transformer, achieving state-of-the-art performance with a fraction of the parameters of larger systems and released under Apache-2.0.

0 favorites 0 likes
#diffusion-transformer

RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation

Hugging Face Daily Papers · 2026-07-07 Cached

Introduces digital teleoperation using action-conditioned world models to generate diverse training data for robotics, decoupling data collection from physical hardware. The system achieves real-time generation and enables zero-shot Sim2Real transfer.

0 favorites 0 likes
#diffusion-transformer

PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

Hugging Face Daily Papers · 2026-07-02 Cached

PointDiT presents a minimalist pixel-space diffusion transformer using a plain ViT architecture for monocular geometry estimation, outperforming complex latent-based models while maintaining simplicity and robustness in ambiguous regions.

0 favorites 0 likes
#diffusion-transformer

Walking in the Implicit: Interactive World Exploration via Neural Scene Representation

Hugging Face Daily Papers · 2026-06-29 Cached

NeuWorld is a new interactive video generation system that uses compact neural implicit scene representations and a transformer VAE with diffusion transformer for trajectory-conditioned rendering, achieving long-horizon consistency.

0 favorites 0 likes
#diffusion-transformer

MirrorPPR: Exemplar-Based Portrait Photo Retouching

Hugging Face Daily Papers · 2026-06-28 Cached

MirrorPPR introduces an exemplar-based portrait retouching framework using Diffusion Transformer with LoRA adaptation and self-augmented training data, achieving superior quality and identity preservation.

0 favorites 0 likes
#diffusion-transformer

EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting

Hugging Face Daily Papers · 2026-06-25 Cached

EO-WM proposes a video diffusion transformer for probabilistic Earth observation forecasting that incorporates physically informed conditioning to capture weather-driven uncertainties, achieving improved prediction of vegetation indices under extreme weather.

0 favorites 0 likes
#diffusion-transformer

@rohanpaul_ai: AI video is moving into its real-time reaction era, with MaineCoon now leading in low-latency AI video. @catnips_ai jus…

X AI KOLs Following · 2026-06-23 Cached

MaineCoon is a 22B real-time text-to-audio-video model that achieves up to 47.5 FPS on a single H100 GPU, enabling low-cost, long-duration streaming with synchronized speech and visuals for live AI characters.

0 favorites 0 likes
#diffusion-transformer

MeshFlow: Mesh Generation with Equivariant Flow Matching

Hugging Face Daily Papers · 2026-06-22 Cached

MeshFlow introduces an equivariant optimal-transport flow matching model for direct triangle mesh generation, achieving state-of-the-art quality while providing approximately 18x inference speedup over autoregressive methods.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback