Tag
This paper introduces Drift Variation autoencoder, which uses conditional posterior flow matching to unify generative and representation learning, achieving high performance in controlled multimodal benchmarks.
seedance2 is an AI model that converts depth-video inputs into full video outputs.
This paper proposes a unified denoising diffusion framework for conditional generation of graph signals, introducing a novel U-GNN architecture that extends U-Net to graph-structured data. The method is demonstrated on stock price forecasting and wireless resource allocation tasks.
FlowBender is a closed-loop framework that improves constraint satisfaction in diffusion and flow models by training networks to correct alignment errors using inference-time feedback, outperforming traditional supervised and guidance-based approaches.
This paper introduces AnyMo, a unified multimodal framework for human motion generation that combines a Residual FSQ-based motion tokenizer with a scalable masked modeling transformer, along with the OmniHuMo dataset of over 5,000 hours of motion data to enable high-quality synthesis under arbitrary modality combinations.
UniSteer introduces a text-guided activation flow matching method to learn a universal conditional velocity field in activation space, enabling versatile LLM behavior control and classification tasks without task-specific intervention modules.
MERIT is a framework that learns disentangled music representations for melody, rhythm, and timbre using conditional audio generation and source-separated stems, enabling nuanced and factor-specific audio similarity queries.