Tag
TransNormal-2 improves monocular normal estimation by correcting VAE reconstruction errors with geometry-aware training losses and a lightweight RGB-guided refinement module, achieving strong results with minimal annotations.
This paper introduces StraightDP, a geometry-aware differential privacy framework for text-conditioned rectified-flow transformers. It partitions the privacy budget to release class-conditional moments and use DP-SGD, improving accuracy and FID over uniform DP training at strong privacy levels.
Mage-Flow is a compact 4B-parameter generative stack for efficient text-to-image generation and instruction-based image editing, featuring a co-designed lightweight tokenizer (Mage-VAE) and a native-resolution multimodal diffusion transformer trained with rectified flow matching. It achieves competitive performance while enabling high-resolution generation at 0.59s on a single A100 GPU.
This paper presents AffectFlow-DINO, a multi-task learning system for the 11th ABAW challenge that uses a conditional rectified-flow head to model uncertainty in in-the-wild facial behavior estimation, achieving substantial improvements over the baseline.
This paper proposes SRT (Super-Resolution for Time Series), a framework that reconstructs high-resolution temporal patterns from low-resolution inputs using a disentangled rectified flow approach. The method decomposes input into trend and seasonal components, applies implicit neural representation for resolution alignment, and introduces cross-resolution attention to generate fine-grained details, achieving state-of-the-art performance on multiple datasets.
ChangeFlow presents a generative framework for remote sensing change detection that synthesizes change masks in latent space using rectified flow, achieving improved accuracy and robustness through sampling-based prediction ensembling, with an average F1 of 80.4% across four benchmarks.
Irodori-TTS-500M-v3 is a Japanese TTS model based on Rectified Flow Diffusion Transformer, supporting zero-shot voice cloning and unique emoji-based style/sound effect control.
This paper introduces PNAPO, an offline preference optimization framework for rectified flow models that augments preference data with noise samples and uses dynamic regularization to improve training efficiency and sample efficiency.