Tag
Self-OPD introduces a teacher-free on-policy distillation framework for flow matching models that uses self-explored stochastic branches and normalized advantages to optimize velocity fields, outperforming prior methods in multi-objective alignment.
Introduces renormalization group flow matching (RGFM), a generative framework that uses renormalization group flow for scalable local generative modeling, improving global coherence and computational efficiency.
This paper introduces untied self-conditioning to correct train-inference mismatch in flow-matching language models, improving few-step generation quality with lower perplexity and no retraining.
GameWAM introduces the first unified world-action model for native video-game control, jointly predicting future visuals and executable actions using block-causal flow matching and mode-specific distributions.
TracingFlow is a simulation-free framework using second-order dynamics for trajectory inference, improving accuracy in capturing high-curvature transitions in single-cell omics data.
ForeTime-VLA introduces a causal future-token distillation method from a world action model to improve dynamic conveyor-belt manipulation, achieving higher grasp success rates compared to existing VLA policies.
This paper investigates unsupervised anomaly detection using flow matching on tabular data, focusing on contaminated training sets and comparing different scoring methods for robustness.
The paper identifies manifold drift as a root cause of reward hacking in flow preference optimization and introduces ThermoDPO, a temperature-controlled method that anchors optimization on preferred samples to preserve data manifold integrity and improve performance.
The paper introduces GALA, a two-stage method for text-to-time-series synthesis that uses generation-aware cross-modal alignment to achieve state-of-the-art results on the TSFragment-600K benchmark, improving both fidelity and caption adherence.
PixRestore is a VAE-free pixel-space diffusion transformer for unified image restoration, achieving high fidelity and efficiency via flow matching and adversarial fine-tuning to a one-step generator.
Proposes ReCoGen, a two-stage framework for multimodal-conditioned time-series generation under irregular missingness, achieving state-of-the-art downstream utility on physiological benchmarks.
MiniMax releases Music 3, a high-performance music generation model that creates complete songs up to five minutes long using lyrics and detailed descriptions, with an 8B global LLM and 0.6B local LLM for long-range coherence and acoustic detail.
The paper introduces Adversarial Fréchet Distance (AdvFD), which adds a learnable adversarial feature space to static Fréchet losses to improve generator post-training, with real-feature whitening to stabilize optimization.
This paper introduces Task-Conditional Flow Matching (TCFM), a framework for adapting multilingual text embedding models that selectively uses flow matching for translation tasks and other objectives for retrieval/classification. It achieves state-of-the-art results on the Indic Massive Text Embedding Benchmark.
Energy-Guided Flow Matching improves generative image quality by using a moving endpoint and adaptive scheduling, achieving state-of-the-art FID scores with reduced training cost.
SimWAM is a simple yet effective World Action Model for end-to-end autonomous driving that uses video generation purely as a training signal, achieving state-of-the-art 91.5 PDMS on NAVSIM while reducing inference latency.
Poly-OPD is a framework for distilling complementary strengths from heterogeneous text-to-image flow models into a single compact flow-matching student, using pixel bridges and gradient-compatible adapters. It improves GenEval and DrawBench scores while consolidating multiple teacher capabilities.
Proposes Quantile Coupling Flow Matching (QC-FM), a lightweight one-sided coupling that constructs source samples from data ranks along random directions without needing pairwise cost matrices or assignment. Achieves up to 12.9% FID improvement over baseline on CIFAR-10, CelebA, FFHQ, and ImageNet-64.
This paper introduces DRIFT, an adversarial patch attack targeting flow-matching vision-language-action models like pi0 and pi0.5, showing that prior robustness claims are illusory and that attacking only the first denoising step is both stronger and cheaper, breaking nearly all solvable tasks in LIBERO suites.
Any-OPD presents the first framework for on-policy distillation between arbitrary latent flow-matching generators, enabling distillation from a 12B FLUX model to a 2.5B SD3.5 model by bridging via a frozen vision representation. It improves the student's PickScore from 0.846 to 0.884, rivaling the teacher at a fifth of its size.