Tag
This paper shows that fine-tuning autoencoders for reconstruction reduces effective dimensionality, making standard velocity prediction inefficient in diffusion models, and proposes using x0-prediction to focus on the signal manifold, consistently improving text-to-image generation.
InternW0 is a foundational physical world model from Shanghai AI Laboratory that jointly learns visual dynamics and robot control for efficient real-world interactions, trained on heterogeneous data and evaluated on scientific tasks.
The paper proposes CorrFlow, a correlation-guided flow matching framework for predicting spatial transcriptomics from histology images, explicitly modeling gene-gene dependencies to improve biological coherence in generated profiles.
The paper introduces SolarFlowRefiner, a refinement-aware flow-matching framework for downscaling surface solar radiation from coarse ERA5 data to high-resolution SolarCube fields, demonstrating consistent improvements over standalone generation and post-hoc refinement methods.
This paper introduces probe guidance, a new method for flow matching models in continuous diffusion language models, which achieves state-of-the-art performance on unconditional generation and improves multiple choice question answering benchmarks while providing insights into autoguidance.
AntennaFlow is a three-stage generative flow model framework that jointly addresses phase acquisition and offset correction challenges in antenna testing, enabling fast, phaseless, and offset-vector-free near-field to far-field reconstruction from sparse amplitude-only measurements.
The paper introduces a flow-matching model for predicting aircraft trajectories using ADS-B data, achieving superior performance over traditional methods in probabilistic trajectory prediction.
Notes on the equivalence between ODE and SDE sampling in diffusion models, discussing conditions and practical changes.
GradRepair-ODE introduces a reliability framework for certifying and repairing gradients in Neural ODE training to address numerical stability issues in scientific machine learning and generative models.
This paper identifies a duality between continuous and discrete flow matching, showing that projecting continuous convex-interpolant paths via argmax yields discrete flows, and explores how different source geometries affect transition timing and generation quality.
StepAudio 3 Music introduces a large-scale, long-form music generation model with explicit musical planning via ABC-CoT, achieving high scores in audio quality and similarity metrics compared to other systems.
Marigold V2 repurposes diffusion transformers for monocular depth estimation via single-step inference and a novel fine-tuning protocol, achieving sharper depth maps and significant improvements on benchmarks like KITTI and ETH3D.
This paper introduces Quantile AlignTree Flow Matching (QAT-FM), a structured coupling method for flow matching that constructs hierarchical couplings using quantile-aligned trees to achieve non-crossing paths and efficient training for high-dimensional generative tasks.
This paper proposes CAT-OV and CAT-OT, two lightweight, training-free algorithms that adapt step-sizes in Flow Matching sampling based on curvature, improving image quality and reducing generation steps by up to 40%.
The paper proposes GeoLAMP, a geometry-aware latent autoregressive generative model for solving multiphysics partial differential equations in complex geometries, using a dual-encoder architecture and causal self-attention transformer with flow matching for stable and scalable predictions.
SimpleMemVLA introduces a simple memory mechanism for Vision-Language-Action models by feeding intact timestamped video history into a pretrained VLM backbone, achieving state-of-the-art results on long-horizon manipulation tasks without dedicated memory modules.
Flow-JEPA introduces a conditional flow matching approach to JEPA world models, improving robustness and accuracy in predicting future latent states under noisy conditions.
This paper introduces Gromov-Monge flow matching for equivariant graph generation, improving sample quality with structure-aware couplings compatible with standard architectures.
GameWAM introduces the first world-action model for native closed-loop gameplay and GUI control in video games, jointly generating visual observations and executable actions with competitive task success and revealing a source-sensitivity failure mode.
This paper introduces Drift Variation autoencoder, which uses conditional posterior flow matching to unify generative and representation learning, achieving high performance in controlled multimodal benchmarks.