Tag
This paper introduces LC-GRPO, a flow-based GRPO framework with Langevin correction that bridges the train-inference gap by aligning stochastic training rollouts with deterministic ODE sampling, improving reward optimization on models like SD3.5, FLUX.1-Dev, and HunyuanVideo.
This paper shows that matching a marginal Gaussian prior in factorized generative models does not prevent conditional style leakage, where style latents carry class information. Multiple remedies are explored, but the authors conclude that marginal statistics alone cannot certify class-invariance.
DiffImaginE is a research paper proposing a diffusion-based verifier for multimodal named entity recognition, replacing deterministic imagination with conditional latent diffusion inference for more robust entity type verification.
Proposes Quantile Coupling Flow Matching (QC-FM), a lightweight one-sided coupling that constructs source samples from data ranks along random directions without needing pairwise cost matrices or assignment. Achieves up to 12.9% FID improvement over baseline on CIFAR-10, CelebA, FFHQ, and ImageNet-64.
This paper introduces StraightDP, a geometry-aware differential privacy framework for text-conditioned rectified-flow transformers. It partitions the privacy budget to release class-conditional moments and use DP-SGD, improving accuracy and FID over uniform DP training at strong privacy levels.
This paper proposes using atom-averaged features from pretrained MLIPs like MACE as coarse coordinates for evaluating and guiding inorganic crystal structure generation, introducing the Coarse-Fine Transport Distance (CFTD) metric that captures both quality and novelty in a distribution-based framework.
This paper introduces PlatformBid, the first comprehensive auto-bidding benchmark designed from a unified advertising platform perspective, along with BidFlow, a novel flow-matching-based auto-bidding method. Experiments show BidFlow improves target cost by +0.68% in online tests on Kuaishou.
This paper introduces SGFlow, a method for learning flow maps for diffusion models that avoids invertibility constraints and backpropagation through model iterations, achieving competitive FID scores on CIFAR with a proven stationary-point guarantee.
Introduces C-VCE, a diffusion framework that builds an interpretable concept bottleneck layer into the generative model, enabling human-guided visual counterfactual explanations without relying on external noise-robust classifiers.
This paper establishes a rigorous quantitative connection between neural network score function approximation and the resulting distribution approximation in score-based diffusion models, proving that accurate score approximation leads to close distribution approximation in KL divergence, with an explicit bound.
A systematic review of diffusion-based methods for medical image inpainting, covering architectures, applications, datasets, and evaluation strategies, with a proposed taxonomy and identification of challenges such as lack of standardized benchmarks.
Chamaileon introduces a framework for multi-target and multi-state protein binder design using contextualized sequence-structure co-modeling and mixed sampling, achieving adaptability across diverse conformational landscapes and multi-target requirements.
This paper investigates the failure of boundary-seeking knowledge distillation (CAKE) when applied to bottlenecked generative autoencoders, showing that the shared latent manifold creates gradient conflicts that prevent effective synthesis of contrastive samples. A simple noise forward pass baseline is proposed instead.
This paper presents a taxonomy-guided evaluation protocol for assessing temporal fidelity in synthetic sequential tabular data, revealing that conventional evaluation overlooks temporal failures and that rankings differ substantially when time is considered.
This position paper argues that probabilistic scaling of generative models is insufficient for quantum circuit generation due to strict mathematical constraints, and proposes verification-aware architectures that integrate formal methods.
Latent actions are gaining traction in robotics as a way to learn from unlabeled video without action labels. Recent papers from DeepMind and FAIR demonstrate progress from controlled game environments to in-the-wild internet video, promising scalable training for imitation learning.
Introduces HEDGEHOG, a hierarchical benchmark for evaluating molecular generative models in drug discovery, revealing that only 0.65% of generated molecules pass all medicinal chemistry and docking filters.
This paper develops conservation laws for diffusion models using generalized extrinsic information transfer (GEXIT) functions, showing that the cross-entropy can be characterized as an integral of local information-theoretic derivatives along the noise path, unifying likelihood characterization for discrete and continuous diffusion.
This paper presents a fair-comparison study of variational quantum circuits in diffusion models, introducing a squeeze-and-excitation scaffold to isolate quantum contributions. It finds functional parity with classical controls and identifies angle-embedding failures in score-based settings, offering a rigorous methodology and mechanistic analysis.
This paper presents a method for fine-grained identity tuning in text-to-image personalization models. It explores the latent space of a frozen encoder to enable localized, semantically coherent facial edits without additional training.