Tag
Latent-Foresight proposes an end-to-end framework that jointly learns a latent tokenizer and a flow-based generative dynamics model, explicitly structuring the latent space for temporal predictability and outperforming two-stage pipelines on future scene understanding tasks.
The paper proposes ReLOBGen, a limit order book message generator that ensures replayability by construction by selecting referenced orders from the current resting book and masking invalid tokens, achieving 100% replayability in 500-message rollouts and a 2.7-3.6x speedup per replayed message over the LOBS5 baseline.
The paper introduces T-RoPE, a time-aware modification to Rotary Position Embedding for sequential recommendation systems, demonstrating significant performance improvements on benchmarks and real-world impact in e-commerce.
This paper introduces Riemannian Wasserstein Entropic Flow Matching (RWEFM), a generative framework for modeling probability distributions on Riemannian manifolds, with applications in scientific domains like single-cell biology and protein conformations.
The paper provides a unified probabilistic framework for large language models, describing them through probability measures, training via maximum-likelihood estimation, and text generation as stochastic simulation, with insights into phenomena like hallucination and the role of diffusion models.
This paper introduces a continuous-time generative dynamics framework that leverages score-based models to generate data at arbitrary timestamps on learned data manifolds, with applications to video and scientific data.
The paper presents VIOT, a variational incompressible optimal transport operator that uses a Fourier Neural Operator to predict divergence-free velocity fields for efficient incompressible density transport, achieving orders of magnitude speedup over traditional optimization methods.
The paper proposes Representation-based Masked Diffusion Model (RMDM), which leverages text representations to improve parallel token updates in masked diffusion models, enhancing generation quality especially in few-step sampling.
The paper proposes CoMA-DiT, a bidirectional cross-modal Diffusion Transformer for latent augmentation in multimodal brain state decoding, which enhances performance in tasks like auditory attention decoding and emotion recognition by using paired modalities as mutual supervisory signals.
This paper identifies a duality between continuous and discrete flow matching, showing that projecting continuous convex-interpolant paths via argmax yields discrete flows, and explores how different source geometries affect transition timing and generation quality.
This paper introduces Newton Matching, a unified framework for fine-tuning and sampling in generative models, which addresses limitations of existing methods by treating learning as an iterative optimization process and leveraging conditional-matching structure.
This paper proposes CAT-OV and CAT-OT, two lightweight, training-free algorithms that adapt step-sizes in Flow Matching sampling based on curvature, improving image quality and reducing generation steps by up to 40%.
PathGuide reformulates classifier-free guidance selection as an on-policy transport problem in flow-based generative models, using the weak form of the continuity equation to dynamically optimize guidance scales for improved sample fidelity.
The article presents BFN-RL, a unified generative modeling framework for offline reinforcement learning based on Bayesian Flow Networks, capable of generating effective trajectories across discrete and continuous state spaces.
This paper isolates the irreducible excess in denoising score matching training loss, showing it equals the trace of the Fisher-Rao metric along diffusion trajectories, connecting variational principles, latent space geometry, and information theory.
FlowNeg is a GFlowNet-based method for diverse hard negative sampling in knowledge graph embedding, improving performance by generating context-conditioned negatives that balance hardness and diversity without treating structural similarity as absolute truth.
Introduces renormalization group flow matching (RGFM), a generative framework that uses renormalization group flow for scalable local generative modeling, improving global coherence and computational efficiency.
A Google DeepMind research paper demonstrates that converting verifiers into generative next-token predictors significantly improves reasoning accuracy, enabling chain-of-thought verification and better performance on math problems through inference-time compute scaling.
Proposes a model-agnostic framework for assortment optimization using guided discrete diffusion, representing assortments as binary vectors and using reward-guided reverse diffusion to avoid combinatorial enumeration. Shows robustness and high-quality solutions in high-dimensional settings.
This paper introduces Simplax, an exact Dirichlet–categorical augmentation for discrete diffusion models that enriches training objectives and reverse transitions while preserving the original categorical corruption process, improving perplexity–entropy tradeoff on OpenWebText and validity on Sudoku.