Tag
The paper introduces GALA, a two-stage method for text-to-time-series synthesis that uses generation-aware cross-modal alignment to achieve state-of-the-art results on the TSFragment-600K benchmark, improving both fidelity and caption adherence.
CALM is a framework for learning interpretable associations between brain regions and genetic pathways from completely unpaired datasets, enabling biomarker discovery for neuropsychiatric disorders like autism without requiring paired multimodal data.
This paper introduces sparse autoencoders to resolve superposition in neural networks, improving interpretability and geometric fidelity of latent spaces, and presents GW-map for cross-modal alignment between image representations and single-cell RNA sequencing data.
FLUX3D introduces a framework for high-fidelity image-to-3D Gaussian Splatting generation by enhancing representation learning and cross-modal alignment with diffusion-aligned structured latents and a sparse-structure-aware diffusion transformer, achieving state-of-the-art results.
HeRA aligns individual attention heads in Multimodal Large Language Models (MLLMs) to preserve local neighborhood relationships across modalities, improving vision-centric task performance and reducing visual hallucinations.
This paper investigates the Platonic Representation Hypothesis, proposing that alignment arises from linear structure in representations, and introduces a statistical framework of signal, bias, and noise.