Tag
This paper analyzes failures in reconstruction-based unsupervised learning through geometric conditions and proposes new methods like Dynamic Push and Pull to enhance anomaly detection performance.
The paper introduces CARE, a condition-aware representation regularization framework for diffusion models that improves sample quality and training efficiency by dynamically modulating feature distributions based on condition similarity. Empirically, it achieves significant reductions in FID and faster convergence for both class-to-image and text-to-image tasks.
This paper challenges the interpretation of the Platonic Representation Hypothesis by distinguishing between relational structure and metric geometry, showing that relational convergence is robust while metric geometry convergence is weaker in various models after calibration.
The paper proposes EviGDA, a framework that enhances graph domain adaptation by combining graph-aware and graph-free experts to improve prediction under structural shifts.
This paper shows that fine-tuning autoencoders for reconstruction reduces effective dimensionality, making standard velocity prediction inefficient in diffusion models, and proposes using x0-prediction to focus on the signal manifold, consistently improving text-to-image generation.
LE4Mob is an inductive, distance-aware, and general-purpose location embedding framework for enhancing human mobility modeling in tasks such as next location prediction and commuter flow generation.
MIRCID is a framework that infers hub-miRNAs to enhance drug mechanism-of-action modeling by comparing gene expression with inferred transcription factor activity and miRNA expression, achieving improvements in pathway classification and similarity-based retrieval.
This paper establishes two limitations of Maximal Coding Rate Reduction (MCR²) for out-of-distribution generalisation, showing that it can fail under distribution shift and that incorporating invariance principles does not eliminate this failure.
Tactile-JEPA is a self-supervised pre-training method for distributed tactile sensors that uses spatial topology to learn representations, improving force estimation and orientation tasks in robotics over prior state-of-the-art.
The paper introduces THAW-VLA, a method that distills world-model representations into Vision-Language-Action models for robotics, enhancing robustness and performance on simulation and real hardware without additional inference overhead.
The paper explores grokking, a delayed generalization phenomenon in neural networks, and introduces Geometric Dimensionality Regularization (GeomDR) to control representation geometry, accelerating grokking by up to 52x.
The paper introduces Lens, a training-free framework for multimodal representation learning that addresses semantic perspective misalignment, achieving significant performance improvements on MMEB datasets without parameter updates.
This paper proposes a new unified pre-training framework for medical code sequences that captures hierarchical structures and complex interactions, demonstrating superior performance in clinical event prediction and drug repositioning case studies for Alzheimer's disease.
FLAT introduces a shared sequence of continuous tokens for images and text, enabling flexible-length representations for multimodal tasks like generation and retrieval, with strong performance on benchmarks.
This paper introduces a mean-field analysis of attention that predicts average representation dynamics and reveals context-specific computation in language models, validated across models like GPT-2, Pythia, and Qwen-3-14B.
The paper proposes Representation-based Masked Diffusion Model (RMDM), which leverages text representations to improve parallel token updates in masked diffusion models, enhancing generation quality especially in few-step sampling.
This paper decomposes transformer representation updates into parallel and perpendicular components to study evolution geometry, linking it to editing robustness, compression diagnosis, and training improvements.
This paper proposes a reference-based method for detecting bias in large language models by analyzing relative representations of hidden states across model variants, introducing Representational Bias Shift (ΔB) that efficiently correlates with output-level bias changes.
Capsule Lens is a framework for mechanistic interpretability that uses geometric capsules to locate and track how concepts are encoded in neural network representations, enabling analysis of both static and dynamic model internals.
This paper proposes a Multi-Granularity Hypergraph Representation Learning (MGHRL) framework that adaptively generates hyperedges at multiple granularities using granular-ball splitting to capture high-order relationships in graphs, outperforming baseline models on benchmark datasets.