Tag
This paper proposes BridgeMIL, a two-stage framework for EEG-based disease diagnosis that decouples instance representation learning from subject-level supervision using multiple instance learning. It achieves state-of-the-art accuracy on three EEG disease datasets, outperforming strong baselines.
This paper proposes CURL, a plug-in adapter that uses estimator uncertainty to allocate pretrained LLM semantic capacity for improving heterogeneous treatment effect (CATE) estimation. It introduces two role-conditioned prompts to construct assignment- and heterogeneity-oriented representations, improving ten host learners on four benchmarks.
This paper argues that concepts in LLMs should be treated as a design axis, mapping the design space along pipeline stage and grounding dimension, and proposes moving from recovering concepts post-hoc to designing explicit conceptual representations.
This paper explores how semantic relations are encoded in the geometry of language model semantic spaces, finding that asymmetric relations occupy distinct regions and that lexical information matters more for causal models while contextual information matters more for masked and diffusion models.
This paper reports that deep reinforcement learning agents using frozen, randomly initialized CNN feature extractors spontaneously develop extremely sparse fully-connected representations, compressing task-relevant information through very few neurons without any sparsity-inducing objective.
IRIS is a training-free framework that uses frozen large language models to construct reusable identity representations for entities in knowledge graphs, enabling efficient entity alignment across different KGs without pair-dependent processing.
Proposes temporal-distance JEPA (TD-JEPA) which mines directed temporal cost from offline trajectories to improve latent world model predictive control, achieving higher success rates on robotic environments.
This paper analyzes why deterministic JEPA-style latent prediction works for images but not for text, attributing the failure to high conditional variance in language where masked contexts admit multiple valid completions whose representations lack a coherent center.
This paper identifies group preference collapse in personalized multimodal large language models and proposes PrefMoE, a preference-centric framework that separates profile information from preferences to improve personalization and reduce collapse.
LC-SEPLM adapts ESM2 with LoRA and long-range residue-pair contact supervision to incorporate structural information into protein sequence representations, achieving significant improvements across eight protein-level tasks including remote-homology recognition.
This paper proposes Music-JEPA, a world model that learns piano sound representations by framing audio as a state and piano roll as an action. It captures action-sound relationships and enables downstream tasks like beat tracking and piano transcription via planning.
Proposes Unbiased Open World Regularization (UOWReg), an encoder-only framework that enforces conditional distribution matching to achieve statistical independence between learned representations and sensitive attributes, reducing bias while maintaining accuracy.
MA-DAR is a plug-and-play framework that addresses representation conflicts in replay-based continual temporal knowledge graph reasoning by aligning replayed and current representations on a shared manifold and using a dynamic gating mechanism for adaptive fusion.
This paper proposes adjustment speed as a safety constraint for nonstationary reinforcement learning, defining safety in terms of adaptation feasibility and using representation learning with context forecasts to proactively regulate behavior when predicted adaptation demand exceeds the system's achievable capacity.
The author revisits influential open-source works in representation learning, listing key papers from MoCo v1 to LeJEPA that advanced vision foundation models and self-supervised learning.
This paper proposes Hyper-Spherical Quantization (HSQ) to address codebook collapse in discretizing visual representations, achieving high-fidelity reconstruction and scalable codebook budgets up to 131,072 with 100% utilization.
This paper introduces Semantic Field Theory (SFT), a computational model for lexical semantics that models meaning through semantic fields, contextual deformation, interaction terms, and energy minimization. It provides formal elements including Gaussian product closure, Möbius inversion for higher-order interactions, and stability conditions.
This paper proposes the Structured Dynamics Model (SDM), which self-supervisedly learns to separate camera and object motion from frozen image features, outperforming baselines on a new evaluation suite.
This paper models Chain-of-Thought reasoning as a switching dynamical system, showing that reasoning fine-tuning globally reorganizes latent policy states, leading to improved multi-step reasoning. The proposed framework combines time-aware contrastive learning with discrete regime discovery, and experiments demonstrate that fine-tuned models exhibit richer latent-policy organization with functional specialization.
This paper identifies and attacks the alignment layer of graph foundation models, showing it is a distinct attack surface vulnerable to perturbations at low budgets, especially for spectral tokenizers, and proposes detection-based defenses.