Tag
RECTOR is a self-supervised framework that learns joint region-channel-temporal representations from EEG/sEEG signals for affective and cognitive state classification, achieving state-of-the-art results on emotion recognition and task-engagement benchmarks.
This paper proposes replacing the inner product scoring in sparse autoencoders with a learned combination of cosine similarity and input magnitude, showing that the resulting features are more interpretable and concept-aligned, with the optimizer consistently preferring cosine over inner product.
This paper explores the use of variational autoencoders to learn latent representations of large-scale X-ray scattering data, enabling efficient data compression and analysis.
RepFusion introduces a method to use pretrained multimodal LLMs as noisy representation encoders in diffusion transformers for text-to-image generation, outperforming baselines with similar compute.
Proposes CoMAG, a unified backbone for multimodal attributed graphs that learns task-adaptive reliable contexts and performs modality-preserving alignment, achieving state-of-the-art results on graph-level prediction, modality matching, and graph-conditioned generation.
Introduces Adelic operation-preserved embeddings (AOE), a training-free representation that encodes numbers by combining real value with p-adic expansions, preserving additive and multiplicative structure. Achieves perfect accuracy on the Weaving Pattern benchmark.
This paper proposes a falsifiable applicability criterion for a training-free, fixed-length descriptor for multivariate time series based on time-lagged spectral embeddings, showing when it can be expected to work and validating it on multiple benchmarks.
RepFusion proposes using multimodal large language models as noisy representation encoders for diffusion transformers in text-to-image generation, outperforming traditional denoising approaches.
This blog post explains the connection between JEPA (Joint Embedding Predictive Architecture) models and Canonical Correlation Analysis (CCA), a statistical method from 1936, arguing that CCA is the conceptual precursor to JEPA and that the idea of maximizing correlation in embedding space dates back to Hotelling.
This paper introduces an approach to map unitary operators into the latent space of an LLM, enabling quantum circuit synthesis and language-conditioned gate constraint specification, achieving competitive results on Clifford+T circuit synthesis.
RepWAM introduces a world action modeling approach using representation visual-action tokenizers, aiming to learn unified visual and action representations for planning and control.
Introduces READER, a lightweight framework for dynamic black-box LLM provenance that uses a frozen proxy LLM to extract authorship evidence from responses and performs Bayesian evidence accumulation across multiple queries, achieving high accuracy on the Agent500 dataset.
This paper proposes an automated hyperparameter optimization framework based on Differential Evolution for Latent Factorization of Tensors (LFT) to improve prediction accuracy on large-scale dynamic weighted directed networks, reducing the need for manual tuning.
The article introduces SynIB, a scalable objective based on the information bottleneck principle that targets synergistic information in multimodal learning by penalizing confident predictions when a modality is masked, improving performance on tasks requiring cross-modal reasoning.
Proposes CF-JEPA, a mask-free self-supervised learning framework for time-series that uses multi-horizon forward prediction from random crops and exploits asymmetry between online and target encoders for improved performance on classification, forecasting, and anomaly detection.
This book presents a mathematical theory of deep representation learning, aiming to demystify the internal mechanisms of large deep networks using optimization and information theory, making architecture design a matter of linear algebra and calculus.
This paper presents a data-efficient anatomy-aware benchmark for cardiac pathology prediction on the ACDC MRI dataset, showing that under limited labels, anatomical representation matters more than model complexity.
TRL-Bench is a unified framework and library for standardizing the evaluation of tabular representation learning models across 20 encoders, 16 tasks, and 87 datasets. It provides a common interface to compare heterogeneous tabular models and reveals that no single encoder is best for all tasks.
A tweet highlights Chris Potts' talk on how large language models learn linguistic structures, reinforcing the view that LLMs capture syntax and semantics.
This paper argues that representation learning, not model-based planning, is the key to scalable multitask deep reinforcement learning. It introduces MR.Q, a simple model-free algorithm with auxiliary predictive objectives that outperforms prior world-model-based methods across diverse continuous control tasks.