Tag
This paper presents a computational framework for steering representational geometry to improve bidirectional alignment between biological and artificial neural networks, showing a 55% relative enhancement in bidirectional predictivity.
NAPE is a self-supervised audio learning framework that uses causal Transformers to predict next spectrogram patch embeddings, achieving state-of-the-art performance on multiple audio and speech benchmarks with a minimalist design.
The paper introduces AC-MTM, a contrastive inverse dynamics method to prevent encoder collapse in JEPA world models, achieving improved performance on multi-object tasks without Gaussian constraints.
The paper establishes a theoretical connection between probabilistic Joint-Embedding Predictive Learning (JEPA) and Hidden Markov Models (HMMs), providing a state-space interpretation and introducing Markov-Chain JEPA for enhanced consistency.
ConceptFormer learns continuous latent concept representations for visual document retrieval, bridging visual evidence and semantic relevance without text intermediates, achieving significant improvements over baselines.
Introduces Action-Conditioned Predictive Consistency (ACPC), a diagnostic for JEPA world models that measures how clean and perturbed observations diverge under action-conditioned rollouts, with theoretical bounds on prediction error and planner cost. Experiments on visual control tasks validate the diagnostic across models like LeWM and PLDM.
This paper presents a controlled study on ECG self-supervised representation learning, examining how temporal context length (16s to 10min) and encoding strategy (continuous patch embeddings vs discretized tokens) affect downstream rhythm detection and patient-level retrieval. Results show longer context and continuous encoders improve performance, motivating extended-context ECG foundation models.
This paper introduces TTARO, an online deep-kernel Bayesian optimization framework that adapts circuit representations at test time using evaluated figure-of-merit labels. It improves sample efficiency for analog circuit topology search, reducing regret AUC by 15.2% over standard BO and 20.7% over fixed deep kernel learning.
This paper introduces low interaction rank as a unified theoretical framework for multiplicative dual-encoder networks, covering approximation, sample complexity, normalization, and identifiability, with experiments on operator learning and CLIP models.
This paper proposes a gloss-free representation learning approach for cross-dataset sign spotting, using weakly aligned broadcast transcripts in Turkish Sign Language. It shows that LLM-assisted pseudo-gloss normalization improves temporal localization and downstream translation quality.
This paper studies how β-VAEs act as effective theories where the KL weight acts as a spectral cutoff, and analyzes how nonlinear interactions and network depth affect the tolerance-dependent effective dimension of representations.
This paper introduces Sheaf-based Federated Representation Learning (SFRL), a framework that aligns heterogeneous local representations via learnable sheaf restriction maps and a quadratic gluing regularizer, without assuming a shared global latent space. A decentralized algorithm (Sheaf-FRL) with convergence guarantees is proposed and shown to outperform baselines in cooperative classification under data and model heterogeneity.
Introduces Rationale-Guided Learning (RGL), a framework that reframes multimodal emotion recognition in conversation as a cognitively-inspired reasoning task using dual-process theory and MLLM-generated rationales, achieving state-of-the-art results on IEMOCAP and MELD.
This paper introduces a geometry-aware adversarial attack framework that targets relational structure in contrastive embedding manifolds, showing that verification systems like Markmatch can be severely degraded by distorting pairwise similarities rather than decision boundaries.
Introduces DALMA, a probabilistic representation learning framework that uses biological supervision to improve cross-center generalization of MALDI-TOF mass spectrometry models for clinical microbiology tasks like microbial identification and antimicrobial resistance prediction.
DoGMA is a central-dogma-guided foundation model for pan-cancer multi-omics analysis, using a Transformer-MoE architecture with directed attention to align DNA-RNA-protein flows and pretraining via masked hierarchical omics reconstruction. It shows strong performance across cancer representation learning, survival prediction, and metastasis prediction tasks.
Introduces TREAT, a benchmark for evaluating whether large language models can recover known theorem identities from equivalence-preserving transformations of mathematical formulas. The best tested model achieves only 60.73% accuracy, showing that theorem knowledge is fragile under representation changes.
The author shares early progress on Leo/PSCLS, an experimental system that learns sequence relationships and improves its story generation and metrics as it is trained on more stories.
This paper investigates how molecular generative models internally organize molecular identity in their latent spaces, revealing piecewise-constant regions and coarse-to-fine boundaries across three architectures.
The paper introduces NysHD, a method that bridges hyperdimensional computing and kernel methods via the Nyström approximation, allowing any positive-semidefinite similarity function to be used as an HDC encoding. It demonstrates improved classification accuracy on graph and string datasets compared to existing HDC encoding methods.