Tag
This paper analytically computes the optimal representations under a contrastive loss for basic augmentations and natural images with stationary statistics, showing that the optimal CNN first-layer filters are sinusoids and that weights can be computed via a waterfilling algorithm.
This paper investigates Rank-Order N-of-M codes for sparse distributed memory, disentangling representation and learning effects to evaluate noise robustness compared to contemporary neuromorphic architectures.
CrossBERT decouples representation learning from token reconstruction, enabling higher masking ratios and better sample efficiency, outperforming BERT on MTEB and GLUE benchmarks.
Proposes a conditional diffusion-guided knowledge transfer framework for multi-domain knowledge graph completion, generating domain-general entity embeddings without suppressing domain-specific information, achieving 4.3% average MRR improvement over state-of-the-art methods.
This paper introduces Mining via Activation Geometry (MAG), an unsupervised framework that extracts reasoning features from LLM activations using natural-language instructions, enabling activation steering and effective training data selection for classifier probes.
This paper proposes SiamJEPA, which uses masked Siamese student encoders with an EMA teacher network in JEPA models, showing improved representation separability and training efficiency compared to single-encoder variants and MAE.
The paper introduces role-aware neural convex divergence heads that apply source and target role projections before evaluating an input-convex neural Bregman divergence, enabling structured and interpretable asymmetric distance learning for tasks like lexical entailment, sentence entailment, and ontology hierarchy. Experiments show consistent improvements in directional accuracy over plain ICNN-Bregman heads across semantic and ontology benchmarks.
This paper investigates learning useful policy representations (embeddings) in two-player zero-sum imperfect-information games, introducing methods for creating policy datasets, learning embeddings, and evaluating them on downstream tasks using Kuhn and Leduc Poker.
This paper introduces a benchmark for predicting taste qualities (sweet, bitter, etc.) from audio embeddings. It evaluates 10 pretrained audio encoders, achieving 0.134 RMSE, outperforming previous state-of-the-art and enabling taste-based music retrieval.
This paper introduces LOPA, a lightweight framework for spoken language assessment that uses latent ordinal prototype alignment and semantic-anchored layer routing on a frozen Whisper encoder, achieving performance comparable to billion-parameter models without LLM fine-tuning.
This paper proposes a self-supervised Mamba-based model to learn effective representations from electronic health records for improved patient subtyping, demonstrating better performance than baseline models on real-world datasets.
This paper introduces textual latent states and factorized GRPO (fGRPO) to enforce strict mediation in text-based world models, addressing the identifiability problem and achieving up to 57% gains in representation quality and 98% improvements in rollout performance.
The Prism Transformer replaces uniform multi-head attention with a progressive head schedule that increases head count across layers, enabling a local-to-global hierarchy without extra parameters or FLOPs. It consistently outperforms standard Transformers on language modeling and zero-shot benchmarks at 124M, 354M, and 757M scales.
A discussion on whether world models, which learn internal environment representations to simulate physics and plan actions, could lead to AGI by overcoming the limitations of reactive predictive text models like LLMs.
This paper presents a general framework for using Graph Neural Networks to learn algebraic properties from Cayley graphs, offering a new approach to algebraic reasoning with GNNs.
This paper introduces three datasets (Hell-Char, PaLit-Char, Med-Char) for diachronic representation learning of ancient Greek letterforms and proposes a similarity-weighted supervised contrastive loss with lacuna-driven augmentation to robustly learn character embeddings across centuries of handwriting variation.
This paper introduces the Swarm-Inspired Emergent Synchronizer (SIES), a graph-dynamical framework that learns generalizable local interaction rules for controllable collective organization, applicable to synchronization control and heterophilous graph representation learning.
This paper studies when conservation laws can be certified in learned latent world models, proposing bounded horizons that guarantee how long rollouts stay on physical invariant level sets using measurable model defects.
This paper presents a spectral phase diagram for binary few-shot classification, analyzing intrinsic dimensionality and geometric saturation for representational diagnosis.
Pangram Labs explores the internal representations of its AI detection model Pangram 3.3.2, analyzing how the model distinguishes human vs AI text across layers using a balanced dataset of 5,000 documents from various sources.