Tag
This paper evaluates how reliably sentiment concept directions extracted from language model representations reproduce across splits, comparing English, Hausa, and Yoruba over four language models. It finds consistent language rank order in direction agreement (English > Hausa > Yoruba) and shows that high probe classification accuracy does not imply directional consistency, though it does not establish tokenizer fertility as the cause.
This NYU/Xi'an Jiaotong paper shows that early representational drift forms a low-dimensional 'scaffold' subspace that networks preferentially reuse when learning later tasks, indicating that early experience leaves a persistent geometric imprint on neural network adaptation.
ChronoSRL gives self-supervised RL critics an explicit temporal geometry, training embeddings so distances match actual goal-reaching time; it outperforms contrastive and survival RL baselines on seven benchmarks and sim-to-real quadruped locomotion tasks.
This paper explores graph neural network encoders for projecting heterogeneously sampled signal sets from sensor networks and radars into fixed-size, sampling-invariant embedding spaces, evaluated via waveform discrimination on synthetic complex-valued RF signals.
This paper evaluates Jev, a fixed general-purpose probabilistic decision model, for training-free human activity recognition on accelerometer data, finding it far below supervised models (macro-F1 0.038–0.118 vs 0.686–0.907) and showing that adding more numerical sensor features actually degrades performance unless paired with deterministic semantic renderings.
This paper introduces Staged Depth Training (SDT), a representation curriculum for Physics-Informed Neural Networks (PINNs) that explicitly learns and transfers hidden representations to improve training performance.
This paper analyzes failures in reconstruction-based unsupervised learning through geometric conditions and proposes new methods like Dynamic Push and Pull to enhance anomaly detection performance.
The paper introduces CARE, a condition-aware representation regularization framework for diffusion models that improves sample quality and training efficiency by dynamically modulating feature distributions based on condition similarity. Empirically, it achieves significant reductions in FID and faster convergence for both class-to-image and text-to-image tasks.
This paper challenges the interpretation of the Platonic Representation Hypothesis by distinguishing between relational structure and metric geometry, showing that relational convergence is robust while metric geometry convergence is weaker in various models after calibration.
The paper proposes EviGDA, a framework that enhances graph domain adaptation by combining graph-aware and graph-free experts to improve prediction under structural shifts.
This paper shows that fine-tuning autoencoders for reconstruction reduces effective dimensionality, making standard velocity prediction inefficient in diffusion models, and proposes using x0-prediction to focus on the signal manifold, consistently improving text-to-image generation.
LE4Mob is an inductive, distance-aware, and general-purpose location embedding framework for enhancing human mobility modeling in tasks such as next location prediction and commuter flow generation.
MIRCID is a framework that infers hub-miRNAs to enhance drug mechanism-of-action modeling by comparing gene expression with inferred transcription factor activity and miRNA expression, achieving improvements in pathway classification and similarity-based retrieval.
This paper establishes two limitations of Maximal Coding Rate Reduction (MCR²) for out-of-distribution generalisation, showing that it can fail under distribution shift and that incorporating invariance principles does not eliminate this failure.
Tactile-JEPA is a self-supervised pre-training method for distributed tactile sensors that uses spatial topology to learn representations, improving force estimation and orientation tasks in robotics over prior state-of-the-art.
The paper introduces THAW-VLA, a method that distills world-model representations into Vision-Language-Action models for robotics, enhancing robustness and performance on simulation and real hardware without additional inference overhead.
The paper explores grokking, a delayed generalization phenomenon in neural networks, and introduces Geometric Dimensionality Regularization (GeomDR) to control representation geometry, accelerating grokking by up to 52x.
The paper introduces Lens, a training-free framework for multimodal representation learning that addresses semantic perspective misalignment, achieving significant performance improvements on MMEB datasets without parameter updates.
This paper proposes a new unified pre-training framework for medical code sequences that captures hierarchical structures and complex interactions, demonstrating superior performance in clinical event prediction and drug repositioning case studies for Alzheimer's disease.
FLAT introduces a shared sequence of continuous tokens for images and text, enabling flexible-length representations for multimodal tasks like generation and retrieval, with strong performance on benchmarks.