representation-learning

Tag

Cards List
#representation-learning

Supervised Latent Restructuring for Small-Data Quantum Learning in Plant Phenomics

arXiv cs.LG ↗ · 2026-05-21 Cached

This paper proposes a hybrid quantum-classical workflow for plant phenomics classification under small-data regimes, using supervised latent restructuring (PCA + LDA) to improve geometric separability before quantum kernel alignment. Experiments show improved separability but highlight compression trade-offs and the difficulty of achieving strong quantum performance.

0 favorites 0 likes
#representation-learning

Neural Collapse by Design: Learning Class Prototypes on the Hypersphere

arXiv cs.LG ↗ · 2026-05-21 Cached

This paper shows that cross-entropy and supervised contrastive learning are both forms of prototype learning on the hypersphere and proposes normalized losses (NTCE and NONL) that achieve Neural Collapse by design, outperforming standard methods.

0 favorites 0 likes
#representation-learning

Representation over Routing: Overcoming Surrogate Hacking in Multi-Timescale PPO

Hugging Face Daily Papers ↗ · 2026-05-21 Cached

This paper identifies surrogate hacking and temporal uncertainty as failure modes in multi-timescale RL, and proposes a Target Decoupling architecture that removes routing from the actor, using the critic for auxiliary representation learning. The method eliminates policy collapse on the LunarLander-v2 benchmark and stably surpasses the 'Environment Solved' threshold without hyperparameter hacking.

0 favorites 0 likes
#representation-learning

RiT: Vanilla Diffusion Transformers Suffice in Representation Space

Hugging Face Daily Papers ↗ · 2026-05-21 Cached

The paper introduces RiT, a vanilla Diffusion Transformer trained on frozen DINOv2 features using flow matching with x-prediction, achieving competitive FID scores on ImageNet 256×256 with fewer parameters and fast sampling without distillation.

0 favorites 0 likes
#representation-learning

VCR: Learning Valid Contextual Representation for Incomplete Wearable Signals

arXiv cs.LG ↗ · 2026-05-20

VCR is a self-supervised framework that learns robust representations from incomplete wearable signals using orthogonal tokenization and missing-aware mixture-of-experts, improving performance under modality missingness.

0 favorites 0 likes
#representation-learning

In-Context Learning Operates as Concept Subspace Learning

arXiv cs.LG ↗ · 2026-05-20

This paper proposes that in-context learning in LLMs operates through low-dimensional concept subspaces, where task-relevant information concentrates in a small fraction of the representation space, supported by experiments on Llama-3-8B and Qwen2.5-7B.

0 favorites 0 likes
#representation-learning

TERGAD: Structure-Aware Text-Enhanced Representations for Graph Anomaly Detection

arXiv cs.CL ↗ · 2026-05-20 Cached

TERGAD is a novel data augmentation framework that uses large language models to translate node-level topological properties into semantic narratives, then fuses these with original node attributes via a gated dual-branch autoencoder for graph anomaly detection, achieving state-of-the-art results on six datasets.

0 favorites 0 likes
#representation-learning

Language models struggle with compartmentalization

arXiv cs.CL ↗ · 2026-05-20 Cached

This paper investigates compartmentalization in LLMs, where models fail to share statistical strength across distinct representations of the same concept, leading to reduced sample efficiency and model capacity. The authors demonstrate this phenomenon in multilingual and multi-format settings and show that synthetic parallel data does not fully resolve it.

0 favorites 0 likes
#representation-learning

Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws

Hugging Face Daily Papers ↗ · 2026-05-20 Cached

This paper shows that different optimizers, particularly Muon versus AdamW, induce distinct spectral scaling behaviors in Transformer models, with Muon achieving significantly better utilization of representation capacity, and suggests that optimizer choice should be a first-class axis in scaling laws.

0 favorites 0 likes
#representation-learning

Geometric Asymmetry in MoE Specialization: Functional Decorrelation and Representational Overlap

arXiv cs.LG ↗ · 2026-05-19 Cached

This paper introduces a Jacobian-PCA-Grassmann framework to analyze the geometric structure of expert specialization in Mixture-of-Experts (MoE) Transformers. It finds that experts exhibit strong functional decorrelation while their representations overlap, and that routing sparsity significantly influences this geometry.

0 favorites 0 likes
#representation-learning

TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens

arXiv cs.AI ↗ · 2026-05-19 Cached

The paper introduces TTE-Flash, a method that replaces explicit chain-of-thought reasoning with latent think tokens to generate reasoning-aware multimodal representations at constant inference cost, outperforming explicit CoT baselines on the MMEB-v2 benchmark.

0 favorites 0 likes
#representation-learning

Representation Without Reward: A JEPA Audit for LLM Fine-Tuning

arXiv cs.LG ↗ · 2026-05-18 Cached

This paper audits Joint-embedding predictive architectures (JEPA) for LLM fine-tuning on a natural-language-to-regex task, testing twenty-two auxiliary objectives. The results show that hidden-state representation improvements are only weakly coupled to decoded-task accuracy, with no auxiliary surviving family-wise correction.

0 favorites 0 likes
#representation-learning

How Data Augmentation Shapes Neural Representations

arXiv cs.LG ↗ · 2026-05-18 Cached

This paper uses shape analysis tools to characterize how different data augmentation strategies reshape the geometry of neural network representations, finding that augmentation strength and type lead to distinct, well-behaved trajectories in shape space.

0 favorites 0 likes
#representation-learning

AudioMosaic: Contrastive Masked Audio Representation Learning

arXiv cs.LG ↗ · 2026-05-15 Cached

AudioMosaic introduces a contrastive learning-based audio encoder that uses structured time-frequency masking on spectrogram patches for efficient large-batch training, achieving state-of-the-art performance on audio benchmarks and improving audio-language models.

0 favorites 0 likes
#representation-learning

CSI-JEPA: Towards Foundation Representations for Ubiquitous Sensing with Minimal Supervision

arXiv cs.LG ↗ · 2026-05-15 Cached

CSI-JEPA is a self-supervised framework for learning reusable representations from unlabeled Wi-Fi channel state information, enabling label-efficient multi-task sensing. It achieves up to 98% label savings and outperforms supervised models.

0 favorites 0 likes
#representation-learning

A Unified Geometric Framework for Weighted Contrastive Learning

arXiv cs.LG ↗ · 2026-05-15 Cached

This paper introduces a unified geometric framework showing that weighted InfoNCE objectives can be interpreted as Distance Geometry Problems, providing exact characterizations of optimal embeddings for supervised and weakly supervised contrastive learning methods and revealing when such embeddings are geometrically realizable, degenerate, or inconsistent.

0 favorites 0 likes
#representation-learning

Rethinking Molecular OOD Generalization via Target-Aware Source Selection

arXiv cs.LG ↗ · 2026-05-15 Cached

This paper introduces SCOPE-Bench, a benchmark for evaluating molecular out-of-distribution generalization, and POMA, a framework using reinforcement learning to select source domains for domain adaptation, achieving significant error reductions on 3D molecular models.

0 favorites 0 likes
#representation-learning

Non-linear Interventions on Large Language Models

arXiv cs.CL ↗ · 2026-05-15 Cached

This paper introduces a general formulation of non-linear intervention for large language models, extending beyond the Linear Representation Hypothesis to manipulate features encoded along non-linear manifolds, and validates the approach on refusal bypass steering.

0 favorites 0 likes
#representation-learning

Network-Aware Bilinear Tokenization for Brain Functional Connectivity Representation Learning

arXiv cs.AI ↗ · 2026-05-15 Cached

NERVE proposes a network-aware bilinear tokenization method for self-supervised learning on brain functional connectivity matrices using masked autoencoders, improving representation learning across developmental cohorts.

0 favorites 0 likes
#representation-learning

Graph-Based Financial Fraud Detection with Calibrated Risk Scoring and Structural Regularization

arXiv cs.LG ↗ · 2026-05-14 Cached

This paper proposes a graph neural network framework for financial fraud detection that integrates transaction records and identity information into node attributes, employs a multi-layer message passing mechanism, and uses weighted supervision and structural consistency regularization to improve risk scoring and probability calibration. Experiments on a public dataset show the method outperforms existing approaches.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback