inductive-bias

Tag

Cards List
#inductive-bias

Graph Machine: Exploring Edge Mechanisms as an Inductive Bias

arXiv cs.LG · 2026-08-10 Cached

This paper introduces Graph Machine, an architecture with explicit edge-based mechanisms (edge-augmented attention and edge-centric referral) to improve iterative relational reasoning. Experiments on Sudoku show it outperforms Transformer baselines, with ablations and mechanistic analysis attributing gains to the edge mechanisms.

0 favorites 0 likes
#inductive-bias

Hierarchical Grading in Large Language Models

arXiv cs.LG · 2026-07-28 Cached

This paper introduces Graded Large Language Models (GLLMs), an algebraic framework that imposes a hierarchical grading on transformer representations, theoretically improving sample efficiency for language hierarchies while preserving inference complexity. It provides geometric and information-theoretic justifications, and outlines a grade-selection procedure validated in a companion manuscript.

0 favorites 0 likes
#inductive-bias

The Information Shadow: Measuring Structural Limits on What Language Models Can Learn

arXiv cs.LG · 2026-07-22 Cached

The paper introduces the 'information shadow' concept—structural limits on what language models can learn from text that persist regardless of scale, data, or architecture. It identifies three distinct types of limits and provides probes to detect them, with implications for benchmark design and capability auditing.

0 favorites 0 likes
#inductive-bias

Language Re-generation: An investigation into information locality effects on reconstruction

arXiv cs.CL · 2026-07-14 Cached

This paper investigates how GPT-2 models pre-trained on impossible languages (with disrupted information locality) can recover natural English, showing a bias toward shorter dependency lengths and dissociation between structural and surface recovery.

0 favorites 0 likes
#inductive-bias

You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences

Hugging Face Daily Papers · 2026-06-14 Cached

The paper introduces Temporal Difference in Vision (TDV), a self-supervised learning method for video that relies only on a causal assumption that past causes future, avoiding strong inductive biases while matching state-of-the-art on dense spatial tasks.

0 favorites 0 likes
#inductive-bias

[R] Measuring the Symmetry--Data Exchange Rate

Reddit r/MachineLearning · 2026-06-04 Cached

This paper empirically measures the symmetry–data exchange rate predicted by equivariance theory, finding that wrong-group symmetry constraints are actively harmful, augmentation with test-time orbit averaging matches equivariant architectures, and the theoretical |G|-fold sample complexity reduction is only weakly confirmed with wide confidence intervals. The study is explicitly exploratory and not pre-registered.

0 favorites 0 likes
#inductive-bias

The Loss Is Not Enough: Sampling Conditions and Inductive Bias in Contrastive Representation Learning

arXiv cs.LG · 2026-06-04 Cached

This paper develops a measure-theoretic framework analyzing when contrastive learning recovers meaningful latent geometry, introducing a 'diversity condition' on positive-pair sampling and a support-corrected InfoNCE variant, with experiments validating that sampling diversity and architectural inductive bias interact critically in contrastive representation learning.

0 favorites 0 likes
#inductive-bias

Measuring the Symmetry--Data Exchange Rate

Hugging Face Daily Papers · 2026-05-31

This exploratory study empirically measures the symmetry–data exchange rate predicted by equivariance theory on controlled C_n-symmetric tasks, finding that wrong-group constraints are actively harmful, augmentation with test-time orbit averaging matches equivariant models exactly, and the empirical exchange rate is broadly consistent with theory but statistically inconclusive. The authors emphasize the study's exploratory nature and call for registered replications.

0 favorites 0 likes
#inductive-bias

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias

arXiv cs.LG · 2026-05-29 Cached

This paper introduces the concept of 'initialization memory' to study how much of the random initialization bias survives training in deep networks, showing that low-learning-rate SGD preserves initialization while Adam-family optimizers erase it, and linking this to forgetting dynamics.

0 favorites 0 likes
#inductive-bias

Energy-Gated Attention and Wavelet Positional Encoding: Complementary Inductive Biases for Transformer Attention

arXiv cs.LG · 2026-05-27 Cached

This paper proposes Energy-Gated Attention (EGA) and Morlet Positional Encoding (MoPE) to address missing inductive biases in transformer attention: token salience and scale-adaptive locality. Experiments on TinyShakespeare show superadditive gains when combined, highlighting complementarity.

0 favorites 0 likes
#inductive-bias

On the Role of Inductive Bias in Time-Series Pretraining: A Case Study in Learning Generalizable Representations for Clinical Time Series

arXiv cs.LG · 2026-05-27 Cached

This paper investigates the role of inductive bias in time-series pretraining for clinical data, proposing PathoFM, an encoder-centric transformer pretrained on multivariate gait windows. The study compares different pretraining objectives and finds that dynamics-centric mixtures yield the most balanced transfer across classification and regression tasks.

0 favorites 0 likes
#inductive-bias

Graph Alignment Topology as an Inductive Bias for Grounding Detection

arXiv cs.CL · 2026-05-25 Cached

This paper introduces Graph Alignment Topology as an inductive bias for grounding detection, using a graph neural network to model alignment structure between reference information and LLM outputs. The method achieves state-of-the-art results on multiple hallucination and question-answering datasets, outperforming GPT-4o.

0 favorites 0 likes
#inductive-bias

When Irregularity Helps: A Subclass Analysis of Inductive Bias in Neural Morphology

arXiv cs.CL · 2026-05-21 Cached

This paper investigates how character-level transformer models generalize to irregular verb subtypes in Japanese past-tense inflection. Controlled experiments show that including irregular examples can improve generalization, challenging the assumption that regularity simplifies learning.

0 favorites 0 likes
← Back to home

Submit Feedback