layer-wise

Tag

Cards List
#layer-wise

DREG: A Layer-Wise Jacobian Regularization as a General-Purpose Penalty

arXiv cs.LG · 2026-06-24 Cached

This paper presents a large-scale empirical study of the Derivative Regularization (DREG) penalty, showing it achieves high accuracy and noise robustness, particularly with GELU activation and data-scarce regimes, positioning it as a general-purpose plug-and-play regularizer for neural networks.

0 favorites 0 likes
#layer-wise

Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs

arXiv cs.LG · 2026-05-14 Cached

This paper introduces Layer-wise Representation Dynamics (LRD), a framework with three measurement families to analyze how hidden states change across layers in language models. Applied to 31 models on 30 MTEB tasks, LRD reveals architectural differences and enables label-free model selection and inference-time layer pruning.

0 favorites 0 likes
← Back to home

Submit Feedback