Tag
SkillEvoReg introduces a regularization framework to prevent overfitting in the skill evolution of language-model agents, combining dropout, local regularization, and causal validation to maintain performance while controlling skill growth.
FuseReg is a regularization method for Representation Autoencoders that mitigates the reconstruction-generation gap by using random layer-subset sampling during training, improving generation performance and decoder robustness across different layer inputs.
The paper proposes a spectral connectivity-regularized graph learning framework (SCoGL) that incorporates Laplacian spectral priors to improve graph recovery and downstream tasks like graph signal denoising when data is scarce.
SPAR is a pretraining objective that uses a gated KL loss to stabilize language model predictions against irrelevant prefix text, improving robustness in long-context scenarios as demonstrated on multiple benchmarks.
This paper introduces Regularized Recursive Self-Improvement (RRSI) for AI agent harnesses, which applies regularization to prevent overfitting during recursive evolution, demonstrating performance gains on multiple benchmarks.
The paper explores grokking, a delayed generalization phenomenon in neural networks, and introduces Geometric Dimensionality Regularization (GeomDR) to control representation geometry, accelerating grokking by up to 52x.
The paper introduces Regularized Emphatic Temporal-Difference Learning (RETD), which ensures stability under constant stepsizes by normalizing the emphatic TD signal, with proofs of convergence and experimental validation.
This paper investigates the limits of proper learning and regularization in multiclass learning, resolving open problems by demonstrating that learning cannot always be reduced to proper learning and that regularization has structural constraints.
This paper proposes a training-time explainability framework for multilingual hate speech detection, aligning model reasoning with human rationales to improve classification performance and interpretability, evaluated on English and Hinglish datasets.
The paper introduces ERPO, a method that moves regularization from the action-side to the input-side by controlling query distribution, addressing the stability-exploration dilemma in LLM policy optimization, and showing improvements on mathematical reasoning benchmarks.
This paper presents a compositional theory of curvature in probabilistic circuits, showing that the Hessian trace factorizes per sum node into circuit flow and local sharpness, and introduces an adaptive sharpness-aware regularizer that preserves closed-form EM updates while improving generalization.
A paper introduces TC-LeWM, applying regularization losses like SIGReg and VISReg to temporal latent residuals to enable multi-task learning in LeWM, significantly improving success rates on the LIBERO robot arm benchmark.
This paper identifies 'suboptimal collapse' in RL post-training of time series foundation models and proposes Ground-Truth Neighborhood Regularization (GTN-R) to keep output distributions near the ground truth, improving forecasting performance.
This paper proposes BRACE, a method that encodes an Ordered Reasoning Chain to detect ever-shifting harmful chat dialogue, achieving high harm-type F1 scores with both encoder and decoder backbones.
PhysAttNet is a physics-informed attention framework that augments lightweight CNN forecasters with domain-guided regularization to improve accuracy and generalization in industrial and astrophysical time series forecasting.
This paper proposes a neural ODE-based regularization method that enforces latent embeddings in reinforcement learning agents to follow consistent ODE flows, aligning representation learning with environment dynamics and yielding performance gains on Atari and gridworld benchmarks.
This paper proposes SCORE, a self-concordance-inspired quasi-Newton method for training physics-informed neural networks (PINNs). It uses a decrement-coupled shifted secant geometry to improve final accuracy on nonlinear PDE benchmarks without requiring Hessian computations.
This paper introduces a novel measure called relative parameter importance for task-agnostic, replay-free continual learning, enabling better balance between stability and plasticity by regularizing only parameters critical for past tasks while allowing others to update for backward knowledge transfer. The method is evaluated on class-incremental and domain-incremental text classification tasks.
This paper introduces Omega-S, a lightweight, data-free regularization penalty for low-rank fine-tuning that improves retention of original model capabilities by penalizing variance in weight-matrix node degrees. Experiments on Llama-3-8B with LoRA show it retains more original capability than no regularization or tuned baselines.
This paper introduces Modality Contribution Drift (MCD) in multimodal continual learning and proposes CMCDR, a regularization method with replay-based and replay-free variants to preserve modality contribution structures across incremental tasks.