NOML-NOML: hierarchical TD3 + anchor policy for flight control [P]
Summary
Introduced NOML, a custom reinforcement learning algorithm for continuous flight control that uses a hierarchical actor, anchor policy, and mirror learning to prevent oscillation and improve stability. The code is open-sourced on GitHub.
Similar Articles
Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization
The paper introduces BATON, a dual-axis policy optimization framework for LLM agents using Bayesian Feedback Attribution and Trajectory Mass Normalization, demonstrating improved performance in reinforcement learning experiments.
DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat
A hierarchical multi-agent reinforcement learning framework combining graph attention and dynamic role assignment improves tactical coordination and win rates in air combat.
Interactive Training 2: Auditable Control Plane for Live Model Training
Interactive Training 2 introduces an auditable control plane for steering live model training through a shared protocol, allowing humans and automated controllers to submit requests that are validated and applied at safe control points. The system is demonstrated across NLP and reinforcement-learning workflows.
Normative Loss Landscape Navigation: A Trajectory-Based Approach to Mitigating Forgetting in Incremental Learning
A new paper proposes TMLN (Trajectory-Modulatory Landscape Navigation), which treats continual learning as an optimal control problem over a curved loss landscape, using a diagonal empirical Fisher Information Matrix and trajectory-based gradient preconditioning to mitigate catastrophic forgetting without additive regularization penalties.
ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents
The paper introduces ATOD, a hybrid online distillation algorithm combining on-policy distillation and reinforcement learning for training small language model agents in multi-turn tasks, featuring an annealed OPD-RL schedule and Turn-level Disagreement-Uncertainty Reweighting to improve dense supervision.