Tag
This paper introduces a predictive formulation for deep reinforcement learning that augments the state space with future reference horizons to enable anticipatory control for trajectory tracking. Simulation results show significant error reduction, though zero-shot transfer to physical hardware reveals a sim-to-real gap.
This work-in-progress paper proposes embedding a differentiable physics model into the PPO actor loss function to penalize anticipated safety violations in reinforcement learning, evaluated on a simulated 1-DoF helicopter system. The physics-informed soft regularizations reduce constraint violations while maintaining reliable target tracking.
Introduces EVOM, an agentic meta-evolution framework using an LLM-based design agent to automatically discover high-performance actor-critic architectures for reinforcement learning, outperforming manual baselines and prior methods on continuous control tasks.
ESPO introduces an early-stopping mechanism for reinforcement learning that detects and terminates failed reasoning trajectories in LLMs, improving mathematical reasoning performance while reducing compute by over 20%.