Tag
This paper introduces Resource-Adaptive Primal-Dual Learning for one-warehouse multi-store inventory systems with censored demand, achieving logarithmic regret by dynamically adapting to changing resource levels.
This paper proposes Constraint-Sensitive Policy Optimization (CSPO), a first-order primal-dual method for safe reinforcement learning that incorporates local constraint sensitivity to improve safety recovery and reduce oscillations near safety boundaries, achieving higher constrained returns on navigation and locomotion benchmarks.