safe-reinforcement-learning

Tag

Cards List
#safe-reinforcement-learning

Adaptive Probabilistic Shielding by Learning MDPs for Safe Reinforcement Learning

arXiv cs.LG · 2026-08-21 Cached

This paper introduces adaptive probabilistic shielding for safe reinforcement learning, where the shield is computed from an online learned MDP model, adapting as the model becomes more accurate.

0 favorites 0 likes
#safe-reinforcement-learning

SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy

arXiv cs.LG · 2026-07-16 Cached

SteinGate introduces a distributional safety certificate using Kernelized Stein Discrepancy to detect rare catastrophic tail events in safe reinforcement learning, dynamically adapting policy updates to reduce constraint violations while maintaining competitive returns.

0 favorites 0 likes
#safe-reinforcement-learning

Integrating Physics-Informed Neural Networks for Safe Reinforcement Learning in a 1-DoF Helicopter System

arXiv cs.LG · 2026-07-07 Cached

This work-in-progress paper proposes embedding a differentiable physics model into the PPO actor loss function to penalize anticipated safety violations in reinforcement learning, evaluated on a simulated 1-DoF helicopter system. The physics-informed soft regularizations reduce constraint violations while maintaining reliable target tracking.

0 favorites 0 likes
#safe-reinforcement-learning

CSPO: Constraint-Sensitive Policy Optimization for Safe Reinforcement Learning

arXiv cs.AI · 2026-06-15 Cached

This paper proposes Constraint-Sensitive Policy Optimization (CSPO), a first-order primal-dual method for safe reinforcement learning that incorporates local constraint sensitivity to improve safety recovery and reduce oscillations near safety boundaries, achieving higher constrained returns on navigation and locomotion benchmarks.

0 favorites 0 likes
#safe-reinforcement-learning

Contract-Based Compositional Shielding for Safe Multi-Agent Reinforcement Learning

arXiv cs.LG · 2026-06-15 Cached

A method for contract-based compositional shielding that ensures global safety in multi-agent reinforcement learning without centralized runtime control, using local LTL obligations and a multi-armed bandit to optimize team reward.

0 favorites 0 likes
#safe-reinforcement-learning

Robust Shielding for Safe Reinforcement Learning

arXiv cs.AI · 2026-06-02 Cached

Introduces a novel shielding framework for robust Markov decision processes (RMDPs) that formally guarantees safety under uncertain transition dynamics, proving soundness and optimality. The approach combines with PAC guarantees for learned models, enabling safe reinforcement learning in unknown environments.

0 favorites 0 likes
← Back to home

Submit Feedback