Delay-corrected Bellman operator + causal attribution for constrained RL contraction proof under unknown stochastic delay [R]

Reddit r/MachineLearning Papers

Summary

Introduces CCPL, a method to address delayed and stochastic consequences in constrained reinforcement learning using a delay-corrected Bellman operator and an Interventional Consequence Net for causal attribution, with a contraction proof under unknown delays.

Standard constrained RL assumes consequences are immediate and attributable to the current action. This breaks down whenever violations are delayed and stochastic, which is most real-world settings you end up penalizing whatever action happened to precede the observed violation, not the action that caused it. Working on CCPL (Causal Consequence-Penalized Learning) to address this: - A delay-corrected Bellman operator using an adaptive effective discount learned from the consequence-delay distribution. Contraction proof holds under unknown stochastic delay. - An Interventional Consequence Net (ICN), pretrained on structural-causal-model labels, estimating marginal causal contribution per action for attribution rather than penalizing based on temporal proximity. Limitations, to be upfront about them: - The ICN currently requires access to the environment's structural causal model to generate pretraining labels it's not learned end-to-end from observational or interventional data alone. That's a real constraint on applicability outside benchmark settings where the SCM is known or can be reasonably specified. Open to contributions and collaborators, especially if you work in constrained/safe RL or causal inference feel free to open an issue or reach out directly.
Original Article

Similar Articles

An Introduction to Causal Reinforcement Learning

arXiv cs.AI

This paper introduces causal reinforcement learning (CRL), unifying causal inference and reinforcement learning under a structural causal model framework, and explores novel learning settings such as generalized policy learning and counterfactual learning.