Delay-corrected Bellman operator + causal attribution for constrained RL contraction proof under unknown stochastic delay [R]
Summary
Introduces CCPL, a method to address delayed and stochastic consequences in constrained reinforcement learning using a delay-corrected Bellman operator and an Interventional Consequence Net for causal attribution, with a contraction proof under unknown delays.
Similar Articles
Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF
This paper introduces Retroactive Advantage Correction (RAC), a closed-form bias correction method for delay-aware RLHF that handles asynchronous reward signals by queuing and reinjecting delayed rewards with a V-trace-style clipped residual update.
An Introduction to Causal Reinforcement Learning
This paper introduces causal reinforcement learning (CRL), unifying causal inference and reinforcement learning under a structural causal model framework, and explores novel learning settings such as generalized policy learning and counterfactual learning.
From Cumulative Constraints to Adaptive Runtime Safety Control for Nonstationary Reinforcement Learning
Proposes CPSS, a runtime safety mechanism that converts cumulative cost constraints into adaptive state-level thresholds for safe reinforcement learning in nonstationary environments, demonstrating reduced violations in highway merging scenarios.
Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents
Proposes CVT-RL, a constrained policy-gradient algorithm with policy-conditioned counterfactual contribution estimation and verifiable rewards, improving long-horizon language agent reliability and reducing reward hacking.
Delayed Verification Destabilizes Multi-Agent LLM Belief: Instability Thresholds and Optimal Corrector Placement
This paper models the impact of delayed verification in multi-agent LLM systems, revealing that delayed correction can destabilize consensus and cause oscillations. It derives closed-form stability thresholds and provides a greedy approximation for optimal corrector placement, validated with experiments on five open models.