Tag
This paper provides high-probability guarantees for an unprojected linear TD(0) algorithm with Polyak–Ruppert averaging under Markovian sampling, using a single stepsize schedule that achieves both robust curvature-free and fast curvature-dependent convergence rates.
This paper proposes behavior-aware auxiliary corrections for off-policy temporal-difference prediction, introducing BA-TDC and BA-TDRC algorithms that replace the auxiliary covariance matrix with the behavior Bellman matrix to improve stability and convergence. Theoretical analysis and experiments on standard benchmarks validate the effectiveness of the proposed methods.