temporal-difference-learning

Tag

Cards List
#temporal-difference-learning

A Single Stepsize Suffices for Unprojected Linear TD(0): Simultaneous Robust and Fast Rates via Polyak--Ruppert Averaging

arXiv cs.LG · 2026-06-25 Cached

This paper provides high-probability guarantees for an unprojected linear TD(0) algorithm with Polyak–Ruppert averaging under Markovian sampling, using a single stepsize schedule that achieves both robust curvature-free and fast curvature-dependent convergence rates.

0 favorites 0 likes
#temporal-difference-learning

Temporal Difference Learning for Diffusion Models

arXiv cs.LG · 2026-06-16 Cached

This paper introduces a temporal difference (TD) learning objective for diffusion models that enforces cross-time consistency along the denoising trajectory. It reformulates denoising as a reinforcement learning policy evaluation problem, showing significant improvements in sample quality (FID), especially for few-step samplers.

0 favorites 0 likes
#temporal-difference-learning

Behavior-Induced Mirror-Prox Temporal-Difference Learning for Faster Off-Policy Prediction

arXiv cs.AI · 2026-05-29 Cached

This paper proposes STHTD-MP, a behavior-induced Mirror-Prox temporal-difference method for faster off-policy prediction in reinforcement learning. It replaces the covariance metric with the behavior-policy Bellman matrix and provides convergence analysis and experimental comparisons.

0 favorites 0 likes
#temporal-difference-learning

On the Divergence of Differential Temporal Difference Learning without Local Clocks

arXiv cs.LG · 2026-05-11 Cached

This paper addresses an open problem in reinforcement learning by providing a counterexample showing that differential temporal difference learning can diverge when using a global clock, despite converging with a local clock, in average-reward settings.

0 favorites 0 likes
← Back to home

Submit Feedback