标签
Introduces CODS, an iterative critic-guided data selection method for offline reinforcement learning that retains task performance at low data budgets by selecting high-residual transitions over multiple rounds.