Tag
This paper proposes environment evolution to incrementally increase task difficulty for terminal agents off-policy, enhancing continuous learning signals and improving performance on benchmarks like Terminal-Bench 2.1.
OpenAI presents Hindsight Experience Replay (HER), a reinforcement learning technique that enables robots to learn from failed attempts by retroactively treating achieved alternative outcomes as successful goals, allowing learning even with sparse reward signals.