Tag
This paper proposes NFTR, a method for offline goal-conditioned reinforcement learning that uses normalizing flows for subgoal policies and a triangle-slack reweighting to address optimistic bias and mode collapse in hierarchical implicit Q-learning.