Tag
This paper introduces a reinforcement learning method for training a meta-reasoning policy that selects between fast reactive control and slower deliberative planning based on uncertainty in the reactive policy, achieving better balance and adaptivity in navigation tasks.