learning-from-preferences

Tag

Cards List
#learning-from-preferences

Learning from human preferences

OpenAI Blog · 2017-06-13 Cached

OpenAI presents a method for training AI agents using human preference feedback, where an agent learns reward functions from human comparisons of behavior trajectories and uses reinforcement learning to optimize for the inferred goals. The approach demonstrates strong sample efficiency, requiring less than 1000 bits of human feedback to train an agent to perform a backflip.

0 favorites 0 likes
← Back to home

Submit Feedback