human-demonstrations

Tag

Cards List
#human-demonstrations

@dair_ai: Highly-recommended read from MIT on the part of RL with verifiable rewards that everyone keeps hitting. RLVR only optim…

X AI KOLs Timeline · 2026-07-03 Cached

This paper from MIT proposes an adversarial generator-discriminator framework that combines verifiable rewards with a learned signal from human demonstrations to address issues like diversity collapse, unnatural responses, and reward hacking in RLVR training of language models.

0 favorites 0 likes
#human-demonstrations

@dair_ai: // Self-play with a pinch of human data // Really cool paper combining human demonstrations and self-play RL. 30 minute…

X AI KOLs Following · 2026-06-20 Cached

A research paper that combines a small amount of human demonstrations as a regularization objective with self-play reinforcement learning, enabling human-compatible driving policies using far less human data (30 minutes vs thousands of hours) and training in 15 hours on a single consumer GPU.

0 favorites 0 likes
#human-demonstrations

Procgen and MineRL Competitions

OpenAI Blog · 2020-06-20 Cached

OpenAI co-organizes the MineRL 2020 Competition to advance sample-efficient reinforcement learning algorithms that leverage human demonstrations. Participants compete to obtain a diamond in Minecraft using only 8 million simulator samples and 4 days of single-GPU training, with access to a 60+ million frame human demonstration dataset.

0 favorites 0 likes
← Back to home

Submit Feedback