deep-rl

Tag

Cards List
#deep-rl

Revisiting Overestimation Bias Problem of Q-learning: Settling Large Discrete Action Space via Action Intersection

arXiv cs.LG ↗ · 2026-08-14 Cached

This paper revisits the overestimation bias in Q-learning under large discrete action spaces, proposing an action intersection strategy that enables semi-decoupling between two Q-functions to balance overestimation and underestimation. Experiments in tabular and deep RL settings show improved performance over several baselines.

0 favorites 0 likes
#deep-rl

Self-Play Reinforcement Learning under Imperfect Information in Big 2

arXiv cs.LG ↗ · 2026-05-29 Cached

This paper presents a self-play reinforcement learning framework for the four-player imperfect-information card game Big 2, comparing policy-gradient and value-based methods and finding that PPO with entropy regularization outperforms others.

0 favorites 0 likes
#deep-rl

Approximate Next Policy Sampling: Replacing Conservative Target Policy Updates in Deep RL

arXiv cs.LG ↗ · 2026-05-08 Cached

This paper introduces Approximate Next Policy Sampling (ANPS) as an alternative to conservative policy updates in deep reinforcement learning. It proposes Stable Value Approximate Policy Iteration (SV-API) and SV-RL, which align training data with the next policy's state distribution to allow for larger and safer policy updates.

0 favorites 0 likes
#deep-rl

What are MDPs? And how can we Solve them?

ML at Berkeley ↗ · 2021-02-23 Cached

This article explains the fundamentals of Markov Decision Processes (MDPs), a core framework in deep reinforcement learning, using an educational example of a student's daily decisions.

0 favorites 0 likes
#deep-rl

OpenAI Five defeats Dota 2 world champions

OpenAI Blog ↗ · 2019-04-15 Cached

OpenAI Five becomes the first AI to defeat world-champion esports professionals in Dota 2, winning two back-to-back matches against OG at the OpenAI Five Finals. The breakthrough was achieved through unprecedented scaling of training compute rather than novel algorithms, and the team is retiring OpenAI Five while announcing plans to deploy it for public internet play.

0 favorites 0 likes
#deep-rl

Variance reduction for policy gradient with action-dependent factorized baselines

OpenAI Blog ↗ · 2018-03-20 Cached

OpenAI researchers derive a bias-free action-dependent baseline for variance reduction in policy gradient methods, demonstrating improved learning efficiency on high-dimensional control tasks, multi-agent, and partially observed environments.

0 favorites 0 likes
#deep-rl

Learning to model other minds

OpenAI Blog ↗ · 2017-09-14 Cached

OpenAI and University of Oxford researchers present LOLA (Learning with Opponent-Learning Awareness), a reinforcement learning method that enables agents to model and account for the learning of other agents, discovering cooperative strategies in multi-agent games like the iterated prisoner's dilemma and coin game.

0 favorites 0 likes
#deep-rl

UCB exploration via Q-ensembles

OpenAI Blog ↗ · 2017-06-05 Cached

OpenAI presents a novel exploration strategy for deep reinforcement learning using ensembles of Q-functions with upper-confidence bounds (UCB), demonstrating significant performance improvements on the Atari benchmark.

0 favorites 0 likes
#deep-rl

Stochastic Neural Networks for hierarchical reinforcement learning

OpenAI Blog ↗ · 2017-04-10 Cached

OpenAI researchers propose a framework using stochastic neural networks for hierarchical reinforcement learning that pre-trains useful skills guided by a proxy reward, then leverages these skills for faster learning in downstream tasks with sparse rewards or long horizons.

0 favorites 0 likes
#deep-rl

#Exploration: A study of count-based exploration for deep reinforcement learning

OpenAI Blog ↗ · 2016-11-15 Cached

OpenAI researchers demonstrate that a simple count-based exploration approach using hash codes can achieve near state-of-the-art performance on high-dimensional deep RL benchmarks, challenging the assumption that count-based methods cannot scale to continuous state spaces.

0 favorites 0 likes
#deep-rl

RL²: Fast reinforcement learning via slow reinforcement learning

OpenAI Blog ↗ · 2016-11-09 Cached

RL² proposes encoding a fast reinforcement learning algorithm as the weights of a recurrent neural network, learned through slow general-purpose RL, enabling agents to adapt to new tasks with few trials similar to biological learning. The method demonstrates strong performance on both small-scale bandit problems and large-scale vision-based navigation tasks.

0 favorites 0 likes
← Back to home

Submit Feedback