Tag
This paper revisits the overestimation bias in Q-learning under large discrete action spaces, proposing an action intersection strategy that enables semi-decoupling between two Q-functions to balance overestimation and underestimation. Experiments in tabular and deep RL settings show improved performance over several baselines.