Tag
This paper systematically explores four deep reinforcement learning solutions (DQN, REINFORCE, PPO, and MuZero) for the asymmetric Nepali board game Baghchal, finding that MuZero achieves the best win rates due to model-based planning via Monte Carlo Tree Search.