Tag
This paper evaluates AlphaZero-inspired reinforcement learning for topological control in power networks, achieving 98.43% survivability and emphasizing the effectiveness of minimalist integration with domain heuristics.
This paper examines the gap between strong play and perfect play in AlphaZero for sparsely rewarded games, using Connect Four and Chomp as testbeds, and proposes an auxiliary loss (AZAL) to improve oracle consistency in optimal play.
This paper presents WallZero, an AlphaZero-based agent for the two-player board game WallGo, which defeats professional Go players and is used to analyze game balance and strategies.
The article analyzes how AlphaZero's value predictions are shaped by self-play training data and noise, questioning whether they reliably estimate win chances against opponents with different play styles despite AlphaZero's strong empirical performance.