Browser demo of our Clash Royale RL environment: a 5.6k-parameter REINFORCE policy learns defensive placement against a brute-force optimum [P]
Summary
The article introduces a browser-based demo for an open-source Clash Royale reinforcement learning environment, where a small REINFORCE policy learns defensive placements with comparisons to brute-force optimal strategies.
Similar Articles
ClashRoyaleAi: an open-source, deterministic Clash Royale simulator for RL, with recurrent PPO, lookahead search and expert iteration [P]
ClashRoyaleAi is an open-source, deterministic Clash Royale simulator designed for reinforcement learning, featuring a PPO agent with lookahead search and expert iteration, showing improved win rates through simple lookahead techniques.
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning
This paper introduces CHERRL, a controllable environment for studying reward hacking in rubric-based reinforcement learning, where LLM-as-a-Judge biases can be injected to reproduce and analyze hacking behaviors. The authors also explore an agent-based system for automatically detecting reward hacking onset from training logs.
SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning
SERPO introduces a self-evolving rubric policy optimization framework for test-time reinforcement learning in open-ended generation, replacing answer voting with a closed loop that co-evolves response evidence, query-specific rubrics, and policy parameters, achieving significant improvements on health and research benchmarks.
I transformed Pokelike.xyz into a LLM and RL benchmark!
A data scientist has created a benchmark for reinforcement learning and large language models by transforming Pokelike.xyz into a game environment where AI bots can be trained and tested. The project is open-source and invites contributions to improve bot performance.
@dair_ai: // Self-play with a pinch of human data // Really cool paper combining human demonstrations and self-play RL. 30 minute…
A research paper that combines a small amount of human demonstrations as a regularization objective with self-play reinforcement learning, enabling human-compatible driving policies using far less human data (30 minutes vs thousands of hours) and training in 15 hours on a single consumer GPU.