Tag
This paper introduces a game-theoretic approach to fine-tuning language models that optimizes the trade-off between reward and deviating from a reference policy, providing a principled method for setting the KL regularization coefficient.
FootsiesGym is an open-source environment for studying non-trivial two-player zero-sum imperfect-information games, built on the minimalist fighting game Footsies. It provides a vectorized simulator for efficient training and benchmarks several reinforcement learning algorithms.