Play Like Champions: Counterfactual Feedback Generation in Latent Space

arXiv cs.LG Papers

Summary

Introduces a framework called Latent Maps of Performance for generating counterfactual feedback in StarCraft II using a Guided Variational Autoencoder trained on professional replays, enabling improvement trajectories for amateur players.

arXiv:2607.00190v1 Announce Type: new Abstract: Recent advances in reinforcement learning have produced superhuman agents across a wide range of competitive games. As a byproduct, researchers have begun studying how these agents play, extracting behavioral representations, analyzing decision structure, and modeling the latent geometry of expert performance. However, this growing body of work has overwhelmingly focused on defeating human players rather than providing feedback, leaving a critical gap in creating model solutions to improve human players. Unlike chess and Go, where AI has become integral to player training, real-time strategy (RTS) games lack principled frameworks for translating expert knowledge into actionable feedback. We introduce Latent Maps of Performance, a framework for counterfactual path generation. We focus on StarCraft~II data to model player improvement as an algorithmic recourse within a learned representation space. As inspiration for our work, we have looked at the championship model used in sports science. We trained a Guided Variational Autoencoder model on 23,305 professional tournament replays, enabling counterfactual traversal between losing and winning gameplay profiles. To fulfill our goal, we have devised and verified four traversal strategies on out-of-distribution (OOD) data randomly sampled from a dataset of amateur replays, namely linear interpolation, iterative optimal transport, density-regularized gradient ascent, and neural flow matching, each designed to generate multi-step improvement trajectories that remain grounded in observed expert behavior while moving a player's profile toward winning configurations. Feedback is extracted at multiple granularities to support players at different stages of improvement. Finally, we conclude that there is a trade-off between the path-finding methods we employ and hope that future research will focus on developing model solutions for human improvement.
Original Article
View Cached Full Text

Cached at: 07/02/26, 05:36 AM

# Play Like Champions: Counterfactual Feedback Generation in Latent Space
Source: [https://arxiv.org/html/2607.00190](https://arxiv.org/html/2607.00190)
Andrzej Białecki111contact \(in\-order\):andrzej\.bialecki94@gmail\.com,adam\.mastalerz@polsl\.pl,hzhou30@student\.ubc\.ca,Adam Mastalerz111contact \(in\-order\):andrzej\.bialecki94@gmail\.com,adam\.mastalerz@polsl\.pl,hzhou30@student\.ubc\.ca,Silesian University of TechnologyHan Zhou111contact \(in\-order\):andrzej\.bialecki94@gmail\.com,adam\.mastalerz@polsl\.pl,hzhou30@student\.ubc\.ca,University of British Columbia

###### Abstract

Recent advances in reinforcement learning have produced superhuman agents across a wide range of competitive games\. As a byproduct, researchers have begun studying how these agents play, extracting behavioral representations, analyzing decision structure, and modeling the latent geometry of expert performance\. However, this growing body of work has overwhelmingly focused on defeating human players rather than providing feedback, leaving a critical gap in creating model solutions to improve human players\. Unlike chess and Go, where AI has become integral to player training, real\-time strategy \(RTS\) games lack principled frameworks for translating expert knowledge into actionable feedback\. We introduce Latent Maps of Performance, a framework for counterfactual path generation\. We focus on StarCraft II data to model player improvement as an algorithmic recourse within a learned representation space\. As inspiration for our work, we have looked at the championship model used in sports science\. We trained a Guided Variational Autoencoder model on 23,305 professional tournament replays, with macro\-economic gameplay structure conditioned on match outcome, enabling counterfactual traversal between losing and winning gameplay profiles\. To fulfill our goal, we have devised and verified four traversal strategies on out\-of\-distribution \(OOD\) data randomly sampled from a dataset of amateur replays, namely linear interpolation, iterative optimal transport, density\-regularized gradient ascent, and neural flow matching, each designed to generate multi\-step improvement trajectories that remain grounded in observed expert behavior while moving a player’s profile toward winning configurations\. Feedback is extracted at multiple granularities to support players at different stages of improvement\. Finally, we conclude that there is a trade\-off between the path\-finding methods we employ and hope that future research will focus on developing model solutions for human improvement\.

*Keywords*generative artificial intelligence⋅\\cdotlatent space traversal⋅\\cdotoptimal transport⋅\\cdotvariational autoencoder⋅\\cdotesports

## 1Introduction

Mastering real\-time strategy \(RTS\) games such as StarCraft II has long stood as a grand challenge for artificial intelligence, one only partially addressed by recent advances in reinforcement learning \(RL\)\[[1](https://arxiv.org/html/2607.00190#bib.bib1),[2](https://arxiv.org/html/2607.00190#bib.bib2)\]\. The difficulty stems from the cognitive demands these games place on their players: precise control of units, careful management of economies, and continuous decision\-making in adversarial settings where even momentary lapses can prove decisive\. Success thus hinges on the interplay of multitasking, strategic foresight, and rapid reaction\[[3](https://arxiv.org/html/2607.00190#bib.bib3)\], making RTS an especially rich testbed for studying intelligent behavior\. These high\-frequency interactions make RTS games especially well\-suited for large\-scale behavioral study through open\-source replay parsers and direct game\-engine access\[[4](https://arxiv.org/html/2607.00190#bib.bib4),[5](https://arxiv.org/html/2607.00190#bib.bib5)\]\. A growing body of work leverages data to surface game information for player decision support, both through digital interfaces and physical prototypes\[[6](https://arxiv.org/html/2607.00190#bib.bib6)\]\. In StarCraft II, community tools such as sc2replaystats\[[7](https://arxiv.org/html/2607.00190#bib.bib7)\]and replayman\[[8](https://arxiv.org/html/2607.00190#bib.bib8)\]have emerged to support replay analysis, alongside real\-time dashboards that contextualize gameplay for spectators and post\-match review\[[9](https://arxiv.org/html/2607.00190#bib.bib9)\]\. Strategic summaries and encounter\-level analysis are highly valued by players across genres\[[10](https://arxiv.org/html/2607.00190#bib.bib10)\]\. Game state retrieval by similarity to estimate win probabilities on demand is a possibility\[[11](https://arxiv.org/html/2607.00190#bib.bib11)\]\. In parallel, AI methods have become deeply embedded in game research and development, powering procedural content generation\[[12](https://arxiv.org/html/2607.00190#bib.bib12)\], voice\-driven agents that deepen immersion\[[13](https://arxiv.org/html/2607.00190#bib.bib13)\], human\-like behavior modeling\[[14](https://arxiv.org/html/2607.00190#bib.bib14)\], and automated quality assurance\[[15](https://arxiv.org/html/2607.00190#bib.bib15)\]\. However, most existing analysis tools remain oriented toward broadcast and streaming audiences rather than the players themselves\[[16](https://arxiv.org/html/2607.00190#bib.bib16)\]\.

This player\-facing gap is not unique to gaming\. In robotics, efficient simulators have driven dramatic breakthroughs\[[17](https://arxiv.org/html/2607.00190#bib.bib17),[18](https://arxiv.org/html/2607.00190#bib.bib18)\], producing systems that now rival or exceed human performance in domains as varied as drone racing\[[19](https://arxiv.org/html/2607.00190#bib.bib19),[20](https://arxiv.org/html/2607.00190#bib.bib20)\], badminton\[[21](https://arxiv.org/html/2607.00190#bib.bib21),[22](https://arxiv.org/html/2607.00190#bib.bib22)\], and table tennis\[[23](https://arxiv.org/html/2607.00190#bib.bib23)\]\. Such interdisciplinary efforts are increasingly recognized as accelerators of research progress\[[24](https://arxiv.org/html/2607.00190#bib.bib24)\]\. However, despite agents and robotic systems consistently surpassing average human ability, few of these works offer mechanisms to translate the resulting expertise back to human practitioners seeking to improve\. The skill translation problem is well understood in the sport sciences, where the championship model describes how athletes shorten the path to performance gains by adopting the training methods, techniques, and tactics of successful peers\[[25](https://arxiv.org/html/2607.00190#bib.bib25),[26](https://arxiv.org/html/2607.00190#bib.bib26)\]\. In domains where the competitive space is naturally digitized, this dynamic increasingly extends to AI\. In chess, AI has reshaped human learning by serving as a scalable training partner\. Analyses show that elite human play has steadily improved across the engine era\[[27](https://arxiv.org/html/2607.00190#bib.bib27),[28](https://arxiv.org/html/2607.00190#bib.bib28),[29](https://arxiv.org/html/2607.00190#bib.bib29)\]\. AlphaZero’s games further illustrate how superhuman agents can surface novel strategic ideas for human study\[[30](https://arxiv.org/html/2607.00190#bib.bib30)\]\. Similar patterns have emerged in Go following the rise of superhuman agents\[[31](https://arxiv.org/html/2607.00190#bib.bib31),[32](https://arxiv.org/html/2607.00190#bib.bib32)\]\.

RTS games share the same computational substrate, but, to our knowledge, no comparable bridge exists between agent expertise, representational learning, and human improvement\. In this work, we introduce a latent\-space feedback system that learns compressed representations of quantitative gameplay features from StarCraft II replays and provides manifold\-aware improvement guidance to players\. Our contribution is a framework for training representational models that recover counterfactual improvement trajectories and reconstruct them back into the original feature space\. Given a well\-trained model, points sampled along an “improvement trajectory” in latent space can be decoded and compared with the player’s input vector, yielding immediate feedback on which gameplay features need to change to improve performance, as determined by the model\. Our work builds directly on “SC2EGSet”, rather than asking “*who will win*”, we ask “*what the player should do differently*”\.

## 2Related Work

Our work sits at the intersection of four research threads: StarCraft II as a machine learning domain; variational autoencoders and disentangled representation learning; latent space traversal; and counterfactual explanations as actionable feedback\. We discuss each in turn and position our contribution relative to prior art\.

##### StarCraft II as a Machine Learning Domain:

StarCraft II has become a canonical benchmark for sequential decision\-making under partial observability\.Vinyals et al\. \[[33](https://arxiv.org/html/2607.00190#bib.bib33)\]demonstrated that a combination of imitation learning, multi\-agent self\-play, and a latent conditioning variable for strategy style can produce grandmaster\-level play, establishing that large\-scale replay data contains rich, learnable structure\. However, AlphaStar is an autonomous agent; it optimizes for winning, not for explaining to human players how to improve\. On the other hand, the simplistic nature of benchmarks geared primarily towards multi\-agent solutions does not fit well in the context of providing feedback to players\[[34](https://arxiv.org/html/2607.00190#bib.bib34),[35](https://arxiv.org/html/2607.00190#bib.bib35)\]\. Work on modeling player skill from replays predates AlphaStar\.Avontuur et al\. \[[36](https://arxiv.org/html/2607.00190#bib.bib36)\]showed that even simple classifiers trained on APM and economy features can predict a player’s league with meaningful accuracy\. Subsequent work demonstrated that macro\-level economic measures, such as the Spending Quotient introduced byBowman et al\. \[[37](https://arxiv.org/html/2607.00190#bib.bib37)\], are among the strongest predictors of both skill and match outcomes\.

##### Variational Autoencoders and Disentangled Representations:

Naturally, the Variational Autoencoder \(VAE\)\[[38](https://arxiv.org/html/2607.00190#bib.bib38)\]acts as the backbone for our work\. By learning a probabilistic encoder and decoder jointly with a Kullback\-Leibler \(KL\) divergence regulariser, the VAE produces a smooth, continuous latent space from which new samples can be reconstructed\. As an extension,Higgins et al\. \[[39](https://arxiv.org/html/2607.00190#bib.bib39)\]introducedβ\\beta\-VAE, and addressed the interpretability by up\-weighting the KL term\. Theβ\\betafactor was applied to force the model to trade reconstruction fidelity for statistical independence between latent dimensions\. Guided VAE addresses this in a direct mannerDing et al\. \[[40](https://arxiv.org/html/2607.00190#bib.bib40)\]\. It attaches a supervised classifier to designated latent dimensions, and uses an adversarial excitation\-inhibition mechanism to concentrate the target factor in that dimension while preventing it from leaking into the remaining dimensions\.Schrum et al\. \[[41](https://arxiv.org/html/2607.00190#bib.bib41)\]presented “SAIL”, which learns persistent skill embeddings from naturalistic behavioral data using expert\-novice basis blending and counterfactual subskill swaps applied to motor tasks such as driving and baseball batting\. We draw inspiration and intuition from these works\.

##### Latent Space Traversal:

Generating semantically meaningful paths through a learned latent space is a non\-trivial problem\. Naive linear interpolation between two latent codes can pass through low\-density regions of the prior, leading to decoded samples that lie outside the data manifold\.Korkmaz et al\. \[[42](https://arxiv.org/html/2607.00190#bib.bib42)\]formalize this distribution mismatch and show that optimal transport \(OT\) maps can correct linear trajectories so that all intermediate points remain consistent with the prior distribution while minimally deviating from a straight line\.Song et al\. \[[43](https://arxiv.org/html/2607.00190#bib.bib43)\]propose a more general framework that models latent structures as learned dynamic potential landscapes, deriving traversal trajectories as the gradient flow of a partial differential equation \(PDE\)\-based potential field\.Yeh et al\. \[[44](https://arxiv.org/html/2607.00190#bib.bib44)\]take a related approach in the explainability domain with partial focus on StarCraft II by leveraging a fixed “SC2 Assault” scenario\. They generate counterfactuals from a jointly trained generative latent space where the traversal is guided toward a target outcome during decoding\.

##### Counterfactual Explanations as Actionable Feedback:

Counterfactual explanations answer the question: “*what is the minimal change to the input that would change the model’s prediction?*” In a performance\-improvement context, this is equivalent to algorithmic recourse\. Providing a ranked list of feature changes that move a player from their current state to a more desirable one\.Crupi et al\. \[[45](https://arxiv.org/html/2607.00190#bib.bib45)\]proposed “CEILS”, generating counterfactuals as interventions in the latent space of a trained VAE\[[45](https://arxiv.org/html/2607.00190#bib.bib45)\]\. Work beyond passive explanation towards actionable coaching was shown byBae et al\. \[[46](https://arxiv.org/html/2607.00190#bib.bib46)\], leveraging counterfactual explanations in racing scenarios with language\-based guidance\.Pegios et al\. \[[47](https://arxiv.org/html/2607.00190#bib.bib47)\]extend latent\-space counterfactual generation by equipping the VAE latent space with a Riemannian metric pulled back through both the decoder and the classifier\.

## 3Material and Methods

##### Replay Preprocessing and Feature Extraction:

To fulfill our goal of providing feedback based on learned representations, we have decided to use a dataset consisting of professional StarCraft II games named “SC2EGSet”\[[4](https://arxiv.org/html/2607.00190#bib.bib4)\]licensed under CC\-BY 4\.0\. Please refer to[AppendixA](https://arxiv.org/html/2607.00190#A1)for a simplified game description\. At the time our work was prepared, the dataset consisted of 23,476 files containing game\-state information sourced from 71 “replaypacks”\. Before training, each replay is converted into a tensor containing information about both players\. For every player, we extract a 196\-dimensional feature vectorℝ196\\mathbb\{R\}^\{196\}\. Selected features primarily include economical game progression statistics that inform various aspects of the game\. Additionally, for more expressive modeling insights, we split the array of player statistics into three windows, each spanning one\-third of the total game duration\. These features are named as “early”, “middle”, and “late game” windows \(3⋅393\\cdot 39features\)\. Tracker statistics averaged within each window\. Finally, we include the final economy state \(3939features\) and the economy difference features \(3939features\), computed as late\-game economy minus early\-game economy\. Besides the economy, we include a scalar value, “supply capped percent”, denoting the percentage of the game duration during which the player was unable to build additional units due to insufficient in\-game infrastructure\. The final replay tensor has shapeℝ2×196\\mathbb\{R\}^\{2\\times 196\}, where the first dimension corresponds to the player\. The label used for disentanglement is the match result from player 0’s perspective, represented as a binary outcomey∈0,1y\\in\{0,1\}, wherey=1y=1indicates that player 0 won\. Prior to training, all features are standardized using per\-feature z\-score normalization\. The normalization statistics \(mean and standard deviation\) are computed exclusively on the training set and subsequently applied to the validation and test sets, preventing any data leakage\. A small epsilon \(ϵ=10−8\\epsilon=10^\{\-8\}\) is added to each standard deviation to avoid division by zero for constant features\.

The initial split between training, validation, and test sets was random \(80%/10%/10%\)\. Replays with missing player statistics, player information, or undecided/draw outcomes were skipped\. Finally, to assure that the training, and validation sets were indeed \(80%/10%\), the missing samples were taken out of the test set without replacement\. After excluding samples, the splits consisted of the following numbers of samples in the training set \(nt​r​a​i​n​i​n​g=18780n\_\{training\}=18780\), the validation set \(nv​a​l​i​d​a​t​i​o​n=2347n\_\{validation\}=2347\), and the test set \(nt​e​s​t=2178n\_\{test\}=2178\)\. To confirm the efficacy of our method on out\-of\-distribution \(OOD\) data, we have randomly sampled an additional \(no​o​d=2178n\_\{ood\}=2178\) samples from an unreleased dataset of players who submitted their replays to the sc2replaystats\[[7](https://arxiv.org/html/2607.00190#bib.bib7)\]in 2016\-2020\. Access to all of the pre\-processed data and code is available\. For more information please see[AppendixG](https://arxiv.org/html/2607.00190#A7)\.

##### Modelling

To ensure the possibility of transitioning between learned latent space representations and reconstruction of the player features, we have decided to use a model adapted from the original Guided VAE\[[40](https://arxiv.org/html/2607.00190#bib.bib40)\]architecture\. We guide the latent space separation by the game outcome\. Therefore, a part of the latent representation is encouraged to contain information about features that are significant for predicting winning or losing outcomes\. The structure of our model consists of a symmetrical encoder\-decoder multilayer perceptron \(MLP\) network with ReLU activations, with an additional supervised classifier attached to the latent space\. In our modified Guided VAE\[[40](https://arxiv.org/html/2607.00190#bib.bib40)\], a selected number of dimensions are set to be supervised\. An adversarial classifier is trained on the remaining latent dimensions, excluding the supervised dimensions\. In training, the supervised classifier receives concatenated supervised latent dimensions of both players,\[z\(0\)​1:k\|z\(1\)​1:k\]\[z^\{\(0\)\}\{1:k\}\|z^\{\(1\)\}\{1:k\}\], and is trained to predictyydirectly\. At inference\-time, when generating improvement paths for the player in seat 1, the input order is swapped, and the output probability is complemented, so that the score always represents the win probability of the player whose path is being improved\. Please refer to[AppendixB](https://arxiv.org/html/2607.00190#A2)for more implementation details\. We use AdamW optimizers\[[48](https://arxiv.org/html/2607.00190#bib.bib48),[49](https://arxiv.org/html/2607.00190#bib.bib49)\], implemented in PyTorch\. The first updates the VAE and its guided classifier jointly\. The second trains an auxiliary adversarial classifier on the non\-guided \(free\) latent dimensions\. The third updates the VAE parameters adversarially, penalising the encoder for producing free\-dimension representations that are predictive of match outcome\. Validation is monitored using the VAE loss, and early stopping is used to avoid overfitting\. We conducted a hyperparameter search to find the best\-performing model, see[AppendixD](https://arxiv.org/html/2607.00190#A4)\. All of the experiments were run on consumer hardware, see[AppendixC](https://arxiv.org/html/2607.00190#A3)\.

##### Latent Space Traversal and Feedback Generation:

After training, a replay is encoded into a latent vector\. We construct a path from the current game state to a winning region in the supervised dimensions of the latent space\. Each point interpolated along the path in the latent space can be reconstructed back to the feature space, where the difference between the current feature values and the reconstructed feature values informs the player about the improvement “trajectory”\. We compute feedback in three ways\. First, we measure the raw change from the start of the path to the end\. Second, we compute a minimum viable change that stops at the first waypoint where the predicted win probability crosses 0\.5, if such a crossing exists\. Third, we compute a win\-probability\-weighted change, where each step’s feature change is weighted by the corresponding increase in predicted win probability; steps where win probability decreases contribute zero weight\. The final output is a ranked list of feature changes\. These changes act as rough suggestions, and should be interpreted as model\-generated hypotheses\. Example generated feedback reports are available in[AppendixF](https://arxiv.org/html/2607.00190#A6)\.

## 4Standardized Path Representation

To ensure compatibility with downstream generation and visualization tasks, the output of the path charting pipeline is strictly standardized, regardless of the specific employed strategy\. The generated path𝒫\\mathcal\{P\}is defined as an ordered sequence ofnndiscrete waypoints in thedd\-dimensional latent space\. This sequence is structured as a matrix𝐏∈ℝn×d\\mathbf\{P\}\\in\\mathbb\{R\}^\{n\\times d\}, where each row𝐩j∈ℝd\\mathbf\{p\}\_\{j\}\\in\\mathbb\{R\}^\{d\}corresponds to a specific waypoint along as seen in[Equation 1](https://arxiv.org/html/2607.00190#S4.E1)\.

𝐏=\[—𝐩0——𝐩1—⋮⋮⋮—𝐩n−1—\]=\[—𝐳s​t​a​r​t—⋮—𝐳e​n​d—\]\\mathbf\{P\}=\\begin\{bmatrix\}\\text\{\-\-\-\}&\\mathbf\{p\}\_\{0\}&\\text\{\-\-\-\}\\\\ \\text\{\-\-\-\}&\\mathbf\{p\}\_\{1\}&\\text\{\-\-\-\}\\\\ \\vdots&\\vdots&\\vdots\\\\ \\text\{\-\-\-\}&\\mathbf\{p\}\_\{n\-1\}&\\text\{\-\-\-\}\\end\{bmatrix\}=\\begin\{bmatrix\}\\text\{\-\-\-\}&\\mathbf\{z\}\_\{start\}&\\text\{\-\-\-\}\\\\ &\\vdots&\\\\ \\text\{\-\-\-\}&\\mathbf\{z\}\_\{end\}&\\text\{\-\-\-\}\\end\{bmatrix\}\(1\)By definition across all strategies, the first waypoint𝐩0\\mathbf\{p\}\_\{0\}corresponds exactly to the initial latent vector𝐳s​t​a​r​t\\mathbf\{z\}\_\{start\}\. The final waypoint𝐩n−1\\mathbf\{p\}\_\{n\-1\}corresponds to the terminal state of the strategy, denoted generally as𝐳e​n​d\\mathbf\{z\}\_\{end\}\(which may represent a predefined target𝐳t​a​r​g​e​t\\mathbf\{z\}\_\{target\}, a successfully converged optimization state, or the integrated endpoint of a velocity field\)\. Traversal methods have their own set of tunable parameters fully described in[AppendixD](https://arxiv.org/html/2607.00190#A4)\.

### 4\.1Linear Strategy

The linear strategy serves as the foundational baseline for counterfactual generation\. It implements a simple Linear Interpolation \(LERP\) to navigate the high\-dimensional latent space, constructing a straight\-line trajectory between the initial losing latent vector and a specified winning target\.Target Selection Strategies:Because the "winning" state is represented by a distribution of points rather than a single vector, the algorithm first collapses the winning latents𝒵w​i​n=\{𝐰1,…,𝐰M\}\\mathcal\{Z\}\_\{win\}=\\\{\\mathbf\{w\}\_\{1\},\\dots,\\mathbf\{w\}\_\{M\}\\\}into a single target vector𝐳t​a​r​g​e​t∈ℝd\\mathbf\{z\}\_\{target\}\\in\\mathbb\{R\}^\{d\}using one of two selectable methods\.Centroid Strategy \(method="centroid"\):The target is defined globally as the unweighted mean position \(centroid\) of all known winning latent vectors\. This provides a robust, generalized direction toward the center of the winning class, see[Equation 2](https://arxiv.org/html/2607.00190#S4.E2)\.Nearest Neighbors Strategy \(method="nearest"\):To preserve local manifold structure and find the "closest" way to win, the target is derived locally\. It calculates the mean of only thekkwinning latents that are closest to the starting sample𝐳s​t​a​r​t\\mathbf\{z\}\_\{start\}\(wherekkis defined byk\_neighbours\), see[Equation 3](https://arxiv.org/html/2607.00190#S4.E3)\.

𝐳t​a​r​g​e​t=1M​∑i=1M𝐰i\\mathbf\{z\}\_\{target\}=\\frac\{1\}\{M\}\\sum\_\{i=1\}^\{M\}\\mathbf\{w\}\_\{i\}\(2\)
𝐳t​a​r​g​e​t=1k​∑𝐰∈k\-NN​\(𝐳s​t​a​r​t,𝒵w​i​n\)𝐰\\mathbf\{z\}\_\{target\}=\\frac\{1\}\{k\}\\sum\_\{\\mathbf\{w\}\\in\\text\{k\-NN\}\(\\mathbf\{z\}\_\{start\},\\mathcal\{Z\}\_\{win\}\)\}\\mathbf\{w\}\(3\)

Linear Interpolation \(LERP\) Formula:Once the target𝐳t​a​r​g​e​t\\mathbf\{z\}\_\{target\}is established, the path is generated as a sequence ofnnwaypoints\{𝐩0,𝐩1,…,𝐩n−1\}\\\{\\mathbf\{p\}\_\{0\},\\mathbf\{p\}\_\{1\},\\dots,\\mathbf\{p\}\_\{n\-1\}\\\}\. The interpolation coefficientαj\\alpha\_\{j\}is defined by an evenly spaced linear progression from0\.00\.0to1\.01\.0, see[Equation 4](https://arxiv.org/html/2607.00190#S4.E4)\. Each intermediate point𝐩j\\mathbf\{p\}\_\{j\}along the trajectory is calculated using a convex combination of the start and target vectors, moving progressively closer to the target asα\\alphaincreases, see[Equation 5](https://arxiv.org/html/2607.00190#S4.E5)\. This straight\-line interpolation assumes a globally Euclidean latent space, transitioning semantic features at a constant velocity without explicit regard for the underlying data density\.

αj=jn−1,for​j∈\{0,1,…,n−1\}\\alpha\_\{j\}=\\frac\{j\}\{n\-1\},\\quad\\text\{for \}j\\in\\\{0,1,\\dots,n\-1\\\}\(4\)
𝐩j=\(1−αj\)​𝐳s​t​a​r​t\+αj​𝐳t​a​r​g​e​t\\mathbf\{p\}\_\{j\}=\(1\-\\alpha\_\{j\}\)\\mathbf\{z\}\_\{start\}\+\\alpha\_\{j\}\\mathbf\{z\}\_\{target\}\(5\)

### 4\.2Iterative Optimal Transport Strategy

The iterative optimal transport \(OT\) strategy simulates a vector field flow toward the winning distribution\. Instead of interpolating toward a single static target \(such as a centroid\), the algorithm dynamically recalculates a local barycentric target at each step by finding the optimal mass transport plan between the current position and the target distribution\.Distance Matrix and Transport Plan:At each steptt, the current latent position𝐳\(t\)\\mathbf\{z\}^\{\(t\)\}is treated as a point mass with weight𝐚=1\\mathbf\{a\}=1\. The set ofNNwinning latents𝒵w​i​n=\{𝐰1,…,𝐰N\}\\mathcal\{Z\}\_\{win\}=\\\{\\mathbf\{w\}\_\{1\},\\dots,\\mathbf\{w\}\_\{N\}\\\}is treated as a uniform target distribution with weights𝐛=1N​𝟏\\mathbf\{b\}=\\frac\{1\}\{N\}\\mathbf\{1\}\. First, the squared Euclidean distance matrix𝐌\(t\)∈ℝ1×N\\mathbf\{M\}^\{\(t\)\}\\in\\mathbb\{R\}^\{1\\times N\}is computed\. To prevent numerical underflow during exponentiation in the regularized transport step, the distance matrix is normalized by its maximum value as seen in[Equation 6](https://arxiv.org/html/2607.00190#S4.E6)\. Next, the optimal transport plan𝐓\(t\)\\mathbf\{T\}^\{\(t\)\}is computed\. If entropic regularizationλ\>0\\lambda\>0\(ot\_reg\) is provided, the mathematically stable log\-domain Sinkhorn algorithm is utilized as seen in[Equation 7](https://arxiv.org/html/2607.00190#S4.E7), whereH​\(𝐓\)H\(\\mathbf\{T\}\)is the entropy of the coupling matrix\. Ifλ=0\\lambda=0, the exact Earth Mover’s Distance \(EMD\) is computed\.Local Barycentric Target:The resulting transport plan𝐓\(t\)\\mathbf\{T\}^\{\(t\)\}dictates the optimal distribution of mass from the current position to the winning points\. These transport weights are normalized to form a localized probability distribution as seen in[Equation 8](https://arxiv.org/html/2607.00190#S4.E8)\. The local target𝐳t​a​r​g​e​t\(t\)\\mathbf\{z\}\_\{target\}^\{\(t\)\}is then defined as the barycenter \(weighted average\) of the winning points, pulled specifically according to the optimal transport plan presented as[Equation 9](https://arxiv.org/html/2607.00190#S4.E9)\.Euler Step \(Vector Flow\):Rather than jumping directly to the target, the algorithm treats the vector\(𝐳t​a​r​g​e​t\(t\)−𝐳\(t\)\)\(\\mathbf\{z\}\_\{target\}^\{\(t\)\}\-\\mathbf\{z\}^\{\(t\)\}\)as a local velocity field\. It takes a small Euler step of sizeη\\eta\(step\_size\) toward the local barycenter, as seen in[Equation 10](https://arxiv.org/html/2607.00190#S4.E10)\. This process is repeated iteratively to trace a smooth trajectory\. Because the transport plan dynamically updates at each spatial step, the resulting path closely mimics a continuous vector flow into the densest regions of the winning distribution\.

Mi\(t\)=‖𝐳\(t\)−𝐰i‖2maxj⁡‖𝐳\(t\)−𝐰j‖2\+ϵM\_\{i\}^\{\(t\)\}=\\frac\{\\\|\\mathbf\{z\}^\{\(t\)\}\-\\mathbf\{w\}\_\{i\}\\\|^\{2\}\}\{\\max\_\{j\}\\\|\\mathbf\{z\}^\{\(t\)\}\-\\mathbf\{w\}\_\{j\}\\\|^\{2\}\+\\epsilon\}\(6\)
𝐓\(t\)=arg⁡min𝐓∈Π​\(𝐚,𝐛\)⁡⟨𝐓,𝐌\(t\)⟩−λ​H​\(𝐓\)\\mathbf\{T\}^\{\(t\)\}=\\arg\\min\_\{\\mathbf\{T\}\\in\\Pi\(\\mathbf\{a\},\\mathbf\{b\}\)\}\\langle\\mathbf\{T\},\\mathbf\{M\}^\{\(t\)\}\\rangle\-\\lambda H\(\\mathbf\{T\}\)\(7\)

τi=Ti\(t\)∑j=1NTj\(t\)\+ϵ\\tau\_\{i\}=\\frac\{T\_\{i\}^\{\(t\)\}\}\{\\sum\_\{j=1\}^\{N\}T\_\{j\}^\{\(t\)\}\+\\epsilon\}\(8\)
𝐳t​a​r​g​e​t\(t\)=∑i=1Nτi​𝐰i\\mathbf\{z\}\_\{target\}^\{\(t\)\}=\\sum\_\{i=1\}^\{N\}\\tau\_\{i\}\\mathbf\{w\}\_\{i\}\(9\)
𝐳\(t\+1\)=𝐳\(t\)\+η​\(𝐳t​a​r​g​e​t\(t\)−𝐳\(t\)\)\\mathbf\{z\}^\{\(t\+1\)\}=\\mathbf\{z\}^\{\(t\)\}\+\\eta\\left\(\\mathbf\{z\}\_\{target\}^\{\(t\)\}\-\\mathbf\{z\}^\{\(t\)\}\\right\)\(10\)

### 4\.3Gradient Ascent Strategy

In the context of a Guided VAE, the gradient ascent strategy actively searches for a counterfactual path\. Starting from a starting latent representation, the goal is to discover the minimal feature changes required to transition into a winning state\. To ensure the generated counterfactuals remain realistic and do not exploit adversarial blind spots in the classifier, the trajectory is explicitly regularized by the data manifold\.Objective Formulation:The optimization seeks to iteratively adjust the latent vector𝐳\\mathbf\{z\}to maximize an opponent\-aware classification scoreS​\(𝐳\)S\(\\mathbf\{z\}\)\(e\.g\., the logit of winning against a specific opponent𝐳o​p​p\\mathbf\{z\}\_\{opp\}\), constrained by a manifold density penaltyD​\(𝐳\)D\(\\mathbf\{z\}\)\. The density is modeled using a fully differentiable Gaussian Kernel Density Estimate \(KDE\) evaluated over the reference dataset of known winning latents𝒵w​i​n=\{𝐰1,…,𝐰N\}\\mathcal\{Z\}\_\{win\}=\\\{\\mathbf\{w\}\_\{1\},\\dots,\\mathbf\{w\}\_\{N\}\\\}with bandwidthhh, as seen in[Equation 11](https://arxiv.org/html/2607.00190#S4.E11)\. The total combined gradient at stepttmerges the direction that increases the likelihood of winning with the direction that points toward denser, realistic regions of the latent space, as shown in[Equation 12](https://arxiv.org/html/2607.00190#S4.E12), whereλ\\lambdarepresents the weighting of the density prior \(density\_weight\)\.Momentum\-Based Optimization Update:To traverse the disentangled latent space smoothly and avoid local minima, the latent vector is updated using gradient ascent with momentum\. Letα\\alphabe the learning rate andβ\\betabe the momentum factor\. The velocity𝐯\\mathbf\{v\}and position𝐳\\mathbf\{z\}are updated as seen in[Equation 13](https://arxiv.org/html/2607.00190#S4.E13), and[Equation 14](https://arxiv.org/html/2607.00190#S4.E14)\.Convergence and Resampling:This iterative process continues until the predicted probability of the winning class exceeds a specifiedconvergence\_thresholdτ\\tauas in[Equation 15](https://arxiv.org/html/2607.00190#S4.E15)\. Because the number of optimization stepsTTrequired to reach this threshold is variable, the resulting sequence\{𝐳\(0\),𝐳\(1\),…,𝐳\(T\)\}\\\{\\mathbf\{z\}^\{\(0\)\},\\mathbf\{z\}^\{\(1\)\},\\dots,\\mathbf\{z\}^\{\(T\)\}\\\}is evenly resampled to extract exactlynnwaypoints\. This final trajectory is then output as the standardized path matrix𝐏∈ℝn×d\\mathbf\{P\}\\in\\mathbb\{R\}^\{n\\times d\}, representing a smooth, realistic counterfactual feature transition\.

D​\(𝐳\)=log​∑i=1Nexp⁡\(−12​h2​‖𝐳−𝐰i‖2\)D\(\\mathbf\{z\}\)=\\log\\sum\_\{i=1\}^\{N\}\\exp\\left\(\-\\frac\{1\}\{2h^\{2\}\}\\left\\\|\\mathbf\{z\}\-\\mathbf\{w\}\_\{i\}\\right\\\|^\{2\}\\right\)\(11\)
𝐠t​o​t​a​l\(t\)=∇𝐳S​\(𝐳\(t\)\)\+λ​∇𝐳D​\(𝐳\(t\)\)\\mathbf\{g\}^\{\(t\)\}\_\{total\}=\\nabla\_\{\\mathbf\{z\}\}S\(\\mathbf\{z\}^\{\(t\)\}\)\+\\lambda\\nabla\_\{\\mathbf\{z\}\}D\(\\mathbf\{z\}^\{\(t\)\}\)\(12\)

𝐯\(t\+1\)=β​𝐯\(t\)\+α​𝐠t​o​t​a​l\(t\)\\mathbf\{v\}^\{\(t\+1\)\}=\\beta\\mathbf\{v\}^\{\(t\)\}\+\\alpha\\mathbf\{g\}^\{\(t\)\}\_\{total\}\(13\)
𝐳\(t\+1\)=𝐳\(t\)\+𝐯\(t\+1\)\\mathbf\{z\}^\{\(t\+1\)\}=\\mathbf\{z\}^\{\(t\)\}\+\\mathbf\{v\}^\{\(t\+1\)\}\(14\)
P​\(win∣𝐳\(t\),𝐳o​p​p\)≥τP\(\\text\{win\}\\mid\\mathbf\{z\}^\{\(t\)\},\\mathbf\{z\}\_\{opp\}\)\\geq\\tau\(15\)

### 4\.4Neural Flow Strategy

The neural flow strategy utilizes continuous normalizing flows via an Optimal Transport \(OT\) Flow Matching framework\. Rather than relying on simple geometric interpolations or local gradient steps, this approach trains a neural network to learn a global velocity field\. This field models the continuous optimal transport of probability mass from the "losing" to the "winning" latent distribution\.Velocity Field Training and OT Pairing:A Multi\-Layer Perceptron \(MLP\) acts as a time\-conditioned velocity field𝐯θ​\(𝐳,t\)\\mathbf\{v\}\_\{\\theta\}\(\\mathbf\{z\},t\), parameterized by weightsθ\\theta\. During training, mini\-batches of losing latents𝐙0\\mathbf\{Z\}\_\{0\}and winning latents𝐙1\\mathbf\{Z\}\_\{1\}are extracted\. To ensure the network learns the most efficient, non\-crossing paths between these distributions, the samples are dynamically paired using exact Earth Mover’s Distance \(EMD\) based on squared Euclidean distance\. For each𝐳0\\mathbf\{z\}\_\{0\}, an optimal𝐳1,paired\\mathbf\{z\}\_\{1,\\text\{paired\}\}is identified\. At a uniformly sampled timet∈\[0,1\]t\\in\[0,1\], the intermediate state is defined by linear interpolation as shown in[Equation 16](https://arxiv.org/html/2607.00190#S4.E16)\. The network is then trained to predict the constant\-velocity vector between paired samples by minimizing the Mean Squared Error \(MSE\), as shown in[Equation 17](https://arxiv.org/html/2607.00190#S4.E17)\.Trajectory Integration \(Euler Method\):To generate a counterfactual path during inference, a starting \(losing\) latent𝐳s​t​a​r​t\\mathbf\{z\}\_\{start\}is integrated through the learned velocity field fromt=0t=0totm​a​x=1\.0t\_\{max\}=1\.0\. Using a discrete number of integration stepsNs​t​e​p​sN\_\{steps\}, the time increment isΔ​t=1\.0Ns​t​e​p​s\\Delta t=\\frac\{1\.0\}\{N\_\{steps\}\}\. The latent position is updated iteratively using the Euler method seen in[Equation 18](https://arxiv.org/html/2607.00190#S4.E18), where the initial condition is𝐳\(0\)=𝐳s​t​a​r​t\\mathbf\{z\}^\{\(0\)\}=\\mathbf\{z\}\_\{start\}\. This produces a smooth flow along the learned data manifold\.Classifier Guidance \(Optional\):To explicitly steer the trajectory toward regions with a higher probability of winning against a specific opponent, optional classifier guidance can be injected into the Euler integration\. At each step, the gradient of the probability scoreS​\(𝐳\)S\(\\mathbf\{z\}\)is computed\. To ensure the guidance scale remains a consistent fraction of the flow step size, regardless of the raw gradient’s magnitude, the gradient is normalized to a unit vector\. The position is updated as shown in[Equation 19](https://arxiv.org/html/2607.00190#S4.E19), whereγ\\gamma\(guidance\_scale\) dictates how strongly the path is pulled toward the classifier’s optimal regions, andϵ\\epsilonprevents division by zero\.

𝐳t=\(1−t\)​𝐳0\+t​𝐳1,paired\\mathbf\{z\}\_\{t\}=\(1\-t\)\\mathbf\{z\}\_\{0\}\+t\\mathbf\{z\}\_\{1,\\text\{paired\}\}\(16\)
ℒ​\(θ\)=𝔼t,𝐳0,𝐳1​\[‖𝐯θ​\(𝐳t,t\)−\(𝐳1,paired−𝐳0\)‖22\]\\mathcal\{L\}\(\\theta\)=\\mathbb\{E\}\_\{t,\\mathbf\{z\}\_\{0\},\\mathbf\{z\}\_\{1\}\}\\left\[\\left\\\|\\mathbf\{v\}\_\{\\theta\}\(\\mathbf\{z\}\_\{t\},t\)\-\(\\mathbf\{z\}\_\{1,\\text\{paired\}\}\-\\mathbf\{z\}\_\{0\}\)\\right\\\|\_\{2\}^\{2\}\\right\]\(17\)

𝐳\(t\+1\)=𝐳\(t\)\+𝐯θ​\(𝐳\(t\),t\)​Δ​t\\mathbf\{z\}^\{\(t\+1\)\}=\\mathbf\{z\}^\{\(t\)\}\+\\mathbf\{v\}\_\{\\theta\}\(\\mathbf\{z\}^\{\(t\)\},t\)\\Delta t\(18\)
𝐳\(t\+1\)=𝐳\(t\)\+Δ​t​\(𝐯θ​\(𝐳\(t\),t\)\+γ​∇𝐳S​\(𝐳\(t\)\)‖∇𝐳S​\(𝐳\(t\)\)‖2\+ϵ\)\\mathbf\{z\}^\{\(t\+1\)\}=\\mathbf\{z\}^\{\(t\)\}\+\\Delta t\\left\(\\mathbf\{v\}\_\{\\theta\}\(\\mathbf\{z\}^\{\(t\)\},t\)\+\\gamma\\frac\{\\nabla\_\{\\mathbf\{z\}\}S\(\\mathbf\{z\}^\{\(t\)\}\)\}\{\\\|\\nabla\_\{\\mathbf\{z\}\}S\(\\mathbf\{z\}^\{\(t\)\}\)\\\|\_\{2\}\+\\epsilon\}\\right\)\(19\)

## 5Experiments

Model Performance:To evaluate the quality of the trained Guided VAE, we assess both its reconstruction capabilities and the effectiveness of the latent space separation\. The model’s performance on the held\-out test set is summarized using several key metrics as seen in[Table 1](https://arxiv.org/html/2607.00190#S5.T1)\.

Table 1:GuidedVAE Test\-Set EvaluationEvaluation of the GuidedVAE demonstrates strong generative and reconstruction fidelity, with the MSE and KL divergence confirming accurate game\-state reconstruction from a well\-regularised latent space\. Furthermore, evaluation of the predictive guidance imposed on the first latent dimension—measured via test accuracy, ROC\-AUC, and Brier score—indicates robust classification performance and highly calibrated win probabilities\. Overall, the model successfully balances precise feature reconstruction with meaningful latent disentanglement, establishing a reliable foundation for generating counterfactual improvement trajectories\.Conterfactual Paths:To evaluate the performance of the final model against our main goal of latent space traversal generating counterfactual “improvement trajectory”, we conduct a comparative assessment of all path generation strategies in[Table 2](https://arxiv.org/html/2607.00190#S5.T2)\. Where the success rate is the fraction of samples for which the path reachesP​\(win\)≥0\.5P\(\\text\{win\}\)\\geq 0\.5at any waypoint along the path\. Crossoverα\\alphais the position along the path at whichP​\(win\)P\(\\text\{win\}\)first crosses 0\.5, whereα=0\\alpha=0is the start andα=1\\alpha=1is the end\. Reported only for successful runs,𝚫\\bm\{\\Delta\}P\(win\) is the absolute gain in predicted win probability from path start to path end, i\.e\.P​\(win\)end−P​\(win\)startP\(\\text\{win\}\)\_\{\\text\{end\}\}\-P\(\\text\{win\}\)\_\{\\text\{start\}\}\., AUC is the area under theP​\(win\)P\(\\text\{win\}\)curve overα∈\[0,1\]\\alpha\\in\[0,1\]\. Monotonicity is the fraction of consecutive waypoint pairs for whichP​\(win\)P\(\\text\{win\}\)is non\-decreasing\. A value of 1\.0 meansP​\(win\)P\(\\text\{win\}\)increases or stays flat at every step\. Lower values indicate oscillation or regression along the path\. Finally, the nearest\-win distance is the Euclidean distance in the supervised latent subspace between the path endpoint and the closest winning latent vector in the training set\. While Gradient Ascent achieves a nominally perfect success rate \(1\.0001\.000\), further inspection of the secondary metrics suggests this performance is largely driven by adversarial off\-manifold drift\. Compared to geometrically grounded strategies such as Optimal Transport, Gradient Ascent tends to explore low\-density regions of the latent space\. This is directly evidenced by a heavily inflated maximum latent norm \(max⁡‖𝐳‖=3\.95±1\.74\\max\\\|\\mathbf\{z\}\\\|=3\.95\\pm 1\.74, compared to just2\.03±0\.582\.03\\pm 0\.58for Optimal Transport\) and a severely degraded path KDE density \(−4\.63±1\.40\-4\.63\\pm 1\.40vs\.−2\.80±0\.50\-2\.80\\pm 0\.50\)\. Furthermore, its significantly reduced monotonicity \(0\.727±0\.2270\.727\\pm 0\.227vs\.0\.998±0\.0410\.998\\pm 0\.041\) and higher nearest\-win distance \(0\.15±0\.120\.15\\pm 0\.12vs\.0\.06±0\.040\.06\\pm 0\.04\) indicate erratic, unconstrained traversal rather than smooth semantic interpolation\. Consequently, the representations generated by this strategy are likely to exploit classifier blind spots and warrant much closer examination, see[AppendixE](https://arxiv.org/html/2607.00190#A5)\. We hypothesize that these failure modes could be addressed in future work by imposing stricter regularization constraints to firmly anchor the trajectory to the learned data prior\.

Table 2:Cross\-Dataset Comparison of Path\-Charting Strategies
The experimental results reveal several key insights regarding the trade\-offs between path reliability and quality:\(1\)Reliability and Success Rates:Gradient Ascent emerges as the strategy maintaining a near\-perfect success rate on both theSC2EGSet\(1\.000\) andOOD Data\(0\.998\)\. In contrast, the Linear \(k\-NN\) baseline exhibits significant fragility under distributional shift, with success rates dropping by over 30% \(Δ=−0\.302\\Delta=\-0\.302\)\.\(2\)Path Efficiency and Crossover:Optimal Transport \(OT\) demonstrates superior efficiency in trajectory charting\. As shown by theCrossoverα\\alpha\(0\.120±0\.0650\.120\\pm 0\.065\), OT\-generated paths transition to a winning state much earlier than Linear methods \(α≈0\.65\\alpha\\approx 0\.65\)\. Furthermore, OT achieves the highestAUC\(0\.752±0\.3010\.752\\pm 0\.301\), suggesting it identifies more direct routes through the latent space\.\(3\)The Success\-Monotonicity Trade\-off :A clear divergence exists between raw success and path smoothness\. While Gradient Ascent is the most successful, it records the lowestMonotonicity\(0\.727±0\.2270\.727\\pm 0\.227\), indicating more erratic trajectories\. Conversely, Neural Flow and Optimal Transport maintain near\-perfect monotonicity \(\>0\.99\>0\.99\) even on OOD data, providing highly stable and interpretable transitions\.The substantial variance inΔ​P​\(win\)\\Delta P\(\\text\{win\}\)across the linear baselines further suggests that simple interpolation is insufficient to capture the model’s complex decision boundaries\. In contrast, neural and transport\-based methods provide more consistent counterfactual evidence, with a meanP​\(win\)P\(\\text\{win\}\)along the counterfactual path shown in[Fig\. 1](https://arxiv.org/html/2607.00190#S5.F1), and directly showcase some of the aforementioned trade\-offs on OOD data\. Further model interpretability is covered in[AppendixE](https://arxiv.org/html/2607.00190#A5)\.

![Refer to caption](https://arxiv.org/html/2607.00190v1/x1.png)Figure 1:Mean OOD dataP​\(win\)P\(\\text\{win\}\)performance of generated paths progress in the latent space\.
## 6Limitations and Future Research

Limitations:Despite the promising results, our approach has several limitations that should be acknowledged\. First, we train our model on a dataset of games spanning multiple years\. We do not explicitly account for the game updates\. Additionally, we do not encode the players’ in\-game race information in any way\. By design, our model is incapable of providing feedback directed towards specific in\-game actions and environment configurations\. The model feedback is additionally constrained by the dataset choice; tournament gameplay samples can be seen as a very specific subset of all of the games\. Our model does not directly convey a more granular approach of jumping between leagues when leveraging its feedback\. Finally, the model and method parameters were not verified with human participants, and aside from expert input from known professional players, we were unable to set up a human\-in\-the\-loop type of experiment\.Future Research:We hope that by extending research efforts in representational learning geared towards providing feedback, we can inspire others to prepare end\-to\-end feedback generation\. Automating ways to improve humans based on deep generative solutions and other computational means\. In the future, we wish to address most of the concerns raised above\. Moving towards optimizing human performance jointly with actions against an environment sounds incredibly exciting, with the potential to uncover environment configurations that promote positive training or learning outcomes\. Given the recent advancements and rapid adoption of AI systems, creating models that provide data\-driven, actionable feedback is crucial\.

## 7Conclusion and Summary

We have demonstrated the computational feasibility and the potential for developing many practical solutions for providing feedback based on learned representations\. Additionally, simplistic methods, while appealing, have drawbacks that become more evident when dealing with a more advanced nonlinear model\. Finally, we have accomplished our goal of bridging the gap between representational learning and generating feedback by extracting actionable, counterfactual improvement trajectories from the latent space, effectively shifting the analytical paradigm from predicting game outcomes to providing players with tangible guidance on what they should do differently\. Based on our ongoing discussions with esports professionals, our solution is proving to be a real asset for future feedback systems\.

## References

- Vinyals et al\. \[2017\]O\. Vinyals, T\. Ewalds, S\. Bartunov, P\. Georgiev, A\. S\. Vezhnevets, M\. Yeo, A\. Makhzani, H\. Küttler, J\. Agapiou, J\. Schrittwieser, J\. Quan, S\. Gaffney, S\. Petersen, K\. Simonyan, T\. Schaul, H\. van Hasselt, D\. Silver, T\. Lillicrap, K\. Calderone, P\. Keet, A\. Brunasso, D\. Lawrence, A\. Ekermo, J\. Repp, and R\. Tsing, “Starcraft ii: A new challenge for reinforcement learning,” 2017\. \[Online\]\. Available:[https://arxiv\.org/abs/1708\.04782](https://arxiv.org/abs/1708.04782)
- Mathieu et al\. \[2023\]M\. Mathieu, S\. Ozair, S\. Srinivasan, C\. Gulcehre, S\. Zhang, R\. Jiang, T\. L\. Paine, R\. Powell, K\. Żołna, J\. Schrittwieser, D\. Choi, P\. Georgiev, D\. Toyama, A\. Huang, R\. Ring, I\. Babuschkin, T\. Ewalds, M\. Bordbar, S\. Henderson, S\. G\. Colmenarejo, A\. van den Oord, W\. M\. Czarnecki, N\. de Freitas, and O\. Vinyals, “Alphastar unplugged: Large\-scale offline reinforcement learning,” 2023\. \[Online\]\. Available:[https://arxiv\.org/abs/2308\.03526](https://arxiv.org/abs/2308.03526)
- Thompson et al\. \[2013\]J\. J\. Thompson, M\. R\. Blair, L\. Chen, and A\. J\. Henrey, “Video game telemetry as a critical tool in the study of complex skill learning,”*PLOS ONE*, vol\. 8, no\. 9, pp\. 1–12, 09 2013\. \[Online\]\. Available:[https://doi\.org/10\.1371/journal\.pone\.0075129](https://doi.org/10.1371/journal.pone.0075129)
- Białecki et al\. \[2023\]A\. Białecki, N\. Jakubowska, P\. Dobrowolski, P\. Białecki, L\. Krupiński, A\. Szczap, R\. Białecki, and J\. Gajewski, “Sc2egset: Starcraft ii esport replay and game\-state dataset,”*Scientific Data*, vol\. 10, no\. 1, p\. 600, Sep 2023\. \[Online\]\. Available:[https://doi\.org/10\.1038/s41597\-023\-02510\-7](https://doi.org/10.1038/s41597-023-02510-7)
- Ferenczi et al\. \[2024\]B\. Ferenczi, R\. Newbury, M\. Burke, and T\. Drummond, “Carefully structured compression: Efficiently managing starcraft ii data,” 2024\. \[Online\]\. Available:[https://arxiv\.org/abs/2410\.08659](https://arxiv.org/abs/2410.08659)
- Rijnders et al\. \[2022\]F\. Rijnders, G\. Wallner, and R\. Bernhaupt, “Live feedback for training through real\-time data visualizations: A study with league of legends,”*Proc\. ACM Hum\.\-Comput\. Interact\.*, vol\. 6, no\. CHI PLAY, oct 2022\. \[Online\]\. Available:[https://doi\.org/10\.1145/3549506](https://doi.org/10.1145/3549506)
- Martin \[2012\]A\. Martin, “sc2replaystats,”[https://sc2replaystats\.com/](https://sc2replaystats.com/), 2012, acessed: 2026\.04\.28\.
- Dibbell \[2026\]B\. Dibbell, “REPLAYMAN — SC2 Replay Analysis & Management – replayman\.com,”[https://replayman\.com/](https://replayman.com/), 2026, \[Accessed 28\-04\-2026\]\.
- Charleer et al\. \[2018\]S\. Charleer, K\. Gerling, F\. Gutiérrez, H\. Cauwenbergh, B\. Luycx, and K\. Verbert, “Real\-time dashboards to support esports spectating,” in*Proceedings of the 2018 Annual Symposium on Computer\-Human Interaction in Play*, ser\. CHI PLAY ’18\. New York, NY, USA: Association for Computing Machinery, 2018, pp\. 59–71\. \[Online\]\. Available:[https://doi\.org/10\.1145/3242671\.3242680](https://doi.org/10.1145/3242671.3242680)
- Wallner and Kriglstein \[2016\]G\. Wallner and S\. Kriglstein, “Visualizations for retrospective analysis of battles in team\-based combat games: A user study,” in*Proceedings of the 2016 Annual Symposium on Computer\-Human Interaction in Play*, ser\. CHI PLAY ’16\. New York, NY, USA: Association for Computing Machinery, 2016, pp\. 22–32\. \[Online\]\. Available:[https://doi\.org/10\.1145/2967934\.2968093](https://doi.org/10.1145/2967934.2968093)
- Xenopoulos et al\. \[2022\]P\. Xenopoulos, J\. a\. Rulff, and C\. Silva, “ggviz: Accelerating large\-scale esports game analysis,”*Proc\. ACM Hum\.\-Comput\. Interact\.*, vol\. 6, no\. CHI PLAY, oct 2022\. \[Online\]\. Available:[https://doi\.org/10\.1145/3549501](https://doi.org/10.1145/3549501)
- Shaker et al\. \[2016\]N\. Shaker, J\. Togelius, and M\. J\. Nelson,*Procedural Content Generation in Games*\. Springer International Publishing, 2016\. \[Online\]\. Available:[http://dx\.doi\.org/10\.1007/978\-3\-319\-42716\-4](http://dx.doi.org/10.1007/978-3-319-42716-4)
- Wei et al\. \[2025\]W\. Wei, S\. Yang, Q\. Zhou, R\. Liu, X\. Zhang, Y\. Yuan, Y\. Jiang, Y\. Luo, H\. Wang, T\. Wang, P\. Jin, W\. Liu, Z\. Zhao, X\. Jin, and E\. S\. Liu, “F\.a\.c\.u\.l\.: Language\-based interaction with ai companions in gaming,” 2025\. \[Online\]\. Available:[https://arxiv\.org/abs/2511\.13112](https://arxiv.org/abs/2511.13112)
- Sestini et al\. \[2025\]A\. Sestini, J\. Bergdahl, J\.\-P\. Barrette\-LaPierre, F\. Fuchs, B\. Chen, M\. Jones, and L\. Gisslén, “Human\-like goalkeeping in a realistic football simulation: a sample\-efficient reinforcement learning approach,” 2025\. \[Online\]\. Available:[https://arxiv\.org/abs/2510\.23216](https://arxiv.org/abs/2510.23216)
- Tufano et al\. \[2022\]R\. Tufano, S\. Scalabrino, L\. Pascarella, E\. Aghajani, R\. Oliveto, and G\. Bavota, “Using reinforcement learning for load testing of video games,” in*Proceedings of the 44th International Conference on Software Engineering*, ser\. ICSE ’22\. New York, NY, USA: Association for Computing Machinery, 2022, pp\. 2303–2314\. \[Online\]\. Available:[https://doi\.org/10\.1145/3510003\.3510625](https://doi.org/10.1145/3510003.3510625)
- Kokkinakis et al\. \[2020\]A\. V\. Kokkinakis, S\. Demediuk, I\. Nölle, O\. Olarewaju, S\. Patra, J\. Robertson, P\. York, A\. P\. Pedrassoli Chitayat, A\. Coates, D\. Slawson, P\. Hughes, N\. Hardie, B\. Kirman, J\. Hook, A\. Drachen, M\. F\. Ursu, and F\. Block, “Dax: Data\-driven audience experiences in esports,” in*Proceedings of the 2020 ACM International Conference on Interactive Media Experiences*, ser\. IMX ’20\. New York, NY, USA: Association for Computing Machinery, 2020, pp\. 94–105\. \[Online\]\. Available:[https://doi\.org/10\.1145/3391614\.3393659](https://doi.org/10.1145/3391614.3393659)
- The Newton Contributors \[2025\]The Newton Contributors, “Newton: GPU\-accelerated physics simulation for robotics and simulation research,” apr 2025\. \[Online\]\. Available:[https://github\.com/newton\-physics/newton](https://github.com/newton-physics/newton)
- Mittal et al\. \[2025\]M\. Mittal, P\. Roth, J\. Tigue, A\. Richard, O\. Zhang, P\. Du, A\. Serrano\-Muñoz, X\. Yao, R\. Zurbrügg, N\. Rudin, L\. Wawrzyniak, M\. Rakhsha, A\. Denzler, E\. Heiden, A\. Borovicka, O\. Ahmed, I\. Akinola, A\. Anwar, M\. T\. Carlson, J\. Y\. Feng, A\. Garg, R\. Gasoto, L\. Gulich, Y\. Guo, M\. Gussert, A\. Hansen, M\. Kulkarni, C\. Li, W\. Liu, V\. Makoviychuk, G\. Malczyk, H\. Mazhar, M\. Moghani, A\. Murali, M\. Noseworthy, A\. Poddubny, N\. Ratliff, W\. Rehberg, C\. Schwarke, R\. Singh, J\. L\. Smith, B\. Tang, R\. Thaker, M\. Trepte, K\. Van Wyk, F\. Yu, A\. Millane, V\. Ramasamy, R\. Steiner, S\. Subramanian, C\. Volk, C\. Chen, N\. Jawale, A\. V\. Kuruttukulam, M\. A\. Lin, A\. Mandlekar, K\. Patzwaldt, J\. Welsh, J\.\-F\. Lafleche, N\. Moënne\-Loccoz, S\. Park, R\. Stepinski, D\. Van Gelder, C\. Amevor, J\. Carius, J\. Chang, A\. He Chen, P\. d\. H\. Ciechomski, G\. Daviet, M\. Mohajerani, J\. von Muralt, V\. Reutskyy, M\. Sauter, S\. Schirm, E\. L\. Shi, P\. Terdiman, K\. Vilella, T\. Widmer, G\. Yeoman, T\. Chen, S\. Grizan, C\. Li, L\. Li, C\. Smith, R\. Wiltz, K\. Alexis, Y\. Chang, L\. J\. Fan, F\. Farshidian, A\. Handa, S\. Huang, M\. Hutter, Y\. Narang, S\. Pouya, S\. Sheng, Y\. Zhu, M\. Macklin, A\. Moravanszky, P\. Reist, Y\. Guo, D\. Hoeller, and G\. State, “Isaac Lab \- A GPU\-Accelerated Simulation Framework for Multi\-Modal Robot Learning,”*arXiv preprint arXiv:2511\.04831*, 2025\. \[Online\]\. Available:[https://arxiv\.org/abs/2511\.04831](https://arxiv.org/abs/2511.04831)
- Kaufmann et al\. \[2023\]E\. Kaufmann, L\. Bauersfeld, A\. Loquercio, M\. Müller, V\. Koltun, and D\. Scaramuzza, “Champion\-level drone racing using deep reinforcement learning,”*Nature*, vol\. 620, no\. 7976, pp\. 982–987, Aug 2023\. \[Online\]\. Available:[https://doi\.org/10\.1038/s41586\-023\-06419\-4](https://doi.org/10.1038/s41586-023-06419-4)
- Lamberti et al\. \[2024\]L\. Lamberti, E\. Cereda, G\. Abbate, L\. Bellone, V\. J\. K\. Morinigo, M\. Barciś, A\. Barciś, A\. Giusti, F\. Conti, and D\. Palossi, “A sim\-to\-real deep learning\-based framework for autonomous nano\-drone racing,”*IEEE Robotics and Automation Letters*, vol\. 9, no\. 2, pp\. 1899–1906, 2024\.
- Ma et al\. \[2025\]Y\. Ma, A\. Cramariuc, F\. Farshidian, and M\. Hutter, “Learning coordinated badminton skills for legged manipulators,”*Science Robotics*, vol\. 10, no\. 102, may 2025\. \[Online\]\. Available:[http://dx\.doi\.org/10\.1126/scirobotics\.adu3922](http://dx.doi.org/10.1126/scirobotics.adu3922)
- Liu et al\. \[2026\]C\. Liu, L\. Jiang, Y\. Wang, K\. Yao, J\. Fu, and X\. Ren, “Humanoid whole\-body badminton via multi\-stage reinforcement learning,” 2026\. \[Online\]\. Available:[https://arxiv\.org/abs/2511\.11218](https://arxiv.org/abs/2511.11218)
- Dürr et al\. \[2026\]P\. Dürr, M\. El Gheche, G\. J\. Maeda, N\. Mukai, N\. Takahashi, S\. Heusser, H\. Sahloul, Y\. Saraiji, P\. Adodin, Y\. Bi, S\. Blakeman, C\. Conti, D\. Fuentes Hitos, Y\. Hu, F\. Khadivar, R\. Kreiser, L\. Martinez, F\. Schilling, R\. Tapiador Morales, G\. Torrente, M\. Ynocente Castro, L\. Abecassis, A\. Giammarino, Y\.\-T\. Huang, Y\. Nagel, A\. Scotti, A\. Sigrist, T\. Silva, E\. Walther, J\. Wong, B\. Yang, A\. Aydin, D\. Grover, A\. Saha, V\. Cavinato, T\. Kakinuma, T\. Kunori, V\. Monferrato, S\. Richter, S\. Charalambous, S\. Guist, M\. A\. Kuhlmann\-Jorgensen, L\. Miele, A\. Politis, M\. Scardecchia, H\. Kitano, P\. R\. Wurman, P\. Stone, and M\. Spranger, “Outplaying elite table tennis players with an autonomous robot,”*Nature*, vol\. 652, no\. 8111, pp\. 886–891, Apr 2026\. \[Online\]\. Available:[https://doi\.org/10\.1038/s41586\-026\-10338\-5](https://doi.org/10.1038/s41586-026-10338-5)
- Leite et al\. \[2025\]I\. Leite, W\. Ahlberg, A\. Pereira, A\. Sestini, L\. Gisslén, and K\. Tollmar, “A call for deeper collaboration between robotics and game development,” in*2025 IEEE Conference on Games \(CoG\)*, 2025, pp\. 1–8\.
- Hancock et al\. \[2011\]D\. J\. Hancock, A\. M\. Rymal, and D\. M\. Ste\-Marie, “A triadic comparison of the use of observational learning amongst team sport athletes, coaches, and officials,”*Psychology of Sport and Exercise*, vol\. 12, no\. 3, pp\. 236–241, 2011\. \[Online\]\. Available:[https://doi\.org/10\.1016/j\.psychsport\.2010\.11\.002](https://doi.org/10.1016/j.psychsport.2010.11.002)
- Sozański et al\. \[2015\]H\. Sozański, J\. Sadowski, and J\. Czerwiński,*Podstawy Teorii i Technologii Treningu Sportowego*\. Akademia Wychowania Fizycznego Józefa Piłsudskiego Filia w Białej Podlaskiej, 2015, vol\. 2\.
- McIlroy\-Young et al\. \[2020\]R\. McIlroy\-Young, S\. Sen, J\. Kleinberg, and A\. Anderson, “Aligning superhuman ai with human behavior: Chess as a model system,” in*Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining*, ser\. KDD ’20\. New York, NY, USA: Association for Computing Machinery, 2020, pp\. 1677–1687\. \[Online\]\. Available:[https://doi\.org/10\.1145/3394486\.3403219](https://doi.org/10.1145/3394486.3403219)
- Gaessler and Piezunka \[2023\]F\. Gaessler and H\. Piezunka, “Training with ai: Evidence from chess computers,”*Strategic Management Journal*, vol\. 44, no\. 11, pp\. 2724–2750, 2023\. \[Online\]\. Available:[https://doi\.org/10\.1002/smj\.3512](https://doi.org/10.1002/smj.3512)
- Bilalić et al\. \[2026\]M\. Bilalić, M\. Graf, and N\. Vaci, “Computers and chess masters: The role of ai in transforming elite human performance,”*British Journal of Psychology*, vol\. 117, no\. 2, pp\. 585–609, 2026\. \[Online\]\. Available:[https://doi\.org/10\.1111/bjop\.12750](https://doi.org/10.1111/bjop.12750)
- Sadler and Regan \[2019\]M\. Sadler and N\. Regan,*Game Changer: AlphaZero’s Groundbreaking Chess Strategies and the Promise of AI*\. Alkmaar, Netherlands: New In Chess, 2019\.
- Kang et al\. \[2022\]J\. Kang, J\. S\. Yoon, and B\. Lee, “How ai\-based training affected the performance of professional go players,” in*Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems*, ser\. CHI ’22\. New York, NY, USA: Association for Computing Machinery, 2022\. \[Online\]\. Available:[https://doi\.org/10\.1145/3491102\.3517540](https://doi.org/10.1145/3491102.3517540)
- Shin et al\. \[2021\]M\. Shin, J\. Kim, and M\. Kim, “Human learning from artificial intelligence: Evidence from human go players’ decisions after alphago,” in*CogSci 2021 \- The 43rd Annual Meeting of the Cognitive Science Society*, 07 2021\. \[Online\]\. Available:[https://doi\.org/10\.5281/zenodo\.5095146](https://doi.org/10.5281/zenodo.5095146)
- Vinyals et al\. \[2019\]O\. Vinyals, I\. Babuschkin, W\. M\. Czarnecki, M\. Mathieu, A\. Dudzik, J\. Chung, D\. H\. Choi, R\. Powell, T\. Ewalds, P\. Georgiev*et al\.*, “Grandmaster level in StarCraft II using multi\-agent reinforcement learning,”*Nature*, vol\. 575, no\. 7782, pp\. 350–354, 2019\.
- Samvelyan et al\. \[2019\]M\. Samvelyan, T\. Rashid, C\. S\. de Witt, G\. Farquhar, N\. Nardelli, T\. G\. J\. Rudner, C\.\-M\. Hung, P\. H\. S\. Torr, J\. Foerster, and S\. Whiteson, “The starcraft multi\-agent challenge,” 2019\. \[Online\]\. Available:[https://arxiv\.org/abs/1902\.04043](https://arxiv.org/abs/1902.04043)
- Ellis et al\. \[2023\]B\. Ellis, J\. Cook, S\. Moalla, M\. Samvelyan, M\. Sun, A\. Mahajan, J\. N\. Foerster, and S\. Whiteson, “Smacv2: An improved benchmark for cooperative multi\-agent reinforcement learning,” 2023\. \[Online\]\. Available:[https://arxiv\.org/abs/2212\.07489](https://arxiv.org/abs/2212.07489)
- Avontuur et al\. \[2013\]T\. Avontuur, P\. Spronck, and M\. van Zaanen, “Player skill modeling in StarCraft II,” in*Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment*, vol\. 9, no\. 1, 2013, pp\. 2–8\.
- Bowman et al\. \[2021\]S\. Bowman, D\. Lux, R\. Vidal, and A\. Drachen, “StarCraft winner prediction,” in*Proceedings of the 16th International Conference on the Foundations of Digital Games*, 2021\.
- Kingma and Welling \[2022\]D\. P\. Kingma and M\. Welling, “Auto\-encoding variational bayes,” 2022\. \[Online\]\. Available:[https://arxiv\.org/abs/1312\.6114](https://arxiv.org/abs/1312.6114)
- Higgins et al\. \[2017\]I\. Higgins, L\. Matthey, A\. Pal, C\. P\. Burgess, X\. Glorot, M\. M\. Botvinick, S\. Mohamed, and A\. Lerchner, “β\\beta\-VAE: Learning basic visual concepts with a constrained variational framework,” in*Proceedings of the 5th International Conference on Learning Representations*, Toulon, France, 2017\. \[Online\]\. Available:[https://openreview\.net/forum?id=Sy2fzU9gl](https://openreview.net/forum?id=Sy2fzU9gl)
- Ding et al\. \[2020\]Z\. Ding, Y\. Xu, W\. Xu, G\. Parmar, Y\. Yang, M\. Welling, and Z\. Tu, “Guided variational autoencoder for disentanglement learning,” in*2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\)*, 2020, pp\. 7917–7926\.
- Schrum et al\. \[2025\]M\. L\. Schrum, S\. Srivatsa, D\. E\. Gopinath, G\. Rosman, and T\. L\. Chen, “Disentangled skill representations for predictive human modeling,” in*ICLR 2026 Conference Withdrawn Submission*, 2025, withdrawn from ICLR 2026\. \[Online\]\. Available:[https://openreview\.net/forum?id=rwvTTjcuHv](https://openreview.net/forum?id=rwvTTjcuHv)
- Korkmaz et al\. \[2018\]E\. Korkmaz, O\. Anil Koyejo, and P\. Smyth, “Optimal transport maps for distribution preserving operations on latent spaces of generative models,” in*ICLR Workshop on Deep Generative Models for Highly Structured Data*, 2018\. \[Online\]\. Available:[https://openreview\.net/forum?id=BklCusRct7](https://openreview.net/forum?id=BklCusRct7)
- Song et al\. \[2023\]Y\. Song, A\. Keller, N\. Sebe, and M\. Welling, “Latent traversals in generative models as potential flows,” in*Proceedings of the 40th International Conference on Machine Learning*, ser\. ICML’23\. JMLR\.org, 2023\.
- Yeh et al\. \[2023\]E\. Yeh, P\. Sequeira, J\. Hostetler, and M\. Gervasio, “Outcome\-guided counterfactuals from a jointly trained generative latent space,” in*Explainable Artificial Intelligence \(xAI 2023\)*, ser\. Communications in Computer and Information Science\. Springer, 2023, pp\. 449–469\. \[Online\]\. Available:[https://arxiv\.org/abs/2207\.07710](https://arxiv.org/abs/2207.07710)
- Crupi et al\. \[2022\]R\. Crupi, A\. Castelnovo, D\. Regoli, and B\. S\. M\. Gonzalez, “Counterfactual explanations as interventions in latent space,”*Data Mining and Knowledge Discovery*, vol\. 38, pp\. 2733–2769, 2022\.
- Bae et al\. \[2025\]J\. Bae, H\. Nam, K\. Ryu, J\. Lee, J\. Kim, H\. Chun, J\. Han, and J\. Choi, “Data\-driven driver training via counterfactual and language\-based guidance in racing scenarios,”*IEEE Access*, vol\. 13, pp\. 170 181–170 199, 2025\.
- Pegios et al\. \[2024\]P\. Pegios, A\. Feragen, A\. A\. Hansen, and G\. Arvanitidis, “Counterfactual explanations via Riemannian latent space traversal,”*arXiv preprint arXiv:2411\.02259*, 2024\. \[Online\]\. Available:[https://arxiv\.org/abs/2411\.02259](https://arxiv.org/abs/2411.02259)
- Kingma and Ba \[2017\]D\. P\. Kingma and J\. Ba, “Adam: A method for stochastic optimization,” 2017\. \[Online\]\. Available:[https://arxiv\.org/abs/1412\.6980](https://arxiv.org/abs/1412.6980)
- Loshchilov and Hutter \[2019\]I\. Loshchilov and F\. Hutter, “Decoupled weight decay regularization,” in*International Conference on Learning Representations*, 2019\. \[Online\]\. Available:[https://openreview\.net/forum?id=Bkg6RiCqY7](https://openreview.net/forum?id=Bkg6RiCqY7)

## Appendix AStarCraft II Game Description

Before diving deeper it is important to state the rules governing StarCraft II competitive gameplay\. The game contains three main “races” of choice for the players\. Each race is differentiated by unit mechanics therefore forcing certain playstyles\. In most cases the game is played in the format of player versus player \(PvP\), one versus one \(1 vs 1\)\. The ultimate goal for competitors is to destroy all of the opponents’ structures, or to force their counterpart to resign\. In many cases tournaments have varying stages, such as the initial group stage, and a subsequent knockout bracket stage\. Depending on the tournament format in most cases the group stages are played out as best of one \(Bo1\), best of two \(Bo2\), or best of three \(Bo3\) matches\. Finally, the knockout bracket features Bo3, best of five \(Bo5\), and best of seven \(Bo7\) matches\.

## Appendix BModel Details

##### Training Objective:

The VAE is trained with the standard reconstruction and KL divergence objective\. The reconstruction term is mean squared error between the original feature vector and the decoded feature vector\. The KL term regularizes the posterior distribution toward a unit Gaussian prior\. The VAE loss is shown in[Equation 20](https://arxiv.org/html/2607.00190#A2.E20)whereℒMSE\\mathcal\{L\}\_\{\\mathrm\{MSE\}\}is the summed mean squared error between the input feature vector and its reconstruction, andℒKL\\mathcal\{L\}\_\{\\mathrm\{KL\}\}regularizes the approximate posterior toward a unit Gaussian prior\. To guide the latent space, we add a binary cross\-entropy loss on the prediction produced from the first latent dimension\. The main Guided VAE update is seen on[Equation 21](https://arxiv.org/html/2607.00190#A2.E21), whereλcls\\lambda\_\{\\mathrm\{cls\}\}controls the strength of supervision and encourages outcome\-relevant information to be concentrated in the guided latent dimension\.

ℒVAE=ℒMSE\+ℒKL,\\mathcal\{L\}\_\{\\mathrm\{VAE\}\}=\\mathcal\{L\}\_\{\\mathrm\{MSE\}\}\+\\mathcal\{L\}\_\{\\mathrm\{KL\}\},\(20\)
ℒmain=ℒVAE\+λcls​ℒBCE,\\mathcal\{L\}\_\{\\mathrm\{main\}\}=\\mathcal\{L\}\_\{\\mathrm\{VAE\}\}\+\\lambda\_\{\\mathrm\{cls\}\}\\mathcal\{L\}\_\{\\mathrm\{BCE\}\},\(21\)

Aside from the typical structure of a Guided VAE model we clamp the log variance output of the encoder to the interval\[−20,2\]\[\-20,2\]before the reparameterization step\. This hard bound constrains the effective standard deviation to the range\[≈4\.5×10−5,≈2\.7\]\[\\approx 4\.5\\times 10^\{\-5\},\\ \\approx 2\.7\], preventing two sources of numerical instability: a near\-zero variance, which causes the KL divergence term in the evidence lower bound \(ELBO\) to diverge and produces exploding gradients; and an excessively large variance, which overwhelms the mean and injects too much noise into the decoder input, destabilizing reconstruction\. These dimensions are used to predict the game outcome\. The point of this is to force at least one direction of the latent space to be directly related to winning and losing\.

## Appendix CHardware

### C\.1Hardware and Computational Requirements

All experiments were conducted on a high\-performance workstation using consumer\-grade hardware\. The specific configuration of the system components is detailed in Table[3](https://arxiv.org/html/2607.00190#A3.T3)\.

Table 3:Hardware specifications used for all experimental runs and hyperparameter sweeps\.The model architectures and optimization strategies were designed for high efficiency\. Consequently, the majority of individual experimental runs, including those within the Guided\-VAE hyperparameter sweep and path generation evaluations, were completed in under 5 minutes\.

## Appendix DHyperparameter Search

### D\.1Guided\-VAE Hyperparameter Search

#### D\.1\.1Search Space

The search space for the Guided\-VAE is defined by a set of hierarchical constraints to ensure a valid bottleneck architecture\. Let𝒲=\{w0,w1,…,wm\}\\mathcal\{W\}=\\\{w\_\{0\},w\_\{1\},\\dots,w\_\{m\}\\\}denote the set of available layer widths in ascending order\.

Architecture Constraints

The encoder configuration is sampled via two primary parameters: the number of layersn∈\{2,3,4\}n\\in\\\{2,3,4\\\}and a categorical start widthws​t​a​r​t∈𝒲w\_\{start\}\\in\\mathcal\{W\}\. To guarantee that the resulting hidden dimensions𝐇e​n​c\\mathbf\{H\}\_\{enc\}are strictly decreasing, we calculate the actual starting indexiias:

i=max⁡\(n−1,index​\(ws​t​a​r​t\)\)i=\\max\(n\-1,\\text\{index\}\(w\_\{start\}\)\)\(22\)The sequence of hidden dimensions is then defined as:

𝐇e​n​c=\{wi−j\}j=0n−1\\mathbf\{H\}\_\{enc\}=\\\{w\_\{i\-j\}\\\}\_\{j=0\}^\{n\-1\}\(23\)
The latent dimensionalitynzn\_\{z\}is coupled to the final encoder widthhl​a​s​t∈𝐇e​n​ch\_\{last\}\\in\\mathbf\{H\}\_\{enc\}through a fractionfz∈\{0\.25,0\.5,1\.0\}f\_\{z\}\\in\\\{0\.25,0\.5,1\.0\\\}, constrained by a minimum floor:

nz=max⁡\(8,⌊hl​a​s​t⋅fz⌋\)n\_\{z\}=\\max\(8,\\lfloor h\_\{last\}\\cdot f\_\{z\}\\rfloor\)\(24\)
Optimization Parameters The remaining parameters are sampled according to the following distributions:

- •Learning Rates:ηv​a​e,ηc​l​s∼LogUniform​\(10−5,10−3\)\\eta\_\{vae\},\\eta\_\{cls\}\\sim\\text\{LogUniform\}\(10^\{\-5\},10^\{\-3\}\)
- •Weight Decays:λv​a​e,λc​l​s∼LogUniform​\(10−6,10−3\)\\lambda\_\{vae\},\\lambda\_\{cls\}\\sim\\text\{LogUniform\}\(10^\{\-6\},10^\{\-3\}\)
- •Classification Weight:α∼LogUniform​\(1,250\)\\alpha\\sim\\text\{LogUniform\}\(1,250\)
- •Supervised Dimensions:ns∈\{1,2,4\}n\_\{s\}\\in\\\{1,2,4\\\}

#### D\.1\.2Hyperparameter Run Configuration

To find the best hyperparameters for our training, we ran the search using Ray \(https://www\.ray\.io/\) and Optuna \(https://optuna\.org/\)\. Upon execution, we have decided on 150 total runs\. The objective function,𝒪H​P​O\\mathcal\{O\}\_\{HPO\}, is defined as a weighted scalar sum of validation metrics logged during the training of the Guided VAE model\. Formally, the minimization objective is expressed as:

minθ⁡𝒪H​P​O=∑i∈ℳwi⋅ℒi​\(θ\)\\min\_\{\\theta\}\\mathcal\{O\}\_\{HPO\}=\\sum\_\{i\\in\\mathcal\{M\}\}w\_\{i\}\\cdot\\mathcal\{L\}\_\{i\}\(\\theta\)\(25\)
whereℳ\\mathcal\{M\}denotes the set of validation metrics,ℒi\\mathcal\{L\}\_\{i\}is the value of theii\-th metric, andwiw\_\{i\}is the user\-defined weight for that metric\. In the configuration utilized for this sweep, the objective was set to equally weight the reconstruction and classification components:

𝒪H​P​O=wvae​ℒval\_vae\+wcls​ℒval\_cls\\mathcal\{O\}\_\{HPO\}=w\_\{\\text\{vae\}\}\\mathcal\{L\}\_\{\\text\{val\\\_vae\}\}\+w\_\{\\text\{cls\}\}\\mathcal\{L\}\_\{\\text\{val\\\_cls\}\}\(26\)
Given our configuration parameterswvae=0\.5w\_\{\\text\{vae\}\}=0\.5andwcls=0\.5w\_\{\\text\{cls\}\}=0\.5, the final objective function simplifies to:

𝒪H​P​O=0\.5​\(ℒval\_vae\)\+0\.5​\(ℒval\_cls\)\\mathcal\{O\}\_\{HPO\}=0\.5\(\\mathcal\{L\}\_\{\\text\{val\\\_vae\}\}\)\+0\.5\(\\mathcal\{L\}\_\{\\text\{val\\\_cls\}\}\)\(27\)

#### D\.1\.3Guided VAE Final Hyperparameters

[Table 4](https://arxiv.org/html/2607.00190#A4.T4)contains the model we have selected for a best performing model\.

Table 4:Final hyperparameter values for the Guided\-VAE model discovered via the Ray/Optuna optimization sweep\.CategoryHyperparameterValueArchitectureInput Dimension196196Encoder Hidden Dimensions \(𝐇e​n​c\\mathbf\{H\}\_\{enc\}\)\[32,16\]\[32,16\]Latent Dimensionality \(nzn\_\{z\}\)1616Supervised Dimensions \(nsn\_\{s\}\)44OptimizationVAE Learning Rate \(ηv​a​e\\eta\_\{vae\}\)1\.7725×10−41\.7725\\times 10^\{\-4\}VAE Weight Decay \(λv​a​e\\lambda\_\{vae\}\)1\.3597×10−51\.3597\\times 10^\{\-5\}Classifier Learning Rate \(ηc​l​s\\eta\_\{cls\}\)4\.1398×10−44\.1398\\times 10^\{\-4\}Classifier Weight Decay \(λc​l​s\\lambda\_\{cls\}\)4\.5500×10−64\.5500\\times 10^\{\-6\}Classification Weight \(α\\alpha\)1\.28241\.2824

### D\.2Path Generation Strategies: Hyperparameter Search

#### D\.2\.1Search Space

The search space for the latent space traversal is structured hierarchically, where the subset of active hyperparameters is conditioned on the chosen strategySS\.

##### Linear Strategy

The linear interpolation strategy relies on neighborhood density constraints:

- •Nearest Neighbors \(kn​e​i​g​h​b​o​r​sk\_\{neighbors\}\):kn​b∼DiscreteUniform​\(3,15\)k\_\{nb\}\\sim\\text\{DiscreteUniform\}\(3,15\)
- •Opponent Constraints \(ko​p​p​o​n​e​n​t​sk\_\{opponents\}\):ko​p​p∼DiscreteUniform​\(10,200\)k\_\{opp\}\\sim\\text\{DiscreteUniform\}\(10,200\)

##### Gradient Ascent Strategy

This strategy utilizes a density\-based optimization approach with fixed stepsT=2000T=2000and a convergence thresholdτ=0\.95\\tau=0\.95:

- •Learning Rate \(η\\eta\):η∼LogUniform​\(10−4,0\.1\)\\eta\\sim\\text\{LogUniform\}\(10^\{\-4\},0\.1\)
- •Momentum \(μ\\mu\):μ∼Uniform​\(0\.0,0\.95\)\\mu\\sim\\text\{Uniform\}\(0\.0,0\.95\)
- •Density Weight \(wρw\_\{\\rho\}\):wρ∼Uniform​\(0\.0,1\.0\)w\_\{\\rho\}\\sim\\text\{Uniform\}\(0\.0,1\.0\)
- •KDE Bandwidth \(hh\):h∼Uniform​\(0\.1,2\.0\)h\\sim\\text\{Uniform\}\(0\.1,2\.0\)

##### Optimal Transport Strategy

The optimal transport strategy balances regularization and geometric constraints:

- •Regularization \(ϵ\\epsilon\):ϵ∼LogUniform​\(0\.01,0\.5\)\\epsilon\\sim\\text\{LogUniform\}\(0\.01,0\.5\)
- •Step Size \(γ\\gamma\):γ∼Uniform​\(0\.05,0\.5\)\\gamma\\sim\\text\{Uniform\}\(0\.05,0\.5\)
- •Opponent Constraints \(ko​p​p​o​n​e​n​t​sk\_\{opponents\}\):ko​p​p∼DiscreteUniform​\(10,200\)k\_\{opp\}\\sim\\text\{DiscreteUniform\}\(10,200\)

##### Neural Flow Strategy

The neural flow strategy utilizes a fixed guidance scale for its transformation:

- •Guidance Scale \(ss\):s=1\.0s=1\.0\(fixed\)

#### D\.2\.2Hyperparameter Run Configuration

The optimization of hyperparameters for the latent space traversal strategies was performed using the Optuna framework\. For each of the strategy we executedN=100N=100independent trials\. In each trial, the performance was evaluated by generatingns​a​m​p​l​e​s=1000n\_\{samples\}=1000latent paths\.

Unlike the model training phase, the objective for path charting is a maximization task\. The objective function𝒥\\mathcal\{J\}is defined as the mean performance of a specified evaluation metricℳ\\mathcal\{M\}: AUC across all generated samples:

maxϕ⁡𝒥​\(ϕ\)=𝔼s∼𝒮​\(ϕ\)​\[ℳ​\(s\)\]\\max\_\{\\phi\}\\mathcal\{J\}\(\\phi\)=\\mathbb\{E\}\_\{s\\sim\\mathcal\{S\}\(\\phi\)\}\[\\mathcal\{M\}\(s\)\]\(28\)
whereϕ\\phirepresents the set of strategy\-specific hyperparameters sampled from the search space, and𝒮​\(ϕ\)\\mathcal\{S\}\(\\phi\)denotes the distribution of paths generated under those parameters\.

#### D\.2\.3Path Generation Strategies: Final Hyperparameters

[Table 5](https://arxiv.org/html/2607.00190#A4.T5)contains the specific hyperparameters used for the final runs of our path generation strategies\.

Table 5:Hyperparameters used for the evaluated methods\. Continuous values discovered via hyperparameter optimization are rounded to four decimal places\.MethodHyperparameterValueNeural FlowGuidance Scale1\.01\.0Gradient AscentSteps20002000Learning Rate \(lr\)0\.03300\.0330Momentum0\.88150\.8815Density Weight0\.84440\.8444KDE Bandwidth0\.65160\.6516Convergence Threshold0\.950\.95Linear CentroidkkNeighbours1414kkOpponents7474Linear NearestkkNeighbours1010kkOpponents1111Optimal TransportRegularization \(reg\)0\.08870\.0887Step Size0\.49990\.4999kkOpponents6464

## Appendix EAdditional Results

Table 6:Cross\-Dataset Comparison of Path\-Charting Strategies with all of the computed metrics\.
### E\.1Model Interpretability

To verify that the Guided VAE’s outcome classifier grounds its predictions in strategically meaningful features, we analyze its decision process using SHapley Additive exPlanations \(SHAP\) as shown in[Fig\. 2](https://arxiv.org/html/2607.00190#A5.F2)and[Fig\. 3](https://arxiv.org/html/2607.00190#A5.F3)\. SHAP values provide a unified measure of feature importance by attributing the log\-odds of the predicted outcome to the individual input features\. The mean absolute SHAP values highlight the top global contributors to the win\-probability prediction, confirming that the model relies on core economic and macroscopic indicators rather than spurious correlations\. Furthermore, the SHAP beeswarm plot reveals the distribution of these impacts across the test\. It visualizes both the magnitude and the direction of the feature effects, illustrating how higher or lower values of specific features correlate with the predicted likelihood of winning\. This transparency is crucial, as it ensures that the counterfactual paths generated by traversing the latent space correspond to interpretable, domain\-consistent shifts in player behavior\.

![Refer to caption](https://arxiv.org/html/2607.00190v1/x2.png)Figure 2:Mean absolute SHAP values for the top\-8 input features of the GuidedVAE win\-probability classifierP​\(w​i​n\)P\(win\)\.![Refer to caption](https://arxiv.org/html/2607.00190v1/x3.png)Figure 3:SHAP dependence plots for the top\-8 features by mean\|S​H​A​P\|\|SHAP\|forP​\(w​i​n\)P\(win\)\.

## Appendix FExample Feedback Reports

[Fig\. 4](https://arxiv.org/html/2607.00190#A6.F4)and[Fig\. 5](https://arxiv.org/html/2607.00190#A6.F5)are examples of the output, informing the user about the counteractual latent space path, and which parameters they should focus on\.

![Refer to caption](https://arxiv.org/html/2607.00190v1/x4.png)Figure 4:UMAP latent space projection with the counterfactual path shown\.![Refer to caption](https://arxiv.org/html/2607.00190v1/x5.png)Figure 5:Feedback report with three distinct user interpretable signals\.
## Appendix GData and Code Repositories

Similar Articles

Coachable agents for interactive gameplay

arXiv cs.AI

This paper presents a framework for training reinforcement learning agents that can be coached in real-time to adopt different styles while performing core tasks, demonstrated in Horizon Forbidden West, Gran Turismo, and a humanoid walking domain.