GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models

arXiv cs.AI Papers

Summary

This paper introduces GENSTRAT, a benchmark that uses procedurally generated strategic environments to evaluate LLMs' strategic reasoning across multiple axes, addressing limitations of fixed game suites.

arXiv:2605.23238v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as economic agents in marketplaces, auctions, and bidding settings. Anticipating their behavior in any specific deployment is hard. Existing strategic-reasoning benchmarks evaluate models on fixed canonical games. These benchmarks may saturate as the frontier improves, and they do not allow evaluators to generalize with confidence from benchmark performance to the varied and messy strategic environments that actual deployments involve. We introduce GENSTRAT, which uses procedurally generated strategic environments to address these challenges. Concretely, we generate a distribution of two-player zero-sum imperfect-information card games. The generator can draw fresh games on demand, allowing for evergreen evaluation and resistance to contamination. We pair the game distribution with a capability-profile methodology that decomposes model competence across six axes (state space, temporal depth, information sensitivity, opponent modeling, risk, and brittleness). We also introduce a jaggedness measure of within-distribution smoothness that detects when a model's advantage jumps unpredictably between strategically similar games. We sample 50 benchmark games from a 2,000-game generated pool and evaluate nine frontier and open-weight LLMs in a head-to-head tournament with over 36,000 matches. Newer frontier-tier models score higher on average. Beyond that average, models with near-identical overall strength show qualitatively different capability profiles, and two of the top three leaderboard models (gpt-5 and claude) are noticeably more locally volatile than the third (gemini-3.1-pro), despite being close in overall strength. Together, the capability profile and the jaggedness measure give a deployment-relevant diagnostic that the overall ranking alone cannot provide.
Original Article
View Cached Full Text

Cached at: 05/25/26, 08:57 AM

# GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models
Source: [https://arxiv.org/html/2605.23238](https://arxiv.org/html/2605.23238)
Vartan Shadarevian Princeton University &Kia Ghods Princeton University &Alex Kenich Google &Anany Kotawala Princeton University

###### Abstract

Large language models \(LLMs\) are increasingly deployed as economic agents in marketplaces, auctions, and bidding settings\. Anticipating their behavior in any specific deployment is hard\. Existing strategic\-reasoning benchmarks evaluate models on fixed canonical games\. These benchmarks may saturate as the frontier improves, and they do not allow evaluators to generalize with confidence from benchmark performance to the varied and messy strategic environments that actual deployments involve\. We introduceGENSTRAT, which uses procedurally generated strategic environments to address these challenges\. Concretely, we generate a distribution of two\-player zero\-sum imperfect\-information card games\. The generator can draw fresh games on demand, allowing for evergreen evaluation and resistance to contamination\. We pair the game distribution with a capability\-profile methodology that decomposes model competence across six axes \(state space, temporal depth, information sensitivity, opponent modeling, risk, and brittleness\)\. We also introduce a jaggedness measure of within\-distribution smoothness that detects when a model’s advantage jumps unpredictably between strategically similar games\. We sample 50 benchmark games from a 2,000\-game generated pool and evaluate nine frontier and open\-weight LLMs in a head\-to\-head tournament with over 36,000 matches\. Newer frontier\-tier models score higher on average\. Beyond that average, models with near\-identical overall strength show qualitatively different capability profiles, and two of the top three leaderboard models \(gpt\-5 and claude\) are noticeably more locally volatile than the third \(gemini\-3\.1\-pro\), despite being close in overall strength\. Together, the capability profile and the jaggedness measure give a deployment\-relevant diagnostic that the overall ranking alone cannot provide\.

## 1Introduction

Frontier LLMs are increasingly placed in economic\-agent roles in controlled experiments, including running small commerce operations\[[3](https://arxiv.org/html/2605.23238#bib.bib38)\]and participating in marketplace simulations\[[4](https://arxiv.org/html/2605.23238#bib.bib39)\], and they exhibit algorithmic\-collusion behavior in LLM\-based pricing studies\[[14](https://arxiv.org/html/2605.23238#bib.bib37)\]\. As LLMs see more use in multi\-agent strategic settings, how well a given model will actually perform once deployed has become hard to anticipate\. AI model performance on canonical games does not transfer cleanly to the specific strategic environment in which a deployer would use the model\.

Existing strategic\-reasoning benchmarks evaluate LLMs on fixed canonical games and suites\. Poker\-based LLM evaluations and agents, including Leduc Hold’em and Texas Hold’em settings\[[15](https://arxiv.org/html/2605.23238#bib.bib10),[16](https://arxiv.org/html/2605.23238#bib.bib11)\], AvalonBench\[[18](https://arxiv.org/html/2605.23238#bib.bib12)\], Diplomacy\[[1](https://arxiv.org/html/2605.23238#bib.bib24)\], and broader game\-theoretic and gameplay suites such as GTBench and GameBench\[[13](https://arxiv.org/html/2605.23238#bib.bib13),[12](https://arxiv.org/html/2605.23238#bib.bib14)\]all fall in this category\. Two limits constrain their ability to serve as evaluations of deployment\-relevant strategic capability\. First, fixed game suites may saturate as the frontier improves, and the closer a benchmark’s contents are to canonical games, the harder it is to rule out corpus contamination from training data\. Second, reducing a model’s strategic competence to performance on a small number of games limits the deployer’s ability to generalize from benchmark performance to novel strategic environments, where variation, real\-world messiness, and shifts in information structure can profoundly affect optimal play\.

We address both limits withGENSTRAT, a procedurally generated distribution of two\-player zero\-sum imperfect\-information card games, which we call generalized betting games \(GBGs\)\. Procedural generation has proven productive for single\-agent reinforcement learning generalization \(ProcGen\[[10](https://arxiv.org/html/2605.23238#bib.bib6)\], MiniGrid\[[9](https://arxiv.org/html/2605.23238#bib.bib7)\]\), but its potential for evaluating multi\-agent strategic reasoning in LLMs has been less explored\. Multi\-agent settings exhibit an*amplification effect*that makes evaluation through procedural generation especially informative: small increases in the complexity of the underlying environment can produce substantial increases in the complexity of the resulting strategic problems agents face\. Each game in our benchmark is played for chips, a numeric stake that accrues over the course of a match and determines the final payoff\. Because the generator can draw freely from the same distribution at any time, the GENSTRAT benchmark cannot be saturated by training on the 50\-game benchmark\. Even if an evaluator that trains directly on those 50 games saturates that fixed subset, a held\-out fresh draw from the same procedural distribution remains uncontaminated\. We pair the distribution with a six\-axis capability\-profile decomposition \(state space, temporal depth, information sensitivity, opponent modeling, risk, brittleness\) so that a model’s performance is reported across strategic dimensions rather than through a single ranking\. We also introduce a jaggedness measure that quantifies how sharply a model’s win\-margin residuals fluctuate between similarly situated games\.

We then run a 9\-model tournament over more than36,00036\{,\}000game matches \(the merged tournament data contains36,93736\{,\}937slot rows\), including both open\-weight and closed\-source models\. Larger, more recent, and reasoning\-capable models score higher on average, with the leaderboard separating models across roughly three chips per game in a clean ordering\. Models with near\-identical overall strength show qualitatively different capability\-profile shapes:gemini\-3\.1\-pro\-previewgains ground on the broadest set of axes, whereasclaude\-sonnet\-4\-6\-maxgains most of its ground on brittleness alone\. The strongest tested model by mean win margin \(gpt\-5\-4\-high\) is also among the most locally jagged \(Section[8](https://arxiv.org/html/2605.23238#S8)\), whereas the second\-strongest \(gemini\-3\.1\-pro\-preview\) is the smoothest among the top\-tier models\. A thinking\-mode ablation, in which the same model plays anchor opponents at low and high reasoning effort across seven of the eight family\-anchor combinations, finds that the chip\-margin return to extra reasoning has positive point estimates of comparable magnitude across all four model families, with two of the four intervals excluding zero and the remaining two underpowered by sample size rather than null in expectation\. The implication for deployment is that a model’s strategic capability is best understood as its full performance profile across different regions of the procedural game space, together with its level of local jaggedness\.

## 2Related work

Strategic games have played a longstanding role in AI research, including specific well\-known cases like AlphaZero\[[24](https://arxiv.org/html/2605.23238#bib.bib20)\]on chess/Go, Libratus\[[7](https://arxiv.org/html/2605.23238#bib.bib21)\]and Pluribus\[[8](https://arxiv.org/html/2605.23238#bib.bib22)\]on poker, DeepNash\[[22](https://arxiv.org/html/2605.23238#bib.bib23)\]on Stratego, and CICERO\[[1](https://arxiv.org/html/2605.23238#bib.bib24)\]on Diplomacy\. These evaluations often focused on isolated, specialized systems meant to play a single well\-known game\. They do not address generalization across novel strategic environments\. More recently, efforts have been made to benchmark general\-purpose LLMs against strategic games\. For example, GTBench\[[13](https://arxiv.org/html/2605.23238#bib.bib13)\], GameBench\[[12](https://arxiv.org/html/2605.23238#bib.bib14)\], and AvalonBench\[[18](https://arxiv.org/html/2605.23238#bib.bib12)\]operate on fixed known games\. Akata et al\.\[[2](https://arxiv.org/html/2605.23238#bib.bib25)\]study LLMs in repeated games\. Lorè and Heydari\[[21](https://arxiv.org/html/2605.23238#bib.bib26)\]disentangle game structure from contextual framing, Collins et al\.\[[11](https://arxiv.org/html/2605.23238#bib.bib15)\]probe LLMs’ abilities to evaluate novel games, and Lin et al\.\[[20](https://arxiv.org/html/2605.23238#bib.bib30)\]study LLMs on professional\-poker\-style tasks with agentic tool use\. The Theory of Mind \(ToM\) literature is closely related\. Strachan et al\.\[[27](https://arxiv.org/html/2605.23238#bib.bib27)\]report human\-level performance for some frontier models on classical false\-belief tasks, while Ullman\[[29](https://arxiv.org/html/2605.23238#bib.bib28)\]shows that small task alterations can sharply reduce apparent ToM performance\. The poker\-ToM coding scheme of\[[19](https://arxiv.org/html/2605.23238#bib.bib29)\]is a closely related effort to read strategic reasoning from model traces\.

Most closely related to our work, gg\-bench\[[33](https://arxiv.org/html/2605.23238#bib.bib16)\]generates novel games via LLM authoring and evaluates LLMs on them by win rate against a self\-play\-trained reinforcement learning \(RL\) agent\.GENSTRATdiffers in four respects: \(i\) a parameterized rule generator \(rather than LLM authoring\), so the game distribution and its complexity are controlled by us directly rather than by an LLM’s design prior; \(ii\) game complexity scales arbitrarily through the same generator, letting the benchmark track the model frontier without being rebuilt; \(iii\) we decompose performance along these axes of complexity and measure the jaggedness of model performance; \(iv\) we run a large\-scale tournament to map the contours of model performance across these axes\.

More broadly, measuring how foundation\-model performance generalizes off the training\-and\-evaluation distribution has motivated a recent line of work on how people expect LLMs to generalize\[[32](https://arxiv.org/html/2605.23238#bib.bib46)\], on the implicit world model of generative models\[[31](https://arxiv.org/html/2605.23238#bib.bib47)\], and on inductive\-bias probes for foundation models\[[30](https://arxiv.org/html/2605.23238#bib.bib48)\]\. That work focuses on single\-agent world modeling; our paper takes the same generalization concern to the multi\-agent strategic setting that economic\-agent deployment lands the model in\.

Other work has examined procedural generation in other contexts as well, though in non\-strategic settings\. ProcGen\[[10](https://arxiv.org/html/2605.23238#bib.bib6)\], MiniGrid\[[9](https://arxiv.org/html/2605.23238#bib.bib7)\], and the broader procedural content generation literature\[[23](https://arxiv.org/html/2605.23238#bib.bib8)\]test single\-agent RL generalization\.

## 3Generalized betting games and GENSTRAT

We define a generalized betting game \(GBG\) as a two\-player zero\-sum extensive\-form game with imperfect information consisting of a deck, private hands, other card piles, structured phases, and conditions that gate branches of the game or otherwise control the occurrence of events\. GBGs generalize games such as Kuhn poker\[[17](https://arxiv.org/html/2605.23238#bib.bib1)\]and Leduc poker\[[26](https://arxiv.org/html/2605.23238#bib.bib2)\]by adding features such as non\-betting actions and rounds, alternative game tree and information structures, and different conditions for accessing branches of the game\.

Structured phases determine how a game unfolds\. They include betting phases, simultaneous\-move phases, auction phases, and observation phases that give players access to signals or other information\. For example, in variations with different levels of observability, the engine controls which observations are visible to each player\.

The game\-building engine randomizes the structural composition of a GBG, not just its surface parameters\. The phase graph itself is sampled, so different draws yield structurally different game forms\. Within that randomized structure, surface features such as ranks, suits, hand sizes, betting order, inclusion of other types of rounds \(such as auction or simultaneous\-move rounds\), observation triggers, position criteria, showdown metrics, side\-bet structures, and conditional\-branch predicates are also drawn at random\. Because randomized configurations can interact in non\-trivial ways, the engine resolves the conditional structures that arise so that the resulting game remains coherent and playable\. Further details on the modular construction are in Appendix[A](https://arxiv.org/html/2605.23238#A1)\.

The complexity of generated games can be scaled by relaxing the generator caps used for the 50\-game benchmark\. The six axes introduced in Section[3\.1](https://arxiv.org/html/2605.23238#S3.SS1)are Monte\-Carlo\-measured diagnostics rather than direct generator controls, but they are used at selection time to target coverage of specific regions of the axis space\. Fresh evaluation games can be drawn from the same procedural distribution at any time, so training on the 50\-game benchmark in Section[4](https://arxiv.org/html/2605.23238#S4)does not exhaust the procedural distribution\. The generator can also produce games at higher complexity than the cap used for the 50\-game benchmark, so the benchmark can scale with the model frontier\. We ensure that every game is a deterministic function of its integer seed for reproducibility\.

When games are generated, we apply multiple quality checks\. A draw is only accepted if it passes three conditions based on Monte Carlo simulations involving random\-playing agents\. First, the average number of moves per player must be no more than ten\. Second, every phase must fire in at least5%5\\%of Monte Carlo episodes, with at most30%30\\%of phases permitted to fall below that threshold before the game is rejected\. Third, in games whose phase graph contains conditional branches, no more than34%34\\%of those branches may remain dead across the Monte Carlo run\. The Monte Carlo budget is2,0002\{,\}000episodes per candidate game with the random agent that selects uniformly at random among legal actions at every decision node\. To collect an accepted pool of2,0002\{,\}000games, the procedural builder sampled12,35112\{,\}351candidate seeds, of which roughly one in six passed the acceptance check to form the candidate pool from which the 50\-game benchmark is then selected following the procedure in §[4](https://arxiv.org/html/2605.23238#S4)\.

### 3\.1Six complexity axes to characterize games

To better understand model performance variation across generated games and to ensure coverage of different notions of complexity, we compute six complexity ‘axes’ measuring the game along different dimensions, constructed from Monte Carlo simulation\. Each axis captures a distinct strategic type of complexity that a player may face\. Together they form the space we use for sampling games and measuring capability profiles \(full formulas are provided in Appendix[C](https://arxiv.org/html/2605.23238#A3)\)\.

- •State space\.The state space axis measures the general combinatorial complexity of the game\. We approximate this aslog10\\log\_\{10\}of the distinct observable information states observed by simulations of random\-agent play\. In this case, for example, small decks with fewer phases score low, whereas larger decks with more phases score high\.
- •Temporal depth\.The temporal depth axis measures how strongly early decisions affect later payoffs\. When early actions are inconsequential, a player can decide myopically\. When they constrain or set up later phases, the player must plan forward\. We measure this as the fraction of total payoff variance that is attributable to decisions taken early in the game, weighted by the number of remaining decisions that still lie ahead\.
- •Information sensitivity\.The information sensitivity axis measures how strongly the best action depends on the player’s private information\. When the optimal move changes with the private hand, the policy must condition on private state\. In low\-information\-sensitivity games, a single near\-best action works regardless\. We measure this as the visit\-weighted fraction of information\-state buckets in which the argmax action depends on the player’s private information\.
- •Opponent modeling\.The opponent modeling axis measures how much the best response shifts when the opponent’s policy changes\. A game scores high when a player must adapt to opponent behavior, and low when the same action works against most opponents\. We measure this as one minus the fraction of opponent policies \(drawn from a Sobol low\-discrepancy quasi\-random sequence\[[25](https://arxiv.org/html/2605.23238#bib.bib45)\]over the probability simplex of mixed strategies, so opponent diversity is covered evenly with few samples\) for which the single most\-common best response remains optimal\.
- •Risk\.The risk axis measures how much a game involves an expected\-value\-versus\-downside tradeoff\. A high\-risk game has large upside and downside swings that force a tradeoff against worst\-case payoff\. A low\-risk game has actions that are safe regardless of opponent play or chance\. We measure this as the visit\-weighted gap between the expected\-value \(EV\) maximizing action and the downside\-safest action at each decision, divided by the standard deviation of payoffs across decision instances so that the gap is dimensionless and comparable across games with different chip stakes\.
- •Brittleness\.Brittleness measures the narrowness of the strategic margins\. A brittle game has a payoff that changes sharply under small \(3%\) policy perturbations, so execution must be precise\. We measure this as follows\. The player’s best\-response action is replaced with a uniformly random alternative at a set of decision contexts whose total visit probability sums to 3% of play\. We then regress per\-trial chip\-margin change on whether the perturbation was applied at decision typed​tdt, and report the ordinary least squares \(OLS\) slope as the brittleness contribution ofd​tdt\.

#### Axis coverage\.

The pairwise Pearson correlation matrix across the 50 benchmark games \(full matrix in Appendix[D](https://arxiv.org/html/2605.23238#A4)\) shows that no two axes are too correlated: all pairwise\|r\|\|r\|are well below the values that would make joint analysis problematic\. The strongest pairs are state\-space×\\timesinformation\-sensitivity \(r=0\.65r\{=\}0\.65\), state\-space×\\timestemporal\-depth \(r=0\.57r\{=\}0\.57\), and information\-sensitivity×\\timesopponent\-modeling \(r=0\.50r\{=\}0\.50\)\. Risk and brittleness are nearly independent of the rest \(\|r\|≤0\.34\|r\|\{\\leq\}0\.34with every other axis\)\. The farthest\-point sampling procedure \(Section[4](https://arxiv.org/html/2605.23238#S4)\) promotes joint Euclidean coverage of the 6\-axis cube subject to staying within the 2,000\-game accepted pool\. Marginal coverage on each axis follows as a consequence of the joint criterion rather than as a separate guarantee, and the realized 50 cover the full observed range on each axis without concentrating in any one corner\. Variance inflation factors \(VIF\) are reported in Appendix[D](https://arxiv.org/html/2605.23238#A4), all below the conservativeVIF=5\\text\{VIF\}\{=\}5rule of thumb \(and well below the looserVIF=10\\text\{VIF\}\{=\}10threshold\)\.

## 4Benchmark construction

From the 2,000\-game accepted pool, we select 50 games via farthest\-point sampling \(FPS\) in the six\-axis embedding\. Each axis is min\-max normalized to\[0,1\]\[0,1\], and FPS greedily picks the next game maximizing minimum Euclidean distance to the already\-selected set, seeded with the centroid game\. Figure[1](https://arxiv.org/html/2605.23238#S4.F1)shows the 50 selected games in a scatter plot with state space and information sensitivity as axes\. A scatter covering the full 2,000\-game pool against the selected 50 is reported in Appendix[E](https://arxiv.org/html/2605.23238#A5)\.

![Refer to caption](https://arxiv.org/html/2605.23238v1/x1.png)Figure 1:The 50\-game benchmark in two diagnostic axes \(state space vs\. information sensitivity\)\.Each point denotes a game, with stars marking the annotated games\.#### Per\-axis coverage\.

Figure[2](https://arxiv.org/html/2605.23238#S4.F2)plots the empirical distribution of each of the six axes across the 50 selected games\. Dashed lines mark the per\-axis 33rd and 67th percentiles, the cuts used for per\-axis tertile splits elsewhere in the analysis\. \(The composite\-complexity tertile split used in Appendix[I](https://arxiv.org/html/2605.23238#A9)is a separate construction based on a principal component of the per\-model slope matrix, not on any single axis\.\) The state\-space axis is reported inlog10\\log\_\{10\}of the raw info\-state count, and the underlying counts span roughly three orders of magnitude across the benchmark \(from about2020on Kuhn\-like games to about1\.6×1041\.6\\times 10^\{4\}on the most complex\), while the other five bounded axes have most of their mass concentrated toward the middle of the observed range with thinner tails on either side; we sample the full observed range on each axis but the bounded axes do not approach a flat distribution\.

![Refer to caption](https://arxiv.org/html/2605.23238v1/x2.png)Figure 2:Per\-axis distributions of the 50 benchmark games\.Kernel\-density estimates with tick\-marked observations and tertile cuts \(dashed\)\. FPS achieves broad per\-axis coverage on all six axes rather than concentrating on the two diagonal axes in Figure[1](https://arxiv.org/html/2605.23238#S4.F1)\.Three distributions appear in the pipeline and should be kept distinct\. The first is the raw generator distribution, defined by the GBG sampling procedure of Section[3](https://arxiv.org/html/2605.23238#S3)before any quality gates apply\. The second is the accepted pool, the subset of raw draws that pass the Monte Carlo acceptance filter described above, of which we collected 2,000 from 12,351 candidate seeds\. The third is the 50\-game evaluation benchmark obtained by farthest\-point sampling within the accepted pool\.

## 5Tournament design

#### Models\.

Nine frontier and open\-weight LLMs participate in the overall tournament:gpt\-5\-4\-high,gemini\-3\.1\-pro\-preview,claude\-sonnet\-4\-6\-max,gemini\-2\.5\-pro,gemma\-4\-31b\-it,deepseek\-v3\.1\-together,gemini\-3\.1\-flash\-lite\-preview,qwen\-3\.5\-together, andllama\-3\.3\-70b\-together\. Per\-model provider parameters are pinned to dated snapshots\. Reasoning\-capable models run with maximum thinking budget \(capability\_tier=max\_thinking\)\. Models without a thinking control run at provider defaults\.

#### Pairing\.

A seat is one of the two player positions in a game\. In each game there is a player seat ‘Alice’ and a player seat ‘Bob’\. Most games are not symmetric across seats, so the seat a model occupies materially affects its payoff\. For each \(model 1, model 2, game\) matchup we run 40 matches, with each model occupying each seat in exactly 20 of them\. A run groups two matches between the same two models, one with each seat assignment, sharing a deterministicplay\_seedderived from\(game seed,run id,matchup\)\(\\text\{game seed\},\\text\{run id\},\\text\{matchup\}\)so that the same chance draws \(dealing, shuffles, etc\.\) are played out under both seat assignments\. A slot is one of those individual matches, i\.e\., a single \(game seed, model 1, model 2, seat assignment, run id\) match in the tournament\. The merged tournament data is a table of slot rows\.

#### Coverage\.

The tested models vary significantly in their per\-token cost\. To keep costs manageable, we vary which subset of possible matchups per game a model faces\. We ensure that each model plays each game against at least two opponents\. The merged tournament data contributes36,93736\{,\}937rows \(some matches hit time\-out issues on multiple attempts and were discarded\)\. Under the additive\-strength specification, the paired\-comparison estimator \(Section[6](https://arxiv.org/html/2605.23238#S6)\) corrects for opponent\-mix imbalance by fitting a single strength indicatorα^m\\hat\{\\alpha\}\_\{m\}for a model jointly across all matchups, so a model that plays a stronger or weaker mix of opponents than another is adjusted for in the strength comparison\.

#### Prompt and parsing\.

Each model receives the auto\-generated natural\-language rulebook in its system prompt and a per\-turn observation prompt composed from its visible history, current phase context, and legal action menu\. Responses must terminate with a JSON object specifying the chosen action\. A lenient parser recovers from minor format violations \(stray whitespace, partial JSON\)\. When both strict and lenient parsing fail, the engine selects a uniformly random legal action so play can continue\. Per\-model fallback rates are uniformly low \(worst case0\.5%0\.5\\%of moves\) and are reported in Appendix[F](https://arxiv.org/html/2605.23238#A6)\.

## 6Overall results

Throughout this paper, payoffs are measured in chips, and all model\-strength scores are reported in chips per game\. For each modelmmwe estimate a single strength scoreα^m\\hat\{\\alpha\}\_\{m\}on a common chips\-per\-game scale\.111α^m\\hat\{\\alpha\}\_\{m\}is a continuous\-margin analogue of a Bradley–Terry\[[6](https://arxiv.org/html/2605.23238#bib.bib40)\]paired\-comparison rating\. It uses the signed win margin in chips at the end of each match rather than the binary win/loss indicator that the standard Bradley–Terry model uses\.

#### Estimator\.

We fit an additive paired\-comparison model on match\-level signed margins\. Letysy\_\{s\}be the win margin \(Alice chips minus Bob chips\) at the end of matchss, and leti\(s\),j\(s\)i^\{\(s\)\},j^\{\(s\)\}be the models seated as Alice and Bob inss\. We fit

ys=αi\(s\)−αj\(s\)\+εssubject to∑mαm=0\.y\_\{s\}\\;=\\;\\alpha\_\{i^\{\(s\)\}\}\-\\alpha\_\{j^\{\(s\)\}\}\+\\varepsilon\_\{s\}\\qquad\\text\{subject to\}\\quad\\textstyle\\sum\_\{m\}\\alpha\_\{m\}=0\.The data identify only strength differences\. An unconstrained fit is degenerate since adding a constantccto everyαm\\alpha\_\{m\}gives the same predictions\. The sum\-to\-zero constraint anchors the level so theα^m\\hat\{\\alpha\}\_\{m\}are uniquely identified, and consistent with the zero\-sum structure of each game\. Confidence intervals come fromB=2,000B=2\{,\}000paired\-cluster bootstrap resamples, with clusters indexed by the four\-tuplec=\(g,m1,m2,r\)c=\(g,m\_\{1\},m\_\{2\},r\), whereggis a game seed,\(m1,m2\)\(m\_\{1\},m\_\{2\}\)is the unordered model pair, andrris the run id that uniquely identifies a paired play seed within that matchup\. The two matches in a cluster \(one for each seat assignment\) share the sameplay\_seed\.BBis the number of bootstrap resamples throughout this paper, fixed unless otherwise noted\. The resultingα^m\\hat\{\\alpha\}\_\{m\}is in chips/game\.

#### Leaderboard\.

Table[1](https://arxiv.org/html/2605.23238#S6.T1)reports overall strengths\.GENSTRATcleanly separates models across roughly three chips/game of range\.

Table 1:Overall leaderboard\.Overall leaderboard —α^\\hat\{\\alpha\}\(chips/game\)Paired\-cluster bootstrap 95% CIs fromB=2,000B\{=\}2\{,\}000resamples, fit on36,93736\{,\}937tournament rows across the 50 benchmark games\.
#### Pairwise margin matrix\.

The full matrix, consisting of pairwise differences in expected chips, is reported in Appendix[J](https://arxiv.org/html/2605.23238#A10)\. Only 2 of the 72 off\-diagonal entries showed a sign reversal from the expected chip margin based on the overall leaderboard\.

#### Robustness checks\.

The overall ordering survives three perturbations of the data\. Underleave\-one\-game\-out, refittingα^\\hat\{\\alpha\}with each of the 50 games dropped in turn preserves the order in 48 of 50 refits \(mean Kendallτ=0\.998\\tau\{=\}0\.998, min0\.9440\.944\)\. Underllama\-excluded, dropping the bottom outlier and refitting on the remaining eight models fully preserves relative order\. Underaxis\-space partition refits, partitioning the 50 games into six clusters by their position in the 6\-axis space and refittingα^\\hat\{\\alpha\}inside each cluster yields per\-cluster rankings that correlate with the overall at Spearmanρ≥0\.95\\rho\{\\geq\}0\.95in five of six clusters \(the sixth is a two\-game cluster withρ=0\.87\\rho\{=\}0\.87\)\. As an additional reference point, we ran an abstracted counterfactual regret minimization \(CFR\)\[[34](https://arxiv.org/html/2605.23238#bib.bib4)\]solver baseline in its CFR\+variant\[[28](https://arxiv.org/html/2605.23238#bib.bib49)\]on 5 tractable seeds against all 9 models \(100 matches each\); model performance against these 5 seeds correlates with the overall leaderboard at Spearmanρ=0\.95\\rho\{=\}0\.95\(Appendix[K](https://arxiv.org/html/2605.23238#A11)\)\. Full tables are in Appendix[H](https://arxiv.org/html/2605.23238#A8)\.

#### Bradley–Terry on win indicator\.

A win\-only Bradley–Terry \(BT\) ranking on the same data gives a noticeably different ordering thanα^\\hat\{\\alpha\}: small\-margin frequent winners \(e\.g\.,gemma\-4\-31b\-it\) score higher under BT, while large\-margin infrequent winners \(e\.g\.,gpt\-5\-4\-high\) score lower\. The two estimators answer different questions\. BT measures frequency of winning, whileα^\\hat\{\\alpha\}measures average win margin\. Since our models were expressly instructed to maximize expected chips rather than expected win probability, we reportα^\\hat\{\\alpha\}in the main text and treat the BT divergence as a complementary stress test rather than a contradiction\. The full BT\-vs\-α^\\hat\{\\alpha\}table is in Appendix[H](https://arxiv.org/html/2605.23238#A8)\.

#### Composite\-complexity tertile refits\.

We also collapse the six axes into a single per\-game*composite\-complexity*score\. The construction proceeds in two steps\. First, we compute the first principal component of the centered nine\-by\-six matrix of per\-model axis\-slopes, yielding a length\-six weight vector with sign fixed so that the entries sum positive\. Second, we project each game’s six\-axiszz\-scored vector onto that weight vector to produce a single per\-game scalar\. The weights and construction details are in Appendix[I](https://arxiv.org/html/2605.23238#A9)\. We then sort the 50 games on this score, split them into low, medium, and high composite\-complexity tertiles, and refitα^\\hat\{\\alpha\}within each tertile\. Two findings emerge\. First, the overall ranking is stable across tertiles:gpt\-5\-4\-highandgemini\-3\.1\-pro\-previewoccupy the top two ranks in every tertile, andllama\-3\.3\-70b\-togetheroccupies the bottom rank in every tertile\. Second, the gap between the top\-three group and llama roughly doubles as composite complexity rises\. The gap between the average top\-three strength and llama’s strength is about1\.751\.75chips per game in the easiest tertile and about4\.54\.5chips per game in the hardest tertile\. Mid\-pack ordering reshuffles modestly in the easiest tertile, where the absolute gaps between mid\-pack models are smallest\. Per\-tertile leaderboards are reported in Appendix[I](https://arxiv.org/html/2605.23238#A9)\.

#### Per\-game strength estimateα^m,g\\hat\{\\alpha\}\_\{m,g\}\.

The capability\-profile and jaggedness analyses below also need a per\-\(model, game\) strength, not just the overallα^m\\hat\{\\alpha\}\_\{m\}\. We refit the additive paired\-comparison model from the previous section separately on each game’s slot rows, with the sum\-to\-zero contrast applied across the models present on that game\. The resultingα^m,g\\hat\{\\alpha\}\_\{m,g\}is modelmm’s win\-margin strength on gameggrelative to the across\-model mean of the nine models ongg, and inherits the same opponent\-mix correction the overallα^m\\hat\{\\alpha\}\_\{m\}applies globally\. A model that drew weaker opponents on gameggis not credited with a higher per\-game strength, because the opponent’s per\-gameα^\\hat\{\\alpha\}on that game is also estimated on the same fit\. Each of the9×50=4509\{\\times\}50\{=\}450\(model, game\) pairs is identified on the merged tournament data\. We useα^m,g\\hat\{\\alpha\}\_\{m,g\}as a primitive throughout the rest of the paper\.

To separate sources of variation in mean chip difference in \(model 1, model 2, game\) matchups, we decompose the variance ofα^m,g\\hat\{\\alpha\}\_\{m,g\}across the 450 \(model, game\) pairs\. The variance ofα^m,g\\hat\{\\alpha\}\_\{m,g\}across the 450 pairs splits into a model main effectσM2\\sigma^\{2\}\_\{M\}, measuring how muchα^m,g\\hat\{\\alpha\}\_\{m,g\}varies from one model to another \(i\.e\. the leaderboard signal\), and a model×\\timesgame interactionσM​G2\\sigma^\{2\}\_\{MG\}, measuring how much a model’s strength on individual games deviates from its own average as the game changes\. Formally,σM2\\sigma^\{2\}\_\{M\}is the variance across models of the per\-model meanα¯m⁣⋅\\bar\{\\alpha\}\_\{m\\cdot\}taken over the 50 games, andσM​G2\\sigma^\{2\}\_\{MG\}is the variance of the residualsα^m,g−α¯m⁣⋅−α¯⋅g\+α¯\\hat\{\\alpha\}\_\{m,g\}\-\\bar\{\\alpha\}\_\{m\\cdot\}\-\\bar\{\\alpha\}\_\{\\cdot g\}\+\\bar\{\\alpha\}across the 450 pairs\. The game main effectα¯⋅g−α¯\\bar\{\\alpha\}\_\{\\cdot g\}\-\\bar\{\\alpha\}is zero by construction becauseα^m,g\\hat\{\\alpha\}\_\{m,g\}is sum\-to\-zero across models on every game \(Section[6](https://arxiv.org/html/2605.23238#S6)\)\. Full definitions, including the within\-cell bootstrap that propagates sampling noise intoα^m,g\\hat\{\\alpha\}\_\{m,g\}, are in Appendix[G](https://arxiv.org/html/2605.23238#A7)\. On the full nine\-model dataset the leaderboard signal dominates but interaction remains substantial: the ratioσM​G2/σM2\\sigma^\{2\}\_\{MG\}/\\sigma^\{2\}\_\{M\}is0\.490\.49, with a 95% paired\-cluster bootstrap interval running from0\.360\.36to0\.650\.65\. If we exclude llama\-3\.3\-70b, whose outlier position inflates the model main effect, the same ratio rises to1\.291\.29\(bootstrap interval0\.640\.64to2\.192\.19\), so model×\\timesgame interaction is roughly comparable in magnitude to the model main effect among the remaining eight non\-outlier models\. A substantial fraction of the variance in per\-\(model, game\) strength therefore reflects which model plays which game, and this is what the capability profiles in Section[7](https://arxiv.org/html/2605.23238#S7)decompose along the six axes\. Full per\-component table in Appendix[G](https://arxiv.org/html/2605.23238#A7)\.

The aggregate leaderboard is a summary across 50 games, and the per\-game ranking induced byα^m,g\\hat\{\\alpha\}\_\{m,g\}need not agree with it\. To test how far apart the two can be, we constructed a per\-game rank\-stability test that, for each game, compares the observed number of pairwise rank reversals against the leaderboard to the number expected under a noise\-only null\. The null fixes each true pairwise gap at the overallα^mi−α^mj\\hat\{\\alpha\}\_\{m\_\{i\}\}\-\\hat\{\\alpha\}\_\{m\_\{j\}\}and treats the per\-game ranking as a noisy draw from it, using the bootstrap\-empirical pairwise standard error on the per\-cellα^mi,g−α^mj,g\\hat\{\\alpha\}\_\{m\_\{i\},g\}\-\\hat\{\\alpha\}\_\{m\_\{j\},g\}as the noise scale\.222The per\-gamepp\-value is computed from a standard\-normal approximation to the sum of pairwise reversal indicators, treating them as independent Bernoullis\. The 36 pairwise reversals on a given game are not strictly independent because the per\-gameα^m,g\\hat\{\\alpha\}\_\{m,g\}estimates share information through the joint per\-game refit\. The independence assumption is the standard construction for a sum of correlated Bernoulli indicators of this form\.We apply the Benjamini–Hochberg false discovery rate \(BH\-FDR\)\[[5](https://arxiv.org/html/2605.23238#bib.bib50)\]procedure across the 50 games\. Two patterns stand out\. First, across the benchmark as a whole the per\-game ranking departs from the overall ranking by more than sampling noise alone would predict on a non\-trivial fraction of games \(15 of 50 atq<0\.05q<0\.05, 19 of 50 atq<0\.10q<0\.10\)\. Hereqqdenotes the Benjamini\-Hochberg\-adjustedpp\-value\. A threshold ofq<0\.05q<0\.05bounds the expected proportion of false discoveries among rejected nulls at5%5\\%\. Second, this dispersion is concentrated almost entirely on the easiest games, with reversal significance falling sharply on more complex ones: 12 of the 16 lowest\-composite\-complexity games are reversal\-significant atq<0\.05q<0\.05, but none of the 17 highest\-complexity games are\. On the hardest games the top models pull away by larger margins, and the per\-game ranking tracks the overall order closely\. We document the construction and per\-game outputs of this rank\-stability test in Appendix[P](https://arxiv.org/html/2605.23238#A16)\.

## 7Capability profiles

The overall ranking treats two models with similar mean win margin as equivalent even when their advantages come from distinct portions of game space\. We want to separate overall model strength, already captured byα^\\hat\{\\alpha\}, from a model’s relative gains and losses in axis space, the profile shape that two models with similar overall rank can still differ on\. Figure[3](https://arxiv.org/html/2605.23238#S7.F3)reports that profile shape via a per\-axis OLS fit of per\-game strengthα^m,g\\hat\{\\alpha\}\_\{m,g\}on the sixzz\-scored axes, in raw chips/game units\.

#### The capability\-profile regression\.

For each modelmmwe fit, on the 50\-game benchmark and with a per\-model intercept:

α^m,g=βm,0\+∑a=16βm,a​za​\(g\)\+εm,g,\\hat\{\\alpha\}\_\{m,g\}\\;=\\;\\beta\_\{m,0\}\+\\sum\_\{a=1\}^\{6\}\\beta\_\{m,a\}\\,z\_\{a\}\(g\)\+\\varepsilon\_\{m,g\},whereza​\(g\)z\_\{a\}\(g\)is thezz\-scored value of axisaaon gamegg\.β^m,a\\hat\{\\beta\}\_\{m,a\}is the change inmm’s chip advantage or deficit versus the across\-model mean perσ\\sigmaof axisaa, controlling for the other five axes and formm’s overall level \(absorbed into the intercept\)\. Becauseα^m,g\\hat\{\\alpha\}\_\{m,g\}is sum\-to\-zero across the nine models on each game by construction \(Section[6](https://arxiv.org/html/2605.23238#S6)\), the slopes automatically sum to zero across models on every axis: if one model gains ground with axisaa, the rest of the model pool collectively falls behind\. Confidence intervals come fromB=500B=500paired\-cluster bootstrap resamples, with clusters indexed by the same \(game seed, run id\) pairing used in Section[6](https://arxiv.org/html/2605.23238#S6)\.

![Refer to caption](https://arxiv.org/html/2605.23238v1/x3.png)Figure 3:Capability profile\.Per\-model OLS slopesβ^m,a\\hat\{\\beta\}\_\{m,a\}of per\-game strengthα^m,g\\hat\{\\alpha\}\_\{m,g\}on the sixzz\-scored axes, fit with a per\-model intercept on the 50 benchmark games\. An outward slope indicates that the model’s chip lead over the across\-model mean grows with that axis, while an inward slope indicates that the lead shrinks\. Units are chips/game perσ\\sigmaof axis\. Slopes sum to zero across the nine models on every axis by construction, since per\-gameα^m,g\\hat\{\\alpha\}\_\{m,g\}is sum\-to\-zero across models on each game\. Shaded bands are 95% paired\-cluster bootstrap CIs \(BB= 500 resamples\)\.Per\-model coefficients with bootstrap CIs and BH\-corrected significance markers are reported in Appendix[L](https://arxiv.org/html/2605.23238#A12); a level\-restored companion plot \(predicted per\-game strength at each axis’s observed maximum, folding overall leaderboard level and profile shape back together\) is in Appendix[M](https://arxiv.org/html/2605.23238#A13)\.

Models with similar overallα^\\hat\{\\alpha\}display structurally different profiles in how their chip advantage or deficit versus the rest of the model pool shifts with axis values\. Brittleness is the single axis on which all three top\-leaderboard models pull further ahead of the rest as the axis grows\.claude\-sonnet\-4\-6\-maxadds\+0\.27\+0\.27chips/game perσ\\sigmaof brittleness \(the largest single\-axis pull\-ahead in the table\),gpt\-5\-4\-highadds\+0\.23\+0\.23, andgemini\-3\.1\-pro\-previewadds\+0\.15\+0\.15, with all three slopes clearing BH atq<0\.05q\{<\}0\.05\.gemini\-3\.1\-pro\-previewpulls ahead on the broadest set of axes among the top three, also gaining BH\-significant ground on state space \(\+0\.22\+0\.22\) and opponent modeling \(\+0\.13\+0\.13\)\.claude\-sonnet\-4\-6\-max’s profile is the most concentrated of the three: brittleness dominates and the other five axes are flat after control\.gpt\-5\-4\-highpulls ahead on several axes \(state space, information sensitivity, opponent modeling\) but only the brittleness slope clears BH after correction\.

At the other end,llama\-3\.3\-70b\-togetherfalls further behind as information sensitivity \(−0\.40\-0\.40\) and brittleness \(−0\.42\-0\.42\) grow, both BH\-significant\. The other four axes are moderately negative but not BH\-significant under the per\-game\-strength fit\.gemini\-3\.1\-flash\-lite\-previewhas the most direction\-splitting profile: it pulls ahead of peers on temporal depth \(\+0\.18\+0\.18\) and information sensitivity \(\+0\.19\+0\.19\) but falls further behind on state space \(−0\.30\-0\.30\), opponent modeling \(−0\.14\-0\.14\), and brittleness \(−0\.15\-0\.15\), all five BH\-significant\.qwen\-3\.5\-togetherloses BH\-significant ground on risk \(−0\.19\-0\.19\) and opponent modeling \(−0\.15\-0\.15\); its temporal\-depth and information\-sensitivity slopes are moderately negative \(−0\.16\-0\.16and−0\.11\-0\.11\) but do not survive BH correction, and the remaining two axes are near zero\. The mid\-packgemini\-2\.5\-pro,gemma\-4\-31b\-it, anddeepseek\-v3\.1\-togetherare largely flat after axis controls, with their relative position not shifting much with any single axis\. This is consistent with capability that scales evenly across axes rather than concentrating on any one\.

#### Robustness to rulebook\-based confounds\.

Because the rulebooks our agents read are themselves auto\-generated, the capability\-profile slopes could in principle pick up rulebook\-presentation effects rather than strategic structure\. The most direct presentation confound is verbosity\. The shortest rulebook in our 50\-game benchmark is roughly 8,600 characters \(approximately 2,000 tokens\) and the longest is just under 39,000 characters \(approximately 9,200 tokens\), a span of roughly fivefold\. We refit the per\-model regression with the base\-ten logarithm of the rulebook character count included as a regressor\. The six axis slopes are stable to this control\. No slope changes sign for any \(model, axis\) pair, and every pair that was BH\-significant in Figure[3](https://arxiv.org/html/2605.23238#S7.F3)remains significant\. The length regressor absorbs only modest variance\.

## 8Local jaggedness

A capability profile reports each model’s average response to a single axis\. Separately, a model’s win margin may vary smoothly between similar games or jump unpredictably between them\. When the per\-game win\-margin surface is smooth, performance on one game extrapolates to nearby games in axis space\. When the surface is jagged, considerable per\-game volatility can be present beneath the overallα^m\\hat\{\\alpha\}\_\{m\}, and deployment outcomes on games that were not in the benchmark become harder to anticipate\.

#### Construction ofJmJ\_\{m\}\.

For each modelmmand gameggwe form the per\-game deviationδm,g=α^m,g−α^m\\delta\_\{m,g\}=\\hat\{\\alpha\}\_\{m,g\}\-\\hat\{\\alpha\}\_\{m\}, the model’s win\-margin strength on gameggminus its across\-model strength\. The deviation averages to approximately zero across the 50 games for each model, and exactly zero whenα^m\\hat\{\\alpha\}\_\{m\}coincides with the unweighted mean ofα^m,g\\hat\{\\alpha\}\_\{m,g\}\. It also inherits the opponent\-mix correction thatα^m,g\\hat\{\\alpha\}\_\{m,g\}already applies\. We then normalize this deviation by the per\-game stakes scaleσg\\sigma\_\{g\}, the standard deviation of all signed match margins played on gamegg\(an intrinsic property of the game rather than of any one model\)\. The studentized deviation is

zm,g=α^m,g−α^mσg\.z\_\{m,g\}\\;=\\;\\frac\{\\hat\{\\alpha\}\_\{m,g\}\-\\hat\{\\alpha\}\_\{m\}\}\{\\sigma\_\{g\}\}\.The denominatorσg\\sigma\_\{g\}is a per\-game empirical scale, computed from the realized match\-margin distribution on gamegg, and is therefore not a property of the rulebook alone\. It depends on the tested model pool, the schedule of matchups, and the retained slot rows\.σg\\sigma\_\{g\}should be read as a common per\-game denominator within this tournament rather than as a game invariant\.333The dependence is less direct than for alternatives that divide explicitly by the model’s overall strength\|α^m\|\|\\hat\{\\alpha\}\_\{m\}\|or by the typical opponent skill gap\.

To turn thezm,gz\_\{m,g\}surface into a per\-model scalar, we average local axis\-space dispersion using aKK\-nearest\-neighbor \(kNN\) aggregator\. For each gamegg, letNK​\(g\)N\_\{K\}\(g\)be its three nearest neighbors in the six\-axis space, with each axis min\-max\-normalized to\[0,1\]\[0,1\]before the Euclidean distance is taken\. Write𝒩​\(g\)=\{g\}∪NK​\(g\)\\mathcal\{N\}\(g\)=\\\{g\\\}\\cup N\_\{K\}\(g\)for the four\-game neighborhood ofgg\(withK=3K=3\), andz¯m,g=1\|𝒩​\(g\)\|​∑g′∈𝒩​\(g\)zm,g′\\bar\{z\}\_\{m,g\}=\\tfrac\{1\}\{\|\\mathcal\{N\}\(g\)\|\}\\sum\_\{g^\{\\prime\}\\in\\mathcal\{N\}\(g\)\}z\_\{m,g^\{\\prime\}\}for the model’s meanzz\-value on that neighborhood\. We compute the population standard deviation of the neighborhood and average over all fifty benchmark games:

Jm=1\|G\|​∑g∈G1\|𝒩​\(g\)\|​∑g′∈𝒩​\(g\)\(zm,g′−z¯m,g\)2\.J\_\{m\}\\;=\\;\\frac\{1\}\{\|G\|\}\\sum\_\{g\\in G\}\\sqrt\{\\frac\{1\}\{\|\\mathcal\{N\}\(g\)\|\}\\sum\_\{g^\{\\prime\}\\in\\mathcal\{N\}\(g\)\}\\big\(z\_\{m,g^\{\\prime\}\}\-\\bar\{z\}\_\{m,g\}\\big\)^\{2\}\}\.We compute 95% confidence intervals from a bias\-corrected paired\-cluster bootstrap\. Robustness to the neighborhood sizeKKand to alternative studentization choices is analyzed in Appendix[N](https://arxiv.org/html/2605.23238#A14)\.

JmJ\_\{m\}as defined does not subtract the fitted capability\-profile surface fromzm,gz\_\{m,g\}before taking the local dispersion, so a smooth but steep trend across the six\-axis space contributes toJmJ\_\{m\}alongside genuine local volatility\.444JmJ\_\{m\}also absorbs sampling noise inα^m,g\\hat\{\\alpha\}\_\{m,g\}in addition to genuine per\-game variation\. The bias correction in the bootstrap removes a component of the inflation, but a fully sampling\-noise\-free analogue would require shrinking eachα^m,g\\hat\{\\alpha\}\_\{m,g\}toward a model\-specific prior, an extension we leave to future work\.

![Refer to caption](https://arxiv.org/html/2605.23238v1/x4.png)Figure 4:Local jaggednessJmJ\_\{m\}, per model\.Bars sorted ascending, with horizontal whiskers showing bias\-corrected paired\-cluster bootstrap 95% confidence intervals \(B=500B=500\)\. HigherJmJ\_\{m\}means the stakes\-normalized per\-game performance surface swings more between axis\-space\-similar games\.
#### InterpretingJmJ\_\{m\}\.

llama\-3\.3\-70b\-togetheris the most locally jagged model, with a centralJmJ\_\{m\}of0\.1520\.152and a 95% confidence interval that runs from0\.1380\.138to0\.1650\.165\. That lower bound exceeds the upper bound of every other model in the table, so the separation is substantive rather than an artifact of interval width\.gpt\-5\-4\-highandclaude\-sonnet\-4\-6\-maxare the next\-most jagged, with centralJmJ\_\{m\}values of0\.0920\.092and0\.0860\.086respectively, followed byqwen\-3\.5\-togetherat0\.0820\.082\. Their CIs overlap one another\. At the smooth end of the table,deepseek\-v3\.1\-togetherhas the lowest centralJmJ\_\{m\}\(0\.0370\.037\), followed bygemini\-2\.5\-pro\(0\.0470\.047\),gemma\-4\-31b\-it\(0\.0520\.052\),gemini\-3\.1\-flash\-lite\-preview\(0\.0570\.057\), andgemini\-3\.1\-pro\-preview\(0\.0620\.062\)\. One caveat is that the stakes\-normalized measure does not condition on the model’s overall strength, so llama’s position at the top of theJmJ\_\{m\}table partly reflects the fact that its absolute chip\-margin swings are larger across most games\. We address this directly when pairingJmJ\_\{m\}withα^m\\hat\{\\alpha\}\_\{m\}below\.

#### PairingJmJ\_\{m\}withα^m\\hat\{\\alpha\}\_\{m\}\.

Read together,JmJ\_\{m\}andα^m\\hat\{\\alpha\}\_\{m\}recover four deployment\-relevant regimes\. HighJmJ\_\{m\}paired with highα^m\\hat\{\\alpha\}\_\{m\}, as forgpt\-5\-4\-highandclaude\-sonnet\-4\-6\-max, describes top\-tier strength with residual local volatility\. These models are strong on average but have axis\-space pockets of much higher or lower edge\. They are also the most informative high\-JmJ\_\{m\}cases, because their elevated jaggedness cannot be explained as a side\-effect of weak play\. LowJmJ\_\{m\}paired with highα^m\\hat\{\\alpha\}\_\{m\}, as forgemini\-3\.1\-pro\-preview, describes the smoothest top\-tier strength regime, where performance on the fifty\-game benchmark is the best guide to performance on neighboring draws\. LowJmJ\_\{m\}paired with low or midα^m\\hat\{\\alpha\}\_\{m\}, as fordeepseek\-v3\.1\-together,gemini\-2\.5\-pro, andgemma\-4\-31b\-it, describes consistently mid\-pack performance with little axis\-space surprise\. The remaining regime is highJmJ\_\{m\}paired with lowα^m\\hat\{\\alpha\}\_\{m\}, as forllama\-3\.3\-70b\-together, a weak model that also moves unpredictably across nearby games\. The\(Jm,α^m\)\(J\_\{m\},\\hat\{\\alpha\}\_\{m\}\)pair carries strictly more information than either alone, and two models with the sameα^m\\hat\{\\alpha\}\_\{m\}but differentJmJ\_\{m\}should be deployed differently\.

## 9Reasoning ablation

We test whether extra thinking budget pays off uniformly across model families\. We consider four families that have both a low\-effort and a high\-effort variant in the model registry:gpt\-5,claude\-sonnet\-4\-6,gemini\-3\.1\-pro, andgemini\-2\.5\-pro\. For each family, we play both variants against two anchor opponents on identical shuffled decks\. The anchors aregemini\-3\.1\-pro\-preview, a top\-tier frontier model, andgemma\-4\-31b\-it, a mid\-pack open\-weight model\. We choose one anchor from each end of the leaderboard so that any thinking\-budget effect we measure is averaged over both stronger and weaker opposition, rather than being tied to a single matchup\. Every low\-effort match has a high\-effort sibling match in which the deal, the deck order, and the opponent are held constant, so any difference in win margin between the two siblings is attributable to the thinking budget rather than to chance variation in the cards\. The estimand is the within\-pair win\-margin gain from going from low to high effort,Δ=𝔼​\[win marginhigh−win marginlow\]\\Delta=\\mathbb\{E\}\[\\text\{win margin\}\_\{\\text\{high\}\}\-\\text\{win margin\}\_\{\\text\{low\}\}\]\. We fitΔ^\\hat\{\\Delta\}on620620such sibling pairs, spread across seven \(family, anchor\) cells, and report 95% confidence intervals from a cluster bootstrap on game seed\.

Table 2:Thinking ablation: paired\-Δ\\Deltawin margin by family\.Families whoseΔ^\\hat\{\\Delta\}is significantly different from zero are starred\.∗95% CI excludes 0\.

The four central estimates are positive and of comparable magnitude, ranging from\+0\.21\+0\.21forgpt\-5to\+0\.53\+0\.53forgemini\-2\.5\-pro\. Two of the four intervals exclude zero, and the other two do not, but the difference between significant and non\-significant rows is largely a difference in interval width rather than in the central effect\. The non\-significant rows are therefore better read as underpowered at the current sample size than as evidence of a null return to extra reasoning\. We note, separately, that each provider implements the low\-effort/high\-effort dial differently\. Theclaude\-sonnet\-4\-6andgpt\-5families expose an explicit reasoning\-token budget\. The Gemini families switch a discrete thinking mode on or off\. The differences in match counts \(50 to 260\) reflect which models we could schedule paired siblings on\.

## 10Limitations and discussion

#### Data, methodology, and scope\.

Axis\-conditional analyses at fifty games are directionally robust, but a larger benchmark would tighten the estimates further\. Even with our farthest\-point sampling, the state\-space\-related axes carry mechanical correlations with the other axes, so the six axes are complementary rather than fully orthogonal\. Provider snapshots are pinned to dated model identifiers, so later snapshots are not guaranteed to reproduce the ordering\.GENSTRATcovers two\-player zero\-sum imperfect\-information English\-language betting games\. Cooperative, multi\-player, non\-betting, post\-training, agentic\-scaffold, and heuristic\-baseline settings are left to future work\.

All reported quantities, including the leaderboard, the capability profiles, and the jaggedness measures, are computed under a sum\-to\-zero contrast across the nine tested models and are therefore comparisons within the model pool\. The CFR baseline on five tractable seeds is the only absolute reference point in our evaluation, and the remaining results should be read as relative rather than as a calibration of absolute strategic competence on the GBG distribution\.

#### Broader impacts and deployment implications\.

Frontier LLMs are deployed as economic agents in marketplace, auction, and bidding settings \(Section[1](https://arxiv.org/html/2605.23238#S1)\)\. The leaderboard ordering is stable across leave\-one\-game\-out and composite\-complexity tertile refits within the 50\-game benchmark, so the relative ranking of models does not depend on any particular subset of games within the procedural distribution that the six axes span\. We are more limited in our ability to evaluate strategic settings whose axis scores fall outside the range covered by the 50\-game benchmark, or whose structural properties differ entirely from GBGs\. Extending the benchmark to higher complexity caps and to non\-GBG strategic families is left to future work\.

Within the span the benchmark covers, the overallα^\\hat\{\\alpha\}ranking on its own is not enough to guide deployment\. A model’s capability profile on the axes that match the deployment, taken together with its local smoothnessJmJ\_\{m\}on the region of game space the deployment is closest to, is the deployment\-relevant summary\. A model with high overallα^\\hat\{\\alpha\}but a flat slope on information sensitivity may be a less suitable candidate than its leaderboard rank suggests for a bidding marketplace in which private\-value revelation matters\. Similarly, a model with highJmJ\_\{m\}may perform less reliably in a deployment whose game distribution lies near but not inside the benchmark\.

## 11Conclusion

GENSTRATreframes strategic\-reasoning evaluation around a procedurally generated game distribution and a six\-axis decomposition of strategic complexity, so that capability profiles, jaggedness, and ablations can be read off the same evaluation\. The generator scales as models improve and extends naturally to richer action primitives, larger state spaces, and multi\-player settings\.

## Acknowledgements

We thank Peter Henderson, Zeyu Shen, and Vikram Kakaria for valuable discussions, assistance, and feedback that helped shape this work\.

## References

- \[1\]M\. F\. A\. R\. D\. T\. \(FAIR\), A\. Bakhtin, N\. Brown, E\. Dinan, G\. Farina, C\. Flaherty, D\. Fried, A\. Goff, J\. Gray, H\. Hu, A\. P\. Jacob, M\. Komeili, K\. Konath, M\. Kwon, A\. Lerer, M\. Lewis, A\. H\. Miller, S\. Mitts, A\. Renduchintala, S\. Roller, D\. Rowe, W\. Shi, J\. Spisak, A\. Wei, D\. Wu, H\. Zhang, and M\. Zijlstra\(2022\)Human\-level play in the game of Diplomacy by combining language models with strategic reasoning\.Science378\(6624\),pp\. 1067–1074\.Cited by:[§1](https://arxiv.org/html/2605.23238#S1.p2.1),[§2](https://arxiv.org/html/2605.23238#S2.p1.1)\.
- \[2\]E\. Akata, L\. Schulz, J\. Coda\-Forno, S\. J\. Oh, M\. Bethge, and E\. Schulz\(2025\)Playing repeated games with large language models\.Nature Human Behaviour9\(7\),pp\. 1380–1390\.Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p1.1)\.
- \[3\]Anthropic\(2025\)Project Vend: Can Claude run a small shop? \(And why does that matter?\)\.Note:Anthropic ResearchExternal Links:[Link](https://www.anthropic.com/research/project-vend-1)Cited by:[§1](https://arxiv.org/html/2605.23238#S1.p1.1)\.
- \[4\]Anthropic\(2026\)Project Deal: our Claude\-run marketplace experiment\.Note:AnthropicExternal Links:[Link](https://www.anthropic.com/features/project-deal)Cited by:[§1](https://arxiv.org/html/2605.23238#S1.p1.1)\.
- \[5\]Y\. Benjamini and Y\. Hochberg\(1995\)Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing\.Journal of the Royal Statistical Society\. Series B \(Methodological\)57\(1\),pp\. 289–300\.Cited by:[§6](https://arxiv.org/html/2605.23238#S6.SS0.SSS0.Px7.p3.10)\.
- \[6\]R\. A\. Bradley and M\. E\. Terry\(1952\)Rank Analysis of Incomplete Block Designs: I\. The Method of Paired Comparisons\.Biometrika39\(3/4\),pp\. 324–345\.Cited by:[footnote 1](https://arxiv.org/html/2605.23238#footnote1)\.
- \[7\]N\. Brown and T\. Sandholm\(2018\)Superhuman AI for heads\-up no\-limit poker: Libratus beats top professionals\.Science359\(6374\),pp\. 418–424\.Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p1.1)\.
- \[8\]N\. Brown and T\. Sandholm\(2019\)Superhuman AI for multiplayer poker\.Science365\(6456\),pp\. 885–890\.Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p1.1)\.
- \[9\]M\. Chevalier\-Boisvert, B\. Dai, M\. Towers, R\. Perez\-Vicente, L\. Willems, S\. Lahlou, S\. Pal, P\. S\. Castro, and J\. K\. Terry\(2023\)Minigrid & Miniworld: Modular & Customizable Reinforcement Learning Environments for Goal\-Oriented Tasks\.InAdvances in Neural Information Processing Systems \(NeurIPS\) Datasets and Benchmarks Track,Vol\.36,pp\. 73383–73394\.Cited by:[§1](https://arxiv.org/html/2605.23238#S1.p3.1),[§2](https://arxiv.org/html/2605.23238#S2.p4.1)\.
- \[10\]K\. Cobbe, C\. Hesse, J\. Hilton, and J\. Schulman\(2020\)Leveraging Procedural Generation to Benchmark Reinforcement Learning\.InInternational Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.119,pp\. 2048–2056\.Cited by:[§1](https://arxiv.org/html/2605.23238#S1.p3.1),[§2](https://arxiv.org/html/2605.23238#S2.p4.1)\.
- \[11\]K\. M\. Collins, C\. E\. Zhang, G\. Todd, L\. Ying, M\. B\. da Costa, R\. Liu, P\. Sharma, A\. Weller, I\. Kuperwajs, L\. Wong, J\. B\. Tenenbaum, and T\. L\. Griffiths\(2026\)Evaluating Language Models’ Evaluations of Games\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p1.1)\.
- \[12\]A\. Costarelli, M\. Allen, R\. Hauksson, G\. Sodunke, S\. Hariharan, C\. Cheng, W\. Li, J\. Clymer, and A\. Yadav\(2024\)GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents\.Note:arXiv:2406\.06613v2Cited by:[§1](https://arxiv.org/html/2605.23238#S1.p2.1),[§2](https://arxiv.org/html/2605.23238#S2.p1.1)\.
- \[13\]J\. Duan, R\. Zhang, J\. Diffenderfer, B\. Kailkhura, L\. Sun, E\. Stengel\-Eskin, M\. Bansal, T\. Chen, and K\. Xu\(2024\)GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game\-Theoretic Evaluations\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.37,pp\. 28219–28253\.Cited by:[§1](https://arxiv.org/html/2605.23238#S1.p2.1),[§2](https://arxiv.org/html/2605.23238#S2.p1.1)\.
- \[14\]S\. Fish, Y\. A\. Gonczarowski, and R\. I\. Shorrer\(2024\)Algorithmic Collusion by Large Language Models\.Note:arXiv:2404\.00806v5, revised 2026Cited by:[§1](https://arxiv.org/html/2605.23238#S1.p1.1)\.
- \[15\]J\. Guo, B\. Yang, P\. Yoo, B\. Y\. Lin, Y\. Iwasawa, and Y\. Matsuo\(2024\)Suspicion Agent: Playing Imperfect Information Games with Theory of Mind Aware GPT\-4\.InFirst Conference on Language Modeling \(COLM\),Cited by:[§1](https://arxiv.org/html/2605.23238#S1.p2.1)\.
- \[16\]C\. Huang, Y\. Cao, Y\. Wen, T\. Zhou, and Y\. Zhang\(2024\)PokerGPT: An End\-to\-End Lightweight Solver for Multi\-Player Texas Hold’em via Large Language Model\.Note:arXiv:2401\.06781Cited by:[§1](https://arxiv.org/html/2605.23238#S1.p2.1)\.
- \[17\]H\. W\. Kuhn\(1950\)A simplified two\-person poker\.InContributions to the Theory of Games, Vol\. I,H\. W\. Kuhn and A\. W\. Tucker \(Eds\.\),Annals of Mathematics Studies,pp\. 97–103\.Cited by:[§3](https://arxiv.org/html/2605.23238#S3.p1.1)\.
- \[18\]J\. Light, M\. Cai, S\. Shen, and Z\. Hu\(2023\)AvalonBench: Evaluating LLMs Playing the Game of Avalon\.Note:arXiv:2310\.05036v3Cited by:[§1](https://arxiv.org/html/2605.23238#S1.p2.1),[§2](https://arxiv.org/html/2605.23238#S2.p1.1)\.
- \[19\]H\. Lin and T\. Hou\(2026\)Readable Minds: Emergent Theory\-of\-Mind\-Like Behavior in LLM Poker Agents\.Note:arXiv:2604\.04157Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p1.1)\.
- \[20\]M\. Lin, E\. Dai, H\. Liu, X\. Tang, Y\. Yan, Z\. Dai, J\. Zeng, Z\. Zhang, F\. Wang, H\. Gao, C\. Luo, X\. Zhang, Q\. He, and S\. Wang\(2026\)How Far Are LLMs from Professional Poker Players? Revisiting Game\-Theoretic Reasoning with Agentic Tool Use\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p1.1)\.
- \[21\]N\. Lorè and B\. Heydari\(2024\)Strategic behavior of large language models and the role of game structure versus contextual framing\.Scientific Reports14\(1\),pp\. 18490\.Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p1.1)\.
- \[22\]J\. Perolat, B\. D\. Vylder, D\. Hennes, E\. Tarassov, F\. Strub, V\. de Boer, P\. Muller, J\. T\. Connor, N\. Burch, T\. Anthony, S\. McAleer, R\. Elie, S\. H\. Cen, Z\. Wang, A\. Gruslys, A\. Malysheva, M\. Khan, S\. Ozair, F\. Timbers, T\. Pohlen, T\. Eccles, M\. Rowland, M\. Lanctot, J\. Lespiau, B\. Piot, S\. Omidshafiei, E\. Lockhart, L\. Sifre, N\. Beauguerlange, R\. Munos, D\. Silver, S\. Singh, D\. Hassabis, and K\. Tuyls\(2022\)Mastering the game of Stratego with model\-free multiagent reinforcement learning\.Science378\(6623\),pp\. 990–996\.Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p1.1)\.
- \[23\]N\. Shaker, J\. Togelius, and M\. J\. Nelson\(2016\)Procedural Content Generation in Games\.Springer\.Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p4.1)\.
- \[24\]D\. Silver, T\. Hubert, J\. Schrittwieser, I\. Antonoglou, M\. Lai, A\. Guez, M\. Lanctot, L\. Sifre, D\. Kumaran, T\. Graepel, T\. Lillicrap, K\. Simonyan, and D\. Hassabis\(2018\)A general reinforcement learning algorithm that masters chess, shogi, and Go through self\-play\.Science362\(6419\),pp\. 1140–1144\.Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p1.1)\.
- \[25\]I\. M\. Sobol’\(1967\)On the distribution of points in a cube and the approximate evaluation of integrals\.USSR Computational Mathematics and Mathematical Physics7\(4\),pp\. 86–112\.Cited by:[4th item](https://arxiv.org/html/2605.23238#S3.I1.i4.p1.1)\.
- \[26\]F\. Southey, M\. Bowling, B\. Larson, C\. Piccione, N\. Burch, D\. Billings, and C\. Rayner\(2005\)Bayes’ Bluff: Opponent Modelling in Poker\.InProceedings of the Twenty\-First Conference on Uncertainty in Artificial Intelligence \(UAI\),pp\. 550–558\.Cited by:[§3](https://arxiv.org/html/2605.23238#S3.p1.1)\.
- \[27\]J\. W\. A\. Strachan, D\. Albergo, G\. Borghini, O\. Pansardi, E\. Scaliti, S\. Gupta, K\. Saxena, A\. Rufo, S\. Panzeri, G\. Manzi, M\. S\. A\. Graziano, and C\. Becchio\(2024\)Testing theory of mind in large language models and humans\.Nature Human Behaviour8\(7\),pp\. 1285–1295\.Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p1.1)\.
- \[28\]O\. Tammelin\(2014\)Solving Large Imperfect Information Games Using CFR\+\.Note:arXiv:1407\.5042Cited by:[§6](https://arxiv.org/html/2605.23238#S6.SS0.SSS0.Px4.p1.7)\.
- \[29\]T\. Ullman\(2023\)Large Language Models Fail on Trivial Alterations to Theory\-of\-Mind Tasks\.Note:arXiv:2302\.08399v5Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p1.1)\.
- \[30\]K\. Vafa, P\. G\. Chang, A\. Rambachan, and S\. Mullainathan\(2025\)What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p3.1)\.
- \[31\]K\. Vafa, J\. Y\. Chen, A\. Rambachan, J\. Kleinberg, and S\. Mullainathan\(2024\)Evaluating the World Model Implicit in a Generative Model\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.37,pp\. 26941–26975\.Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p3.1)\.
- \[32\]K\. Vafa, A\. Rambachan, and S\. Mullainathan\(2024\)Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function\.InInternational Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.235,pp\. 48919–48937\.Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p3.1)\.
- \[33\]V\. Verma, D\. Huang, W\. Chen, D\. Klein, and N\. Tomlin\(2025\)Measuring General Intelligence with Generated Games\.Note:arXiv:2505\.07215Cited by:[§2](https://arxiv.org/html/2605.23238#S2.p2.1)\.
- \[34\]M\. Zinkevich, M\. Johanson, M\. Bowling, and C\. Piccione\(2007\)Regret minimization in games with incomplete information\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.20,pp\. 1729–1736\.Cited by:[§6](https://arxiv.org/html/2605.23238#S6.SS0.SSS0.Px4.p1.7)\.

## Appendix AModular game design

Each GBG is assembled from modular components composed by the parameterized GBG builder:

- •Core layer\.GameState\(players, variables, piles\),Rulebook\(phase graph\),Phase\(ordered action lists\)\.
- •Action subclasses\.Every game operation is a typedActionsubclass \(Deal,ChipTransfer,Shuffle,PeekAtHand,SwapWithOpponent,StealCardMove,Wager,ScoreAdjustment, …\), enabling structured logging and analysis\.
- •Phase blocks\.Reusablemake\_action\_round\(\),make\_observation\_round\(\),make\_simultaneous\_round\(\), andmake\_position\_assignment\_round\(\)compose phases from parameterized templates\.
- •Phase graph\.Conditional transitions between phases gated by chip counts, card comparisons, and round counters enable nonlinear game flow \(branches, bounded loops, conditional sub\-phases\)\.
- •Complexity dial\.A single parameterc∈\[0,1\]c\\in\[0,1\]modulates the draw probabilities for more complex structural and surface features, so increasingccsmoothly shifts the distribution from Kuhn\-like simplicity toward multi\-phase complexity\.
- •Deterministic reconstruction\.G=reconstruct​\(s\)G=\\texttt\{reconstruct\}\(s\)exactly reproduces any game from its seedssgiven the builder version hash\.

Natural\-language rulebooks \(Appendix[B](https://arxiv.org/html/2605.23238#A2)\) are auto\-generated from the phase graph and served as LLM system prompts\. The rendering preserves all conditional structure, visibility rules, and chip\-transfer semantics\.

## Appendix BSample game rulebook

Relative to the 50 benchmark games, seed 4643 scores high on state\-space \(90th percentile\), temporal\-depth \(98th\), and risk \(94th\); low\-to\-mid on information\-sensitivity \(32nd\) and opponent\-modeling \(30th\); and mid\-high on brittleness \(72nd\)\. The rulebook below is the verbatim system prompt the LLM agents receive at the start of play\. It is constructed deterministically from the phase graph\.

[⬇](data:text/plain;base64,IyMgT3ZlcnZpZXcNCioqUGxheWVyczoqKiBBbGljZSwgQm9iDQoqKlBsYXllciBvcmRlcjoqKiBQbGF5ZXIgMSA9IEFsaWNlIChhY3RzIGZpcnN0IGluIGJldHRpbmcpLCBQbGF5ZXIgMiA9IEJvYi4NCioqRGVjazoqKiAzIGNhcmRzIChKLCBRLCBLIG9mIGhlYXJ0cykNCioqU3RhcnRpbmcgY2hpcHM6KiogMTAgZWFjaCAodmlzaWJsZSB0byBhbGwpLiBDaGlwcyBtb3ZlIGJldHdlZW4gcGxheWVycycgc3RhY2tzIGFuZCB0aGUgcG90OyB0aGUgdG90YWwgb2YgYWxsIGNoaXBzIChzdGFja3MgcGx1cyBwb3QpIGlzIGNvbnNlcnZlZCBhY3Jvc3MgdGhlIGdhbWUuDQoqKlBvdDoqKiBzdGFydHMgYXQgMA0KKipHZW5lcmFsIHJ1bGVzOioqDQotICoqSW5mb3JtYXRpb246KiogVGhlIG51bWJlciBvZiBjYXJkcyByZW1haW5pbmcgaW4gdGhlIERlY2sgaXMgcHVibGljIGluZm9ybWF0aW9uLiBIYW5kIHNpemVzIGFyZSBwcml2YXRlLiBQbGF5ZXJzIGRvIG5vdCBrbm93IGhvdyBtYW55IGNhcmRzIG9wcG9uZW50cyBob2xkIHVubGVzcyByZXZlYWxlZCBieSBhbiBhY3Rpb24uIElmIHRoZSBEZWNrIGlzIGVtcHR5IHdoZW4gYSBkcmF3IGFjdGlvbiBvY2N1cnMsIG5vIGNhcmRzIGFyZSBkcmF3biBhbmQgdGhlIGFjdGlvbiBoYXMgbm8gZWZmZWN0LiBQdWJsaWMgZHJhd3MgZW1pdCBhIHB1YmxpYyBza2lwcGVkLWRyYXcgbm90aWNlOyBwcml2YXRlIGRyYXdzIGVtaXQgdGhhdCBub3RpY2Ugb25seSB0byB0aGUgZHJhd2luZyBwbGF5ZXIg4oCUIGJ1dCBwZXIgdGhlIGluZmVyYWJpbGl0eSBwcmluY2lwbGUsIHRoZSBEZWNrLXNpemUgY2hhbmdlIGZyb20gYW55IHN1Y2Nlc3NmdWwgZHJhdyBpcyBwdWJsaWNseSB2aXNpYmxlLg0KLSAqKkxlZ2FsIGFjdGlvbnM6KiogUGxheWVycyBtYXkgY2hvb3NlIGFueSBsaXN0ZWQgb3B0aW9uLiBJZiBpdHMgcHJlY29uZGl0aW9ucyBhcmUgdW5tZXQgKGVtcHR5IGhhbmQpLCB0aGUgYWN0aW9uIGhhcyBubyBlZmZlY3QuDQotICoqUmVzb2x1dGlvbjoqKiBJbiBhbnkgbXVsdGktc3RlcCBlZmZlY3QsIGVhY2ggc3RlcCByZXNvbHZlcyBpbiBvcmRlcjsgaWYgb25lIHN0ZXAncyBwcmVjb25kaXRpb24gaXMgdW5tZXQsIHRoYXQgc3RlcCBpcyBza2lwcGVkIGFuZCBsYXRlciBzdGVwcyBzdGlsbCByZXNvbHZlLg0KLSAqKkdhbWUgZW5kOioqIElmIHRoZSBnYW1lIGVuZHMgZHVyaW5nIGFueSBwaGFzZSwgYWxsIHJlbWFpbmluZyBwaGFzZXMgYXJlIHNraXBwZWQuDQotICoqQ29zdHMgJiBwb3Q6KiogQ29zdHMgYXJlIHBhaWQgdG8gdGhlIHBvdCBiZWZvcmUgdGhlIGVmZmVjdCByZXNvbHZlcywgdW5sZXNzIG90aGVyd2lzZSBzcGVjaWZpZWQuIFdoZW4gYW4gYWN0aW9uIGxpc3RzIGEgY29zdDogaWYgdGhlIHBsYXllciBjYW5ub3QgYWZmb3JkIHRoZSBzdGF0ZWQgY29zdCBhdCB0aGUgbW9tZW50IHRoZXkgc2VsZWN0IGl0LCB0aGUgYWN0aW9uIGlzIG5vdCBvZmZlcmVkIGFzIGFuIG9wdGlvbi4gSWYgdGhlIGNvc3QgaXMgcHJvYmFiaWxpc3RpYyAoZS5nLiBhIDUwJSBjaGFuY2Ugb2YgcGF5aW5nIDEgY2hpcCksIHRoZSBwbGF5ZXIgY29tbWl0cyB0byB0aGUgYWN0aW9uIGZpcnN0IGFuZCB0aGUgY29zdCBpcyByZXNvbHZlZCBhZnRlcndhcmRzOyBpZiB0aGV5IGNhbm5vdCBhZmZvcmQgdGhlIHJlc29sdmVkIGNvc3QsIHRoZSBhY3Rpb24gaXMgc2tpcHBlZCBhbmQgdGhlaXIgY2hvaWNlIHNsb3QgaXMgY29uc3VtZWQgKHRoZXkgZG8gbm90IGdldCB0byBwaWNrIGFnYWluKS4NCi0gKipTaG93ZG93biAmIG1pc2MuOioqIFRoZXJlIGlzIG5vIGhhbmQgc2l6ZSBsaW1pdDsgaGFuZHMgZ3JvdyBhbmQgc2hyaW5rIGZyZWVseSB2aWEgZHJhd3MuIFVubGVzcyBzdGF0ZWQgb3RoZXJ3aXNlLCBhbGwgcmFuZG9tIHNlbGVjdGlvbnMgYXJlIHVuaWZvcm0uIEF0IFNob3dkb3duLCBlYWNoIHBsYXllcidzIGhhbmQgaXMgcmV2ZWFsZWQgYXMgb2YgdGhlIG1vbWVudCB0aGUgU2hvd2Rvd24gcGhhc2UgYmVnaW5zLCBpbmNsdWRpbmcgYW55IGNhcmRzIHN3YXBwZWQsIGRyYXduLCBvciBzdG9sZW4gZHVyaW5nIGVhcmxpZXIgcGhhc2VzLg0KIyMgVmlzaWJpbGl0eQ0KLSAqKkhhbmRzKio6IEVhY2ggcGxheWVyJ3MgY2FyZHMgYXJlIHByaXZhdGUgdG8gdGhhdCBwbGF5ZXIuIE9wcG9uZW50cyBjYW5ub3Qgc2VlIGhhbmQgY29udGVudHMgdW5sZXNzIGV4cGxpY2l0bHkgcmV2ZWFsZWQuDQotICoqQWN0aW9ucyoqOiBBbGwgcGxheWVyIGFjdGlvbnMgKGJldHMsIGNoZWNrcywgZm9sZHMpIGFuZCB0aGVpciBlZmZlY3RzIGFyZSBwdWJsaWNseSBhbm5vdW5jZWQgdW5sZXNzIHN0YXRlZCBvdGhlcndpc2UuIEFjdGlvbnMgbWFya2VkIGFzIHByaXZhdGUgYXJlIGV4ZWN1dGVkIHNpbGVudGx5Lg0KLSAqKkluZmVyYWJpbGl0eSoqOiAiUHJpdmF0ZSIgaGlkZXMgKndoaWNoIGJyYW5jaCBmaXJlZCwgd2hhdCBjYXJkIHdhcyBkcmF3bi9wZWVrZWQsIG9yIHdobyBjaG9zZSB3aGF0KiDigJQgbm90IHRoZSBvYnNlcnZhYmxlIHNpZGUtZWZmZWN0cy4gRGVjayBzaXplcywgcG90IHRvdGFscywgY2hpcCBzdGFja3MsIGFuZCBwb3NpdGlvbiBob2xkaW5ncyBhcmUgYWx3YXlzIHB1YmxpYyBhbmQgcmVjb25jaWxlIGF0IHNldHRsZW1lbnQuIElmIG9uZSBwcml2YXRlIGJyYW5jaCB3b3VsZCBjaGFuZ2UgYSBwdWJsaWMgbnVtYmVyIChlLmcuIHNocmluayB0aGUgRGVjayBieSAxKSBhbmQgYW5vdGhlciB3b3VsZCBub3QsIHRoZSBvcHBvbmVudCBjYW4gaW5mZXIgd2hpY2ggYnJhbmNoIHJhbiBmcm9tIHRoYXQgbnVtYmVyLiBUcmVhdCB0aGUgcHJpdmFjeSBsYWJlbCBhcyBjb25jZWFsaW5nIHRoZSAqaWRlbnRpdHkvY2hvaWNlKiwgbm90IHRoZSAqb2NjdXJyZW5jZSouDQotICoqQ29uZGl0aW9uIGNoZWNrcyoqOiBDaGVja3MgbWFya2VkICIocHVibGljKSIgYXJlIGFubm91bmNlZCB0byBhbGwgcGxheWVycyB3aGVuIGV2YWx1YXRlZC4NCi0gKipIaWRkZW4gaW5mb3JtYXRpb24qKjogQWxsIGhpZGRlbiBpbmZvcm1hdGlvbiBpcyB0cmFja2VkIGFuZCBlbmZvcmNlZDsgcGxheWVycyBhcmUgdG9sZCBvbmx5IHdoYXQgdGhlIHJ1bGVzIHNwZWNpZnkuDQoNCiMjIFNlcXVlbmNlIG9mIFBsYXkNCjEuICoqU2V0dXAqKjogc2h1ZmZsZSBkZWNrLCBkZWFsIGNhcmRzDQoyLiAqKkFudGUgUGhhc2UqKjogZWFjaCBwbGF5ZXIgYW50ZXMgaW50byB0aGUgcG90DQozLiAqKkJldHRpbmcgUGhhc2UqKjogcGxheWVycyBtYXkgYmV0LCBjaGVjaywgZm9sZCwgb3IgY2FsbA0KNC4gKipQb3N0LUJldHRpbmcgUGhhc2UqKg0KNS4gKipTaG93ZG93bioqOiByZXZlYWwgaGFuZHMgYW5kIHNldHRsZSB0aGUgcG90DQoNCiMjIFBoYXNlIERldGFpbHMNCg0KIyMjIFNldHVwDQotIFNodWZmbGUgdGhlIERlY2suDQotIERlYWwgMSBjYXJkIGZhY2UtZG93biB0byBlYWNoIHBsYXllciAoQWxpY2UsIEJvYikgKG9yIGFzIG1hbnkgYXMgcmVtYWluIGlmIHRoZSBkZWNrIGlzIHNob3J0KS4gRWFjaCBwbGF5ZXIgc2VlcyBvbmx5IHRoZWlyIG93biBjYXJkcy4NCg0KIyMjIEFudGUgUGhhc2UNCi0gQWxpY2UgYW50ZXMgMiBjaGlwcyAobW92ZWQgZnJvbSBBbGljZSdzIGNoaXAgc3RhY2sgaW50byB0aGUgcG90KS4NCi0gQm9iIGFudGVzIDIgY2hpcHMgKG1vdmVkIGZyb20gQm9iJ3MgY2hpcCBzdGFjayBpbnRvIHRoZSBwb3QpLg0KDQojIyMgQmV0dGluZyBQaGFzZQ0KLSBCZXR0aW5nIHByb2NlZWRzIGZvciB1cCB0byAyIHJvdW5kcy4NCi0gRWFjaCByb3VuZCwgQmV0dGluZyBSb3VuZCBpcyBwbGF5ZWQuIFRoaXMgY29udGludWVzIHVudGlsIGEgcGxheWVyIGhhcyAyIG9yIGZld2VyIGNoaXBzLCB0aGUgcm91bmQgY2FwICgyKSBoYXMgYmVlbiByZWFjaGVkLCBvciB0aGUgZ2FtZSBoYXMgZW5kZWQuDQoNCiMjIyMgU3ViLXBoYXNlOiBCZXR0aW5nIFJvdW5kICooY2FsbGVkIGZyb20gQmV0dGluZyBQaGFzZSkqDQotIMKrYmV0IHJvdW5kwrsgc3RhcnRzIGF0IDAgYW5kIGluY3JlYXNlcyBieSAxIGFmdGVyIGVhY2ggcm91bmQuDQotIElmIMKrYmV0IHJvdW5kwrsgaXMgZ3JlYXRlciB0aGFuIDAgYW5kIHRoZSBnYW1lIGlzIHN0aWxsIGluIHByb2dyZXNzIChwdWJsaWMpOiBQZXJmb3JtICoqSW50ZXItUm91bmQgRHJhdyAoQWxsIFBsYXllcnMsIGludGVyLXJvdW5kIGV2ZW50IDEpKiogc3RlcHMsIHRoZW4gY29udGludWUuDQotIElmIMKrYmV0IHJvdW5kwrsgaXMgZ3JlYXRlciB0aGFuIDAgYW5kIHRoZSBnYW1lIGlzIHN0aWxsIGluIHByb2dyZXNzIChwdWJsaWMpOiBQZXJmb3JtICoqU2ltdWx0YW5lb3VzIFJvdW5kIChpbnRlci1yb3VuZCkqKiBzdGVwcyAoYm90aCBwbGF5ZXJzIHBhcnRpY2lwYXRlIHBlciB0aGlzIHN1Yi1waGFzZSdzIG93biBydWxlcyksIHRoZW4gY29udGludWUuDQotIElmIMKrYmV0IHJvdW5kwrsgaXMgZXF1YWwgdG8gMCAocHVibGljKToNCiAgICAtIElmIEFsaWNlIGhhcyBjaGlwcyByZW1haW5pbmcgYW5kIEJvYiBoYXMgY2hpcHMgcmVtYWluaW5nIChwdWJsaWMpOg0KICAgICAgICBSb3VuZCAxOiB2YXJpYWJsZSBiZXR0aW5nLCB3YWdlciBhbnkgYW1vdW50IGZyb20gMSBjaGlwIHVwIHRvIHlvdXIgcmVtYWluaW5nIHN0YWNrLg0KICAgICAgICBUaGVuLCBwZXJmb3JtICoqVmFyaWFibGUgQmV0dGluZyBSb3VuZCAoUm91bmQgMSkqKiBzdGVwcywgdGhlbiBjb250aW51ZS4NCiAgICAtIE90aGVyd2lzZToNCiAgICAgICAgUm91bmQgMTogdmFyaWFibGUgYmV0dGluZyB1bmFmZm9yZGFibGUsIGZhbGxpbmcgYmFjayB0byBmaXhlZCAxLWNoaXAgYmV0dGluZyBmb3IgdGhpcyByb3VuZC4NCiAgICAgICAgVGhlbiwgQWxpY2UgcGlja3Mgb25lIG9mOiBiZXQgb3IgY2hlY2suIChjaG9pY2UgaXMgYW5ub3VuY2VkIHRvIGFsbCBwbGF5ZXJzKQ0KICAgICAgICAtIElmIEFsaWNlIGNob29zZXMgJ2NoZWNrJyAocHVibGljKTogUGVyZm9ybSAqKkJvYjogQmV0IG9yIENoZWNrIChSb3VuZCAxLCBGYWxsYmFjaykqKiBzdGVwcywgdGhlbiBjb250aW51ZS4NCiAgICAgICAgLSBPdGhlcndpc2UsIGlmIEFsaWNlIGNob29zZXMgJ2JldCc6DQogICAgICAgICAgICBBbGljZSBiZXRzIDEgY2hpcCAobW92ZWQgZnJvbSBBbGljZSdzIGNoaXAgc3RhY2sgaW50byB0aGUgcG90KS4NCiAgICAgICAgICAgIFRoZW4sIHBlcmZvcm0gKipCb2I6IEZvbGQgb3IgQ2FsbCAoUm91bmQgMSwgRmFsbGJhY2spKiogc3RlcHMsIHRoZW4gY29udGludWUuDQogIC0gT3RoZXJ3aXNlOg0KICAgIFJvdW5kIDI6IGZpeGVkIGJldHRpbmcsIGVhY2ggYmV0IG9yIGNhbGwgaXMgMSBjaGlwLg0KICAgIFRoZW4sIEFsaWNlIHBpY2tzIG9uZSBvZjogYmV0IG9yIGNoZWNrLiAoY2hvaWNlIGlzIGFubm91bmNlZCB0byBhbGwgcGxheWVycykNCiAgICAtIElmIEFsaWNlIGNob29zZXMgJ2NoZWNrJyAocHVibGljKTogUGVyZm9ybSAqKkJvYjogQmV0IG9yIENoZWNrIChSb3VuZCAyKSoqIHN0ZXBzLCB0aGVuIGNvbnRpbnVlLg0KICAgIC0gT3RoZXJ3aXNlLCBpZiBBbGljZSBjaG9vc2VzICdiZXQnOg0KICAgICAgICBBbGljZSBiZXRzIDEgY2hpcCAobW92ZWQgZnJvbSBBbGljZSdzIGNoaXAgc3RhY2sgaW50byB0aGUgcG90KS4NCiAgICAgICAgVGhlbiwgcGVyZm9ybSAqKkJvYjogRm9sZCBvciBDYWxsIChSb3VuZCAyKSoqIHN0ZXBzLCB0aGVuIGNvbnRpbnVlLg0KKihQbGF5IHJldHVybnMgdG8gKipCZXR0aW5nIFBoYXNlKiouKSoNCg0KIyMjIyBTdWItcGhhc2U6IEludGVyLVJvdW5kIERyYXcgKEFsbCBQbGF5ZXJzLCBpbnRlci1yb3VuZCBldmVudCAxKSAqKGNhbGxlZCBmcm9tIEJldHRpbmcgUm91bmQpKg0KLSBJZiB0aGUgRGVjayBoYXMgYXQgbGVhc3QgMiBjYXJkcyAocHVibGljKTogSW50ZXItcm91bmQgZHJhdzogZWFjaCBwbGF5ZXIgZHJhd3MgMSBjYXJkIGZyb20gdGhlIERlY2sgKHByaXZhdGUpLg0KICAtIE90aGVyd2lzZTogSW50ZXItcm91bmQgZHJhdzogdGhlIERlY2sgaGFzIGZld2VyIHRoYW4gMiBjYXJkcyByZW1haW5pbmc7IHRoZSBib3RoLXBsYXllcnMgZHJhdyBpcyBza2lwcGVkLg0KKihQbGF5IHJldHVybnMgdG8gKipCZXR0aW5nIFJvdW5kKiouKSoNCg0KIyMjIyBTdWItcGhhc2U6IFNpbXVsdGFuZW91cyBSb3VuZCAoaW50ZXItcm91bmQpICooY2FsbGVkIGZyb20gQmV0dGluZyBSb3VuZCkqDQotIEFsbCBwbGF5ZXJzIGNob29zZSBzaW11bHRhbmVvdXNseSBhbmQgaW4gc2VjcmV0LiBObyBwbGF5ZXIga25vd3MgdGhlIG90aGVycycgY2hvaWNlcyB1bnRpbCBhbGwgYXJlIGxvY2tlZCBpbi4gT25jZSBhbGwgY2hvaWNlcyBhcmUgc3VibWl0dGVkLCBlYWNoIHBsYXllcidzIGNob2ljZSBpcyBwdWJsaWNseSByZXZlYWxlZCwgdGhlbiB0aGUgY29tYmluZWQgb3V0Y29tZSBpcyByZXNvbHZlZCBhbmQgYW5ub3VuY2VkIHB1YmxpY2x5LiBJZiBhIHBsYXllciBmYWlscyB0byBjaG9vc2UsIGEgdW5pZm9ybWx5IHJhbmRvbSBvcHRpb24gaXMgc2VsZWN0ZWQgZm9yIHRoZW0uDQotIEFsaWNlIGNob29zZXMgb25lIG9mIHRoZSBmb2xsb3dpbmcgb3B0aW9uczogSG9sZCBvciBQYXJyeS4NCi0gQm9iIGNob29zZXMgb25lIG9mIHRoZSBmb2xsb3dpbmcgb3B0aW9uczogSG9sZCBvciBQYXJyeS4NCi0gT3V0Y29tZXMgYmFzZWQgb24gY2hvaWNlczoNCiAgLSBJZiBBbGljZSBjaG9vc2VzICdIb2xkJyBBTkQgQm9iIGNob29zZXMgJ1BhcnJ5JzoNCiAgICAgIC0gSWYgQm9iJ3MgaGlnaGVzdC1yYW5rZWQgY2FyZCBoYXMgYSBoaWdoZXIgcmFuayB0aGFuIEFsaWNlJ3MgaGlnaGVzdC1yYW5rZWQgY2FyZDogQWxpY2UgcGF5cyAzIGNoaXBzIHRvIEJvYi4gSWYgQWxpY2UgaGFzIGZld2VyIHRoYW4gMyBjaGlwcywgb25seSB0aGUgYXZhaWxhYmxlIGFtb3VudCBpcyB0cmFuc2ZlcnJlZC4NCiAgICAgIC0gT3RoZXJ3aXNlOiBCb2IgcGF5cyAxIGNoaXAgdG8gQWxpY2UuIElmIEJvYiBoYXMgZmV3ZXIgdGhhbiAxIGNoaXAsIG9ubHkgdGhlIGF2YWlsYWJsZSBhbW91bnQgaXMgdHJhbnNmZXJyZWQuDQogICAgICAgICAgKihUaGUgYnJhbmNoIGNvbmRpdGlvbiBpcyBldmFsdWF0ZWQgcHJpdmF0ZWx5OyB3aGljaCBicmFuY2ggZmlyZXMgaXMgbm90IGFubm91bmNlZCBkaXJlY3RseS4pKi4NCiAgICAgICAgICAqKE5vdGU6IGRpZmZlcmVudCBicmFuY2hlcyBwcm9kdWNlIGRpZmZlcmVudCBvYnNlcnZhYmxlIGNoaXAgZWZmZWN0czsgdGhlIGJyYW5jaCB0YWtlbiBtYXkgYmUgcGFydGlhbGx5IGluZmVyYWJsZSBmcm9tIGNoaXAgY2hhbmdlcy4pKi4NCiAgLSBJZiBBbGljZSBjaG9vc2VzICdQYXJyeScgQU5EIEJvYiBjaG9vc2VzICdIb2xkJzogQm9iIHBheXMgMSBjaGlwIHRvIEFsaWNlLiBJZiBCb2IgaGFzIGZld2VyIHRoYW4gMSBjaGlwLCBvbmx5IHRoZSBhdmFpbGFibGUgYW1vdW50IGlzIHRyYW5zZmVycmVkLg0KICAtIElmIEFsaWNlIGNob29zZXMgJ1BhcnJ5JyBBTkQgQm9iIGNob29zZXMgJ1BhcnJ5JzogTm8gZWZmZWN0Lg0KICAtIElmIEFsaWNlIGNob29zZXMgJ0hvbGQnIEFORCBCb2IgY2hvb3NlcyAnSG9sZCc6DQogICAgICAtIElmIEJvYiBob2xkcyBhdCBsZWFzdCBvbmUgY2FyZCBvZiByYW5rIEsgb3IgSjoNCiAgICAgICAgICBBbGljZSBwZWVrcyBhdCAxIGNhcmQgaW4gQm9iJ3MgaGFuZCAoc2VsZWN0ZWQgdW5pZm9ybWx5IGF0IHJhbmRvbSkuIFRoZSBjYXJkIHJlbWFpbnMgaW4gQm9iJ3MgaGFuZC4gVGhlIHBlZWsgaXMgcHJpdmF0ZTogb25seSBBbGljZSBzZWVzIHRoZSBjYXJkLiBJZiBCb2IncyBoYW5kIGlzIGVtcHR5LCBBbGljZSBpcyBpbmZvcm1lZCB0aGF0IGl0IGlzIGVtcHR5Lg0KICAgICAgICAgIFRoZW4sIEFsaWNlIHBheXMgMSBjaGlwIHRvIEJvYi4gSWYgQWxpY2UgaGFzIGZld2VyIHRoYW4gMSBjaGlwLCBvbmx5IHRoZSBhdmFpbGFibGUgYW1vdW50IGlzIHRyYW5zZmVycmVkLg0KICAgICAgLSBPdGhlcndpc2U6DQogICAgICAgICAgQWxpY2UgcGVla3MgYXQgMSBjYXJkIGluIEJvYidzIGhhbmQgKHNlbGVjdGVkIHVuaWZvcm1seSBhdCByYW5kb20pLiBUaGUgY2FyZCByZW1haW5zIGluIEJvYidzIGhhbmQuIFRoZSBwZWVrIGlzIHByaXZhdGU6IG9ubHkgQWxpY2Ugc2VlcyB0aGUgY2FyZC4gSWYgQm9iJ3MgaGFuZCBpcyBlbXB0eSwgQWxpY2UgaXMgaW5mb3JtZWQgdGhhdCBpdCBpcyBlbXB0eS4NCiAgICAgICAgICBUaGVuLCBCb2IgcGF5cyAxIGNoaXAgdG8gQWxpY2UuIElmIEJvYiBoYXMgZmV3ZXIgdGhhbiAxIGNoaXAsIG9ubHkgdGhlIGF2YWlsYWJsZSBhbW91bnQgaXMgdHJhbnNmZXJyZWQuDQogICAgICAgICAgKihUaGUgYnJhbmNoIGNvbmRpdGlvbiBpcyBldmFsdWF0ZWQgcHJpdmF0ZWx5OyB3aGljaCBicmFuY2ggZmlyZXMgaXMgbm90IGFubm91bmNlZCBkaXJlY3RseS4pKi4NCiAgICAgICAgICAqKE5vdGU6IGRpZmZlcmVudCBicmFuY2hlcyBwcm9kdWNlIGRpZmZlcmVudCBvYnNlcnZhYmxlIGNoaXAgZWZmZWN0czsgdGhlIGJyYW5jaCB0YWtlbiBtYXkgYmUgcGFydGlhbGx5IGluZmVyYWJsZSBmcm9tIGNoaXAgY2hhbmdlcy4pKi4NCiooUGxheSByZXR1cm5zIHRvICoqQmV0dGluZyBSb3VuZCoqLikqDQoNCiMjIyMgU3ViLXBoYXNlOiBCb2I6IEZvbGQgb3IgQ2FsbCAoUm91bmQgMikgKihjYWxsZWQgZnJvbSBCZXR0aW5nIFJvdW5kKSoNCi0gQm9iIHBpY2tzIG9uZSBvZjogZm9sZCBvciBjYWxsLiAoY2hvaWNlIGlzIGFubm91bmNlZCB0byBhbGwgcGxheWVycykNCiAgLSBJZiBCb2IgY2hvb3NlcyAnY2FsbCcgKHB1YmxpYyk6IEJvYiBiZXRzIDEgY2hpcCAobW92ZWQgZnJvbSBCb2IncyBjaGlwIHN0YWNrIGludG8gdGhlIHBvdCkuDQogIC0gT3RoZXJ3aXNlLCBpZiBCb2IgY2hvb3NlcyAnZm9sZCc6DQogICAgICBUaGUgZW50aXJlIHBvdCBpcyBnaXZlbiB0byBBbGljZS4NCiAgICAgIFRoZW4sIHRoZSBnYW1lIGVuZHM6IEFsaWNlIHdpbnMgYWZ0ZXIgQm9iIGZvbGRzLg0KDQojIyMjIFN1Yi1waGFzZTogQm9iOiBCZXQgb3IgQ2hlY2sgKFJvdW5kIDIpICooY2FsbGVkIGZyb20gQmV0dGluZyBSb3VuZCkqDQotIEJvYiBwaWNrcyBvbmUgb2Y6IGJldCBvciBjaGVjay4gKGNob2ljZSBpcyBhbm5vdW5jZWQgdG8gYWxsIHBsYXllcnMpDQogIC0gSWYgQm9iIGNob29zZXMgJ2NoZWNrJyAocHVibGljKTogTm8gYWN0aW9uLg0KICAtIE90aGVyd2lzZSwgaWYgQm9iIGNob29zZXMgJ2JldCc6DQogICAgICBCb2IgYmV0cyAxIGNoaXAgKG1vdmVkIGZyb20gQm9iJ3MgY2hpcCBzdGFjayBpbnRvIHRoZSBwb3QpLg0KICAgICAgVGhlbiwgcGVyZm9ybSAqKkFsaWNlOiBGb2xkIG9yIENhbGwgKHJlc3BvbmRpbmcsIFJvdW5kIDIpKiogc3RlcHMsIHRoZW4gY29udGludWUuDQoqKFBsYXkgcmV0dXJucyB0byAqKkJldHRpbmcgUm91bmQqKi4pKg0KDQojIyMjIFN1Yi1waGFzZTogQWxpY2U6IEZvbGQgb3IgQ2FsbCAocmVzcG9uZGluZywgUm91bmQgMikgKihjYWxsZWQgZnJvbSBCb2I6IEJldCBvciBDaGVjayAoUm91bmQgMikpKg0KLSBBbGljZSBwaWNrcyBvbmUgb2Y6IGZvbGQgb3IgY2FsbC4gKGNob2ljZSBpcyBhbm5vdW5jZWQgdG8gYWxsIHBsYXllcnMpDQogIC0gSWYgQWxpY2UgY2hvb3NlcyAnY2FsbCcgKHB1YmxpYyk6IEFsaWNlIGJldHMgMSBjaGlwIChtb3ZlZCBmcm9tIEFsaWNlJ3MgY2hpcCBzdGFjayBpbnRvIHRoZSBwb3QpLg0KICAtIE90aGVyd2lzZSwgaWYgQWxpY2UgY2hvb3NlcyAnZm9sZCc6DQogICAgICBUaGUgZW50aXJlIHBvdCBpcyBnaXZlbiB0byBCb2IuDQogICAgICBUaGVuLCB0aGUgZ2FtZSBlbmRzOiBCb2Igd2lucyBhZnRlciBBbGljZSBmb2xkcy4NCg0KIyMjIyBTdWItcGhhc2U6IEJvYjogRm9sZCBvciBDYWxsIChSb3VuZCAxLCBGYWxsYmFjaykgKihjYWxsZWQgZnJvbSBCZXR0aW5nIFJvdW5kKSoNCi0gQm9iIHBpY2tzIG9uZSBvZjogZm9sZCBvciBjYWxsLiAoY2hvaWNlIGlzIGFubm91bmNlZCB0byBhbGwgcGxheWVycykNCiAgLSBJZiBCb2IgY2hvb3NlcyAnY2FsbCcgKHB1YmxpYyk6IEJvYiBiZXRzIDEgY2hpcCAobW92ZWQgZnJvbSBCb2IncyBjaGlwIHN0YWNrIGludG8gdGhlIHBvdCkuDQogIC0gT3RoZXJ3aXNlLCBpZiBCb2IgY2hvb3NlcyAnZm9sZCc6DQogICAgICBUaGUgZW50aXJlIHBvdCBpcyBnaXZlbiB0byBBbGljZS4NCiAgICAgIFRoZW4sIHRoZSBnYW1lIGVuZHM6IEFsaWNlIHdpbnMgYWZ0ZXIgQm9iIGZvbGRzLg0KDQojIyMjIFN1Yi1waGFzZTogQm9iOiBCZXQgb3IgQ2hlY2sgKFJvdW5kIDEsIEZhbGxiYWNrKSAqKGNhbGxlZCBmcm9tIEJldHRpbmcgUm91bmQpKg0KLSBCb2IgcGlja3Mgb25lIG9mOiBiZXQgb3IgY2hlY2suIChjaG9pY2UgaXMgYW5ub3VuY2VkIHRvIGFsbCBwbGF5ZXJzKQ0KICAtIElmIEJvYiBjaG9vc2VzICdjaGVjaycgKHB1YmxpYyk6IE5vIGFjdGlvbi4NCiAgLSBPdGhlcndpc2UsIGlmIEJvYiBjaG9vc2VzICdiZXQnOg0KICAgICAgQm9iIGJldHMgMSBjaGlwIChtb3ZlZCBmcm9tIEJvYidzIGNoaXAgc3RhY2sgaW50byB0aGUgcG90KS4NCiAgICAgIFRoZW4sIHBlcmZvcm0gKipBbGljZTogRm9sZCBvciBDYWxsIChyZXNwb25kaW5nLCBSb3VuZCAxLCBGYWxsYmFjaykqKiBzdGVwcywgdGhlbiBjb250aW51ZS4NCiooUGxheSByZXR1cm5zIHRvICoqQmV0dGluZyBSb3VuZCoqLikqDQoNCiMjIyMgU3ViLXBoYXNlOiBBbGljZTogRm9sZCBvciBDYWxsIChyZXNwb25kaW5nLCBSb3VuZCAxLCBGYWxsYmFjaykgKihjYWxsZWQgZnJvbSBCb2I6IEJldCBvciBDaGVjayAoUm91bmQgMSwgRmFsbGJhY2spKSoNCi0gQWxpY2UgcGlja3Mgb25lIG9mOiBmb2xkIG9yIGNhbGwuIChjaG9pY2UgaXMgYW5ub3VuY2VkIHRvIGFsbCBwbGF5ZXJzKQ0KICAtIElmIEFsaWNlIGNob29zZXMgJ2NhbGwnIChwdWJsaWMpOiBBbGljZSBiZXRzIDEgY2hpcCAobW92ZWQgZnJvbSBBbGljZSdzIGNoaXAgc3RhY2sgaW50byB0aGUgcG90KS4NCiAgLSBPdGhlcndpc2UsIGlmIEFsaWNlIGNob29zZXMgJ2ZvbGQnOg0KICAgICAgVGhlIGVudGlyZSBwb3QgaXMgZ2l2ZW4gdG8gQm9iLg0KICAgICAgVGhlbiwgdGhlIGdhbWUgZW5kczogQm9iIHdpbnMgYWZ0ZXIgQWxpY2UgZm9sZHMuDQoNCiMjIyMgU3ViLXBoYXNlOiBWYXJpYWJsZSBCZXR0aW5nIFJvdW5kIChSb3VuZCAxKSAqKGNhbGxlZCBmcm9tIEJldHRpbmcgUm91bmQpKg0KLSBBbGljZSBhbmQgQm9iIG1heSBlYWNoIGJldCBhbnkgd2hvbGUgbnVtYmVyIG9mIGNoaXBzIChtaW5pbXVtIDEpLiBXaGVuIGEgcGxheWVyIGJldHMsIHRoZWlyIGNob3NlbiBhbW91bnQgaXMgdHJhbnNmZXJyZWQgZnJvbSB0aGVpciBjaGlwIHN0YWNrIGludG8gdGhlIHBvdCBpbW1lZGlhdGVseS4gV2hlbiBhIHBsYXllciBjYWxscywgdGhlaXIgbWF0Y2hlZCBhbW91bnQgaXMgdHJhbnNmZXJyZWQgdG8gdGhlIHBvdDsgaWYgdGhleSBjYW5ub3QgbWF0Y2ggdGhlIGZ1bGwgYmV0LCB0aGV5IGdvIGFsbC1pbiB3aXRoIHRoZWlyIHJlbWFpbmluZyBjaGlwcyBhbmQgdGhlIGV4Y2VzcyBmcm9tIHRoZSBvcmlnaW5hbCBiZXR0b3IgaXMgcmV0dXJuZWQgdG8gdGhhdCBiZXR0b3IuIElmIGJvdGggcGxheWVycyBjaGVjayBpbiBzZXF1ZW5jZSwgdGhlIHJvdW5kIGVuZHMgd2l0aCBubyBhZGRpdGlvbmFsIGNoaXBzIG1vdmluZy4NCi0gQWxpY2UgcGlja3Mgb25lIG9mOiBiZXQgb3IgY2hlY2suIChjaG9pY2UgaXMgYW5ub3VuY2VkIHRvIGFsbCBwbGF5ZXJzKQ0KICAtIElmIEFsaWNlIGNob29zZXMgJ2NoZWNrJyAocHVibGljKTogUGVyZm9ybSAqKkJvYjogQmV0IG9yIENoZWNrIChSb3VuZCAxKSoqIHN0ZXBzLCB0aGVuIGNvbnRpbnVlLg0KICAtIE90aGVyd2lzZSwgaWYgQWxpY2UgY2hvb3NlcyAnYmV0JzoNCiAgICAgIEFsaWNlIGNob29zZXMgYSB3YWdlciBhbW91bnQgKG1pbmltdW0gMSBjaGlwLCB1cCB0byBBbGljZSdzIGN1cnJlbnQgY2hpcCBjb3VudCkuIFdoZW4gQWxpY2UgY29tbWl0cyB0aGUgY2hvc2VuIGFtb3VudCwgaXQgaXMgdHJhbnNmZXJyZWQgZnJvbSBBbGljZSdzIGNoaXAgc3RhY2sgaW50byB0aGUgcG90IHZpYSB0aGUgbm9ybWFsIHB1YmxpYyBjaGlwIGFubm91bmNlbWVudC4gSWYgdGhlIHN1Ym1pdHRlZCBhbW91bnQgZXhjZWVkcyB0aGUgbGVnYWwgbWF4aW11bSwgdGhlIGJpZCBpcyBjbGFtcGVkIHRvIHRoZSBtYXhpbXVtIChhbGwtaW4gb3IgZGVjbGFyZWQgY2FwKS4gSWYgdGhlIHN1Ym1pdHRlZCBhbW91bnQgaXMgbWFsZm9ybWVkIG9yIGJlbG93IHRoZSBtaW5pbXVtLCB0aGUgZW5naW5lIHN1YnN0aXR1dGVzIGEgdW5pZm9ybS1yYW5kb20gbGVnYWwgYW1vdW50LiBJbiBlaXRoZXIgY2FzZSB0aGUgc3VibWl0dGluZyBwbGF5ZXIgaXMgbm90aWZpZWQgcHJpdmF0ZWx5IGFuZCBwbGF5IGNvbnRpbnVlcyB3aXRoIHRoZSByZXNvbHZlZCBhbW91bnQuDQogICAgICBUaGVuLCBwZXJmb3JtICoqQm9iOiBGb2xkIG9yIENhbGwgKFJvdW5kIDEpKiogc3RlcHMsIHRoZW4gY29udGludWUuDQoqKFBsYXkgcmV0dXJucyB0byAqKkJldHRpbmcgUm91bmQqKi4pKg0KDQojIyMjIFN1Yi1waGFzZTogQm9iOiBGb2xkIG9yIENhbGwgKFJvdW5kIDEpICooY2FsbGVkIGZyb20gVmFyaWFibGUgQmV0dGluZyBSb3VuZCAoUm91bmQgMSkpKg0KLSBCb2IgcGlja3Mgb25lIG9mOiBmb2xkIG9yIGNhbGwuIChjaG9pY2UgaXMgYW5ub3VuY2VkIHRvIGFsbCBwbGF5ZXJzKQ0KICAtIElmIEJvYiBjaG9vc2VzICdjYWxsJyAocHVibGljKTogQm9iIG1hdGNoZXMgdGhlIGJldC4gSWYgdGhleSBjYW5ub3QgbWF0Y2ggdGhlIGZ1bGwgYW1vdW50LCB0aGV5IGdvIGFsbC1pbiBhbmQgb25seSB0aGUgbWF0Y2hlZCBwb3J0aW9uIGZyb20gZWFjaCBwbGF5ZXIgcmVtYWlucyBpbiB0aGUgcG90Lg0KICAtIE90aGVyd2lzZSwgaWYgQm9iIGNob29zZXMgJ2ZvbGQnOg0KICAgICAgVGhlIGVudGlyZSBwb3QgaXMgZ2l2ZW4gdG8gQWxpY2UuDQogICAgICBUaGVuLCB0aGUgZ2FtZSBlbmRzOiBBbGljZSB3aW5zIGFmdGVyIEJvYiBmb2xkcy4NCg0KIyMjIyBTdWItcGhhc2U6IEJvYjogQmV0IG9yIENoZWNrIChSb3VuZCAxKSAqKGNhbGxlZCBmcm9tIFZhcmlhYmxlIEJldHRpbmcgUm91bmQgKFJvdW5kIDEpKSoNCi0gQm9iIHBpY2tzIG9uZSBvZjogYmV0IG9yIGNoZWNrLiAoY2hvaWNlIGlzIGFubm91bmNlZCB0byBhbGwgcGxheWVycykNCiAgLSBJZiBCb2IgY2hvb3NlcyAnY2hlY2snIChwdWJsaWMpOiBObyBhY3Rpb24uDQogIC0gT3RoZXJ3aXNlLCBpZiBCb2IgY2hvb3NlcyAnYmV0JzoNCiAgICAgIEJvYiBjaG9vc2VzIGEgd2FnZXIgYW1vdW50IChtaW5pbXVtIDEgY2hpcCwgdXAgdG8gQm9iJ3MgY3VycmVudCBjaGlwIGNvdW50KS4gV2hlbiBCb2IgY29tbWl0cyB0aGUgY2hvc2VuIGFtb3VudCwgaXQgaXMgdHJhbnNmZXJyZWQgZnJvbSBCb2IncyBjaGlwIHN0YWNrIGludG8gdGhlIHBvdCB2aWEgdGhlIG5vcm1hbCBwdWJsaWMgY2hpcCBhbm5vdW5jZW1lbnQuIElmIHRoZSBzdWJtaXR0ZWQgYW1vdW50IGV4Y2VlZHMgdGhlIGxlZ2FsIG1heGltdW0sIHRoZSBiaWQgaXMgY2xhbXBlZCB0byB0aGUgbWF4aW11bSAoYWxsLWluIG9yIGRlY2xhcmVkIGNhcCkuIElmIHRoZSBzdWJtaXR0ZWQgYW1vdW50IGlzIG1hbGZvcm1lZCBvciBiZWxvdyB0aGUgbWluaW11bSwgdGhlIGVuZ2luZSBzdWJzdGl0dXRlcyBhIHVuaWZvcm0tcmFuZG9tIGxlZ2FsIGFtb3VudC4gSW4gZWl0aGVyIGNhc2UgdGhlIHN1Ym1pdHRpbmcgcGxheWVyIGlzIG5vdGlmaWVkIHByaXZhdGVseSBhbmQgcGxheSBjb250aW51ZXMgd2l0aCB0aGUgcmVzb2x2ZWQgYW1vdW50Lg0KICAgICAgVGhlbiwgcGVyZm9ybSAqKkFsaWNlOiBGb2xkIG9yIENhbGwgKHJlc3BvbmRpbmcsIFJvdW5kIDEpKiogc3RlcHMsIHRoZW4gY29udGludWUuDQoqKFBsYXkgcmV0dXJucyB0byAqKlZhcmlhYmxlIEJldHRpbmcgUm91bmQgKFJvdW5kIDEpKiouKSoNCg0KIyMjIyBTdWItcGhhc2U6IEFsaWNlOiBGb2xkIG9yIENhbGwgKHJlc3BvbmRpbmcsIFJvdW5kIDEpICooY2FsbGVkIGZyb20gQm9iOiBCZXQgb3IgQ2hlY2sgKFJvdW5kIDEpKSoNCi0gQWxpY2UgcGlja3Mgb25lIG9mOiBmb2xkIG9yIGNhbGwuIChjaG9pY2UgaXMgYW5ub3VuY2VkIHRvIGFsbCBwbGF5ZXJzKQ0KICAtIElmIEFsaWNlIGNob29zZXMgJ2NhbGwnIChwdWJsaWMpOiBBbGljZSBtYXRjaGVzIHRoZSBiZXQuIElmIHRoZXkgY2Fubm90IG1hdGNoIHRoZSBmdWxsIGFtb3VudCwgdGhleSBnbyBhbGwtaW4gYW5kIG9ubHkgdGhlIG1hdGNoZWQgcG9ydGlvbiBmcm9tIGVhY2ggcGxheWVyIHJlbWFpbnMgaW4gdGhlIHBvdC4NCiAgLSBPdGhlcndpc2UsIGlmIEFsaWNlIGNob29zZXMgJ2ZvbGQnOg0KICAgICAgVGhlIGVudGlyZSBwb3QgaXMgZ2l2ZW4gdG8gQm9iLg0KICAgICAgVGhlbiwgdGhlIGdhbWUgZW5kczogQm9iIHdpbnMgYWZ0ZXIgQWxpY2UgZm9sZHMuDQoNCiMjIyBQb3N0LUJldHRpbmcgUGhhc2UNCi0gUG9zdC1iZXR0aW5nIGV2ZW50cyBtYXkgZmlyZSBiYXNlZCBvbiBnYW1lIHN0YXRlLiBFYWNoIG9mIHRoZSBmb2xsb3dpbmcgY29uZGl0aW9ucyBpcyBldmFsdWF0ZWQgdG9wLXRvLWJvdHRvbSBhdCBpdHMgb3duIHN0ZXA7IGV2ZXJ5IGNvbmRpdGlvbiB3aG9zZSBjaGVjayBwYXNzZXMgYXQgdGhhdCBtb21lbnQgcGVyZm9ybXMgaXRzIGFzc29jaWF0ZWQgYWN0aW9uLiBUaGUgY29uZGl0aW9ucyBhcmUgbm90IG11dHVhbGx5IGV4Y2x1c2l2ZTogbW9yZSB0aGFuIG9uZSBtYXkgZmlyZSBpbiB0aGUgc2FtZSBwaGFzZS4NCi0gSWYgdGhlIGN1cnJlbnQgcG90IGlzIG1vcmUgdGhhbiAyIGNoaXBzIChwdWJsaWMpOiBQZXJmb3JtICoqU2ltdWx0YW5lb3VzIFJvdW5kIChwb3N0LWJldHRpbmcpKiogc3RlcHMgKGJvdGggcGxheWVycyBwYXJ0aWNpcGF0ZSBwZXIgdGhpcyBzdWItcGhhc2UncyBvd24gcnVsZXMpLCB0aGVuIGNvbnRpbnVlLg0KLSBJZiBBbGljZSdzIGhhbmQgY29udGFpbnMgYSBLaW5nICh0aGUgZGVjaydzIHRvcCByYW5rKSAocHVibGljKTogUGVyZm9ybSAqKkF1Y3Rpb24gUm91bmQ6IENhcmQgRHJhdyAocG9zdC1iZXR0aW5nKSoqIHN0ZXBzIChib3RoIHBsYXllcnMgcGFydGljaXBhdGUgcGVyIHRoaXMgc3ViLXBoYXNlJ3Mgb3duIHJ1bGVzKSwgdGhlbiBjb250aW51ZS4NCg0KIyMjIyBTdWItcGhhc2U6IFNpbXVsdGFuZW91cyBSb3VuZCAocG9zdC1iZXR0aW5nKSAqKGNhbGxlZCBmcm9tIFBvc3QtQmV0dGluZyBQaGFzZSkqDQotIEFsbCBwbGF5ZXJzIGNob29zZSBzaW11bHRhbmVvdXNseSBhbmQgaW4gc2VjcmV0LiBObyBwbGF5ZXIga25vd3MgdGhlIG90aGVycycgY2hvaWNlcyB1bnRpbCBhbGwgYXJlIGxvY2tlZCBpbi4gT25jZSBhbGwgY2hvaWNlcyBhcmUgc3VibWl0dGVkLCBlYWNoIHBsYXllcidzIGNob2ljZSBpcyBwdWJsaWNseSByZXZlYWxlZCwgdGhlbiB0aGUgY29tYmluZWQgb3V0Y29tZSBpcyByZXNvbHZlZCBhbmQgYW5ub3VuY2VkIHB1YmxpY2x5LiBJZiBhIHBsYXllciBmYWlscyB0byBjaG9vc2UsIGEgdW5pZm9ybWx5IHJhbmRvbSBvcHRpb24gaXMgc2VsZWN0ZWQgZm9yIHRoZW0uDQotIEFsaWNlIGNob29zZXMgb25lIG9mIHRoZSBmb2xsb3dpbmcgb3B0aW9uczogUHJlc3Mgb3IgUmV0cmVhdC4NCi0gQm9iIGNob29zZXMgb25lIG9mIHRoZSBmb2xsb3dpbmcgb3B0aW9uczogUHJlc3Mgb3IgUmV0cmVhdC4NCi0gT3V0Y29tZXMgYmFzZWQgb24gY2hvaWNlczoNCiAgLSBJZiBBbGljZSBjaG9vc2VzICdQcmVzcycgQU5EIEJvYiBjaG9vc2VzICdSZXRyZWF0JzogQm9iIHBlZWtzIGF0IDEgY2FyZCBpbiBBbGljZSdzIGhhbmQgKHNlbGVjdGVkIHVuaWZvcm1seSBhdCByYW5kb20pLiBUaGUgY2FyZCByZW1haW5zIGluIEFsaWNlJ3MgaGFuZC4gVGhlIHBlZWsgaXRzZWxmIGlzIGFubm91bmNlZCB0byBhbGwgcGxheWVycyAob2NjdXJyZW5jZSBpcyBwdWJsaWMpOyBvbmx5IEJvYiBzZWVzIHdoaWNoIGNhcmQgd2FzIG9ic2VydmVkIChjb250ZW50IGlzIHByaXZhdGUpLiBJZiBBbGljZSdzIGhhbmQgaXMgZW1wdHksIEJvYiBpcyBpbmZvcm1lZCB0aGF0IGl0IGlzIGVtcHR5Lg0KICAtIElmIEFsaWNlIGNob29zZXMgJ1JldHJlYXQnIEFORCBCb2IgY2hvb3NlcyAnUHJlc3MnOg0KICAgICAgLSBJZiBBbGljZSdzIGhpZ2hlc3QtcmFua2VkIGNhcmQgaGFzIGEgaGlnaGVyIHJhbmsgdGhhbiBCb2IncyBoaWdoZXN0LXJhbmtlZCBjYXJkOiBCb2IgcGF5cyAyIGNoaXBzIHRvIEFsaWNlLiBJZiBCb2IgaGFzIGZld2VyIHRoYW4gMiBjaGlwcywgb25seSB0aGUgYXZhaWxhYmxlIGFtb3VudCBpcyB0cmFuc2ZlcnJlZC4NCiAgICAgIC0gT3RoZXJ3aXNlOiBBbGljZSBwYXlzIDEgY2hpcCB0byBCb2IuIElmIEFsaWNlIGhhcyBmZXdlciB0aGFuIDEgY2hpcCwgb25seSB0aGUgYXZhaWxhYmxlIGFtb3VudCBpcyB0cmFuc2ZlcnJlZC4NCiAgICAgICAgICAqKFRoZSBicmFuY2ggY29uZGl0aW9uIGlzIGV2YWx1YXRlZCBwcml2YXRlbHk7IHdoaWNoIGJyYW5jaCBmaXJlcyBpcyBub3QgYW5ub3VuY2VkIGRpcmVjdGx5LikqLg0KICAgICAgICAgICooTm90ZTogZGlmZmVyZW50IGJyYW5jaGVzIHByb2R1Y2UgZGlmZmVyZW50IG9ic2VydmFibGUgY2hpcCBlZmZlY3RzOyB0aGUgYnJhbmNoIHRha2VuIG1heSBiZSBwYXJ0aWFsbHkgaW5mZXJhYmxlIGZyb20gY2hpcCBjaGFuZ2VzLikqLg0KICAtIEZvciBhbGwgb3RoZXIgY29tYmluYXRpb25zIChQcmVzcy9QcmVzcywgUmV0cmVhdC9SZXRyZWF0KTogQm9iIGRyYXdzIDEgY2FyZCBmcm9tIHRoZSBEZWNrLiBCb2Igc2VlcyB0aGUgY2FyZCBpZGVudGl0eTsgdGhlIERlY2stc2l6ZSBkZWNyZW1lbnQgaXMgcHVibGljLCBzaWduYWxpbmcgdG8gYm90aCBwbGF5ZXJzIHRoYXQgYSBkcmF3IG9jY3VycmVkLg0KKihQbGF5IHJldHVybnMgdG8gKipQb3N0LUJldHRpbmcgUGhhc2UqKi4pKg0KDQojIyMjIFN1Yi1waGFzZTogQXVjdGlvbiBSb3VuZDogQ2FyZCBEcmF3IChwb3N0LWJldHRpbmcpICooY2FsbGVkIGZyb20gUG9zdC1CZXR0aW5nIFBoYXNlKSoNCi0gU2VhbGVkLWJpZCBhdWN0aW9uLiBFYWNoIHBsYXllciBzZWNyZXRseSBwaWNrcyBhIGJpZCBmcm9tIDAgdG8gMyBjaGlwczsgYSBiaWQgbGFyZ2VyIHRoYW4gdGhlIHBsYXllcidzIGN1cnJlbnQgY2hpcCBzdGFjayBpcyBmaXJzdCBjbGFtcGVkIGRvd24gdG8gdGhlIHN0YWNrIHNpemUsIHNvIHRoZWlyIGVmZmVjdGl2ZSBiaWQgKGFuZCB0aGUgYW1vdW50IHRoZXkgcGF5IGlmIHRoZXkgd2luKSBuZXZlciBleGNlZWRzIHdoYXQgdGhleSBob2xkIChhbGwtaW4pLiBCb3RoIGJpZHMgYXJlIHRoZW4gcmV2ZWFsZWQgYXQgdGhlIHNhbWUgbW9tZW50LiBJZiBib3RoIGVmZmVjdGl2ZSBiaWRzIGFyZSAwLCBubyBvbmUgd2lucyBhbmQgbm8gY2hpcHMgbW92ZS4gT3RoZXJ3aXNlIHRoZSBoaWdoZXIgZWZmZWN0aXZlIGJpZCB3aW5zLCB3aXRoIHRpZXMgYmV0d2VlbiBub256ZXJvIGJpZHMgYnJva2VuIGJ5IGEgY29pbiBmbGlwOyB0aGUgd2lubmVyIHBheXMgdGhlaXIgZWZmZWN0aXZlIGJpZCBpbnRvIHRoZSBwb3QgYW5kIGRyYXdzIDEgY2FyZCBmcm9tIHRoZSBkZWNrLiBJZiB0aGUgZGVjayBpcyBlbXB0eSBhdCB0aGUgbW9tZW50IHRoZSBwcml6ZSBpcyB0byBiZSBwZXJmb3JtZWQsIHRoZSBhdWN0aW9uIGlzIGNhbmNlbGxlZDogdGhlIHdpbm5lciBkb2VzIG5vdCBwYXkgdGhlaXIgYmlkIGFuZCBubyBjaGlwcyBtb3ZlLiBBdWN0aW9ucyBkbyBub3QgY29sbGVjdCBhbiBhbnRlLg0KLSBBbGljZSBjaG9vc2VzIGEgYmlkIGFtb3VudCAobWluaW11bSAwIGNoaXBzLCB1cCB0byAzIGNoaXBzKS4gVGhlIGNob3NlbiBhbW91bnQgaXMgcmVjb3JkZWQgYXMgQWxpY2UncyBiaWQgYnV0IGlzIG5vdCB0cmFuc2ZlcnJlZCB5ZXQ7IGEgc3Vic2VxdWVudCByZXNvbHV0aW9uIHN0ZXAgZGVjaWRlcyB3aGV0aGVyIGNoaXBzIGFjdHVhbGx5IG1vdmUgKGZvciBleGFtcGxlLCBhIHNlYWxlZC1iaWQgYXVjdGlvbiBvbmx5IGNoYXJnZXMgdGhlIHdpbm5lcikuIElmIHRoZSBzdWJtaXR0ZWQgYW1vdW50IGV4Y2VlZHMgdGhlIGxlZ2FsIG1heGltdW0sIHRoZSBiaWQgaXMgY2xhbXBlZCB0byB0aGUgbWF4aW11bSAoYWxsLWluIG9yIGRlY2xhcmVkIGNhcCkuIElmIHRoZSBzdWJtaXR0ZWQgYW1vdW50IGlzIG1hbGZvcm1lZCBvciBiZWxvdyB0aGUgbWluaW11bSwgdGhlIGVuZ2luZSBzdWJzdGl0dXRlcyBhIHVuaWZvcm0tcmFuZG9tIGxlZ2FsIGFtb3VudC4gSW4gZWl0aGVyIGNhc2UgdGhlIHN1Ym1pdHRpbmcgcGxheWVyIGlzIG5vdGlmaWVkIHByaXZhdGVseSBhbmQgcGxheSBjb250aW51ZXMgd2l0aCB0aGUgcmVzb2x2ZWQgYW1vdW50Lg0KLSBCb2IgY2hvb3NlcyBhIGJpZCBhbW91bnQgKG1pbmltdW0gMCBjaGlwcywgdXAgdG8gMyBjaGlwcykuIFRoZSBjaG9zZW4gYW1vdW50IGlzIHJlY29yZGVkIGFzIEJvYidzIGJpZCBidXQgaXMgbm90IHRyYW5zZmVycmVkIHlldDsgYSBzdWJzZXF1ZW50IHJlc29sdXRpb24gc3RlcCBkZWNpZGVzIHdoZXRoZXIgY2hpcHMgYWN0dWFsbHkgbW92ZSAoZm9yIGV4YW1wbGUsIGEgc2VhbGVkLWJpZCBhdWN0aW9uIG9ubHkgY2hhcmdlcyB0aGUgd2lubmVyKS4gSWYgdGhlIHN1Ym1pdHRlZCBhbW91bnQgZXhjZWVkcyB0aGUgbGVnYWwgbWF4aW11bSwgdGhlIGJpZCBpcyBjbGFtcGVkIHRvIHRoZSBtYXhpbXVtIChhbGwtaW4gb3IgZGVjbGFyZWQgY2FwKS4gSWYgdGhlIHN1Ym1pdHRlZCBhbW91bnQgaXMgbWFsZm9ybWVkIG9yIGJlbG93IHRoZSBtaW5pbXVtLCB0aGUgZW5naW5lIHN1YnN0aXR1dGVzIGEgdW5pZm9ybS1yYW5kb20gbGVnYWwgYW1vdW50LiBJbiBlaXRoZXIgY2FzZSB0aGUgc3VibWl0dGluZyBwbGF5ZXIgaXMgbm90aWZpZWQgcHJpdmF0ZWx5IGFuZCBwbGF5IGNvbnRpbnVlcyB3aXRoIHRoZSByZXNvbHZlZCBhbW91bnQuDQoqKFBsYXkgcmV0dXJucyB0byAqKlBvc3QtQmV0dGluZyBQaGFzZSoqLikqDQoNCiMjIyBTaG93ZG93bg0KLSBCb3RoIHBsYXllcnMnIGhhbmRzIGFyZSByZXZlYWxlZCB0byBhbGwgcGxheWVycy4gU2hvd2Rvd24gcnVsZXM6DQogIC0gKipDb21wYXJpc29uOioqIGNvbXBhcmUgY2FyZHMgYnkgcmFuaywgaGlnaGVzdCBmaXJzdC4gVGhlIGZpcnN0IHJhbmsgdGhhdCBkaWZmZXJzIGRldGVybWluZXMgdGhlIHdpbm5lci4NCiAgLSAqKlRpZXM6KiogaWYgaGlnaGVzdCBjYXJkcyBtYXRjaCwgY29tcGFyZSBuZXh0LWhpZ2hlc3QgKGFuZCBzbyBvbikuIElmIGFsbCBjb21wYXJlZCByYW5rcyB0aWUgYW5kIGJvdGggaGFuZHMgcnVuIG91dCBhdCB0aGUgc2FtZSB0aW1lLCB0aGUgcG90IGlzIHNwbGl0Lg0KICAtICoqVW5lcXVhbCBoYW5kIHNpemVzOioqIGlmIG9uZSBwbGF5ZXIgcnVucyBvdXQgb2YgY2FyZHMgdG8gY29tcGFyZSBmaXJzdCwgdGhlIHBsYXllciB3aXRoIHRoZSByZW1haW5pbmcgY2FyZCB3aW5zIHRoYXQgc3RlcC4NCiAgLSAqKkVtcHR5IGhhbmQ6KiogYW4gZW1wdHkgaGFuZCBsb3NlcyB0byBhbnkgbm9uLWVtcHR5IGhhbmQ7IHR3byBlbXB0eSBoYW5kcyB0aWUuDQogIC0gKFJhbmsgdmFsdWVzOiBKPTExLCBRPTEyLCBLPTEzLikNCi0gVGhlIHNob3dkb3duIHdpbm5lciB0YWtlcyB0aGUgZW50aXJlIHBvdC4gT24gYSB0aWUsIGVhY2ggcGxheWVyIHJlY2VpdmVzIGhhbGYgdGhlIHBvdCAocm91bmRlZCBkb3duKTsgYW55IHJlbWFpbmluZyBvZGQgY2hpcCBpcyBhd2FyZGVkIHVuaWZvcm1seSBhdCByYW5kb20gdG8gb25lIHRpZWQgcGxheWVyLg0KLSBUaGUgZ2FtZSBlbmRzOiBTaG93ZG93biBjb21wbGV0ZS4NCg==)\#\#Overview\*\*Players:\*\*Alice,Bob\*\*Playerorder:\*\*Player1=Alice\(actsfirstinbetting\),Player2=Bob\.\*\*Deck:\*\*3cards\(J,Q,Kofhearts\)\*\*Startingchips:\*\*10each\(visibletoall\)\.Chipsmovebetweenplayers'stacksandthepot;thetotalofallchips\(stackspluspot\)isconservedacrossthegame\.\*\*Pot:\*\*startsat0\*\*Generalrules:\*\*\-\*\*Information:\*\*ThenumberofcardsremainingintheDeckispublicinformation\.Handsizesareprivate\.Playersdonotknowhowmanycardsopponentsholdunlessrevealedbyanaction\.IftheDeckisemptywhenadrawactionoccurs,nocardsaredrawnandtheactionhasnoeffect\.Publicdrawsemitapublicskipped\-drawnotice;privatedrawsemitthatnoticeonlytothedrawingplayer—butpertheinferabilityprinciple,theDeck\-sizechangefromanysuccessfuldrawispubliclyvisible\.\-\*\*Legalactions:\*\*Playersmaychooseanylistedoption\.Ifitspreconditionsareunmet\(emptyhand\),theactionhasnoeffect\.\-\*\*Resolution:\*\*Inanymulti\-stepeffect,eachstepresolvesinorder;ifonestep'spreconditionisunmet,thatstepisskippedandlaterstepsstillresolve\.\-\*\*Gameend:\*\*Ifthegameendsduringanyphase,allremainingphasesareskipped\.\-\*\*Costs&pot:\*\*Costsarepaidtothepotbeforetheeffectresolves,unlessotherwisespecified\.Whenanactionlistsacost:iftheplayercannotaffordthestatedcostatthemomenttheyselectit,theactionisnotofferedasanoption\.Ifthecostisprobabilistic\(e\.g\.a50%chanceofpaying1chip\),theplayercommitstotheactionfirstandthecostisresolvedafterwards;iftheycannotaffordtheresolvedcost,theactionisskippedandtheirchoiceslotisconsumed\(theydonotgettopickagain\)\.\-\*\*Showdown&misc\.:\*\*Thereisnohandsizelimit;handsgrowandshrinkfreelyviadraws\.Unlessstatedotherwise,allrandomselectionsareuniform\.AtShowdown,eachplayer'shandisrevealedasofthemomenttheShowdownphasebegins,includinganycardsswapped,drawn,orstolenduringearlierphases\.\#\#Visibility\-\*\*Hands\*\*:Eachplayer'scardsareprivatetothatplayer\.Opponentscannotseehandcontentsunlessexplicitlyrevealed\.\-\*\*Actions\*\*:Allplayeractions\(bets,checks,folds\)andtheireffectsarepubliclyannouncedunlessstatedotherwise\.Actionsmarkedasprivateareexecutedsilently\.\-\*\*Inferability\*\*:"Private"hides\*whichbranchfired,whatcardwasdrawn/peeked,orwhochosewhat\*—nottheobservableside\-effects\.Decksizes,pottotals,chipstacks,andpositionholdingsarealwayspublicandreconcileatsettlement\.Ifoneprivatebranchwouldchangeapublicnumber\(e\.g\.shrinktheDeckby1\)andanotherwouldnot,theopponentcaninferwhichbranchranfromthatnumber\.Treattheprivacylabelasconcealingthe\*identity/choice\*,notthe\*occurrence\*\.\-\*\*Conditionchecks\*\*:Checksmarked"\(public\)"areannouncedtoallplayerswhenevaluated\.\-\*\*Hiddeninformation\*\*:Allhiddeninformationistrackedandenforced;playersaretoldonlywhattherulesspecify\.\#\#SequenceofPlay1\.\*\*Setup\*\*:shuffledeck,dealcards2\.\*\*AntePhase\*\*:eachplayerantesintothepot3\.\*\*BettingPhase\*\*:playersmaybet,check,fold,orcall4\.\*\*Post\-BettingPhase\*\*5\.\*\*Showdown\*\*:revealhandsandsettlethepot\#\#PhaseDetails\#\#\#Setup\-ShuffletheDeck\.\-Deal1cardface\-downtoeachplayer\(Alice,Bob\)\(orasmanyasremainifthedeckisshort\)\.Eachplayerseesonlytheirowncards\.\#\#\#AntePhase\-Aliceantes2chips\(movedfromAlice'schipstackintothepot\)\.\-Bobantes2chips\(movedfromBob'schipstackintothepot\)\.\#\#\#BettingPhase\-Bettingproceedsforupto2rounds\.\-Eachround,BettingRoundisplayed\.Thiscontinuesuntilaplayerhas2orfewerchips,theroundcap\(2\)hasbeenreached,orthegamehasended\.\#\#\#\#Sub\-phase:BettingRound\*\(calledfromBettingPhase\)\*\-«betround»startsat0andincreasesby1aftereachround\.\-If«betround»isgreaterthan0andthegameisstillinprogress\(public\):Perform\*\*Inter\-RoundDraw\(AllPlayers,inter\-roundevent1\)\*\*steps,thencontinue\.\-If«betround»isgreaterthan0andthegameisstillinprogress\(public\):Perform\*\*SimultaneousRound\(inter\-round\)\*\*steps\(bothplayersparticipateperthissub\-phase'sownrules\),thencontinue\.\-If«betround»isequalto0\(public\):\-IfAlicehaschipsremainingandBobhaschipsremaining\(public\):Round1:variablebetting,wageranyamountfrom1chipuptoyourremainingstack\.Then,perform\*\*VariableBettingRound\(Round1\)\*\*steps,thencontinue\.\-Otherwise:Round1:variablebettingunaffordable,fallingbacktofixed1\-chipbettingforthisround\.Then,Alicepicksoneof:betorcheck\.\(choiceisannouncedtoallplayers\)\-IfAlicechooses'check'\(public\):Perform\*\*Bob:BetorCheck\(Round1,Fallback\)\*\*steps,thencontinue\.\-Otherwise,ifAlicechooses'bet':Alicebets1chip\(movedfromAlice'schipstackintothepot\)\.Then,perform\*\*Bob:FoldorCall\(Round1,Fallback\)\*\*steps,thencontinue\.\-Otherwise:Round2:fixedbetting,eachbetorcallis1chip\.Then,Alicepicksoneof:betorcheck\.\(choiceisannouncedtoallplayers\)\-IfAlicechooses'check'\(public\):Perform\*\*Bob:BetorCheck\(Round2\)\*\*steps,thencontinue\.\-Otherwise,ifAlicechooses'bet':Alicebets1chip\(movedfromAlice'schipstackintothepot\)\.Then,perform\*\*Bob:FoldorCall\(Round2\)\*\*steps,thencontinue\.\*\(Playreturnsto\*\*BettingPhase\*\*\.\)\*\#\#\#\#Sub\-phase:Inter\-RoundDraw\(AllPlayers,inter\-roundevent1\)\*\(calledfromBettingRound\)\*\-IftheDeckhasatleast2cards\(public\):Inter\-rounddraw:eachplayerdraws1cardfromtheDeck\(private\)\.\-Otherwise:Inter\-rounddraw:theDeckhasfewerthan2cardsremaining;theboth\-playersdrawisskipped\.\*\(Playreturnsto\*\*BettingRound\*\*\.\)\*\#\#\#\#Sub\-phase:SimultaneousRound\(inter\-round\)\*\(calledfromBettingRound\)\*\-Allplayerschoosesimultaneouslyandinsecret\.Noplayerknowstheothers'choicesuntilallarelockedin\.Onceallchoicesaresubmitted,eachplayer'schoiceispubliclyrevealed,thenthecombinedoutcomeisresolvedandannouncedpublicly\.Ifaplayerfailstochoose,auniformlyrandomoptionisselectedforthem\.\-Alicechoosesoneofthefollowingoptions:HoldorParry\.\-Bobchoosesoneofthefollowingoptions:HoldorParry\.\-Outcomesbasedonchoices:\-IfAlicechooses'Hold'ANDBobchooses'Parry':\-IfBob'shighest\-rankedcardhasahigherrankthanAlice'shighest\-rankedcard:Alicepays3chipstoBob\.IfAlicehasfewerthan3chips,onlytheavailableamountistransferred\.\-Otherwise:Bobpays1chiptoAlice\.IfBobhasfewerthan1chip,onlytheavailableamountistransferred\.\*\(Thebranchconditionisevaluatedprivately;whichbranchfiresisnotannounceddirectly\.\)\*\.\*\(Note:differentbranchesproducedifferentobservablechipeffects;thebranchtakenmaybepartiallyinferablefromchipchanges\.\)\*\.\-IfAlicechooses'Parry'ANDBobchooses'Hold':Bobpays1chiptoAlice\.IfBobhasfewerthan1chip,onlytheavailableamountistransferred\.\-IfAlicechooses'Parry'ANDBobchooses'Parry':Noeffect\.\-IfAlicechooses'Hold'ANDBobchooses'Hold':\-IfBobholdsatleastonecardofrankKorJ:Alicepeeksat1cardinBob'shand\(selecteduniformlyatrandom\)\.ThecardremainsinBob'shand\.Thepeekisprivate:onlyAliceseesthecard\.IfBob'shandisempty,Aliceisinformedthatitisempty\.Then,Alicepays1chiptoBob\.IfAlicehasfewerthan1chip,onlytheavailableamountistransferred\.\-Otherwise:Alicepeeksat1cardinBob'shand\(selecteduniformlyatrandom\)\.ThecardremainsinBob'shand\.Thepeekisprivate:onlyAliceseesthecard\.IfBob'shandisempty,Aliceisinformedthatitisempty\.Then,Bobpays1chiptoAlice\.IfBobhasfewerthan1chip,onlytheavailableamountistransferred\.\*\(Thebranchconditionisevaluatedprivately;whichbranchfiresisnotannounceddirectly\.\)\*\.\*\(Note:differentbranchesproducedifferentobservablechipeffects;thebranchtakenmaybepartiallyinferablefromchipchanges\.\)\*\.\*\(Playreturnsto\*\*BettingRound\*\*\.\)\*\#\#\#\#Sub\-phase:Bob:FoldorCall\(Round2\)\*\(calledfromBettingRound\)\*\-Bobpicksoneof:foldorcall\.\(choiceisannouncedtoallplayers\)\-IfBobchooses'call'\(public\):Bobbets1chip\(movedfromBob'schipstackintothepot\)\.\-Otherwise,ifBobchooses'fold':TheentirepotisgiventoAlice\.Then,thegameends:AlicewinsafterBobfolds\.\#\#\#\#Sub\-phase:Bob:BetorCheck\(Round2\)\*\(calledfromBettingRound\)\*\-Bobpicksoneof:betorcheck\.\(choiceisannouncedtoallplayers\)\-IfBobchooses'check'\(public\):Noaction\.\-Otherwise,ifBobchooses'bet':Bobbets1chip\(movedfromBob'schipstackintothepot\)\.Then,perform\*\*Alice:FoldorCall\(responding,Round2\)\*\*steps,thencontinue\.\*\(Playreturnsto\*\*BettingRound\*\*\.\)\*\#\#\#\#Sub\-phase:Alice:FoldorCall\(responding,Round2\)\*\(calledfromBob:BetorCheck\(Round2\)\)\*\-Alicepicksoneof:foldorcall\.\(choiceisannouncedtoallplayers\)\-IfAlicechooses'call'\(public\):Alicebets1chip\(movedfromAlice'schipstackintothepot\)\.\-Otherwise,ifAlicechooses'fold':TheentirepotisgiventoBob\.Then,thegameends:BobwinsafterAlicefolds\.\#\#\#\#Sub\-phase:Bob:FoldorCall\(Round1,Fallback\)\*\(calledfromBettingRound\)\*\-Bobpicksoneof:foldorcall\.\(choiceisannouncedtoallplayers\)\-IfBobchooses'call'\(public\):Bobbets1chip\(movedfromBob'schipstackintothepot\)\.\-Otherwise,ifBobchooses'fold':TheentirepotisgiventoAlice\.Then,thegameends:AlicewinsafterBobfolds\.\#\#\#\#Sub\-phase:Bob:BetorCheck\(Round1,Fallback\)\*\(calledfromBettingRound\)\*\-Bobpicksoneof:betorcheck\.\(choiceisannouncedtoallplayers\)\-IfBobchooses'check'\(public\):Noaction\.\-Otherwise,ifBobchooses'bet':Bobbets1chip\(movedfromBob'schipstackintothepot\)\.Then,perform\*\*Alice:FoldorCall\(responding,Round1,Fallback\)\*\*steps,thencontinue\.\*\(Playreturnsto\*\*BettingRound\*\*\.\)\*\#\#\#\#Sub\-phase:Alice:FoldorCall\(responding,Round1,Fallback\)\*\(calledfromBob:BetorCheck\(Round1,Fallback\)\)\*\-Alicepicksoneof:foldorcall\.\(choiceisannouncedtoallplayers\)\-IfAlicechooses'call'\(public\):Alicebets1chip\(movedfromAlice'schipstackintothepot\)\.\-Otherwise,ifAlicechooses'fold':TheentirepotisgiventoBob\.Then,thegameends:BobwinsafterAlicefolds\.\#\#\#\#Sub\-phase:VariableBettingRound\(Round1\)\*\(calledfromBettingRound\)\*\-AliceandBobmayeachbetanywholenumberofchips\(minimum1\)\.Whenaplayerbets,theirchosenamountistransferredfromtheirchipstackintothepotimmediately\.Whenaplayercalls,theirmatchedamountistransferredtothepot;iftheycannotmatchthefullbet,theygoall\-inwiththeirremainingchipsandtheexcessfromtheoriginalbettorisreturnedtothatbettor\.Ifbothplayerscheckinsequence,theroundendswithnoadditionalchipsmoving\.\-Alicepicksoneof:betorcheck\.\(choiceisannouncedtoallplayers\)\-IfAlicechooses'check'\(public\):Perform\*\*Bob:BetorCheck\(Round1\)\*\*steps,thencontinue\.\-Otherwise,ifAlicechooses'bet':Alicechoosesawageramount\(minimum1chip,uptoAlice'scurrentchipcount\)\.WhenAlicecommitsthechosenamount,itistransferredfromAlice'schipstackintothepotviathenormalpublicchipannouncement\.Ifthesubmittedamountexceedsthelegalmaximum,thebidisclampedtothemaximum\(all\-inordeclaredcap\)\.Ifthesubmittedamountismalformedorbelowtheminimum,theenginesubstitutesauniform\-randomlegalamount\.Ineithercasethesubmittingplayerisnotifiedprivatelyandplaycontinueswiththeresolvedamount\.Then,perform\*\*Bob:FoldorCall\(Round1\)\*\*steps,thencontinue\.\*\(Playreturnsto\*\*BettingRound\*\*\.\)\*\#\#\#\#Sub\-phase:Bob:FoldorCall\(Round1\)\*\(calledfromVariableBettingRound\(Round1\)\)\*\-Bobpicksoneof:foldorcall\.\(choiceisannouncedtoallplayers\)\-IfBobchooses'call'\(public\):Bobmatchesthebet\.Iftheycannotmatchthefullamount,theygoall\-inandonlythematchedportionfromeachplayerremainsinthepot\.\-Otherwise,ifBobchooses'fold':TheentirepotisgiventoAlice\.Then,thegameends:AlicewinsafterBobfolds\.\#\#\#\#Sub\-phase:Bob:BetorCheck\(Round1\)\*\(calledfromVariableBettingRound\(Round1\)\)\*\-Bobpicksoneof:betorcheck\.\(choiceisannouncedtoallplayers\)\-IfBobchooses'check'\(public\):Noaction\.\-Otherwise,ifBobchooses'bet':Bobchoosesawageramount\(minimum1chip,uptoBob'scurrentchipcount\)\.WhenBobcommitsthechosenamount,itistransferredfromBob'schipstackintothepotviathenormalpublicchipannouncement\.Ifthesubmittedamountexceedsthelegalmaximum,thebidisclampedtothemaximum\(all\-inordeclaredcap\)\.Ifthesubmittedamountismalformedorbelowtheminimum,theenginesubstitutesauniform\-randomlegalamount\.Ineithercasethesubmittingplayerisnotifiedprivatelyandplaycontinueswiththeresolvedamount\.Then,perform\*\*Alice:FoldorCall\(responding,Round1\)\*\*steps,thencontinue\.\*\(Playreturnsto\*\*VariableBettingRound\(Round1\)\*\*\.\)\*\#\#\#\#Sub\-phase:Alice:FoldorCall\(responding,Round1\)\*\(calledfromBob:BetorCheck\(Round1\)\)\*\-Alicepicksoneof:foldorcall\.\(choiceisannouncedtoallplayers\)\-IfAlicechooses'call'\(public\):Alicematchesthebet\.Iftheycannotmatchthefullamount,theygoall\-inandonlythematchedportionfromeachplayerremainsinthepot\.\-Otherwise,ifAlicechooses'fold':TheentirepotisgiventoBob\.Then,thegameends:BobwinsafterAlicefolds\.\#\#\#Post\-BettingPhase\-Post\-bettingeventsmayfirebasedongamestate\.Eachofthefollowingconditionsisevaluatedtop\-to\-bottomatitsownstep;everyconditionwhosecheckpassesatthatmomentperformsitsassociatedaction\.Theconditionsarenotmutuallyexclusive:morethanonemayfireinthesamephase\.\-Ifthecurrentpotismorethan2chips\(public\):Perform\*\*SimultaneousRound\(post\-betting\)\*\*steps\(bothplayersparticipateperthissub\-phase'sownrules\),thencontinue\.\-IfAlice'shandcontainsaKing\(thedeck'stoprank\)\(public\):Perform\*\*AuctionRound:CardDraw\(post\-betting\)\*\*steps\(bothplayersparticipateperthissub\-phase'sownrules\),thencontinue\.\#\#\#\#Sub\-phase:SimultaneousRound\(post\-betting\)\*\(calledfromPost\-BettingPhase\)\*\-Allplayerschoosesimultaneouslyandinsecret\.Noplayerknowstheothers'choicesuntilallarelockedin\.Onceallchoicesaresubmitted,eachplayer'schoiceispubliclyrevealed,thenthecombinedoutcomeisresolvedandannouncedpublicly\.Ifaplayerfailstochoose,auniformlyrandomoptionisselectedforthem\.\-Alicechoosesoneofthefollowingoptions:PressorRetreat\.\-Bobchoosesoneofthefollowingoptions:PressorRetreat\.\-Outcomesbasedonchoices:\-IfAlicechooses'Press'ANDBobchooses'Retreat':Bobpeeksat1cardinAlice'shand\(selecteduniformlyatrandom\)\.ThecardremainsinAlice'shand\.Thepeekitselfisannouncedtoallplayers\(occurrenceispublic\);onlyBobseeswhichcardwasobserved\(contentisprivate\)\.IfAlice'shandisempty,Bobisinformedthatitisempty\.\-IfAlicechooses'Retreat'ANDBobchooses'Press':\-IfAlice'shighest\-rankedcardhasahigherrankthanBob'shighest\-rankedcard:Bobpays2chipstoAlice\.IfBobhasfewerthan2chips,onlytheavailableamountistransferred\.\-Otherwise:Alicepays1chiptoBob\.IfAlicehasfewerthan1chip,onlytheavailableamountistransferred\.\*\(Thebranchconditionisevaluatedprivately;whichbranchfiresisnotannounceddirectly\.\)\*\.\*\(Note:differentbranchesproducedifferentobservablechipeffects;thebranchtakenmaybepartiallyinferablefromchipchanges\.\)\*\.\-Forallothercombinations\(Press/Press,Retreat/Retreat\):Bobdraws1cardfromtheDeck\.Bobseesthecardidentity;theDeck\-sizedecrementispublic,signalingtobothplayersthatadrawoccurred\.\*\(Playreturnsto\*\*Post\-BettingPhase\*\*\.\)\*\#\#\#\#Sub\-phase:AuctionRound:CardDraw\(post\-betting\)\*\(calledfromPost\-BettingPhase\)\*\-Sealed\-bidauction\.Eachplayersecretlypicksabidfrom0to3chips;abidlargerthantheplayer'scurrentchipstackisfirstclampeddowntothestacksize,sotheireffectivebid\(andtheamounttheypayiftheywin\)neverexceedswhattheyhold\(all\-in\)\.Bothbidsarethenrevealedatthesamemoment\.Ifbotheffectivebidsare0,noonewinsandnochipsmove\.Otherwisethehighereffectivebidwins,withtiesbetweennonzerobidsbrokenbyacoinflip;thewinnerpaystheireffectivebidintothepotanddraws1cardfromthedeck\.Ifthedeckisemptyatthemomenttheprizeistobeperformed,theauctioniscancelled:thewinnerdoesnotpaytheirbidandnochipsmove\.Auctionsdonotcollectanante\.\-Alicechoosesabidamount\(minimum0chips,upto3chips\)\.ThechosenamountisrecordedasAlice'sbidbutisnottransferredyet;asubsequentresolutionstepdecideswhetherchipsactuallymove\(forexample,asealed\-bidauctiononlychargesthewinner\)\.Ifthesubmittedamountexceedsthelegalmaximum,thebidisclampedtothemaximum\(all\-inordeclaredcap\)\.Ifthesubmittedamountismalformedorbelowtheminimum,theenginesubstitutesauniform\-randomlegalamount\.Ineithercasethesubmittingplayerisnotifiedprivatelyandplaycontinueswiththeresolvedamount\.\-Bobchoosesabidamount\(minimum0chips,upto3chips\)\.ThechosenamountisrecordedasBob'sbidbutisnottransferredyet;asubsequentresolutionstepdecideswhetherchipsactuallymove\(forexample,asealed\-bidauctiononlychargesthewinner\)\.Ifthesubmittedamountexceedsthelegalmaximum,thebidisclampedtothemaximum\(all\-inordeclaredcap\)\.Ifthesubmittedamountismalformedorbelowtheminimum,theenginesubstitutesauniform\-randomlegalamount\.Ineithercasethesubmittingplayerisnotifiedprivatelyandplaycontinueswiththeresolvedamount\.\*\(Playreturnsto\*\*Post\-BettingPhase\*\*\.\)\*\#\#\#Showdown\-Bothplayers'handsarerevealedtoallplayers\.Showdownrules:\-\*\*Comparison:\*\*comparecardsbyrank,highestfirst\.Thefirstrankthatdiffersdeterminesthewinner\.\-\*\*Ties:\*\*ifhighestcardsmatch,comparenext\-highest\(andsoon\)\.Ifallcomparedrankstieandbothhandsrunoutatthesametime,thepotissplit\.\-\*\*Unequalhandsizes:\*\*ifoneplayerrunsoutofcardstocomparefirst,theplayerwiththeremainingcardwinsthatstep\.\-\*\*Emptyhand:\*\*anemptyhandlosestoanynon\-emptyhand;twoemptyhandstie\.\-\(Rankvalues:J=11,Q=12,K=13\.\)\-Theshowdownwinnertakestheentirepot\.Onatie,eachplayerreceiveshalfthepot\(roundeddown\);anyremainingoddchipisawardeduniformlyatrandomtoonetiedplayer\.\-Thegameends:Showdowncomplete\.

#### Per\-turn observation prompt\.

Each time it is an agent’s turn, the engine sends the model a single user message that contains three pieces of information\. The first describes the current state of the game: which phase we are in, the agent’s own hand, any public cards on the board, every player’s chip stack, the current pot, role or position assignments, and who the opponents are\. The second is a chronological log of everything that has happened in the hand from the agent’s point of view: every public event plus every private observation the agent was entitled to see\. The third describes the decision the agent has to make: the move type, the menu of legal actions allowed by the current phase, and the format the response should take\. The agent is asked to reply with a single line of JSON of the form`\{"action": <chosen\_option\>\}`\. Handling of malformed responses and the resulting fallback rate are described in more detail in Appendix[F](https://arxiv.org/html/2605.23238#A6)\.

## Appendix CComplexity axes \(formal\)

Each axis is a single scalar computed from a fixed precise\-tier measurement budget per game\. We run three thousand random\-play episodes, denoted L0, and one thousand five hundred episodes of an L1 best\-response policy to L0, used for the policy\-perturbation axes\. The opponent\-modeling axis uses three hundred twenty opponent policies drawn from a low\-discrepancy Sobol sequence on the strategy simplex, which we refer to as the*Sobol budget*\.555A Sobol sequence is a low\-discrepancy alternative to uniform random sampling that gives more even coverage of the simplex at a smaller sample size\. The 2,000\-game candidate pool is scored at a faster tier \(one thousand L0 episodes per game and sixty\-four Sobol opponent policies instead of the precise tier’s three hundred twenty\) so that farthest\-point sampling can run at scale, and the 50 benchmark games are then re\-scored at the precise tier\. The precise values are the ones used throughout\. Pool selection is unaffected because benchmark seeds were already chosen under fast\-tier coordinates and the precise\-tier re\-score only refines them\.For each axis below we give the closed\-form score plus one sentence on what it measures and how to read it, with symbols defined where they appear\.

We use a single notational convention throughout\. An information stateIp=\(d​t,hand,action path,chip bin,signals,roles\)I\_\{p\}=\(dt,\\text\{hand\},\\text\{action path\},\\text\{chip bin\},\\text\{signals\},\\text\{roles\}\)encodes a player’s full visible context, with decision typed​t=\(phase,move type,move name\)dt=\(\\text\{phase\},\\text\{move type\},\\text\{move name\}\)as its first component, sod​t​\(Ip\)dt\(I\_\{p\}\)recovers the decision type\. Where an axis aggregates at the coarserd​tdtlevel \(because its inner quantity is defined per decision type\), we sum overd​tdt\. Where an axis is naturally defined per information state \(because the inner quantity, e\.g\. EV\-vs\-floor, is meaningful only at the situation level\), we sum over visitedIpI\_\{p\}\.

#### State space \(log10\\log\_\{10\}\)\.

A measure of how many distinct observable contexts a player can reach\. We approximate it by a Monte\-Carlo draw rather than by exhaustive enumeration\. Across the three thousand random\-play episodes that make up the L0 budget, we record every information stateIpI\_\{p\}that each player actually visits and count the number of distinct ones\.

state​\_​spacelog10=log10⁡\(∑p\|ℐp\|\),\\mathrm\{state\\\_space\}\_\{\\log\_\{10\}\}\\;=\\;\\log\_\{10\}\\\!\\left\(\\textstyle\\sum\_\{p\}\|\\mathcal\{I\}\_\{p\}\|\\right\),whereℐp\\mathcal\{I\}\_\{p\}is the set of distinct information states visited by playerppacross the L0 episode log\. The count is a lower bound on the true reachable state space, since states the random\-play distribution does not reach within three thousand episodes are undercounted\. In practice the random\-play distribution covers the contexts that an LLM is likely to face in play\. Larger values correspond to more distinct contexts visited\. The base\-ten logarithm compresses a hundred\-fold range in the raw count into one unit on the reported axis\.

#### Temporal depth\.

Whether decisions early in an episode have downstream consequences, i\.e\. whether the player must plan forward rather than play myopically\. Indexing decision types byd=\(phase,move type,move name\)d=\(\\text\{phase\},\\text\{move type\},\\text\{move name\}\), we measure three per\-type quantities on the L0 episode log:

- •fd=nd/nepsf\_\{d\}=n\_\{d\}/n\_\{\\mathrm\{eps\}\}, the average number of times decision typeddfires per episode, wherendn\_\{d\}counts visits andnepsn\_\{\\mathrm\{eps\}\}is the episode count\.
- •ηd2\\eta\_\{d\}^\{2\}, the share of final\-payoff variance attributable to which action is chosen atdd\. WritingUeU\_\{e\}for the focal player’s final payoff on episodeee,U¯a,d\\bar\{U\}\_\{a,d\}for the mean payoff over episodes that took actionaaatdd, andU¯d\\bar\{U\}\_\{d\}for the overall mean atdd, the standard one\-way\-ANOVA explained\-variance ratio is ηd2=∑ana,d​\(U¯a,d−U¯d\)2∑e\(Ue−U¯d\)2\.\\eta\_\{d\}^\{2\}\\;=\\;\\frac\{\\sum\_\{a\}n\_\{a,d\}\\,\(\\bar\{U\}\_\{a,d\}\-\\bar\{U\}\_\{d\}\)^\{2\}\}\{\\sum\_\{e\}\(U\_\{e\}\-\\bar\{U\}\_\{d\}\)^\{2\}\}\.
- •rdr\_\{d\}, the mean number of decisions the same player faces afterddwithin the same episode\.

The game\-level temporal\-depth score sums the product across decision types,

temporal​\_​depth=∑dfd​ηd2​rd,\\mathrm\{temporal\\\_depth\}\\;=\\;\\sum\_\{d\}f\_\{d\}\\,\\eta\_\{d\}^\{2\}\\,r\_\{d\},so a decision type contributes more when it fires often, has a strong action\-choice signal, and leaves many subsequent decisions to follow\. The reported axis is the raw score, notzz\-scored\. It is non\-negative and grows with both signal strength and lookahead horizon\. Across the 50 benchmark games the raw score ranges from below0\.010\.01, on myopic games such as canonical Kuhn poker, to approximately0\.760\.76on the deepest multi\-round games in the collection, with a median of about0\.110\.11\.

#### Information sensitivity\.

How often the player’s optimal action depends on their private information\.

info​\_​sens=∑IpfIp∑Ip′fIp′​1​\[arg⁡maxa⁡𝔼​\[U∣Ip,a\]≠arg⁡maxa⁡𝔼​\[U∣d​t​\(Ip\),a\]\],\\mathrm\{info\\\_sens\}\\;=\\;\\sum\_\{I\_\{p\}\}\\frac\{f\_\{I\_\{p\}\}\}\{\\sum\_\{I^\{\\prime\}\_\{p\}\}f\_\{I^\{\\prime\}\_\{p\}\}\}\\;\\mathbf\{1\}\\\!\\Big\[\\,\\arg\\max\_\{a\}\\mathbb\{E\}\[U\\mid I\_\{p\},a\]\\;\\neq\\;\\arg\\max\_\{a\}\\mathbb\{E\}\[U\\mid dt\(I\_\{p\}\),a\]\\,\\Big\],where the sum is over visited information statesIpI\_\{p\}andfIpf\_\{I\_\{p\}\}is the number of L0 episodes in whichIpI\_\{p\}was reached \(so visited states are weighted in proportion to how often they actually arise in play\)\. The indicator inside the sum compares the argmax action when conditioning on the full information stateIpI\_\{p\}to the argmax action when conditioning only on the decision typed​t​\(Ip\)dt\(I\_\{p\}\)\. A high score therefore means the player must condition their action on private information, while a low score means the same best action works across most information states at the same decision type\. The score is bounded between zero and one\.

#### Opponent modeling\.

How often the player’s optimal action depends on which opponent they face, i\.e\., whether best\-response choice is stable across opponent policies or flips\.

opp​\_​mod=∑d​twd​t​\(1−modal​\_​shareπ​\[arg⁡maxa⁡𝔼​\[U∣d​t,a,π\]\]\),\\mathrm\{opp\\\_mod\}\\;=\\;\\sum\_\{dt\}w\_\{dt\}\\,\\Big\(1\\;\-\\;\\mathrm\{modal\\\_share\}\_\{\\pi\}\\big\[\\,\\arg\\max\_\{a\}\\mathbb\{E\}\[U\\mid dt,a,\\pi\]\\,\\big\]\\Big\),wherewd​t=n​\(d​t\)/∑d​t′n​\(d​t′\)w\_\{dt\}=n\(dt\)/\\sum\_\{dt^\{\\prime\}\}n\(dt^\{\\prime\}\)is the visit fraction of decision typed​tdtandπ\\piranges over three hundred twenty opponent policies drawn from a Sobol sequence on the strategy simplex\. Of these, sixty\-four are drawn from a single Sobol pass that gives even global coverage of the simplex\. The remaining two hundred fifty\-six are drawn from a second Sobol pass concentrated on regions of the simplex where the modal best response from the first sixty\-four policies was unstable, so that more opponent diversity is used to test best\-response stability where it appears most fragile\. We run thirty\-two playout episodes against each sampled opponent\. The functionmodal​\_​share\\mathrm\{modal\\\_share\}is the share of opponents for whom the same action is the argmax\. The score is bounded between zero and one\. Zero means one action is the best response across every opponent, so the player has no need to model the opponent, and higher values mean the player must condition their strategy on opponent behavior\.

#### Risk\.

The visit\-weighted cost, expressed in payoff standard deviations, of switching from the expected\-value\-maximizing action to the action that maximizes the worst\-decile payoff floor:

risk=∑IpwIp​EV​\(aIp∗\)−EV​\(asafe,Ip\)σU,\\mathrm\{risk\}\\;=\\;\\sum\_\{I\_\{p\}\}w\_\{I\_\{p\}\}\\,\\frac\{\\mathrm\{EV\}\(a^\{\*\}\_\{I\_\{p\}\}\)\-\\mathrm\{EV\}\(a\_\{\\mathrm\{safe\},\\,I\_\{p\}\}\)\}\{\\sigma\_\{U\}\},where the sum runs over visited information statesIpI\_\{p\}with at least twenty visits and at least two actions each tried at least five times,wIpw\_\{I\_\{p\}\}is the visit fraction ofIpI\_\{p\}, and the score is then averaged across both player seats\.666For visit\-density reasons, the risk\-axis aggregator coarsensIpI\_\{p\}to its strategic\-context subset\(d​t,hand,action path,chip bin\)\(dt,\\text\{hand\},\\text\{action path\},\\text\{chip bin\}\), dropping signals and roles\. The fullIpI\_\{p\}would fragment the per\-bucket sample below the minimum\-visit threshold on most games\.Within anIpI\_\{p\},aIp∗=arg⁡maxa⁡EV​\(a\)a^\{\*\}\_\{I\_\{p\}\}=\\arg\\max\_\{a\}\\mathrm\{EV\}\(a\)is the expected\-value\-maximizing action andasafe,Ip=arg⁡maxa⁡q0\.10​\(a\)a\_\{\\mathrm\{safe\},\\,I\_\{p\}\}=\\arg\\max\_\{a\}q\_\{0\.10\}\(a\)is the action that maximizes the tenth\-percentile payoff floor\. TheIpI\_\{p\}contributes zero whenevera∗=asafea^\{\*\}=a\_\{\\mathrm\{safe\}\}, i\.e\., when no expected\-value\-versus\-floor tradeoff is present\. The denominatorσU\\sigma\_\{U\}is the standard deviation of the realized utilityUUacross L0 episodes\. Values below0\.010\.01are reported as zero\. The score is non\-negative, with larger values indicating that the expected\-value\-maximizing action exposes the player to substantially worse worst\-decile outcomes than the safest alternative\. We aggregate atIpI\_\{p\}rather than atd​tdtbecause the EV\-vs\-floor tradeoff is only meaningful conditional on a specific hand and action path\. Averaging EVs across heterogeneous information sets at the samed​tdtwould mix qualitatively different decisions into a single bucket\.

#### Brittleness \(log10\\log\_\{10\}\)\.

How much a small perturbation to the focal player’s L1 policy at a single decision type swings realized payoff\.

brittlenesslog10=log10⁡\(∑d​twd​t​\|β^d​t\|σU\),\\mathrm\{brittleness\}\_\{\\log\_\{10\}\}\\;=\\;\\log\_\{10\}\\\!\\Bigg\(\\sum\_\{dt\}w\_\{dt\}\\,\\frac\{\|\\hat\{\\beta\}\_\{dt\}\|\}\{\\sigma\_\{U\}\}\\Bigg\),whereβ^d​t\\hat\{\\beta\}\_\{dt\}is the OLS slope of the per\-trial payoff changeΔ​U\\Delta Uregressed on a perturbation indicator, equal to one ifd​tdtwas the perturbed decision type and zero otherwise\. We run twenty such trials per game\. In each trial, three percent of the focal player’s L1 policy mass at one decision type is shifted to a uniform\-random alternative action, and the opponent is held fixed at ten Dirichlet\-random policies \(uniform random points on the probability simplex\) with fifteen playout episodes per opponent\. The score is then averaged across the two focal seats\. Larger values correspond to small policy perturbations producing larger payoff swings\. Values near zero correspond to a payoff swing of one standard deviation per unit of policy\-mass perturbation\.

## Appendix DVariance\-inflation factors for the six axes

The variance inflation factor, or VIF \(see Section[3\.1](https://arxiv.org/html/2605.23238#S3.SS1)\), is a standard collinearity diagnostic\. For each axis, we regress it on the other five \(across the 50 benchmark games\) and reportVIF=1/\(1−R2\)\\mathrm\{VIF\}=1/\(1\-R^\{2\}\), whereR2R^\{2\}is the coefficient of determination of that regression\. A VIF of one means the axis is linearly independent of the others, and a VIF of five is the conventional threshold above which collinearity is severe enough to inflate joint\-regression standard errors\. The per\-axis values on the 50\-game benchmark are as follows\.

All six VIFs are below the conventional collinearity threshold of five\. The two largest are state space, at3\.313\.31, and information sensitivity, at2\.702\.70\. This is consistent with the positive correlation between those two axes: richer state spaces tend to create more contexts in which private information shifts the best action\. Risk and brittleness are near1\.21\.2and are nearly independent of the other axes\.

![Refer to caption](https://arxiv.org/html/2605.23238v1/x5.png)Figure 5:Pairwise Pearson correlation of the six complexity axes across the 50 benchmark games\.The strongest positive pair is state space and information sensitivity, with Pearsonr=0\.65r=0\.65, followed by state space and temporal depth atr=0\.57r=0\.57and information sensitivity and opponent modeling atr=0\.50r=0\.50\. Risk and brittleness are nearly independent of the remaining axes\. In both cases the absolute correlation with every other axis is at most0\.340\.34\. All correlations stay below the conventional\|r\|≈0\.7\|r\|\\approx 0\.7threshold above which joint\-regression standard errors begin to inflate substantially, and the per\-axis VIFs reported in the table above are all below four\.
## Appendix EFPS coverage scatter

![Refer to caption](https://arxiv.org/html/2605.23238v1/x6.png)Figure 6:The 2,000\-game accepted pool vs\. the 50 FPS\-selected benchmark games in two diagnostic axes \(information sensitivity×\\timesopponent modeling\)\.
## Appendix FFallback\-rate diagnostics

A*fallback*is a per\-move event in which the engine could not recover a legal action from the model’s reply\. Each model response is first passed through a strict JSON parser\. If that fails, a lenient parser tries to recover the chosen action from common malformations \(stray whitespace, trailing commentary, partial JSON\)\. If both parsers fail, the engine substitutes a uniformly random legal action so play does not stall, and the move is recorded as a fallback\. Each tournament slot records per\-seat counts of fallbacks and total moves\. We aggregate to model\-level rates over the merged 9\-model tournament data\.

Move rate is the number of fallback\-triggered moves divided by total moves\. Slot rate is the fraction of slots with at least one fallback move\.

The overall observation is that fallback rates are uniformly low: the worst offender \(gpt\-5\) triggers the parser fallback on0\.5%0\.5\\%of moves, and four of nine models record zero fallbacks across the set of 50 games\. The leaderboard ordering is therefore not driven by parser\-format compliance differences\.

## Appendix GVariance decomposition

#### Per\-game strength estimates\.

We decompose variance across \(model, game\) pairs using per\-game strength estimatesα^m,g\\hat\{\\alpha\}\_\{m,g\}from a per\-game refit of the additive paired\-comparison model, rather than the raw mean win margin of modelmmon gamegg\. The refit subtracts off the opponent’s strength ongg, so a model that happened to draw weaker opponents on gameggis not credited with a higher per\-game strength on that account\. The raw mean would carry that opponent\-mix bias straight into the decomposition\. For each gamegg, we restrict the slot\-level rows to matches played onggand refit

ys=αi\(s\),g−αj\(s\),g\+εs,∑mαm,g=0,y\_\{s\}=\\alpha\_\{i^\{\(s\)\},g\}\-\\alpha\_\{j^\{\(s\)\},g\}\+\\varepsilon\_\{s\},\\qquad\\sum\_\{m\}\\alpha\_\{m,g\}=0,by OLS on the matches played ongg\. The resultingα^m,g\\hat\{\\alpha\}\_\{m,g\}is modelmm’s strength on gameggadjusted for which opponents it faced on that game \(the same adjustment the overallα^m\\hat\{\\alpha\}\_\{m\}makes across the full tournament data\)\. All9×50=4509\{\\times\}50=450\(model, game\) pairs are identified on the full tournament data; in the 8\-model \(llama\-excluded\) subset394394of400400\(model, game\) pairs are identified, with the 6 unidentified pairs being \(model, game\) pairs whose only matches were against llama\.

#### Decomposition\.

Letα¯m⁣⋅\\bar\{\\alpha\}\_\{m\\cdot\}be the per\-model average ofα^m,g\\hat\{\\alpha\}\_\{m,g\}across games,α¯\\bar\{\\alpha\}the grand mean, andα¯⋅g\\bar\{\\alpha\}\_\{\\cdot g\}the per\-game average across models \(which is identically zero by the sum\-to\-zero contrast, so the game main effectσG2=1\|G\|​∑g\(α¯⋅g−α¯\)2\\sigma^\{2\}\_\{G\}=\\tfrac\{1\}\{\|G\|\}\\sum\_\{g\}\(\\bar\{\\alpha\}\_\{\\cdot g\}\-\\bar\{\\alpha\}\)^\{2\}is zero by construction and we omit it\)\. The two non\-trivial components are777The technique is standard two\-way variance decomposition into row main effect, column main effect, and row×\\timescolumn interaction\. The per\-model row average around the grand mean givesσM2\\sigma^\{2\}\_\{M\}, and the residual cell variation after subtracting both row and column averages givesσM​G2\\sigma^\{2\}\_\{MG\}\.

σM2=1\|M\|​∑m\(α¯m⁣⋅−α¯\)2,σM​G2=1\|M\|​\|G\|​∑m,g\(α^m,g−α¯m⁣⋅−α¯⋅g\+α¯\)2\.\\sigma^\{2\}\_\{M\}=\\frac\{1\}\{\|M\|\}\\sum\_\{m\}\(\\bar\{\\alpha\}\_\{m\\cdot\}\-\\bar\{\\alpha\}\)^\{2\},\\qquad\\sigma^\{2\}\_\{MG\}=\\frac\{1\}\{\|M\|\|G\|\}\\sum\_\{m,g\}\(\\hat\{\\alpha\}\_\{m,g\}\-\\bar\{\\alpha\}\_\{m\\cdot\}\-\\bar\{\\alpha\}\_\{\\cdot g\}\+\\bar\{\\alpha\}\)^\{2\}\.All9×50=4509\\times 50=450\(model, game\) pairs are identified on the full tournament data, so the decomposition is computed on the balanced grid\. In the 8\-model llama\-excluded subset,394394of400400pairs are identified \(the six unidentified pairs are \(model, game\) pairs whose only matches on that game were against llama\)\.

We compute 95% confidence intervals from a within\-cell bootstrap of two thousand replicates\. A cell is a \(model, game\) pair; for each cell we resample its contributing slot edges with replacement to the original cell size, refit the per\-\(model, game\) cell mean, and re\-decompose the resulting variance components\. This differs from the paired\-cluster bootstrap on \(game seed, run id\) clusters that we use elsewhere in the paper forα^\\hat\{\\alpha\}and the capability\-profile slopes, and we use the within\-cell design here specifically because the paired\-cluster bootstrap is biased upward for variance\-components targets: replicates in which some cells end up with fewer\-than\-original slot counts inflate apparent model×\\timesgame interaction variance\. The within\-cell design preserves cell coverage exactly on each replicate\. We report the bias\-corrected percentile interval\.

The interaction\-vs\-model\-main\-effect ratio is0\.490\.49on the full 9\-model tournament data and1\.291\.29when llama is excluded\. With llama the leaderboard signal clearly dominates \(CI well below11\); without llama, interaction is roughly comparable to the model main effect \(CI straddles11\), so we cannot determine which is larger at the available power\. See Section[6](https://arxiv.org/html/2605.23238#S6)\(*Variance decomposition*\) for discussion in main text\.

## Appendix HRobustness checks

#### Leave\-one\-game\-out\.

We refitα^\\hat\{\\alpha\}dropping each game in turn\. Kendallτ\\taubetween the leave\-one\-out \(LOO\) ranking and the full\-data ranking has mean0\.9980\.998, minimum0\.9440\.944, with48/5048/50LOO refits preserving the overall ranking exactly\.

#### α^\\hat\{\\alpha\}vs\. coverage\.

To check whether models with more slots receive systematically higher or lowerα^\\hat\{\\alpha\}, we plot per\-modelα^\\hat\{\\alpha\}against the number of slotsnslotsn\_\{\\text\{slots\}\}contributed by that model\. We find no monotone relationship, so coverage imbalance does not predict rank\.

#### Llama\-excluded refit\.

llama\-3\.3\-70b\-togetherhas outlier performance, so we exclude it \(the bottom outlier\) and refit on the remaining 8 models\. The order of the 8 surviving models is preserved\.

Table 3:Llama\-excluded leaderboard\.Llama\-excluded refit —α^\\hat\{\\alpha\}\(chips/game\)Paired\-cluster bootstrap 95% CI \(B=2,000B\{=\}2\{,\}000\), sum\-to\-zero contrast\. Refit on the 8 non\-llama models\.The relative ranking of the 8 surviving models is identical to their relative position in the overall \(with\-llama\) leaderboard \(Table[1](https://arxiv.org/html/2605.23238#S6.T1)\); the level shifts because the sum\-to\-zero contrast is now centered on a different population \(no llama at−2\.37\-2\.37to anchor the bottom\), so several mid\-pack models flip sign relative to the new mean\. The qualitative claim thatgpt\-5andgemini\-3\.1\-proare at the top, andqwen\-3\.5andgemini\-3\.1\-flash\-liteare at the bottom, holds in both specifications\.

#### Bradley–Terry on win indicator\.

A logistic win\-probability model on the indicator𝟏​\[edge\>0\]\\mathbf\{1\}\[\\text\{edge\}\>0\]yields per\-model Bradley–Terry \(BT\) scores\. Unlike the win\-marginα^\\hat\{\\alpha\}, BT measures frequency of winning rather than magnitude of margin, so the two estimators answer different questions\. On our tournament data the BT ranking diverges noticeably from the win\-margin ranking: several mid\-pack models that win frequently with small margins \(gemini\-2\.5\-pro,gemini\-3\.1\-flash\-lite\) score high on BT, while models that win less often but with large margins \(gpt\-5,gemini\-3\.1\-pro\) score lower on BT than on win\-marginα^\\hat\{\\alpha\}\. We report win\-marginα^\\hat\{\\alpha\}as the overall estimator because \(i\) we explicitly instructed models to maximize expected chips and \(ii\) win margin contains more information about model play during the game than the binary win indicator\. The divergence from BT is itself a finding\. The way a model considers the tradeoff between winning frequently with small margins and winning less often with large margins is worth reporting for follow\-up work\.

## Appendix IPer\-tertile leaderboards \(composite\-complexity split\)

The composite\-complexity axis is the first principal component of the9×69\{\\times\}6matrix of per\-model axis\-slopesβ^m,a\\hat\{\\beta\}\_\{m,a\}from the capability\-profile regression in Section[7](https://arxiv.org/html/2605.23238#S7), a multivariate OLS of per\-game strengthα^m,g\\hat\{\\alpha\}\_\{m,g\}on the sixzz\-scored axes\. The per\-axis weights are\+0\.72\+0\.72for brittleness\-log10,\+0\.45\+0\.45for state\-space\-log10,\+0\.41\+0\.41for information\-sensitivity,\+0\.29\+0\.29for opponent\-modeling,\+0\.15\+0\.15for temporal\-depth, and\+0\.09\+0\.09for risk\. Per\-game composite scores sort the 50\-game benchmark into three tertiles of1616,1717, and1717games\. Within each tertile we refitα^\\hat\{\\alpha\}on the slot rows restricted to that tertile’s games, with paired\-cluster bootstrap 95% CIs \(B=500B\{=\}500\)\.

Table 4:Per\-tertile additive paired\-comparison leaderboards\.Per\-tertile leaderboards —α^\\hat\{\\alpha\}by composite\-complexity tertileT1 = easiest 16 games, T2 = middle 17, T3 = hardest 17, by composite complexity\. Paired\-cluster bootstrap 95% CI widths \(B=500B\{=\}500\) reach at most0\.530\.53chips/game across all 27 cells \(widest onclaude\-sonnet\-4\-6\-max, T3\), with a median width of0\.280\.28\. The leaderboard is broadly stable across tertiles, but the top\-3 and bottom\-3 are not preserved exactly\. On T1, claude drops out of the top\-3 \(gpt\-5, gemini\-3\.1\-pro, and gemini\-2\.5\-pro are the top three at\+0\.35\+0\.35,\+0\.35\+0\.35, and\+0\.26\+0\.26respectively\) and deepseek replaces qwen in the bottom\-3\. On T2 and T3 the top\-3 ordering settles into gpt\-5, gemini\-3\.1\-pro, claude, and win\-margin gaps widen with complexity\.
## Appendix JPairwise head\-to\-head table \(full9×99\{\\times\}9\)

Table 5:Pairwise mean win margin: row−\-column, across the 50 benchmark games with both seats balanced\. Bold = 95% paired\-cluster bootstrap CI excludes 0 \(B=500B\{=\}500\)\. Rows and columns sorted by leaderboard rank\.Fifty\-eight of the 72 off\-diagonal entries reach significance at the 95% level \(paired\-cluster bootstrap,B=500B\{=\}500\)\.

## Appendix KSolver\-baseline reference matchups \(full\)

For five benchmark seeds, we obtained reasonable tabular CFR\+abstractions and solved to convergence within them\. The resulting abstracted\-solver policies are not Nash policies of the original games \(finite abstraction introduces residual exploitable error\), so we treat them only as a third\-perspective sanity check alongside the LLM\-vs\-LLM tournament\. We play each of the 9 LLMs against the abstracted CFR\+solver on each of the 5 seeds, targetingnslots=100n\_\{\\text\{slots\}\}\{=\}100paired slots per \(model, seed\) matchup; the merged dataset is 4,494 slots \(45 \(model, seed\) matchups×\\times100 nominal, minus 6claude\-sonnet\-4\-6\-maxslots that failed at the provider and were not retried: 5 missing on seed 933 and 1 on seed 10137\)\. Top\-ranked models post small positive average margins \(\+0\.30\+0\.30to\+0\.58\+0\.58chips/game\), consistent with finding and exploiting residual abstraction error;llama\-3\.3\-70bis the only model that the solver consistently exploits, with an average of−1\.89\-1\.89chips per game across the five seeds and a low of−3\.18\-3\.18chips per game on the most complex seed\. The per\-model averages rank\-correlate strongly with the overall LLM\-vs\-LLM leaderboard, with Spearmanρ=0\.95\\rho=0\.95andp=0\.0001p=0\.0001\.

We report each model’s edge against the abstracted solver asy¯±SE\\bar\{y\}\\pm\\mathrm\{SE\}, the mean chip margin per paired\-play\-seed group with paired\-seat standard error\. The game is zero\-sum, so the model edgey¯\\bar\{y\}already pins down the solver’s loss as−y¯\-\\bar\{y\}\. The CFR\+solver policy is the average strategy of CFR\+run to convergence within the abstraction \(default lumping for four seeds, finer abstraction for seed 11520\)\. Decisions are sampled stochastically from the mixed strategy\. We target one hundred paired slots per \(seed, model\) matchup\. The achieved count is one hundred in 43 of 45 matchups, with 95 in the \(seed 933,claude\-sonnet\-4\-6\-max\) matchup and 99 in the \(seed 10137,claude\-sonnet\-4\-6\-max\) matchup, in both cases due to provider\-side failures that were not retried\.

Table 6:LLM\-vs\-abstracted\-CFR\+reference matchups, 5 seeds with tractable solver abstractions×\\times9 models, 50 paired\-play\-seed runs \(≤100\\leq 100slots\) per \(model, seed\) matchup\.Each cell is the mean chip margin of the model against the solver on that seed, averaged over the paired\-play\-seed groups \(the two seat assignments within a group are collapsed by within\-group mean\),±\\pmthe paired\-play\-seed standard error \(standard deviation of the group means divided byngroups\\sqrt\{n\_\{\\text\{groups\}\}\}, wherengroupsn\_\{\\text\{groups\}\}is the number of distinct play\-seed groups for that \(model, seed\) cell\)\. The final column is the per\-model average across the 5 seeds\. The CFR\+policy is the solver output on a finite abstraction, so any residual exploitable error in the abstracted policy can be picked up by the model and shows up here as a positive edge\.Per\-model averages across the 5 seeds rank\-correlate strongly with the LLM\-vs\-LLM overall leaderboard \(Spearmanρ=0\.95\\rho\{=\}0\.95,p=0\.0001p\{=\}0\.0001; Pearsonr=0\.98r\{=\}0\.98,p<0\.0001p\{<\}0\.0001\), indicating the solver baseline and the peer tournament are measuring essentially the same underlying axis of strategic competence on this distribution\.

## Appendix LFull per\-axis regression table

Table[7](https://arxiv.org/html/2605.23238#A12.T7)reports the same per\-model multivariate OLS slopes plotted in Figure[3](https://arxiv.org/html/2605.23238#S7.F3), in tabular form\. CIs are paired\-cluster bootstrap \(B=500B=500\);pp\-values are Benjamini–Hochberg corrected at FDR0\.050\.05across the full family of9×6=549\\times 6=54\(model, axis\) coefficients\. Bold entries clear BH\.

Table 7:Per\-model multivariate OLS slope ofα^m,g\\hat\{\\alpha\}\_\{m,g\}on eachzz\-scored axis, controlling for the other five axes\. Bold = BH\-adjustedq<0\.05q<0\.05across the 54 coefficients\.
## Appendix MCapability profile in absolute units

The main\-text radar \(Figure[3](https://arxiv.org/html/2605.23238#S7.F3)\) plots the per\-\(model, axis\) slopeβ^m,a\\hat\{\\beta\}\_\{m,a\}in chips/game perσ\\sigmaof axis, which is the local sensitivity of modelmm’s edge to axisaa\. The companion plot below uses the same fitted regression, but evaluates each model’s predicted per\-game strength on a synthetic game where axisaais at its observed maximum z\-score and the other five axes are at their observed medians:

α^m,ga∗=β^m,0\+β^m,a​zamax\+∑a′≠aβ^m,a′​za′med\.\\hat\{\\alpha\}\_\{m,g\_\{a\}^\{\*\}\}\\;=\\;\\hat\{\\beta\}\_\{m,0\}\+\\hat\{\\beta\}\_\{m,a\}\\,z\_\{a\}^\{\\max\}\+\\\!\\\!\\sum\_\{a^\{\\prime\}\\neq a\}\\\!\\\!\\hat\{\\beta\}\_\{m,a^\{\\prime\}\}\\,z\_\{a^\{\\prime\}\}^\{\\text\{med\}\}\.Each spoke is then in chips/game on a benchmark\-extreme game for that axis, rather than a per\-σ\\sigmaslope\.

![Refer to caption](https://arxiv.org/html/2605.23238v1/figures/radar_predicted_alpha.png)Figure 7:Capability profile in absolute units\.Predicted per\-game strengthα^m,ga∗\\hat\{\\alpha\}\_\{m,g\_\{a\}^\{\*\}\}from the same multivariate OLS used for Figure[3](https://arxiv.org/html/2605.23238#S7.F3), evaluated at the synthetic input defined above \(each axis at its 50\-game maximumzz\-score in turn, with the other five axes held at their 50\-game medians\)\. Units are chips/game versus the across\-model mean\. Shaded bands are 95% paired\-cluster bootstrap CIs \(B=500B\{=\}500\) propagated through the linear predictor\.Two contrasts come out of the absolute\-units plot that the slope\-only radar cannot show\. First, the three top\-leaderboard models are clearly above the across\-model mean on the most\-brittleness\-heavy benchmark game \(gpt\-5\-4\-highat\+1\.49\+1\.49chips/game,claude\-sonnet\-4\-6\-maxat\+1\.43\+1\.43,gemini\-3\.1\-pro\-previewat\+1\.26\+1\.26\), withgpt\-5\-4\-highthe flattest of the three across spokes \(\+0\.91\+0\.91to\+1\.49\+1\.49\) andclaude\-sonnet\-4\-6\-maxthe most peaked\. Second,gemini\-3\.1\-flash\-lite\-previewis the only direction\-splitting profile: predicted above the mean on temporal\-depth \(\+0\.23\+0\.23\) and information\-sensitivity \(\+0\.13\+0\.13\) but below it on the other four axes, so on the most\-temporal\-depth game in the benchmark flash\-lite is a real \(small\) winner rather than just a gap\-narrower\. The main\-text radar suppresses both contrasts because it plots only the slopes\.

## Appendix NLocal jaggedness: estimation, inference, and subset robustness

This section gives the full construction of the stakes\-normalized local\-jaggedness measureJmJ\_\{m\}used in Section[8](https://arxiv.org/html/2605.23238#S8), reports its sensitivity to the neighborhood sizeKK, and reports the tournament\-subset robustness check that motivates the choice ofσg\\sigma\_\{g\}\-only studentization\.

#### Construction\.

The full formula forJmJ\_\{m\}is given in Section[8](https://arxiv.org/html/2605.23238#S8)\. The kNN aggregator uses the population \(biased\) standard deviation of the four neighborhood values, i\.e\. divides by\|𝒩​\(g\)\|\|\\mathcal\{N\}\(g\)\|rather than\|𝒩​\(g\)\|−1\|\\mathcal\{N\}\(g\)\|\-1\. Switching the convention to the sample standard deviation would multiply everyJmJ\_\{m\}by\|𝒩​\(g\)\|/\(\|𝒩​\(g\)\|−1\)≈1\.155\\sqrt\{\|\\mathcal\{N\}\(g\)\|/\(\|\\mathcal\{N\}\(g\)\|\-1\)\}\\approx 1\.155uniformly, so the choice of convention cancels out of any ranking, ratio, or bootstrap CI; we use the population convention for notational simplicity\. We compute 95% confidence intervals from a bias\-corrected paired\-cluster bootstrap of500500replicates, where clusters are indexed by \(game seed, run id\)\. At each replicate we resample clusters with replacement, re\-fit the global additive paired\-comparisonα^m\\hat\{\\alpha\}\_\{m\}and the per\-game refitα^m,g\\hat\{\\alpha\}\_\{m,g\}, recomputeσg\\sigma\_\{g\}on the resampled rows, and recomputeJmJ\_\{m\}\. We report the bias\-corrected percentile interval\.

#### Sensitivity to the neighborhood sizeKK\.

The main\-text construction usesK=3K=3\. We re\-run theσg\\sigma\_\{g\}\-onlyJmJ\_\{m\}pipeline atK∈\{2,4,5,7,10,15\}K\\in\\\{2,4,5,7,10,15\\\}\. The nine\-model mean ofJmJ\_\{m\}rises smoothly withKKas larger neighborhoods incorporate more cross\-game variance, from0\.0670\.067atK=2K=2to0\.0990\.099atK=15K=15, but the per\-model ranking is extraordinarily stable\. The Spearman rank correlation against theK=3K=3ordering is\+1\.000\+1\.000atK=2K=2and atK=4K=4, and\+0\.983\+0\.983at every largerKKin the sweep\. A single position swap, in whichclaude\-sonnet\-4\-6\-maxandqwen\-3\.5\-togetherswap ranks three and four, accounts for the entire departure from perfect agreement\.llama\-3\.3\-70b\-togetheris the most locally volatile model at everyKKin the sweep, anddeepseek\-v3\.1remains the smoothest\.

#### Tournament\-subset robustness\.

A concern is that a local\-jaggedness measure of the formJm=meang​std​\{zm,g′\}J\_\{m\}=\\mathrm\{mean\}\_\{g\}\\mathrm\{std\}\\\{z\_\{m,g^\{\\prime\}\}\\\}depends on which models share the tournament pool, because bothα^m,g\\hat\{\\alpha\}\_\{m,g\}\(the per\-game refit\) andα^m\\hat\{\\alpha\}\_\{m\}\(the global strength\) are estimated on the shared pool\. We test this directly by rerunning theσg\\sigma\_\{g\}\-onlyJmJ\_\{m\}pipeline under four subset conditions: \(i\) leave\-one\-model\-out, dropping each of the nine models in turn and recomputingJmJ\_\{m\}on the remaining eight; \(ii\) drop top three \(gpt\-5,gemini\-3\.1\-pro,claude\); \(iii\) drop bottom three \(llama\-3\.3\-70b,qwen\-3\.5,gemini\-3\.1\-flash\-lite\); and \(iv\) mid\-pack only \(gemini\-2\.5\-pro,gemma\-4\-31b\-it,deepseek\-v3\.1\)\.

Underσg\\sigma\_\{g\}\-only studentization,JmJ\_\{m\}is robust to which models share the tournament pool\. For the high\-JmJ\_\{m\}models the leave\-one\-out range across the eight subset replicates is tight, on the order of1010to20%20\\%of the baseline value \(gpt\-5,claude\-sonnet\-4\-6,llama\-3\.3\-70b\)\. The low\-JmJ\_\{m\}models have wider relative LOO ranges \(up to roughly4040–60%60\\%of baseline fordeepseek\-v3\.1,qwen\-3\.5, andgemini\-3\.1\-flash\-lite\) because their small denominators amplify any absolute LOO shift; their absolute LOO shifts are themselves small\. The qualitative ordering, withllama\-3\.3\-70bmost jagged, thengpt\-5,claude\-sonnet\-4\-6, andqwen\-3\.5, then the mid\-pack anddeepseek\-v3\.1smoothest, is preserved under every leave\-one\-out replicate\. The drop\-top\-three, drop\-bottom\-three, and mid\-three\-only conditions move the absolute values somewhat more, but the inferential picture is unchanged\.

## Appendix OPer\-game results

The full per\-\(model, game\) pair\-mean win margin across the 50\-game benchmark, sorted ascending by composite complexity score\. Thecmplxcolumn is the simple mean of the sixzz\-scored axis values for each game, so it is centered at zero across the 50\-game benchmark by construction\. Negative values denote games whose axis values are below the benchmark mean \(Kuhn\-like\) and positive values denote above\-mean games\. This is a lightweight summary used only for sorting the rows of this table; the PC1\-based composite used for the tertile leaderboards in Appendix[I](https://arxiv.org/html/2605.23238#A9)is a separate construction\. All 9 models reach per\-\(model, game\) coverage on all 50 seeds via the rotating\-matchup schedule \(Section[5](https://arxiv.org/html/2605.23238#S5)\)\.

Table 8:Per\-\(model, game\) pair\-mean win margin on the 50\-game benchmark, sorted ascending by composite complexity\.cmplx= mean of the sixzz\-scored axes\.
## Appendix PPer\-cellα^m,g\\hat\{\\alpha\}\_\{m,g\}matrix and rank\-stability against a noise null

This appendix reports the full matrix of per\-game strength estimatesα^m,g\\hat\{\\alpha\}\_\{m,g\}and tests whether the per\-game variation in model rankings exceeds what sampling noise alone would produce\. The per\-cell point estimates are the same per\-game additive paired\-comparison refits used throughout the paper \(Section[6](https://arxiv.org/html/2605.23238#S6)\)\. For each \(model, game\) pair we additionally compute a 95% bootstrap confidence interval fromB=500B=500paired\-cluster bootstrap replicates on \(game seed, run id\) clusters\.

Table[9](https://arxiv.org/html/2605.23238#A16.T9)shows the point estimates only, sorted ascending by composite\-complexity score\. Models are column\-ordered by the overall leaderboard\.

Table 9:Per\-game strength estimatesα^m,g\\hat\{\\alpha\}\_\{m,g\}\(chips/game\), per model\. Rows are the 50 benchmark games sorted ascending by composite\-complexity score, and columns are the models in overall leaderboard order\.#### Rank stability\.

The per\-game rankings induced byα^m,g\\hat\{\\alpha\}\_\{m,g\}are not identical to the overall leaderboard, and some pairs flip\. We decompose this disagreement into a component attributable to sampling noise and a residual through three measurements\.

The Kendall rank correlation between the per\-game ranking and the overall ranking has a mean of0\.650\.65and a median of0\.720\.72across the 50 benchmark games\. Three\-quarters of games have aτ\\tauof at least0\.60\.6, and a quarter haveτ≥0\.8\\tau\\geq 0\.8\. The worst\-case game hasτ=0\.11\\tau=0\.11, the bestτ=0\.89\\tau=0\.89\. The overall ordering is therefore a strong but imperfect predictor of the per\-game ordering\.

The average number of pairwise rank reversals per game, against the overall leaderboard, is6\.226\.22out of\(92\)=36\\binom\{9\}\{2\}=36possible pairwise comparisons\. The distribution is skewed, with median55reversals and maximum1616\.

The noise\-baseline comparison computes the expected number of those reversals under a null where the true per\-game ranking equals the overall ranking and the only source of per\-game disagreement is sampling noise on the per\-cellα^m,g\\hat\{\\alpha\}\_\{m,g\}estimates themselves\. For each pair\(mi,mj\)\(m\_\{i\},m\_\{j\}\)on each gamegg, letSE^pair​\(g\)\\widehat\{\\mathrm\{SE\}\}\_\{\\mathrm\{pair\}\}\(g\)be the bootstrap\-empirical standard deviation ofα^mi,g\(b\)−α^mj,g\(b\)\\hat\{\\alpha\}^\{\(b\)\}\_\{m\_\{i\},g\}\-\\hat\{\\alpha\}^\{\(b\)\}\_\{m\_\{j\},g\}across the 500 replicates\. Under the null that the true difference equals the overall differenceα^mi−α^mj\\hat\{\\alpha\}\_\{m\_\{i\}\}\-\\hat\{\\alpha\}\_\{m\_\{j\}\}, the probability of an observed reversal on that pair\-game is

Prev​\(i,j,g\)=Φ​\(−\|α^mi−α^mj\|SE^pair​\(g\)\),P\_\{\\mathrm\{rev\}\}\(i,j,g\)\\;=\\;\\Phi\\\!\\left\(\-\\frac\{\|\\hat\{\\alpha\}\_\{m\_\{i\}\}\-\\hat\{\\alpha\}\_\{m\_\{j\}\}\|\}\{\\widehat\{\\mathrm\{SE\}\}\_\{\\mathrm\{pair\}\}\(g\)\}\\right\),whereΦ\\Phiis the standard normal CDF\. LetEg=∑i<jPrev​\(i,j,g\)E\_\{g\}=\\sum\_\{i<j\}P\_\{\\mathrm\{rev\}\}\(i,j,g\)denote the expected reversal count on gameggunder this null \(the sum across all 36 unordered pairs\), and letVg=∑i<jPrev​\(i,j,g\)​\(1−Prev​\(i,j,g\)\)V\_\{g\}=\\sum\_\{i<j\}P\_\{\\mathrm\{rev\}\}\(i,j,g\)\\bigl\(1\-P\_\{\\mathrm\{rev\}\}\(i,j,g\)\\bigr\)denote its variance\.888Treating the 36 per\-pair reversal indicators as independent Bernoullis gives the Poisson\-binomial expectationEgE\_\{g\}and varianceVgV\_\{g\}; we use the standard\-normal approximation tozgz\_\{g\}for the one\-sidedpp\-valuep^g=1−Φ​\(zg\)\\hat\{p\}\_\{g\}=1\-\\Phi\(z\_\{g\}\)\. The independence assumption is not strict, since the per\-gameα^m,g\\hat\{\\alpha\}\_\{m,g\}for different models share information through the joint per\-game refit, but this is the standard construction for a sum of correlated Bernoulli indicators\.

Averaged across the 50 benchmark games, the expected reversal count under the null is3\.713\.71, compared to the observed6\.226\.22\. The observed\-to\-expected ratio is1\.681\.68\. For each gameggdefinezg=\(Nobs​\(g\)−Eg\)/Vgz\_\{g\}=\(N\_\{\\text\{obs\}\}\(g\)\-E\_\{g\}\)/\\sqrt\{V\_\{g\}\}, the number of noise\-null standard deviations by which the observed reversal countNobs​\(g\)N\_\{\\text\{obs\}\}\(g\)exceeds its null expectation\. The mean ofzgz\_\{g\}across the 50 benchmark games is\+3\.2\+3\.2and the median is\+0\.97\+0\.97\. Thirty percent of games havezg≥2z\_\{g\}\\geq 2, and thirty\-four percent havezg≤0z\_\{g\}\\leq 0\. The right tail is what drives the average\. A substantial subset of the benchmark games carries significantly more per\-game disagreement than sampling noise can produce, while another sizeable subset is consistent with the noise null\.

To assess whether significant parts of the game space carry reversals beyond what sampling noise alone can produce, we apply the Benjamini–Hochberg procedure to the one\-sided excess\-reversalpp\-valuesp^g=1−Φ​\(zg\)\\hat\{p\}\_\{g\}=1\-\\Phi\(z\_\{g\}\)across the 50 benchmark games\. At a false\-discovery rate ofq=0\.05q=0\.05,15 of 50 games \(30%30\\%\)carry statistically more rank reversals than the noise\-only null predicts\. Atq=0\.10q=0\.10,19 of 50 games \(38%38\\%\)do\. Under the null that no game in the benchmark has any per\-game variation beyond sampling noise, the expected number of false positives atq=0\.05q=0\.05is at most2\.52\.5\. The observed 15 therefore constitutes clear evidence of per\-game variation beyond the noise null on a non\-trivial slice of the benchmark\.

At the finer per\-\(pair, game\) level, we identify311311candidate reversal cells \(game seeds, model pairs where the per\-game point estimate reverses the overall leaderboard\)\. For each candidate cell we compute a one\-sided bootstrappp\-value \(fraction of bootstrap replicates in which the reversal does not hold\)\. After BH\-FDR correction across the 311 candidates,five individual reversal cells surviveq<0\.05q<0\.05and ten surviveq<0\.10q<0\.10\. The fiveq<0\.05q<0\.05cells are: seed23852385wheregemini\-3\.1\-flash\-liteoutperformsgemma\-4\-31b; seed34833483whereclaude\-sonnet\-4\-6outperformsgemini\-3\.1\-pro; seed66356635wheredeepseek\-v3\.1outperformsgemma\-4\-31b; and seeds82978297and1013710137wheregemma\-4\-31boutperformsgemini\-2\.5\-pro\. These are mid\-pack rearrangements rather than upsets of the top three by the bottom of the table\.

#### Reversal significance across axis space\.

Figure[8](https://arxiv.org/html/2605.23238#A16.F8)plots the 50 benchmark games in the canonical state\-space versus information\-sensitivity plane, with each point coloured by its BH\-FDRqq\-value on the excess\-reversal test and the marker shape distinguishing the three significance bands\. On no single axis taken alone do the reversal\-significant games cluster cleanly into one tertile, andqq\-values vary across the full observed range of each axis\.

#### Reversal significance by composite\-complexity tertile\.

The composite\-complexity score \(PC1 of the per\-model slope matrix; see Appendix[I](https://arxiv.org/html/2605.23238#A9)\) combines the six axes into a single per\-game difficulty proxy\. Cross\-tabulating the rank\-reversal test against the composite\-complexity tertiles shows that reversal significance is concentrated almost entirely on the easiest tertile\.

The mean obs column is the average per\-game count of observed pairwise rank reversals within the tertile, and mean exp is the average per\-game count expected under the noise\-only null defined above\. The mean\-zzcolumn reports the per\-tertile average ofzgz\_\{g\}, the number of noise\-null standard deviations by which the observed reversal count exceeds its null expectation\.

Twelve of the sixteen easiest games carry BH\-significant excess reversals atq<0\.05q<0\.05, with a meanzz\-score of\+9\+9\. The seventeen hardest games have zero BH\-significant reversals and a meanzzclose to zero\. On hard games, the top models pull away from the rest with larger absolute margins \(Section[6](https://arxiv.org/html/2605.23238#S6)\), so the rank ordering is stable and matches the overall leaderboard\. On easy games the absolute gaps are small, and the per\-game ranking reorders the field beyond what the per\-pair sampling noise alone can produce\. The per\-game capability variation we surface in this appendix exceeds the noise null but is concentrated on easy games\. Hard games are where the overall leaderboard is most trustworthy, and easy games are where it most under\-represents per\-game disagreement\.

![Refer to caption](https://arxiv.org/html/2605.23238v1/x7.png)Figure 8:Reversal significance in axis space\.The 50 benchmark games plotted in the state\-space versus information\-sensitivity plane \(matching Figure[1](https://arxiv.org/html/2605.23238#S4.F1)\), colour\-coded by BH\-FDRqq\-value on the one\-sided excess\-reversal test\. Stars mark games withq<0\.05q<0\.05, indicating per\-game rank reversals that are statistically significant at FDR5%5\\%after correcting across the 50 games\. Squares mark games with0\.05≤q<0\.100\.05\\leq q<0\.10\(marginal significance\)\. Circles mark games withq≥0\.10q\\geq 0\.10\(consistent with the noise null\)\.

Similar Articles

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play

Hugging Face Daily Papers

STRATAGEM is a new framework for improving reasoning transferability in language models by using game self-play with a Reasoning Transferability Coefficient and Reasoning Evolution Reward to reinforce abstract, domain-agnostic reasoning patterns over game-specific heuristics. Experiments show strong improvements on mathematical reasoning, general reasoning, and code generation benchmarks.

Evaluating Large Language Models in a Complex Hidden Role Game

arXiv cs.CL

This paper introduces an open-source framework to evaluate LLMs' reasoning, persuasion, and deception capabilities in the hidden role game Secret Hitler, finding that current models fail at sustained multi-turn manipulation while rule-based agents outperform them.