Tag
This paper introduces an entropy-regularized reinforcement learning framework for zero-sum stochastic differential games in regime-switching jump-diffusion processes, deriving HJBI equations and an actor-critic algorithm with applications to investment games.
Introduces SidConArena, a benchmark framework for evaluating LLM agents in open-ended, positive-sum bargaining games, combining negotiation, production, and auctions to assess mixed-motive interaction and economic planning.
EMAgnet introduces parameter-space exponential moving average regularization for policy gradient self-play in large two-player zero-sum games, achieving lower exploitability compared to uniform regularization targets.
A 58-page paper from Google DeepMind on building agents specialized in game theory, highlighting key insights from the research.
This paper studies decentralized coalition formation as a dynamical process driven by unilateral exit-and-join decisions, using the Aumann-Dreze value for local payoff evaluation. It establishes equilibrium characterizations, Lyapunov and potential representations, and analyzes the impact of switching/acceptance costs on stability.
A blog post discussing how increased granularity in systems, such as tick sizes in financial markets and time slots for booking sports courts, can introduce strategic gaming and inefficiencies, arguing that finer choices are not always beneficial.
MIT researchers co-authored a paper showing that general-purpose policy gradient algorithms can outperform specialized game-theoretic algorithms in imperfect-information games, challenging long-held assumptions in the field.
Introduces the concept of synthetic counteradaptation, where humans and AI systems co-evolve by adapting to each other's strategies, illustrated through examples from Go, social interactions, and geopolitical simulations.
Poker Arena is a new benchmark using no-limit Texas Hold'em to evaluate LLMs' strategic reasoning and memory across multiple cognitive axes. The platform reveals that multi-axis evaluation exposes capability structures that scalar leaderboards misrank.
This paper introduces RogueAI, a reverse Turing test implemented as an interactive webapp where human players interrogate two LLM agents to identify which one is licensed to deceive within a shared fictional scenario. A pilot deployment shows a gap between heuristic detection (75.6% accuracy) and human performance (56.6%), highlighting the potential of the system as a data-collection and teaching tool for AI deception and honesty.
This paper presents a unified multi-modal framework integrating reinforcement learning, high-frequency trading, game-theoretic approaches, and cross-modal sentiment analysis for intelligent financial systems, claiming significant improvements over single-domain systems.
This paper introduces a framework for two-sided matching with temporally extended feedback, formulating it as a partially observable Markov game with costly screening, noisy observations, and evolving latent profiles. The authors present Learn2Match, a multi-agent reinforcement learning benchmark, and show that independent PPO outperforms bandit baselines in social welfare but incurs higher information-friction loss.
A systematic exploration of all possible strategies in a repeated two-player game using computational methods, analyzing cumulative payoffs and winning strategies through ruliology.
Stephen Wolfram explores what happens when agents with all possible strategies compete, using ruliological methods to systematically analyze strategies in a match-or-not game.
This paper introduces Repeated Policy Regret (RP-Regret), a game-theoretic metric for regret minimization in repeated games with adaptive opponents, and proposes three algorithms to minimize it, showing that doing so can lead to cooperative equilibria like in Stag-Hunt.
This paper argues that superintelligent AI systems designed under a solipsistic paradigm that treats the world as stationary will be self-undermining and uncooperative, leading to collective failures. The authors call for a new research paradigm that treats interdependence and cooperation as core design principles.
The article explores extending rock-paper-scissors to more than three options by allowing ties, revealing richer game dynamics and strategies through graph theory.
Proposes a truthful online preference aggregation mechanism for LLM fine-tuning in mobile crowdsourcing, addressing strategic worker misreporting and achieving sublinear regret.
This paper introduces GENSTRAT, a benchmark that uses procedurally generated strategic environments to evaluate LLMs' strategic reasoning across multiple axes, addressing limitations of fixed game suites.
This paper identifies a threshold in decision capacity that determines whether self-play reinforcement learning agents collapse under asymmetric rule perturbations, showing that eliminating all positive-reach contingent decisions leads to rapid convergence to a deterministic exploitation attractor.