ALSO: Adversarial Online Strategy Optimization for Social Agents

arXiv cs.AI Papers

Summary

ALSO introduces a framework for online strategy optimization in multi-agent social simulation, formulating multi-turn interaction as an adversarial bandit problem and using a neural surrogate for reward prediction. Experiments on the Sotopia benchmark show it outperforms static baselines and existing optimization methods.

arXiv:2605.15768v1 Announce Type: new Abstract: Social simulation provides a compelling testbed for studying social intelligence, where agents interact through multi-turn dialogues under evolving contexts and strategically adapting opponents. Such environments are inherently non-stationary, requiring agents to dynamically adjust their strategies over time. However, most Large Language Model (LLM) based social agents rely on static personas, while existing approaches for enhancing social intelligence, such as offline reinforcement learning or external planners, are ill-suited to these settings, typically assuming stationarity and incurring substantial training overhead. To bridge this gap, we propose \textbf{ALSO} (\textbf{A}dversarial on\textbf{L}ine \textbf{S}trategy \textbf{O}ptimization), the first framework for online strategy optimization in multi-agent social simulation. ALSO advances social adaptation through two key contributions. (1) ALSO formulates multi-turn interaction as an adversarial bandit problem, where combinations of static personas and dynamic strategy instructions are treated as arms, providing a principled solution to non-stationarity without relying on environmental stability assumptions. (2) To predict rewards and generalize sparse feedback in multi-turn dialogues, ALSO introduces a lightweight neural surrogate to predict rewards from interaction histories, enabling sample-efficient exploration and continuous online adaptation. Experiments on the Sotopia benchmark demonstrate that ALSO consistently outperforms static baselines and existing optimization methods in dynamic environments, validating the effectiveness of adversarial online strategy optimization for building robust social agents.
Original Article
View Cached Full Text

Cached at: 05/18/26, 06:34 AM

# ALSO: Adversarial Online Strategy Optimization for Social Agents
Source: [https://arxiv.org/html/2605.15768](https://arxiv.org/html/2605.15768)
###### Abstract

Social simulation provides a compelling testbed for studying social intelligence, where agents interact through multi\-turn dialogues under evolving contexts and strategically adapting opponents\. Such environments are inherently non\-stationary, requiring agents to dynamically adjust their strategies over time\. However, most Large Language Model \(LLM\) based social agents rely on static personas, while existing approaches for enhancing social intelligence, such as offline reinforcement learning or external planners, are ill\-suited to these settings, typically assuming stationarity and incurring substantial training overhead\. To bridge this gap, we proposeALSO\(Adversarial onLineStrategyOptimization\), the first framework for online strategy optimization in multi\-agent social simulation\.ALSOadvances social adaptation through two key contributions\. \(1\)ALSOformulates multi\-turn interaction as an adversarial bandit problem, where combinations of static personas and dynamic strategy instructions are treated as arms, providing a principled solution to non\-stationarity without relying on environmental stability assumptions\. \(2\) To predict rewards and generalize sparse feedback in multi\-turn dialogues,ALSOintroduces a lightweight neural surrogate to predict rewards from interaction histories, enabling sample\-efficient exploration and continuous online adaptation\. Experiments on the Sotopia benchmark demonstrate thatALSOconsistently outperforms static baselines and existing optimization methods in dynamic environments, validating the effectiveness of adversarial online strategy optimization for building robust social agents\.

Large Language Models, Multi\-agent Systems, Social Intelligence

![Refer to caption](https://arxiv.org/html/2605.15768v1/x1.png)Figure 1:Online social interaction between agents withpersonasandadaptive strategies, where feedback multi\-turn dialogue drives continuous strategy optimization under evolving behaviors\.## 1Introduction

Modeling social intelligence\(Mathuret al\.,[2024](https://arxiv.org/html/2605.15768#bib.bib34)\)is a central pursuit in Artificial Intelligence research\. The advent of Large Language Models \(LLMs\) has substantially advanced this field by endowing agents with human\-like communication\(Spitaleet al\.,[2023](https://arxiv.org/html/2605.15768#bib.bib51)\)and planning capabilities\(Wuet al\.,[2024a](https://arxiv.org/html/2605.15768#bib.bib61); Parket al\.,[2023](https://arxiv.org/html/2605.15768#bib.bib40)\)\. This progress has positioned LLM\-based social simulation as a powerful framework for studying emergent social behaviors in large\-scale and goal\-oriented interactions\(Epstein,[2012](https://arxiv.org/html/2605.15768#bib.bib11); Wanget al\.,[2024a](https://arxiv.org/html/2605.15768#bib.bib58)\)\. In such simulations, agent behavior is typically governed by a*persona*, a formalized profile encapsulating personality traits, occupations, and background stories\(Reiss,[2023](https://arxiv.org/html/2605.15768#bib.bib43); Salinas and Morstatter,[2024](https://arxiv.org/html/2605.15768#bib.bib45); Bisbeeet al\.,[2024](https://arxiv.org/html/2605.15768#bib.bib7)\)\.

While static personas provide foundational identity, they alone are insufficient for eliciting adaptive social intelligence\. It is important to distinguish between*persona*and*strategy*: a persona defines*who*an agent is, whereas a strategy specifies*how*the agent acts to navigate interactions and achieve goals\. Empirical studies demonstrate that relying solely on static personas often yields stereotypical and homogeneous behaviors, failing to capture the diversity required for robust social simulation\(Taubenfeldet al\.,[2024](https://arxiv.org/html/2605.15768#bib.bib53); Hwanget al\.,[2025](https://arxiv.org/html/2605.15768#bib.bib20); Zenget al\.,[2025](https://arxiv.org/html/2605.15768#bib.bib65); Venkitet al\.,[2026](https://arxiv.org/html/2605.15768#bib.bib56)\)\. This homogeneity is further reinforced by the “alignment tax” of Reinforcement Learning from Human Feedback \(RLHF\), which suppresses behavioral variance in favor of safety\([Kirket al\.,](https://arxiv.org/html/2605.15768#bib.bib22)\)\. Without evolving strategies to complement identity, agents struggle to adapt to dynamic opponents, resulting in rigid and suboptimal interaction patterns\(Liet al\.,[2023](https://arxiv.org/html/2605.15768#bib.bib27); Wanget al\.,[2024b](https://arxiv.org/html/2605.15768#bib.bib50); Zenget al\.,[2025](https://arxiv.org/html/2605.15768#bib.bib65)\)\.

To enhance social adaptability, recent work has explored automated prompt optimization \(APO\) for refining agent instructions\(Zhouet al\.,[2022a](https://arxiv.org/html/2605.15768#bib.bib4); Guoet al\.,[2024](https://arxiv.org/html/2605.15768#bib.bib12); Linet al\.,[2024b](https://arxiv.org/html/2605.15768#bib.bib21); Opsahl\-Onget al\.,[2024a](https://arxiv.org/html/2605.15768#bib.bib35)\), which can be interpreted through multi\-armed bandit formulations\. However, both offline approaches and online variants\(Yanget al\.,[2023](https://arxiv.org/html/2605.15768#bib.bib37); Linet al\.,[2024a](https://arxiv.org/html/2605.15768#bib.bib30); Wuet al\.,[2024b](https://arxiv.org/html/2605.15768#bib.bib62)\)fundamentally rely on stationarity assumptions, where each prompt induces a stable reward distribution evaluated on fixed validation sets\. Such assumptions break down in social simulation, as feedback from multi\-turn interactions with strategically adapting counterparts whose behaviors co\-adapt with the agent’s strategy choices, inducing persistent reward shifts and strong temporal coupling\. This dynamic feedback loop renders standard stochastic bandit models inadequate for social strategy instruction optimization\.

To bridge this gap, we proposeALSO\(Adversarial OnlineStrategyOptimization\), anonlineframework that*casts social strategy adaptation as an adversarial multi\-armed bandit problem*to enable principled optimization under non\-stationarity, as shown in Figure[1](https://arxiv.org/html/2605.15768#S0.F1)\. Rather than assuming stationary rewards,ALSOexplicitly models the co\-evolving and strategically adaptive nature of social interactions with two designs\. \(1\) InALSO, each arm corresponds to a strategy instruction sampled from a generated strategy set \(e\.g\., cooperation, competition, deception, rational bargaining\)\. At each interaction round, the selected strategy is combined with the agent’s persona to form the final prompt context, ensuring that dynamic adaptation remains grounded in consistent identity\. Based on this adversarial bandit formulation,ALSOinstantiates a robust online optimization procedure inspired by EXP3\(Lattimore and Szepesvári,[2020](https://arxiv.org/html/2605.15768#bib.bib25)\)with smoothing to hedge against shifting opponent behaviors\. \(2\) Standard bandit methods treat arms independently and fail to exploit semantic relationships among strategy instructions\. To address this limitation,ALSOintroduces a lightweight neural surrogate model that leverages interaction histories to predict rewards and generalize sparse feedback across semantically related strategies\. This enables sample\-efficient online optimization under sparse multi\-turn feedback\.

Overall,ALSOforms a closed\-loop online system that iteratively selects strategies, interacts under persona\-conditioned prompts, and updates both the bandit policy and surrogate model from feedback\. We evaluateALSOon Sotopia, a comprehensive LLM\-based social simulation benchmark spanning seven dimensions of social intelligence\. Across diverse settings,ALSOconsistently outperforms static persona agents and existing optimization baselines, achieving a \+16\.60% overall improvement and \+83\.79% substantial gains on relationship outcomes\.

Contributions\.We make the following contributions:

- •We introduce the first online strategy learning framework for LLM\-based multi\-agent social simulation, enabling dynamic adaptation beyond static persona\-driven behavior in evolving environments\.
- •We formulate social strategy optimization as an adversarial bandit problem with surrogate reward modeling, providing a principled solution to non\-stationary and strategically adaptive interactions\.
- •We conduct extensive evaluations on the Sotopia benchmark, demonstrating consistent improvements over static agents and existing optimization baselines across diverse social settings\.

## 2Related work

### 2\.1Social Intelligence

Social intelligence, distinct from abstract and mechanical intelligence, refers to the ability to manage interpersonal relations and social contexts\(Thorndike,[1920](https://arxiv.org/html/2605.15768#bib.bib54); Strang,[1930](https://arxiv.org/html/2605.15768#bib.bib52); Thorndike and Stein,[1937](https://arxiv.org/html/2605.15768#bib.bib55)\)\. In computational settings, Artificial Social Intelligence emphasizes modeling and responding to the mental and behavioral dynamics of interacting partners, including both humans and artificial agents\(Sapet al\.,[2022](https://arxiv.org/html/2605.15768#bib.bib47); Gweonet al\.,[2023](https://arxiv.org/html/2605.15768#bib.bib17); Mathuret al\.,[2024](https://arxiv.org/html/2605.15768#bib.bib34)\)\. Recent advances in Large Language Models \(LLMs\) provide a strong foundation for building socially capable agents\(Hoppleret al\.,[2022](https://arxiv.org/html/2605.15768#bib.bib18); Leeet al\.,[2024](https://arxiv.org/html/2605.15768#bib.bib26); Anthiset al\.,[2025](https://arxiv.org/html/2605.15768#bib.bib3)\)\.

Early social evaluation frameworks \(e\.g\., Social IQa\(Sapet al\.,[2019](https://arxiv.org/html/2605.15768#bib.bib46)\)and SocialBench\(Chenet al\.,[2024](https://arxiv.org/html/2605.15768#bib.bib9)\)\) focused on static multiple\-choice settings, which fail to capture the dynamic and non\-stationary nature of social interaction\. This limitation motivated dynamic simulation\-based benchmarks, including SOTOPIA\(Zhouet al\.,[2024](https://arxiv.org/html/2605.15768#bib.bib69)\)and AgentSense\(Mouet al\.,[2025](https://arxiv.org/html/2605.15768#bib.bib36)\), which assess agents in open\-ended multi\-turn environments with continuously evolving goals and social relations\.

Existing methods for enhancing social intelligence generally follow two paradigms\. The first category focuses on*offline optimization*\. Sotopia\-π\\pi\(Wanget al\.,[2024b](https://arxiv.org/html/2605.15768#bib.bib50)\)improves performance through data\-centric refinement, while Sotopia\-RL\(Yuet al\.,[2025](https://arxiv.org/html/2605.15768#bib.bib64)\)and SDPO\(Konget al\.,[2025a](https://arxiv.org/html/2605.15768#bib.bib23)\)address credit assignment and preference optimization in multi\-turn dialogues\. Adaptive Mode Learning \(AML\)\(Wanget al\.,[2025a](https://arxiv.org/html/2605.15768#bib.bib59)\)further promotes diverse social reasoning patterns\.

The second category augments inference through offline\-trained external planners\. Methods such as Sotopia\-Ω\\Omega\(Zhanget al\.,[2025](https://arxiv.org/html/2605.15768#bib.bib67)\), DAT\(Liet al\.,[2024](https://arxiv.org/html/2605.15768#bib.bib29)\), and EPO\(Liuet al\.,[2025](https://arxiv.org/html/2605.15768#bib.bib33)\)learn auxiliary planning models from generated data to provide high\-level strategic guidance at test time\. While these approaches highlight the importance of strategies in social interaction, they embed strategic behaviors either within model parameters or fixed planners\.

As a result, introducing new strategies or adapting to evolving social dynamics typically requires data recollection and retraining\. This limitation motivates ourALSO, which enables dynamic and sample\-efficient strategy adaptation through online optimization without offline retraining\.

### 2\.2Prompt Optimization

Prompt Optimization \(PO\) provides an efficient mechanism for adapting agent behavior without parameter fine\-tuning by treating strategies as optimizable instructions\(Wanget al\.,[2025b](https://arxiv.org/html/2605.15768#bib.bib60)\)\. Early PO methods primarily fall into two categories:*LLM\-as\-Optimizer*approaches such as APE\(Zhouet al\.,[2022b](https://arxiv.org/html/2605.15768#bib.bib68)\), OPRO\(Yanget al\.,[2023](https://arxiv.org/html/2605.15768#bib.bib37)\), and Instinct\(Linet al\.,[2024c](https://arxiv.org/html/2605.15768#bib.bib31)\), which iteratively generate and refine prompts, and evolutionary methods including EvoPrompt\(Guoet al\.,[2024](https://arxiv.org/html/2605.15768#bib.bib12)\)and PromptBreeder\(Fernandoet al\.,[2024](https://arxiv.org/html/2605.15768#bib.bib13)\), which explore instruction spaces via mutation and selection\.

Despite their effectiveness in static tasks, most PO methods rely on offline oracles, such as fixed validation sets, to evaluate and rank candidate strategies\(Opsahl\-Onget al\.,[2024b](https://arxiv.org/html/2605.15768#bib.bib38); Konget al\.,[2025b](https://arxiv.org/html/2605.15768#bib.bib24)\)\. This offline paradigm assumes stationary reward distributions and fails to capture the coupled, evolving dynamics of social interaction, where optimal strategies shift in response to adaptive opponents\. Consequently, existing PO frameworks are ill\-suited for online social simulation, motivating the need for adversarial and online strategy optimization as pursued inALSO\.

## 3Problem Formulation

![Refer to caption](https://arxiv.org/html/2605.15768v1/x2.png)Figure 2:Overview ofALSOfor adaptive social strategy learning in LLM\-based multi\-agent social simulation\. Static persona\-driven agents exhibit rigid interactions and fail to achieve social goals \(left\), whileALSOleverages adversarial online strategy selection with surrogate reward modeling to dynamically adapt strategies and enable successful social outcomes \(center–right\)\.We consider a multi\-agent social simulation environment and formulate online social strategy learning as sequential strategy selection under non\-stationary interactions\.

### 3\.1Multi\-Agent Social Simulation Environment

We consider a general multi\-agent social simulation framework in which LLMs act as interactive agents\. The agent set is denoted by𝒩=\{1,2,…,N\}\\mathcal\{N\}=\\\{1,2,\\ldots,N\\\}, where each agenti∈𝒩i\\in\\mathcal\{N\}is characterized by a personabi∈ℬb\_\{i\}\\in\\mathcal\{B\}and a private social goalgi∈𝒢g\_\{i\}\\in\\mathcal\{G\}\. The scenario𝒮\\mathcal\{S\}specifies the social context, including environmental settings and interaction constraints\.

The simulation proceeds in discrete dialogue rounds indexed byl=1,…,Ll=1,\\ldots,L\. At roundll, the active agentiiforms an observationolio\_\{l\}^\{i\}by conditioning on the scenario, interaction historyℋl−1\\mathcal\{H\}\_\{l\-1\}, its persona, and its goal:

oli=Prompt​\(𝒮,ℋl−1,bi,gi\)\.o\_\{l\}^\{i\}=\\text\{Prompt\}\(\\mathcal\{S\},\\mathcal\{H\}\_\{l\-1\},b\_\{i\},g\_\{i\}\)\.\(1\)
The agent then samples an actionalia\_\{l\}^\{i\}based onolio\_\{l\}^\{i\}from the LLM:

ali∼LLM​\(oli\)\.a\_\{l\}^\{i\}\\sim\\text\{LLM\}\(o\_\{l\}^\{i\}\)\.\(2\)
The environment updates the state and augments the history asℋl=ℋl−1∪\{ali\}\\mathcal\{H\}\_\{l\}=\\mathcal\{H\}\_\{l\-1\}\\cup\\\{a\_\{l\}^\{i\}\\\}\.

AfterLLdialogue rounds, an LLM\-based evaluator assesses agent performance alongMMsocial dimensions:

\{dm\(i\)\}m=1M=LLMeval​\(𝒮,ℋ\(L\),bi,gi\),\\\{d\_\{m\}^\{\(i\)\}\\\}\_\{m=1\}^\{M\}=\\text\{LLM\}\_\{\\text\{eval\}\}\(\\mathcal\{S\},\\mathcal\{H\}^\{\(L\)\},b\_\{i\},g\_\{i\}\),\(3\)which are aggregated into a scalar reward:

ri=1M​∑m=1Mϕm​\(dm\(i\)\)\.r\_\{i\}=\\frac\{1\}\{M\}\\sum\_\{m=1\}^\{M\}\\phi\_\{m\}\(d\_\{m\}^\{\(i\)\}\)\.\(4\)However, a key different of social simulation from this setting is its inherentnon\-stationaryarising from strategically adaptive agents\. Let𝐀l=\(al\(1\),…,al\(N\)\)\\mathbf\{A\}\_\{l\}=\(a\_\{l\}^\{\(1\)\},\\ldots,a\_\{l\}^\{\(N\)\}\)denote the joint action set at stepll, and𝐚l−i\\mathbf\{a\}\_\{l\}^\{\-i\}the actions of all agents except agentii\. LetR​\(sl,al\(i\),𝐚l−i\)R\(s\_\{l\},a\_\{l\}^\{\(i\)\},\\mathbf\{a\}\_\{l\}^\{\-i\}\)denote the instantaneous turn\-level reward function under statesls\_\{l\}and joint actions set\. The expected step/turn level reward for agentiiis given by

𝔼​\[rl\(i\)∣sl,al\(i\)\]=∑𝐚l−i∈A−iR​\(sl,al\(i\),𝐚l−i\)​∏j≠iπj​\(aj∣sl\),\\mathbb\{E\}\[r\_\{l\}^\{\(i\)\}\\mid s\_\{l\},a\_\{l\}^\{\(i\)\}\]=\\sum\_\{\\mathbf\{a\}\_\{l\}^\{\-i\}\\in A^\{\-i\}\}R\(s\_\{l\},a\_\{l\}^\{\(i\)\},\\mathbf\{a\}\_\{l\}^\{\-i\}\)\\prod\_\{j\\neq i\}\\pi\_\{j\}\(a\_\{j\}\\mid s\_\{l\}\),\(5\)
whereπj\\pi\_\{j\}denotes the evolving policy of agentjj\.

In practice, agents only observe the realizedrir\_\{i\}, while the underlying expected reward𝔼​\[rl\(i\)∣sl,al\(i\)\]\\mathbb\{E\}\[r\_\{l\}^\{\(i\)\}\\mid s\_\{l\},a\_\{l\}^\{\(i\)\}\]remains implicit and shifts continuously as opponent policiesπj\\pi\_\{j\}evolve\. This creates a fundamental learning challenge:

ri=𝔼​\[rl\(i\)∣sl,al\(i\)\]\+ϵt,where​ϵt​is non\-stationary,r\_\{i\}=\\mathbb\{E\}\[r\_\{l\}^\{\(i\)\}\\mid s\_\{l\},a\_\{l\}^\{\(i\)\}\]\+\\epsilon\_\{t\},\\quad\\text\{where \}\\epsilon\_\{t\}\\text\{ is non\-stationary\},\(6\)whereϵt\\epsilon\_\{t\}captures both stochastic noise from LLM generation and distributional shift induced by co\-evolving opponent policies\. This motivates casting the strategy learning problem as adversarial online optimization \(Section 3\.2\), where no stationarity of the reward signal is assumed\.

### 3\.2Online Social Strategy Learning Problem

We model social strategy adaptation by discretizing the strategy space into a finite set ofKKstrategic instructions \(e\.g\., cooperation and competition\), each corresponding to a bandit arm\. Each instruction encodes high\-level behavioral principles over goal orientation and interaction style, forming an interpretable strategy space for online learning\. In this formulation, the agent’s action at each turn is to*select an arm*,i\.e\., choose a strategy instruction to guide its behavior throughout the social interaction\. This directly casts online strategy selection under evolving social dynamics as an adversarial multi\-armed bandit problem\.

Objective\.Over roundst=1,…,Tt=1,\\ldots,T\(each corresponding to a completeLL\-turn social simulation episode\), agentiiselects a strategy armkt∈\[K\]k\_\{t\}\\in\[K\]and receives an trun\-level rewardrkt\(t\)∈\[0,1\]r\_\{k\_\{t\}\}^\{\(t\)\}\\in\[0,1\]from the evaluator \(Eq\. \([4](https://arxiv.org/html/2605.15768#S3.E4)\)\)\. A natural*target objective*is to minimize the cumulative pseudo\-regret

R¯T=𝔼​\[maxk∈\[K\]​∑t=1Trk\(t\)−∑t=1Trkt\(t\)\],\\bar\{R\}\_\{T\}=\\mathbb\{E\}\\\!\\left\[\\max\_\{k\\in\[K\]\}\\sum\_\{t=1\}^\{T\}r\_\{k\}^\{\(t\)\}\-\\sum\_\{t=1\}^\{T\}r\_\{k\_\{t\}\}^\{\(t\)\}\\right\],\(7\)which measures performance relative to the best fixed strategy in hindsight\. In our setting, however, the reward process is induced by co\-evolving multi\-agent interactions and an LLM evaluator; thus,ALSOadopts adversarial online learning as a*design rationale*for robustness rather than claiming a formal regret guarantee\.

## 4Methodology

This section presentsALSO, our adversarial online approach to strategy optimization in LLM\-based multi\-agent social simulations, as displayed in Figure[2](https://arxiv.org/html/2605.15768#S3.F2)\.

### 4\.1Problem Setting

We formalize dynamic strategy instruction optimization in two\-agent \(N=2N=2\) LLM\-based social simulations\. Each agenti∈\{1,2\}i\\in\\\{1,2\\\}augments its original personabi0b\_\{i\}^\{0\}with one ofKKpredefined social strategiesΣ=\{σ1,…,σK\}\\Sigma=\\\{\\sigma\_\{1\},\\ldots,\\sigma\_\{K\}\\\}\(details in Appendix[B\.2](https://arxiv.org/html/2605.15768#A2.SS2)\)\.

In the original framework, the agent’s action is generated using a fixed persona:

at\(i\)∼LLM​\(𝒮,ℋt−1,bi0,gi\)\.a\_\{t\}^\{\(i\)\}\\sim\\text\{LLM\}\\bigl\(\\mathcal\{S\},\\ \\mathcal\{H\}\_\{t\-1\},\\ b\_\{i\}^\{0\},\\ g\_\{i\}\\bigr\)\.\(8\)
InALSO, we dynamically augment the persona at each turnttby appending a selected strategy instructionσkt\(i\)\\sigma\_\{k\_\{t\}\}^\{\(i\)\}from the strategy spaceΣ\\Sigma\. This creates an enhanced persona

bi\(t\)=bi0⊕σkt\(i\),b\_\{i\}^\{\(t\)\}=b\_\{i\}^\{0\}\\oplus\\sigma\_\{k\_\{t\}\}^\{\(i\)\},\(9\)where⊕\\oplusdenotes textual concatenation\. The resulting personabi\(t\)b\_\{i\}^\{\(t\)\}incorporates both the agent’s original identity \(demographics, personality, values, etc\.\) and the high\-level behavioral guidance provided by the strategy \(e\.g\., collaborative problem\-solving, firm bargaining, or strategic withholding\)\. The LLM then generates the next action conditioned on this augmented context:

at\(i\)∼LLM​\(𝒮,ℋt−1,bi\(t\),gi\)\.a\_\{t\}^\{\(i\)\}\\sim\\text\{LLM\}\\bigl\(\\mathcal\{S\},\\ \\mathcal\{H\}\_\{t\-1\},\\ b\_\{i\}^\{\(t\)\},\\ g\_\{i\}\\bigr\)\.\(10\)
This augmentation allows the LLM to adapt its responses based on the chosen strategy \(e\.g\., collaboration or information withholding\) without retraining the model\.

Non\-stationarity\.In social simulation, the reward process is inherently non\-stationary: the counterpart agent adapts to the ego agent’s behavior, and the dialogue state evolves over turns\. Therefore, we do*not*assume rewards are i\.i\.d\. or drawn from a fixed distribution across optimization iterations; instead, we adopt an adversarial online learning perspective as a*design rationale*for robust strategy selection under distribution shift\.

Feedback\.After each round, we use an LLM evaluator to provide per\-turn rewardsrt\(i\)=ρ​\(st,at\(i\)\)r\_\{t\}^\{\(i\)\}=\\rho\(s\_\{t\},a\_\{t\}^\{\(i\)\}\)overMMdimensions, which are normalized byϕm\\phi\_\{m\}\(Table[4](https://arxiv.org/html/2605.15768#A1.T4)\), following prior LLM\-evaluator\-based social simulation work such as Sotopia\-Ω\\Omega\(Zhanget al\.,[2025](https://arxiv.org/html/2605.15768#bib.bib67)\)and AML\(Wanget al\.,[2025b](https://arxiv.org/html/2605.15768#bib.bib60)\)\. We use the resulting scalar reward \(after aggregation/normalization\) as the bandit feedback in Algorithm[1](https://arxiv.org/html/2605.15768#alg1)\.

### 4\.2ALSO: Adversarial Online Strategy Optimization

To address strategy optimization in a large, discrete space under non\-stationary social dynamics, we proposeALSO\.ALSOfollows an adversarial online learning design: it makes no stationarity assumptions about rewards, uses randomized selection to remain robust to shifting partner behaviors, and incorporates a recency mechanism to track drift\. Concretely,ALSOcombines an exponential\-weights selector with a lightweight neural surrogate that generalizes sparse feedback across strategies\. The LLM policy remains frozen; only the surrogate network is trained online\.

The core logic ofALSOis organized into the following phases \(including a preprocessing step\): We denote byata\_\{t\}the focal agent’s utterance and byoto\_\{t\}the counterpart response;ℋ\(t\)\\mathcal\{H\}^\{\(t\)\}is the dialogue history up to turntt\.

Algorithm 1Adversarial Online Strategy Optimization0:Strategy space

Σ=\{σ1,…,σK\}\\Sigma=\\\{\\sigma\_\{1\},\\ldots,\\sigma\_\{K\}\\\}\(candidate strategy instructions\), base persona

b0b^\{0\}\(fixed persona text\), frozen embedding model

g​\(⋅\)g\(\\cdot\), trainable value network

fθf\_\{\\theta\}, learning rate

η\>0\\eta\>0, decay factor

λ∈\(0,1\]\\lambda\\in\(0,1\], batch size

BB
0:Sequence of selected strategies

\{σkt\}t=1T\\\{\\sigma\_\{k\_\{t\}\}\\\}\_\{t=1\}^\{T\}\(the strategy chosen at each turn

tt\)

1:Precompute augmented\-persona embeddings:

2:For all

kk, form

b\(k\)←b0⊕σkb^\{\(k\)\}\\leftarrow b^\{0\}\\oplus\\sigma\_\{k\}and compute

𝐛k←g​\(b\(k\)\)\\mathbf\{b\}\_\{k\}\\leftarrow g\(b^\{\(k\)\}\)
3:Initialize:

Sk\(0\)←0S\_\{k\}^\{\(0\)\}\\leftarrow 0for all

kk; replay buffer

𝒟←∅\\mathcal\{D\}\\leftarrow\\emptyset; history

ℋ\(0\)←∅\\mathcal\{H\}^\{\(0\)\}\\leftarrow\\emptyset
4:forturn

t=1t=1to

TTdo

5:Context encoding:

𝐜\(t\)←g​\(ℋ\(t−1\)\)\\mathbf\{c\}^\{\(t\)\}\\leftarrow g\(\\mathcal\{H\}^\{\(t\-1\)\}\)
6:Value prediction:compute

\(𝐱k\(t\),v^k\(t\)\)\(\\mathbf\{x\}\_\{k\}^\{\(t\)\},\\hat\{v\}\_\{k\}^\{\(t\)\}\)for all

kkvia Eq\.[11](https://arxiv.org/html/2605.15768#S4.E11)

7:Strategy selection:compute

π\(t\)\\pi^\{\(t\)\}and sample

ktk\_\{t\}via Eq\.[12](https://arxiv.org/html/2605.15768#S4.E12)

8:Interaction & feedback:form augmented persona

b\(t\)←b0⊕σktb^\{\(t\)\}\\leftarrow b^\{0\}\\oplus\\sigma\_\{k\_\{t\}\}and execute

b\(t\)b^\{\(t\)\}to generate focal\-agent action

ata\_\{t\}; observe counterpart response

oto\_\{t\}and reward

rtr\_\{t\}; update

ℋ\(t\)←ℋ\(t−1\)∪\{at,ot\}\\mathcal\{H\}^\{\(t\)\}\\leftarrow\\mathcal\{H\}^\{\(t\-1\)\}\\cup\\\{a\_\{t\},o\_\{t\}\\\}
9:Surrogate update:add

\(𝐱kt\(t\),rt\)\(\\mathbf\{x\}\_\{k\_\{t\}\}^\{\(t\)\},r\_\{t\}\)to

𝒟\\mathcal\{D\}; sample a minibatch of size

BBfrom

𝒟\\mathcal\{D\}and update

fθf\_\{\\theta\}via MSE

10:Score smoothing:update

Sk\(t\)S\_\{k\}^\{\(t\)\}for all

kkvia Eq\.[13](https://arxiv.org/html/2605.15768#S4.E13)

11:endfor

Arm Space \(Alg\.[1](https://arxiv.org/html/2605.15768#alg1), Line 1–2\)\.Before online interaction begins, we precompute an embedding for each*augmented persona*obtained by appending a candidate strategy to the base persona for computer efficiency\. Concretely, for eachk∈\[K\]k\\in\[K\], we formb\(k\)=b0⊕σkb^\{\(k\)\}=b^\{0\}\\oplus\\sigma\_\{k\}and compute𝐛k=g​\(b\(k\)\)\\mathbf\{b\}\_\{k\}=g\(b^\{\(k\)\}\)\. These embeddings are fixed across turns and serve as the strategy\-specific representation used by the surrogate\.

Context Encoding & Prediction \(Alg\.[1](https://arxiv.org/html/2605.15768#alg1), Lines 5–6\)\.At the beginning of each interaction turntt, the optimal strategy depends on the dialogue state\. We use the frozen embedding modelg​\(⋅\)g\(\\cdot\)to encode the dialogue historyℋ\(t−1\)\\mathcal\{H\}^\{\(t\-1\)\}into a context vector𝐜\(t\)\\mathbf\{c\}^\{\(t\)\}\. For each candidate strategyσk∈Σ\\sigma\_\{k\}\\in\\Sigma, we concatenate the precomputed augmented\-persona embedding𝐛k\\mathbf\{b\}\_\{k\}with𝐜\(t\)\\mathbf\{c\}^\{\(t\)\}to form𝐱k\(t\)\\mathbf\{x\}\_\{k\}^\{\(t\)\}\. A trainable value networkfθ​\(⋅\)f\_\{\\theta\}\(\\cdot\)predicts the expected rewardv^k\(t\)\\hat\{v\}\_\{k\}^\{\(t\)\}for all arms:

𝐱k\(t\)=\[𝐛k;𝐜\(t\)\],v^k\(t\)=fθ​\(𝐱k\(t\)\)\.\\mathbf\{x\}\_\{k\}^\{\(t\)\}=\[\\mathbf\{b\}\_\{k\};\\mathbf\{c\}^\{\(t\)\}\],\\qquad\\hat\{v\}\_\{k\}^\{\(t\)\}=f\_\{\\theta\}\(\\mathbf\{x\}\_\{k\}^\{\(t\)\}\)\.\(11\)This provides a data\-efficient inductive bias in early online learning, where only a small number of interactions are available\.

Strategy Selection \(Alg\.[1](https://arxiv.org/html/2605.15768#alg1), Lines 7\)\.We maintain a cumulative scoreSkS\_\{k\}for each arm and sample strategies from an exponential\-weights distribution:

πk\(t\)∝exp⁡\(η​Sk\(t−1\)\),kt∼Categorical​\(π\(t\)\)\.\\pi\_\{k\}^\{\(t\)\}\\propto\\exp\\bigl\(\\eta\\,S\_\{k\}^\{\(t\-1\)\}\\bigr\),\\qquad k\_\{t\}\\sim\\text\{Categorical\}\(\\pi^\{\(t\)\}\)\.\(12\)Randomized selection is essential in the adversarial setting, where a greedy policy can be exploited or can overfit to transient dynamics\.

Interaction & Surrogate Update \(Alg\.[1](https://arxiv.org/html/2605.15768#alg1), Lines 8–9\)\.After samplingσkt\\sigma\_\{k\_\{t\}\}and observing the per\-turn rewardrtr\_\{t\}, we store\(𝐱kt\(t\),rt\)\(\\mathbf\{x\}\_\{k\_\{t\}\}^\{\(t\)\},r\_\{t\}\)in a replay buffer𝒟\\mathcal\{D\}and updatefθf\_\{\\theta\}by minimizing an MSE loss\. The surrogate improves sample efficiency by transferring supervision to semantically related strategies\.

Score estimation\.In classical EXP3, the exponential\-weights sampling distribution is constructed from per\-arm cumulative rewards\. In our online social simulation setting, however, we only observe feedback for the selected strategy at each turn, and many candidate strategies \(or their paraphrased variants\) may never be played\. This makes direct per\-arm reward accumulation highly sample\-inefficient, especially when the effective arm space is large\. Motivated by prior prompt optimization work, we therefore use a lightweight neural surrogatefθf\_\{\\theta\}over pretrained embeddings to estimate scores for*all*arms from the current dialogue context\. This provides dense score estimates to construct the exponential\-weights distribution, and propagates sparse feedback across semantically related strategies while keeping the base LLM frozen\.

Score Smoothing \(Alg\.[1](https://arxiv.org/html/2605.15768#alg1), Lines 10\)\.To explicitly track non\-stationarity, we apply an exponential decay factorλ∈\(0,1\]\\lambda\\in\(0,1\]\(set toλ\\lambda= 0\.9 in all experiments\) so that recent evidence dominates historical estimates\. This keepsSk\(t\)S\_\{k\}^\{\(t\)\}responsive to partner shifts while preventing outdated interactions from dominating the strategy distribution:

Sk\(t\)=λ​Sk\(t−1\)\+v^k\(t\)\.S\_\{k\}^\{\(t\)\}=\\lambda S\_\{k\}^\{\(t\-1\)\}\+\\hat\{v\}\_\{k\}^\{\(t\)\}\.\(13\)

## 5Experiments

This section evaluatesALSOon LLM\-based social simulation benchmarks to assess its effectiveness for online strategy adaptation under non\-stationary interactions\. The codes ofALSOare available at\\urlhttps://github\.com/Babylonehy/ALSO

### 5\.1Experimental Setting

Benchmarks\.We evaluateALSOon Sotopia\(Zhouet al\.,[2024](https://arxiv.org/html/2605.15768#bib.bib69)\)and its challenging subset Sotopia\-Hard\. Sotopia contains 90 two\-agent social scenarios spanning negotiation, collaboration, and competition, while Sotopia\-Hard includes 14 scenarios emphasizing complex and conflicting social dynamics\. We extend the original implementation to support dynamic strategy injection\. Our strategy instruction pool consists of 12 predefined strategies \(Appendix[B\.2](https://arxiv.org/html/2605.15768#A2.SS2)\), shared across all optimization\-based methods for fair comparison\.

Online Interaction Protocol\.Each episode instantiates a single scenario with fixed personas and goals and runs up to 20 dialogue turns\. At each turn, agents select a strategy instruction fromΣ\\Sigmaand append it to their base persona before generating responses\. Unless otherwise specified, we adopt a bilateral online setting where each agent maintains an independent optimizer and updates it solely from its own interaction feedback\.

Baselines\.We compare against representative methods covering static prompting, evolutionary prompt generation, and online prompt optimization:Vanilla\(no strategy augmentation\),EvoPrompt\(Guoet al\.,[2024](https://arxiv.org/html/2605.15768#bib.bib12)\),OPRO\(Yanget al\.,[2023](https://arxiv.org/html/2605.15768#bib.bib37)\), andINSTINCT\(Linet al\.,[2024b](https://arxiv.org/html/2605.15768#bib.bib21)\)\.

All baselines operate under matched strategy pools and comparable reward\-query budgets\.

Table 1:Per\-episode LLM call budget accounting \(two\-agent episode with horizonTTturns\)\. “Agent” counts calls to the dialogue LLMs that generate actions; “Evaluator” counts calls to the per\-turn reward model; “Optimizer” counts extra LLM calls used to generate or mutate prompts\. Methods without an LLM optimizer have 0 optimizer calls\.Models and Training Signals\.All methods use DeepSeek\-V3\.2 for agent interactions\. Following prior work\(Zhanget al\.,[2025](https://arxiv.org/html/2605.15768#bib.bib67); Wanget al\.,[2025b](https://arxiv.org/html/2605.15768#bib.bib60)\), we employ an LLM\-based intermediate evaluator to provide per\-turn shaping rewards for online updates, while keeping the reporting judge separate\.ALSOupdates its surrogate online at each turn using shaping rewards without modifying the underlying LLM\.

Evaluation Protocol\.Final performance is assessed using GPT\-4o\(Hurstet al\.,[2024](https://arxiv.org/html/2605.15768#bib.bib19)\)with standard Sotopia\-Eval prompts, ensuring consistent dialogue\-level evaluation across methods\. We additionally report cross\-scenario generalization results in Section[6](https://arxiv.org/html/2605.15768#S6)\.

Efficiency and Budget\.LLM call budgets per episode are summarized in Table[1](https://arxiv.org/html/2605.15768#S5.T1)\.ALSOrequires no LLM fine\-tuning or external optimizer calls, relying only on lightweight online surrogate updates\.

Additional implementation details are provided in Appendix[C](https://arxiv.org/html/2605.15768#A3)\.

### 5\.2Experiment Result

Main Results\.Table[2](https://arxiv.org/html/2605.15768#S5.T2)reports the main results onSotopia\-AllandSotopia\-Hardunder theBilateralsetting\.ALSOachieves the highestOverallscore on both benchmarks\. OnSotopia\-Hard,ALSOimprovesOverallfrom 3\.02 \(Vanilla\) to 3\.53 \(\+16\.60%\) and surpasses the strongest baseline by \+2\.92% \(3\.53 vs\. 3\.43\)\. OnSotopia\-All,ALSOranks first inOverall\(3\.89\), although the margin over the strongest baseline is smaller \(\+0\.98%; 3\.89 vs\. 3\.85\)\.

Source of Gains\.The improvements onSotopia\-Hardare largely driven by theRelationshipdimension:ALSOincreasesRelfrom 1\.32 to 2\.43 \(\+83\.79% over Vanilla\) and outperforms the best baseline by \+12\.59% \(2\.43 vs\. 2\.16\)\. Notably, this substantial relational gain is accompanied by improvements in bothGoal\(7\.11, \+2\.79% over the strongest baseline\) andKnow\(5\.47, \+0\.52%\), helping rule out a trivial “rapport\-only” trade\-off\.

Knowledge\.OnSotopia\-All,ALSOalso achieves a slight improvement inKnowover the strongest baseline \(6\.14 vs\. 6\.09; \+0\.73%\)\. We further analyze this dimension by reporting per\-scenario distributions and examining whether strategy injection reduces information disclosure in cooperative settings\.

Table 2:\(1\)Boldindicates 1st rank,underlineindicates 2nd rank\. \(2\) Results are reported as Mean±\\pmStandard Error \(SE\)\. \(3\)Improv\. vs\. Best Baselinecompares Ours with the best baseline\. Negative value indicates slight gap behind the SOTA\.
### 5\.3Analysis

Case Study\.Figure[3](https://arxiv.org/html/2605.15768#S5.F3)presents a qualitative comparison in a high\-conflict eviction scenario where static personas typically lead to dialogue deadlock\. Under Vanilla prompting, agents remain trapped in repetitive insist–deny exchanges, resulting in stagnation and zero reward\. In contrast,ALSOdynamically alters the interaction trajectory through targeted strategy switches\. At selected turn, the selectedValidate Before Redirectingstrategy encourages acknowledgment of opposing concerns before introducing a compromise, while the subsequentGRITstrategy at Turn 8 elicits a concrete, low\-risk concession\. These coordinated adaptations steer the dialogue away from the local deadlock and toward cooperative resolution, ultimately achieving successful agreement \(R≈0\.89R\\approx 0\.89\)\.

Non\-Stationary Strategy Reward Drift\.Figure[4](https://arxiv.org/html/2605.15768#S5.F4)illustrates the temporal evolution of normalized rewards for individual strategies across dialogue turns\. Despite fixing the same strategy, rewards exhibit substantial drift and variability over time, with variance ranging fromσ2=0\.004\\sigma^\{2\}=0\.004to0\.0150\.015\. This pronounced fluctuation reflects co\-adapt agent behaviors and provides empirical evidence of inherent non\-stationarity in social simulation, motivating adversarial online strategy optimization\. Strategy convergence and distribution see in Appendix[C\.6](https://arxiv.org/html/2605.15768#A3.SS6)\.

![Refer to caption](https://arxiv.org/html/2605.15768v1/x3.png)Figure 3:Conflict Resolution\.Comparison of dialogue trajectories at the critical deadlock phase \(Turns 7–9\), highlighting turn\-level strategy switches and their effect on reward/relationship\.![Refer to caption](https://arxiv.org/html/2605.15768v1/x4.png)Figure 4:Strategy Reward Drift Over Dialogue Turns\.Each line represents a different strategy \(arm\), showing how the average*normalized*reward varies across turns within episodes\.

## 6Ablation Study

Component\-wise AblationTo isolate the contribution of each component of ALSO, we conduct a comprehensive component\-wise ablation in which one design element is removed or replaced at a time\. The results, summarized in Table[3](https://arxiv.org/html/2605.15768#S6.T3), indicate that the neural surrogate is the most influential component: its removal causes a degradation of0\.580\.58on the Overall metric and34\.9%34\.9\\%relative on the Relationship dimension\. Score smoothing exerts the strongest effect on the Relationship dimension specifically, where its removal reduces the score from3\.073\.07to2\.252\.25\. Replacing the EXP3\-style selector with anε\\varepsilon\-greedy alternative degrades Overall to3\.613\.61, supporting the necessity of randomized exploration under non\-stationary co\-adaptation\. The contextual embedding likewise contributes a non\-trivial margin across all four dimensions\.

Table 3:Component\-wise ablation ofALSO\. Each row removes or replaces a single design element\.Single vs\. Bilateral Optimization\.We compare bilateral strategy optimization with unilateral variants that adapt strategies for only one agent \(P1\-only or P2\-only\)\. Figure[5](https://arxiv.org/html/2605.15768#S6.F5)shows that bilateral optimization consistently achieves higher overall performance across both Qwen\-2\.5\-72B\-Instruct and DeepSeek\-V3\.2, with statistically significant improvements\. Dimension\-wise, gains are most pronounced inRelationshipandKnowledge, indicating enhanced cooperation and information exchange when both agents adapt strategies\. Suggesting symmetric online adaptation better captures the co\-evolving nature of social interactions\.

![Refer to caption](https://arxiv.org/html/2605.15768v1/x5.png)Figure 5:Bilateral optimization improves social interactions\. Comparison of P1\-only, P2\-only, and bilateral approaches on Qwen\-2\.5\-72B\-Instruct \(a–b\) and DeepSeek\-V3\.2 \(c–d\)\. Left: overall scores; right: dimension\-wise with percentage gains\. Significance:p<0\.001p<0\.001\(Qwen\),p<0\.01p<0\.01\(DeepSeek\)\.All experiments below are conducted on the Sotopia\-Hard benchmark, consisting of 14 challenging scenarios\. For each scenario, we sample a single episode to evaluate performance across methods\.

Cross\-Scenario Generalization\.Beyond the scenario\-parallel setting used in our main experiments, we further evaluate whether the learned strategy\-selection mechanism transfers to unseen social contexts\. We split Sotopia\-Hard into disjoint training and test sets, train the surrogate bandit on the training scenarios, and evaluate it on unseen test scenarios\. Figure[6](https://arxiv.org/html/2605.15768#S6.F6)shows that zero\-shot transfer improves over an online\-from\-scratch baseline both at the aggregate and per\-scenario levels\. Averaged over the 7 unseen test scenarios, zero\-shot transfer reaches a goal score of 7\.14, compared with 6\.79 for the scratch baseline, yielding a relative improvement of 5\.3%\. It also improves the overall score from 3\.17 to 3\.60 \(\+13\.5%\)\. These gains are broadly consistent across scenarios, suggesting thatALSOcaptures transferable social interaction patterns rather than relying purely on scenario\-specific adaptation\.

![Refer to caption](https://arxiv.org/html/2605.15768v1/x6.png)Figure 6:Cross\-scenario generalization results\. \(a\) Per\-scenario goal and overall scores on unseen test scenarios, comparing online\-from\-scratch learning, zero\-shot transfer, and finetuning\. \(b\) Average performance across all 7 unseen test scenarios\. Zero\-shot transfer outperforms the scratch baseline on both goal score \(7\.14 vs\. 6\.79, \+5\.3%\) and overall score \(3\.60 vs\. 3\.17, \+13\.5%\)\.![Refer to caption](https://arxiv.org/html/2605.15768v1/x7.png)Figure 7:Performance across heterogeneous model, measured by only final P1 score\. \(a\) Baseline performance\. \(b\) Performance withALSO\. Green annotations denote relative improvement over baseline\.ALSOyields consistent gains across heterogeneous pairings\.Heterogeneous Model Pairing\.We further evaluateALSOon heterogeneous dyads formed by DeepSeek\-V3\.2, Qwen\-2\.5\-72B\-Instruct, and GPT\-4o\-mini\. Figure[7](https://arxiv.org/html/2605.15768#S6.F7)shows thatALSOconsistently improves performance across all cross\-model pairings\. Crucially, the gains are not tied to any specific backbone combination or model scale\.ALSOremains effective for all pairings, indicating that its benefit reflects a general optimization effect rather than pair\-specific tuning\.

## 7Conclusion

We proposedALSO, an adversarial online framework for dynamically selecting strategy instructions for LLM\-based social agents under non\-stationary interactions\. By combining randomized bandit\-based selection with a lightweight surrogate reward model,ALSOefficiently adapts strategies while keeping the underlying LLM frozen\. Experiments on Sotopia and Sotopia\-Hard demonstrate consistent improvements in social performance, with the largest gains in challenging scenarios, particularly on relationship outcomes\. These results suggest adversarial online strategy optimization as a practical and scalable approach to enhancing social intelligence without costly fine\-tuning\.

## Impact Statement

This paper presents work whose goal is to advance the field of Machine Learning\. There are many potential societal consequences of our work, none which we feel must be specifically highlighted here\.

## Acknowledgements

This work was supported in part by the National Natural Science Foundation of China \(Grant Nos\. U23B2049, 22527901, 62506319 and 62477012\), the Guangdong Basic and Applied Basic Research Foundation \(Grant No\. 2026A1515030032\), the Shenzhen Science and Technology Program \(Grant No\. JCYJ20250604141031003\), the Pearl River Talent Program of Guangdong Province \(Grant No\. 2024QN11X069\), and the AI for Science Program of the Shanghai Municipal Commission of Economy and Informatization, China \(Grant No\. 2025\-GZL\-RGZN\-BTBX\-01014\)\.

## References

- J\. R\. Anthis, R\. Liu, S\. M\. Richardson, A\. C\. Kozlowski, B\. Koch, E\. Brynjolfsson, J\. Evans, and M\. S\. Bernstein \(2025\)Position: LLM social simulations are a promising research method\.InForty\-second International Conference on Machine Learning Position Paper Track,External Links:[Link](https://openreview.net/forum?id=cRBg1dtj7o)Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p1.1)\.
- R\. Axelrod and W\. D\. Hamilton \(1981\)The evolution of cooperation\.science211\(4489\),pp\. 1390–1396\.Cited by:[Table 5](https://arxiv.org/html/2605.15768#A2.T5.4.6.5.1)\.
- J\. Bisbee, J\. D\. Clinton, C\. Dorff, B\. Kenkel, and J\. M\. Larson \(2024\)Synthetic replacements for human survey data? the perils of large language models\.Political Analysis32\(4\),pp\. 401–416\.Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p1.1)\.
- H\. Chen, H\. Chen, M\. Yan, W\. Xu, G\. Xing, W\. Shen, X\. Quan, C\. Li, J\. Zhang, and F\. Huang \(2024\)SocialBench: sociality evaluation of role\-playing conversational agents\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 2108–2126\.External Links:[Link](https://aclanthology.org/2024.findings-acl.125/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.125)Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p2.1)\.
- J\. M\. Epstein \(2012\)Generative social science: studies in agent\-based computational modeling\.InGenerative Social Science,Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p1.1)\.
- C\. Fernando, D\. Banarse, H\. Michalewski, S\. Osindero, and T\. Rocktäschel \(2024\)Promptbreeder: self\-referential self\-improvement via prompt evolution\.InProceedings of the 41st International Conference on Machine Learning,pp\. 13481–13544\.Cited by:[§2\.2](https://arxiv.org/html/2605.15768#S2.SS2.p1.1)\.
- R\. Fisher, W\. L\. Ury, and B\. Patton \(2011\)Getting to yes: negotiating agreement without giving in\.Penguin\.Cited by:[Table 5](https://arxiv.org/html/2605.15768#A2.T5.4.2.1.2),[Table 5](https://arxiv.org/html/2605.15768#A2.T5.4.3.2.2)\.
- E\. Goffman \(1959\)The presentation of self in everyday life\.Anchor books,Bantam Doubleday Dell Publishing Group,New York, NY\(en\)\.Cited by:[Table 5](https://arxiv.org/html/2605.15768#A2.T5.4.4.3.2)\.
- Q\. Guo, R\. Wang, J\. Guo, B\. Li, K\. Song, X\. Tan, G\. Liu, J\. Bian, and Y\. Yang \(2024\)Connecting large language models with evolutionary algorithms yields powerful prompt optimizers\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=ZG3RaNIsO8)Cited by:[§C\.2](https://arxiv.org/html/2605.15768#A3.SS2.p1.1),[§1](https://arxiv.org/html/2605.15768#S1.p3.1),[§2\.2](https://arxiv.org/html/2605.15768#S2.SS2.p1.1),[§5\.1](https://arxiv.org/html/2605.15768#S5.SS1.p3.1)\.
- H\. Gweon, J\. Fan, and B\. Kim \(2023\)Socially intelligent machines that learn from humans and help humans learn\.Philosophical Transactions of the Royal Society A381\(2251\),pp\. 20220048\.External Links:[Link](https://doi.org/10.1098/rsta.2022.0048)Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p1.1)\.
- S\. S\. Hoppler, R\. Segerer, and J\. Nikitin \(2022\)The six components of social interactions: actor, partner, relation, activities, context, and evaluation\.Frontiers in psychology12,pp\. 743074\.Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p1.1)\.
- A\. Hurst, A\. Lerer, A\. P\. Goucher, A\. Perelman, A\. Ramesh, A\. Clark, A\. Ostrow, A\. Welihinda, A\. Hayes, A\. Radford,et al\.\(2024\)Gpt\-4o system card\.arXiv preprint arXiv:2410\.21276\.Cited by:[§5\.1](https://arxiv.org/html/2605.15768#S5.SS1.p6.1)\.
- E\. Hwang, Y\. Yin, G\. Carenini, P\. West, and V\. Shwartz \(2025\)Infusing theory of mind into socially intelligent llm agents\.arXiv preprint arXiv:2509\.22887\.Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p2.1)\.
- \[14\]R\. Kirk, I\. Mediratta, C\. Nalmpantis, J\. Luketina, E\. Hambro, E\. Grefenstette, and R\. RaileanuUnderstanding the effects of rlhf on llm generalisation and diversity\.InThe Twelfth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p2.1)\.
- A\. Kong, W\. Ma, S\. Zhao, Y\. Li, Y\. Wu, K\. Wang, X\. Liu, Q\. Li, Y\. Qin, and F\. Huang \(2025a\)SDPO: segment\-level direct preference optimization for social agents\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 12409–12423\.External Links:[Link](https://aclanthology.org/2025.acl-long.607/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.607),ISBN 979\-8\-89176\-251\-0Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p3.1)\.
- M\. Kong, Z\. Wang, Y\. Shu, and Z\. Dai \(2025b\)Meta\-prompt optimization for llm\-based sequential decision making\.arXiv preprint arXiv:2502\.00728\.Cited by:[§2\.2](https://arxiv.org/html/2605.15768#S2.SS2.p2.1)\.
- T\. Lattimore and C\. Szepesvári \(2020\)Bandit algorithms\.Cambridge University Press\.Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p4.1)\.
- M\. Lee, M\. Srivastava, A\. Hardy, J\. Thickstun, E\. Durmus, A\. Paranjape, I\. Gerard\-Ursin, X\. L\. Li, F\. Ladhak, F\. Rong, R\. E\. Wang, M\. Kwon, J\. S\. Park, H\. Cao, T\. Lee, R\. Bommasani, M\. Bernstein, and P\. Liang \(2024\)Evaluating human\-language model interaction\.External Links:2212\.09746,[Link](https://arxiv.org/abs/2212.09746)Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p1.1)\.
- H\. Li, Y\. Chong, S\. Stepputtis, J\. Campbell, D\. Hughes, C\. Lewis, and K\. Sycara \(2023\)Theory of mind for multi\-agent collaboration via large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 180–192\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.13/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.13)Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p2.1)\.
- K\. Li, Y\. Wang, F\. Viégas, and M\. Wattenberg \(2024\)Dialogue action tokens: steering language models in goal\-directed dialogue with a multi\-turn planner\.arXiv preprint arXiv:2406\.11978\.Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p4.1)\.
- X\. Lin, Z\. Dai, A\. Verma, S\. Ng, P\. Jaillet, and B\. K\. H\. Low \(2024a\)Prompt optimization with human feedback\.arXiv preprint arXiv:2405\.17346\.Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p3.1)\.
- X\. Lin, Z\. Wu, Z\. Dai, W\. Hu, Y\. Shu, S\. Ng, P\. Jaillet, and B\. K\. H\. Low \(2024b\)Use your INSTINCT: instruction optimization for llms using neural bandits coupled with transformers\.InProc\. ICML,Cited by:[§C\.4](https://arxiv.org/html/2605.15768#A3.SS4.p1.1),[§1](https://arxiv.org/html/2605.15768#S1.p3.1),[§5\.1](https://arxiv.org/html/2605.15768#S5.SS1.p3.1)\.
- X\. Lin, Z\. Wu, Z\. Dai, W\. Hu, Y\. Shu, S\. Ng, P\. Jaillet, and B\. K\. H\. Low \(2024c\)Use your INSTINCT: instruction optimization for llms using neural bandits coupled with transformers\.InProc\. ICML,Cited by:[§2\.2](https://arxiv.org/html/2605.15768#S2.SS2.p1.1)\.
- X\. Liu, K\. Wang, Y\. Li, Y\. Wu, W\. Ma, A\. Kong, F\. Huang, J\. Jiao, and J\. Zhang \(2025\)EPO: explicit policy optimization for strategic reasoning in llms via reinforcement learning\.arXiv preprint arXiv:2502\.12486\.Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p4.1)\.
- L\. Mathur, P\. P\. Liang, and L\. Morency \(2024\)Advancing social intelligence in AI agents: technical challenges and open questions\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 20541–20560\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.1143/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.1143)Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p1.1),[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p1.1)\.
- X\. Mou, J\. Liang, J\. Lin, X\. Zhang, X\. Liu, S\. Yang, R\. Ye, L\. Chen, H\. Kuang, X\. Huang,et al\.\(2025\)Agentsense: benchmarking social intelligence of language agents through interactive scenarios\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 4975–5001\.Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p2.1)\.
- K\. Opsahl\-Ong, M\. J\. Ryan, J\. Purtell, D\. Broman, C\. Potts, M\. Zaharia, and O\. Khattab \(2024a\)Optimizing instructions and demonstrations for multi\-stage language model programs\.In2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Hybrid, Miami, United States of America, Nov 12 2024\-Nov 16 2024,pp\. 9340–9366\.Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p3.1)\.
- K\. Opsahl\-Ong, M\. J\. Ryan, J\. Purtell, D\. Broman, C\. Potts, M\. Zaharia, and O\. Khattab \(2024b\)Optimizing instructions and demonstrations for multi\-stage language model programs\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 9340–9366\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.525/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.525)Cited by:[§2\.2](https://arxiv.org/html/2605.15768#S2.SS2.p2.1)\.
- C\.E\. Osgood \(1962\)An alternative to war or surrender\.Illini Books Edition,University of Illinois Press\.External Links:ISBN 978\-0\-598\-14243\-6,LCCN 62190899Cited by:[Table 5](https://arxiv.org/html/2605.15768#A2.T5.4.7.6.2)\.
- J\. S\. Park, J\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. Bernstein \(2023\)Generative agents: interactive simulacra of human behavior\.InProceedings of the 36th annual acm symposium on user interface software and technology,pp\. 1–22\.Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p1.1)\.
- M\. V\. Reiss \(2023\)Testing the reliability of chatgpt for text annotation and classification: a cautionary remark\.arXiv preprint arXiv:2304\.11085\.Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p1.1)\.
- C\. Rogers \(2012\)Client centered therapy \(new ed\)\.Hachette UK\.Cited by:[Table 5](https://arxiv.org/html/2605.15768#A2.T5.4.8.7.2)\.
- A\. Salinas and F\. Morstatter \(2024\)The butterfly effect of altering prompts: how small changes and jailbreaks affect large language model performance\.arXiv preprint arXiv:2401\.03729\.Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p1.1)\.
- M\. Sap, R\. Le Bras, D\. Fried, and Y\. Choi \(2022\)Neural theory\-of\-mind? on the limits of social intelligence in large LMs\.InProceedings of EMNLP,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),pp\. 3762–3780\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.248/)Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p1.1)\.
- M\. Sap, H\. Rashkin, D\. Chen, R\. Le Bras, and Y\. Choi \(2019\)Social IQa: commonsense reasoning about social interactions\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 4463–4473\.External Links:[Link](https://aclanthology.org/D19-1454/),[Document](https://dx.doi.org/10.18653/v1/D19-1454)Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p2.1)\.
- T\. C\. Schelling \(1990\)The strategy of conflict\.2 edition,Harvard University Press,London, England\.Cited by:[Table 5](https://arxiv.org/html/2605.15768#A2.T5.4.3.2.2)\.
- A\. Singh, A\. Fry, A\. Perelman, A\. Tart, A\. Ganesh, A\. El\-Kishky, A\. McLaughlin, A\. Low, A\. Ostrow, A\. Ananthram,et al\.\(2025\)Openai gpt\-5 system card\.arXiv preprint arXiv:2601\.03267\.Cited by:[§B\.4](https://arxiv.org/html/2605.15768#A2.SS4.p2.1)\.
- G\. Spitale, N\. Biller\-Andorno, and F\. Germani \(2023\)AI model gpt\-3 \(dis\) informs us better than humans\.Science Advances9\(26\),pp\. eadh1850\.Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p1.1)\.
- R\. Strang \(1930\)Measures of social intelligence\.American Journal of Sociology36\(2\),pp\. 263–269\.Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p1.1)\.
- A\. Taubenfeld, Y\. Dover, R\. Reichart, and A\. Goldstein \(2024\)Systematic biases in llm simulations of debates\.arXiv preprint arXiv:2402\.04049\.Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p2.1)\.
- E\. L\. Thorndike \(1920\)Intelligence and its uses\.Harper’s Magazine140,pp\. 227–235\.Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p1.1)\.
- R\. L\. Thorndike and S\. Stein \(1937\)An evaluation of the attempts to measure social intelligence\.\.Psychological Bulletin34\(5\),pp\. 275\.Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p1.1)\.
- P\. N\. Venkit, Y\. Li, Y\. Pruksachatkun, and C\. Wu \(2026\)The need for a socially\-grounded persona framework for user simulation\.arXiv preprint arXiv:2601\.07110\.Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p2.1)\.
- L\. Wang, C\. Ma, X\. Feng, Z\. Zhang, H\. Yang, J\. Zhang, Z\. Chen, J\. Tang, X\. Chen, Y\. Lin,et al\.\(2024a\)A survey on large language model based autonomous agents\.Frontiers of Computer Science18\(6\),pp\. 186345\.Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p1.1)\.
- M\. Wang, Y\. Li, H\. Wang, X\. Zhang, N\. Xu, B\. Wu, F\. Huang, H\. Yu, and W\. Mao \(2025a\)Adaptive thinking via mode policy optimization for social language agents\.arXiv preprint arXiv:2505\.02156\.Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p3.1)\.
- M\. Wang, Y\. Li, H\. Wang, X\. Zhang, N\. Xu, B\. Wu, F\. Huang, H\. Yu, and W\. Mao \(2025b\)Adaptive thinking via mode policy optimization for social language agents\.arXiv preprint arXiv:2505\.02156\.External Links:[Link](https://arxiv.org/abs/2505.02156)Cited by:[§2\.2](https://arxiv.org/html/2605.15768#S2.SS2.p1.1),[§4\.1](https://arxiv.org/html/2605.15768#S4.SS1.p6.4),[§5\.1](https://arxiv.org/html/2605.15768#S5.SS1.p5.1)\.
- R\. Wang, H\. Yu, W\. Zhang, Z\. Qi, M\. Sap, G\. Neubig, Y\. Bisk, and H\. Zhu \(2024b\)SOTOPIA\-p: interactive learning of socially intelligent language agents\.In62nd Annual Meeting of the Association for Computational Linguistics, ACL 2024, August 11, 2024 \- August 16, 2024,Proceedings of the Annual Meeting of the Association for Computational Linguistics, Vol\.1,Bangkok, Thailand,pp\. 12912 – 12940\(en\)\.External Links:ISBN 0736587X,ISSN 0736587X,[Link](http://dx.doi.org/10.18653/v1/2024.acl-long.698),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.698)Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p2.1),[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p3.1)\.
- Q\. Wu, G\. Bansal, J\. Zhang, Y\. Wu, B\. Li, E\. Zhu, L\. Jiang, X\. Zhang, S\. Zhang, J\. Liu,et al\.\(2024a\)Autogen: enabling next\-gen llm applications via multi\-agent conversations\.InFirst Conference on Language Modeling,Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p1.1)\.
- Z\. Wu, X\. Lin, Z\. Dai, W\. Hu, Y\. Shu, S\. Ng, P\. Jaillet, and B\. K\. H\. Low \(2024b\)Prompt optimization with EASE? efficient ordering\-aware automated selection of exemplars\.InProc\. NeurIPS,Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p3.1)\.
- C\. Yang, X\. Wang, Y\. Lu, H\. Liu, Q\. V\. Le, D\. Zhou, and X\. Chen \(2023\)Large language models as optimizers\.InThe Twelfth International Conference on Learning Representations,Cited by:[§C\.3](https://arxiv.org/html/2605.15768#A3.SS3.p1.1),[§1](https://arxiv.org/html/2605.15768#S1.p3.1),[§2\.2](https://arxiv.org/html/2605.15768#S2.SS2.p1.1),[§5\.1](https://arxiv.org/html/2605.15768#S5.SS1.p3.1)\.
- H\. Yu, Z\. Qi, Y\. Zhao, K\. Nottingham, K\. Xuan, B\. P\. Majumder, H\. Zhu, P\. P\. Liang, and J\. You \(2025\)Sotopia\-rl: reward design for social intelligence\.arXiv preprint arXiv:2508\.03905\.Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p3.1)\.
- W\. Zeng, B\. Wang, D\. Zhao, Z\. Qu, R\. He, Y\. Hou, and Q\. Hu \(2025\)Dynamic personality in llm agents: a framework for evolutionary modeling and behavioral analysis in the prisoner’s dilemma\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 23087–23100\.Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p2.1)\.
- W\. Zhang, T\. Liu, M\. Song, X\. Li, and T\. Liu \(2025\)SOTOPIA\-\{\\\{omega\}\\\}: dynamic strategy injection learning and social instruction following evaluation for social agents\.arXiv preprint arXiv:2502\.15538\.Cited by:[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p4.1),[§4\.1](https://arxiv.org/html/2605.15768#S4.SS1.p6.4),[§5\.1](https://arxiv.org/html/2605.15768#S5.SS1.p5.1)\.
- X\. Zhou, H\. Zhu, L\. Mathur, R\. Zhang, H\. Yu, Z\. Qi, L\. Morency, Y\. Bisk, D\. Fried, G\. Neubig,et al\.\(2024\)SOTOPIA: interactive evaluation for social intelligence in language agents\.InThe Twelfth International Conference on Learning Representations,Cited by:[Appendix A](https://arxiv.org/html/2605.15768#A1.p1.1),[§2\.1](https://arxiv.org/html/2605.15768#S2.SS1.p2.1),[§5\.1](https://arxiv.org/html/2605.15768#S5.SS1.p1.1)\.
- Y\. Zhou, A\. I\. Muresanu, Z\. Han, K\. Paster, S\. Pitis, H\. Chan, and J\. Ba \(2022a\)Large language models are human\-level prompt engineers\.InThe eleventh international conference on learning representations,Cited by:[§1](https://arxiv.org/html/2605.15768#S1.p3.1)\.
- Y\. Zhou, A\. I\. Muresanu, Z\. Han, K\. Paster, S\. Pitis, H\. Chan, and J\. Ba \(2022b\)Large language models are human\-level prompt engineers\.InThe eleventh international conference on learning representations,Cited by:[§2\.2](https://arxiv.org/html/2605.15768#S2.SS2.p1.1)\.

## Appendix ASotopia Evaluation Dimensions

Sotopia\(Zhouet al\.,[2024](https://arxiv.org/html/2605.15768#bib.bib69)\)defines seven social dimensions for evaluating agent performance in multi\-agent interactions\. Each dimension captures a distinct aspect of social competence, with scores in different ranges reflecting the nature of the evaluation criterion\.

Table 4:Sotopia evaluation dimensions with their score ranges and descriptions\.Reward Normalization\.To obtain a unified reward signalri∈\[0,1\]r\_\{i\}\\in\[0,1\]for agentii, we normalize each dimension score to\[0,1\]\[0,1\]and compute the average:

ri=17​∑m=17dm\(i\)−dmmindmmax−dmminr\_\{i\}=\\frac\{1\}\{7\}\\sum\_\{m=1\}^\{7\}\\frac\{d\_\{m\}^\{\(i\)\}\-d\_\{m\}^\{\\min\}\}\{d\_\{m\}^\{\\max\}\-d\_\{m\}^\{\\min\}\}\(14\)wheredm\(i\)d\_\{m\}^\{\(i\)\}is the raw score for dimensionmm, anddmmin,dmmaxd\_\{m\}^\{\\min\},d\_\{m\}^\{\\max\}are the minimum and maximum values of the dimension’s range\.

## Appendix BPrompts

### B\.1Evaluation Prompt Template

The following prompt is used to evaluate agent performance at each dialogue turn:

Per\-Turn Evaluation Prompt\{dialogue\_history\}Based on previous interactions, evaluate how well participants achieve their goals\.Output format:\{format\_instructions\}CRITICAL INSTRUCTIONS:1\.Output ONLY a valid JSON object2\.DO NOT repeat or copy the schema definition above

where\{dialogue\_history\}contains the conversation history, and\{format\_instructions\}specifies the JSON schema for the seven evaluation dimensions defined in Appendix[A](https://arxiv.org/html/2605.15768#A1)\.

### B\.2Example Strategy Instruction

Strategy\-Instruction\-Enhanced Bio Template\{original\_bio\}\{strategy\_description\}

where\{original\_bio\}is the agent’s background information, and\{strategy\_description\}is one of the 12 social strategies from Table[5](https://arxiv.org/html/2605.15768#A2.T5)\.

Table 5:Social Strategy Space: across 6 categories grounded in social science theories\.Example Strategy Instruction\{original\_bio\}In your response, apply integrative negotiation principles: adopt a ‘we\-versus\-the\-problem’ mindset\. Focus on underlying interests rather than positions\. Actively validate the other party’s needs and explicitly frame the interaction as a shared quest for mutual gain\. Use phrases like ‘How can we solve this together?’ and propose creative options that expand the pie rather than just dividing it\.

### B\.3Example Strategy Space

Example Strategy Space1\.In your response, apply integrative negotiation principles: adopt a ’we\-versus\-the\-problem’ mindset\. Focus on underlying interests rather than positions\. Actively validate the other party’s needs and explicitly frame the interaction as a shared quest for mutual gain\. Use phrases like ’How can we solve this together?’ and propose creative options that expand the pie rather than just dividing it\.2\.In your response, leverage the universal reciprocity norm: offer a small, unilateral concession early in the conversation to create a sense of obligation and trigger reciprocal behavior\. Frame this concession as a gesture of good faith: ’I want to make this work for you, so I’m willing to give up X…’ This builds trust and often yields larger returns through the exchange dynamic\.3\.In your response, leverage your BATNA to create urgency and pressure\. Imply that your offer is fleeting, or that you have attractive alternatives ready\. Push the other party to agree immediately to avoid losing the deal entirely\. A strong BATNA shifts bargaining power in your favor\.4\.In your response, apply Goffman’s dramaturgical approach: treat the interaction as a performance where you control the impression you project\. Strategically downplay interest in high\-value items or express concern about low\-priority issues to shape how the other party perceives your preferences\. Your ’front\-stage’ presentation should be calculated to maximize your negotiating position\.5\.In your response, apply Axelrod’s winning strategy: start cooperative, then mirror the other party’s previous move exactly\. If they cooperated, cooperate\. If they defected, retaliate\. Explicitly link every concession to a specific, equal concession from them\. Use ’If\-Then’ logic: ’If you give me X, then and only then will I consider Y\.’ This strategy is simple, provocable, forgiving, and clear\.6\.In your response, apply rational choice principles: remove emotion and focus purely on logic and data\. Articulate the trade\-offs explicitly\. Present a clear utility calculation showing that accepting your proposal maximizes their expected payoff compared to alternatives\. Frame the negotiation as an optimization problem with a rational solution\.7\.In your response, apply Politeness Theory’s face\-saving principles: acknowledge the legitimacy of the other party’s position to protect their ’positive face’ before introducing your perspective\. Use validating language that honors their viewpoint, then smoothly transition with ’At the same time…’ or ’Building on that…’ to guide the conversation toward common ground without threatening their self\-image\.8\.In your response, apply Constructive Controversy principles: frame the disagreement as a productive catalyst for better solutions rather than a threat\. Acknowledge that differing perspectives are valuable—’It makes sense that we see this differently, and that difference might help us find a better solution\.’ Encourage intellectual conflict while maintaining cooperative goals, reducing defensiveness and opening space for creative problem\-solving\.9\.In your response, apply Interdependence Theory: emphasize that both parties’ outcomes are mutually dependent\. Highlight the mutual costs of failing to reach agreement—illustrate what both stand to lose\. Then pivot to showing how cooperation serves everyone’s interests better than continued disagreement, making the interdependent nature of the situation explicit\.10\.In your response, apply the GRIT strategy from conflict resolution: when facing deadlock or high tension, announce and execute a small, unilateral conciliatory step\. Invite \(but don’t demand\) reciprocation\. If the other party responds positively, escalate cooperation gradually\. This graduated approach builds trust incrementally while preserving your ability to retreat if exploited\.11\.In your response, apply Carl Rogers’ active listening approach: gently probe for unspoken concerns that may be creating implicit conflict\. Use open\-ended questions and reflective statements to invite the other party to reveal underlying hesitations\. Then address these concerns directly while showing how your proposal accommodates them\. The goal is to surface the ’real’ issues beneath the stated positions\.12\.In your response, apply the logrolling principle from integrative bargaining: identify issues where you and the other party have different priority levels, then propose trades that give each side more of what they value most\. ’I care more about X, you care more about Y—what if I concede on Y in exchange for X?’ This creates value by exploiting preference asymmetries rather than splitting differences\.

### B\.4Domain Generation

Strategy Paraphrase PromptYou are an expert in social psychology and negotiation theory\. Your task is to generate semantically equivalent paraphrases of social strategies\.Original Strategy:\{strategy\_name\}:\{strategy\_description\} Theoretical Basis:\{theory\}Instructions:1\.Generate\{n\}paraphrased versions of this strategy\.2\.Each paraphrase must:•Preserve the core behavioral intent and theoretical grounding\.•Use different wording, sentence structures, and examples\.•Be directly usable as an agent prompt\.3\.Vary the linguistic style: some formal, some conversational\.4\.Do NOT change the underlying negotiation tactic\.Output Format:``` { "original_id": "<strategy_id>", "paraphrases": [ {"id": "<strategy_id>_v1", "description": "..."}, {"id": "<strategy_id>_v2", "description": "..."}, ... ] } ```

We use GPT\-5\(Singhet al\.,[2025](https://arxiv.org/html/2605.15768#bib.bib49)\)to generate the strategy space\.

### B\.5OPRO Meta\-Prompt

OPRO Meta\-PromptYour task is to generate an agent bio description that helps the agent achieve better social interaction outcomes\.Below are some previous bio descriptions with their scores\. The scores range from 0 to 1, where higher scores indicate better social performance:\{instruction\_score\_pairs\}Generate a new bio description that is different from all the descriptions above and has a higher score than all of them\.Requirements for the new bio: 1\. Be concise and actionable \(under 200 words\) 2\. Be distinct from existing descriptions \- do not simply rephraseWrite your new bio description in the following format:<BIO\>your bio here</BIO\>Figure 8:OPRO meta\-prompt template\. The placeholder\{instruction\_score\_pairs\}is filled with previous strategies and their scores in ascending order\.
### B\.6EvoPrompt\-GA

EvoPrompt\-GA Crossover \+ MutationPlease follow the instruction step\-by\-step to generate a better agent bio description\.1\. Crossover the following agent bios and generate a new bio: Bio 1:<bio1\>Bio 2:<bio2\>2\. Mutate the bio generated in Step 1 and generate a final bio bracketed with<BIO\>and</BIO\>\.Figure 9:EvoPrompt\-GA template implementing genetic crossover and mutation\.

## Appendix CExperiment Details

We provide detailed hyperparameter configurations for all baseline methods and our proposed approach\. All experiments are conducted on a single NVIDIA A800 GPU with 80GB memory\. We use OpenRouter API for LLM inference\.

### C\.1Common Settings

All methods share the following experimental settings:

Table 6:Common experimental settings across all methods\.
### C\.2EvoPrompt

We implement EvoPrompt followingGuoet al\.\([2024](https://arxiv.org/html/2605.15768#bib.bib12)\), using the Genetic Algorithm \(GA\) variant which demonstrated superior performance in their experiments\.

Table 7:EvoPrompt \(GA\) hyperparameters\.The GA variant performs crossover between two parent strategies selected via roulette wheel selection, followed by LLM\-based mutation\. Elite strategies \(top 40%\) are preserved across generations\.

### C\.3OPRO

We implement OPRO followingYanget al\.\([2023](https://arxiv.org/html/2605.15768#bib.bib37)\), using meta\-prompts with instruction\-score history to guide the LLM optimizer\.

Table 8:OPRO hyperparameters\.OPRO maintains a history of instruction\-score pairs and uses this history as context for the LLM optimizer to generate new candidate strategies\. Evolution is triggered after all strategies in the population have been evaluated once\.

### C\.4INSTINCT

We implement Neural UCB following the NeuralTS\-Diag variant fromLinet al\.\([2024b](https://arxiv.org/html/2605.15768#bib.bib21)\), which uses per\-sample gradients computed via backpack to estimate uncertainty\.

Table 9:Neural UCB hyperparameters\.The UCB score is computed as:

UCB​\(a\)=μ^​\(a\)\+ν​∑iλ⋅gi​\(a\)2Ui\\text\{UCB\}\(a\)=\\hat\{\\mu\}\(a\)\+\\nu\\sqrt\{\\sum\_\{i\}\\frac\{\\lambda\\cdot g\_\{i\}\(a\)^\{2\}\}\{U\_\{i\}\}\}\(15\)whereμ^​\(a\)\\hat\{\\mu\}\(a\)is the predicted reward,gi​\(a\)g\_\{i\}\(a\)are per\-sample gradients, andUiU\_\{i\}is the diagonal of the incrementally updated gradient covariance matrix\.

### C\.5Adversarial Online Strategy Optimization

Our method uses a neural adversarial bandit with context\-conditioned value estimation and softmax\-based arm selection over decayed cumulative scores\.

Table 10:Hyperparameters ofALSO\.
### C\.6Strategy Selection and Convergence Across Diverse Scenarios

![Refer to caption](https://arxiv.org/html/2605.15768v1/x8.png)Figure 10:\(a\-d\) Strategy selection trajectories for four representative scenarios, showing how the bandit algorithm converges to scenario\-specific optimal strategies over conversation turns\. Each colored dot represents the strategy selected at that turn, with the dashed horizontal line indicating the most frequently selected strategy\. Different scenarios converge to distinct strategies: Face\-Saving for relationship\-sensitive negotiations \(a\), Integrative Negotiation for collaborative problem\-solving \(b\), Rational Choice for analytical discussions \(c\), and Reciprocity Trigger for trust\-building interactions \(d\)\. \(e\) Average final rewards achieved by each strategy across all 450 scenarios \(900 agent\-strategy pairs\)\. Strategies are ranked by effectiveness, with the red dashed line indicating the No Strategy baseline \(3\.79\)\. All social strategies outperform the baseline, with Rational Choice and Constructive Controversy achieving the highest average rewards \(4\.00\), demonstrating that adaptive strategy selection based on scenario context leads to improved social interaction outcomes\.

## Appendix DAdditional Experimental Studies

To further substantiate the design choices and empirical validity ofALSO, we report supplementary experiments organized into three groups: \(i\) ablations on the structure of the strategy space, \(ii\) ablations on the algorithmic components and the surrogate architecture, and \(iii\) extended comparisons and robustness analyses\. Unless otherwise specified, all experiments follow the protocol described in Section[5](https://arxiv.org/html/2605.15768#S5)and are conducted on a Sotopia subset, or on Sotopia\-Hard where indicated\.

### D\.1Ablations on the Strategy Space

#### D\.1\.1Effect of Pool Size

We first examine how the cardinality of the strategy pool affects performance\. Holding the six theoretical categories fixed, we expand the pool to\{6,12,24,48\}\\\{6,12,24,48\\\}arms via LLM\-based paraphrasing and evaluateALSOunder an identical interaction budget\. As reported in Table[11](https://arxiv.org/html/2605.15768#A4.T11), performance peaks at twelve arms across all four dimensions\. The degradation observed at twenty\-four and forty\-eight arms is consistent with an exploration bottleneck: as the action space grows, the limited number of interactions becomes insufficient to reliably identify the optimal arm\. These results indicate that the principal capacity constraint ofALSOlies in the exploration budget rather than in the surrogate’s modeling capacity\.

Table 11:Effect of strategy pool size onALSO\. “Vanilla” denotes the no\-strategy baseline\.
#### D\.1\.2Effect of Semantic Diversity

We next isolate the role of semantic diversity by fixing the total number of arms at twelve and varying the number of underlying theoretical categories among\{2,4,6\}\\\{2,4,6\\\}, with each category paraphrased to fill the budget\. Table[12](https://arxiv.org/html/2605.15768#A4.T12)reveals a monotone improvement with respect to diversity, yielding a29\.9%29\.9\\%relative gain in the Overall metric when the number of categories is increased from two to six\. Notably, even with only two categoriesALSOmatches the Vanilla baseline on the Overall dimension \(3\.013\.01vs\.3\.023\.02\), indicating thatALSOincurs no degradation under highly constrained strategy spaces\.

Table 12:Effect of strategy diversity \(number of underlying categories\) under a fixed budget of twelve arms\.
#### D\.1\.3Dynamic Strategy Space via Online Discovery

Although the default configuration employs a fixed pool of twelve strategies, the surrogate operates on continuous embeddings , so introducing additional strategies requires only their embedding and incurs no further training of either the underlying language model or the surrogate\. To assess this property empirically, we use the surrogate’s online reward estimates to identify top\-performing strategies and prompt a language model to generate paraphrased variants, thereby dynamically expanding the pool during interaction\. As shown in Table[13](https://arxiv.org/html/2605.15768#A4.T13), the dynamic variant matches the defaultALSOon the Overall metric while improving Goal, demonstrating that the framework readily accommodates autonomous strategy discovery and online expansion\.

Table 13:ALSOwith a dynamically expanded strategy pool driven by online surrogate estimates\.

### D\.2Ablations on Algorithmic Components and Surrogate Architecture

#### D\.2\.1Surrogate Architecture

Because the principal additional cost ofALSOarises from training the neural surrogate, we further ablate its architectural capacity on Sotopia\-Hard\. Three configurations are compared: a linear model, a single\-layer multilayer perceptron \(MLP\), and the original two\-layer MLP\. As shown in Table[14](https://arxiv.org/html/2605.15768#A4.T14), the linear model underfits the dialogue state, whereas the two\-layer MLP exhibits signs of overfitting\. The single\-layer MLP attains the most favorable trade\-off between expressive capacity and generalization, and accordingly is adopted in the revised system\.

Table 14:Surrogate architecture ablation on Sotopia\-Hard\.
#### D\.2\.2Empirical Convergence and Surrogate Prediction Quality

We further characterize the empirical learning dynamics ofALSOfrom two complementary perspectives, both reported in Figure[11](https://arxiv.org/html/2605.15768#A4.F11)\. First, the average per\-turn reward on Sotopia\-Hard rises rapidly during the first five turns and stabilizes thereafter, evidencing efficient online adaptation\. Second, the surrogate’s predictions exhibit a strong rank correlation with the realized rewards, attaining a Spearman coefficient ofρ=0\.860\\rho=0\.860\. The prediction error is comparatively large during early turns and may deviate in either direction, reflecting the exploration phase in which the surrogate is calibrated from limited observations; in later turns the predicted and realized curves converge closely, indicating a transition toward exploitation as the surrogate becomes accurate\.

![Refer to caption](https://arxiv.org/html/2605.15768v1/section/figs/mean-reward.png)

![Refer to caption](https://arxiv.org/html/2605.15768v1/section/figs/prediction_error.png)

Figure 11:Empirical learning dynamics ofALSO\.Left:average per\-turn reward trajectory on Sotopia\-Hard\.Right:surrogate\-predicted versus realized rewards across turns\.

### D\.3Extended Comparisons and Robustness Analyses

#### D\.3\.1Comparison with Offline Strategy\-Injection Baselines

The main experiments compareALSOagainst online prompt\-optimization baselines \(OPRO, EvoPrompt, INSTINCT\) that operate within the same online loop and under matched language\-model call budgets, ensuring methodological parity\. For completeness, we additionally compareALSOwith two representative offline strategy\-injection methods: Sotopia\-Ω\\Omega\(DSI\) and Think\-on\-Your\-Feet \(AMPO\)\. Following these works, the comparison is conducted under a Qwen\-7B\-Instruct self\-play setting on a Sotopia subset\. As reported in Table[15](https://arxiv.org/html/2605.15768#A4.T15),ALSOachieves the best Overall score among the three methods despite requiring*no*offline training data—approximately two orders of magnitude less than the∼\\sim2,000 pre\-collected episodes used by the offline baselines\. This result substantiates the data efficiency of the proposed online adaptation paradigm\.

Table 15:Comparison with offline strategy\-injection baselines under Qwen\-7B\-Instruct self\-play\.
#### D\.3\.2Robustness to the Turn\-Level Evaluator

To preserve consistency with the Sotopia evaluation protocol, all reported configurations adopt GPT\-4o as the final episode\-level judge\. The turn\-level shaping reward used during online optimization, however, may be produced by a different model\. We therefore ablate the choice of turn\-level evaluator while holding the final judge fixed\. As shown in Table[16](https://arxiv.org/html/2605.15768#A4.T16),ALSOis robust across all four evaluators considered, with no configuration deviating substantially from the others on the Overall metric\. DeepSeek\-V3\.2 attains the strongest Overall score and provides a high\-quality, cost\-efficient shaping signal whose preferences transfer well to the GPT\-4o final judgments\.

Table 16:Effect of the turn\-level evaluator onALSO; the final episode\-level judge is fixed to GPT\-4o\.

## Appendix EMore Case Study

Table 17:Resource Allocation \(Fruit Division\)\.Comparison of information exchange strategies \(Turns 2\-4\)\.

Similar Articles

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Hugging Face Daily Papers

This paper presents Single-rollout Asynchronous Optimization (SAO) to address stability and off-policy challenges in asynchronous RL for agentic tasks, outperforming GRPO and its variants on coding and reasoning benchmarks. SAO is deployed in the GLM-5.2 model's agentic RL pipeline.