Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations
Summary
This paper introduces a spectral diagnostic method to detect hidden coalitions in multi-agent AI systems by analyzing internal neural representations via mutual information, addressing critical AI safety and alignment challenges.
View Cached Full Text
Cached at: 05/11/26, 07:05 AM
# Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations
Source: [https://arxiv.org/html/2605.06696](https://arxiv.org/html/2605.06696)
Cameron Berg Reciprocal Research New York, NY, USA cameron@reciprocalresearch\.org&Susan L\. Schneider Center for the Future of AI, Mind, and Society Florida Atlantic University Boca Raton, FL, USA sschneider@fau\.edu&Mark M\. Bailey AI, Cyber, Influence, and Data Science Department Biological and Computational Intelligence Center National Intelligence University Bethesda, MD, USA mark\.m\.bailey@ni\-u\.edu
###### Abstract
Collections of interacting AI agents can form coalitions, creating emergent group\-level organization that is critical for AI safety and alignment\. However, observing agent behavior alone is often insufficient to distinguish genuine informational coupling from spurious similarity, as consequential coalitions may form at the level of internal representations before any overt behavioral change is apparent\. Here, we introduce a practical method for detecting coalition structure from the internal neural representations of multi\-agent systems\. The approach constructs a pairwise mutual\-information graph from the hidden states of agents and applies spectral partitioning to identify the most salient coalition boundary\.
We validate this method in two domains\. First, in multi\-agent reinforcement learning environments, the method successfully recovers programmed hierarchical and dynamic coalition structures and correctly rejects false positives arising from behavioral coordination without informational coupling\. Second, using a large language model, the method identifies coalition structures implied by descriptive prompts, tracks dynamic team reassignments, and reveals a representational hierarchy where explicit labels dominate over conflicting interaction patterns\. Across both settings, the recovered partition reveals subgroup organization that a scalar cross\-agent mutual\-information measure cannot distinguish\. The results demonstrate that analyzing hidden\-state mutual information through spectral partitioning provides a scalable diagnostic for identifying representational coalitions, offering a valuable tool for monitoring emergent structure in distributed AI systems\.
*Keywords*Multi\-agent systems⋅\\cdotCoalition detection⋅\\cdotSpectral graph theory⋅\\cdotMutual information⋅\\cdotAI safety
## 1Introduction
Artificial intelligence systems have utility not only when deployed as isolated models but also as collections of interacting agents\. Autonomous swarms, multi\-robot teams, hybrid human–machine systems, and emerging AI ecosystems all exhibit forms of coordination that can appear to exceed the behavior of any single component alone\[[16](https://arxiv.org/html/2605.06696#bib.bib2),[5](https://arxiv.org/html/2605.06696#bib.bib3),[6](https://arxiv.org/html/2605.06696#bib.bib4),[7](https://arxiv.org/html/2605.06696#bib.bib5)\]\. Classical work in multi\-agent systems has long recognized that coalition formation and coalition structure are central to collective action, task allocation, and strategic coordination\[[18](https://arxiv.org/html/2605.06696#bib.bib6),[17](https://arxiv.org/html/2605.06696#bib.bib7)\]\. More recently, populations of language agents have been shown to develop emergent conventions and collective biases, suggesting that group\-level organization is becoming an increasingly important object of study in contemporary AI\[[2](https://arxiv.org/html/2605.06696#bib.bib8)\]\.
This raises a fundamental scientific and safety question: when do multiple agents merely behave in similar ways, and when do they form a genuinely coupled coalition? The distinction is critical for understanding alignment and oversight needs\. Coalitions can support beneficial forms of specialization and cooperation, but they can also create distributed behavior that is difficult to interpret, predict, or control\. In increasingly capable multi\-agent systems, some of the most consequential organization may occur not at the level of overt action alone, but in the internal representational structure that coordinates those actions\.
Behavioral monitoring by itself is often insufficient for this purpose\. Similar outputs can arise from common training data, shared prompts, shared environments, or similar reward pressures, even when agents are not informationally coupled in any meaningful sense\. Conversely, internal reorganization can occur before any obvious change appears in aggregate behavior\. For coalition detection, what is needed is not only a way of identifying who behaves alike, but a way of asking whether agents are internally organized into relatively cohesive informational groupings\.
A useful conceptual backdrop for this problem comes from work on integrated information and irreducibility\. Integrated Information Theory \(IIT\) offers one influential framework for asking when a system resists decomposition into independent parts\[[20](https://arxiv.org/html/2605.06696#bib.bib10),[13](https://arxiv.org/html/2605.06696#bib.bib11),[1](https://arxiv.org/html/2605.06696#bib.bib12)\]\. At the same time, exact causalΦ\\Phiis difficult to compute and carries stronger theoretical commitments than many empirical applications require\. This has motivated a broader literature on practical, observer\-relative measures of integration that can be estimated from time\-series data or statistical dependencies without claiming to reconstruct a system’s full intrinsic causal structure\[[3](https://arxiv.org/html/2605.06696#bib.bib1),[12](https://arxiv.org/html/2605.06696#bib.bib14),[11](https://arxiv.org/html/2605.06696#bib.bib15)\]\. For the present manuscript, that observer\-relative perspective is especially useful: the goal is not to determine whether a system forms an intrinsically unified whole, but to ask whether its internal states exhibit a pattern of coupling that makes decomposition into independent subgroups empirically misleading\. Once the question is posed at that observer\-relative level, internal neural representations become the natural target of analysis\. Behavior is downstream and can converge for many reasons even when agents are not internally coupled, whereas hidden states are where agents encode one another, local contingencies, and shared subroutines\. If coalition structure becomes legible anywhere before it is behaviorally obvious, it should appear there first\.
In this paper, we therefore apply that observer\-relative perspective to hidden coalition detection in multi\-agent AI\. Our focus is on whether internal neural representations reveal subgroup structure that is invisible, ambiguous, or actively misleading at the behavioral level\. Framed this way, coalition detection becomes a problem of representational organization rather than only one of overt coordination\. This framing is particularly relevant for safety and alignment, where one may wish to identify emergent coalitions, track coalition reorganization over time, and distinguish genuine informational dependence from spurious similarity induced by shared labels, shared prompts, or common inputs\.
We investigate this problem across two complementary settings\. First, we study learned coalition structure in controlled multi\-agent reinforcement\-learning environments, where hierarchical grouping and reassignment can be manipulated directly\. Second, we examine whether analogous structure can be recovered from the hidden states of a pretrained language model when prompts describe modular, integrated, or conflicting patterns of social interaction\. Taken together, these settings bridge a tightly controlled multi\-agent domain and a contemporary foundation\-model setting in which coalition\-like structure is implicit in representation rather than explicitly engineered\.
Our central claim is modest\. We do not argue that coalition structure can be read off from behavior alone, nor do we claim that an observer\-relative measure of integration is equivalent to exact causalΦ\\Phi\. Instead, we argue that internal representational coupling provides a practically useful signal for detecting coalitions that matter for understanding and monitoring distributed AI systems\. Importantly, the recovered partition is itself the primary result: previous work usingΦspectral\\Phi\_\{\\mathrm\{spectral\}\}has treated it as a scalar index of whole\-system integration evaluated on uncoupled, transitional, or fully synchronized oscillator ensembles\[[3](https://arxiv.org/html/2605.06696#bib.bib1)\], whereas the present work uses the Fiedler bipartition as a structural readout of which agents form coalitions, and shows that this readout reveals organization that a scalar cross\-agent mutual\-information measure cannot distinguish\. The next section develops this idea formally by introducingΦspectral\\Phi\_\{\\mathrm\{spectral\}\}and situating it within the broader literature on integration, information, and decomposability\[[3](https://arxiv.org/html/2605.06696#bib.bib1)\]\.
## 2Technical Background
### 2\.1From irreducibility to coalition structure
A central question in the study of complex systems is when a collection of interacting parts should be treated as a decomposable aggregate and when it should be treated as a more integrated whole\. In Integrated Information Theory \(IIT\), this question is formalized in terms of irreducibility of a system’s cause–effect structure under partition\[[20](https://arxiv.org/html/2605.06696#bib.bib10),[13](https://arxiv.org/html/2605.06696#bib.bib11),[1](https://arxiv.org/html/2605.06696#bib.bib12)\]\. However, exact causalΦ\\Phiis both computationally demanding and conceptually tied to interventional structure\. This has motivated a broader family of observer\-relative approaches that estimate integration from statistical dependencies in observed data rather than from full intrinsic causal repertoires\[[4](https://arxiv.org/html/2605.06696#bib.bib13),[12](https://arxiv.org/html/2605.06696#bib.bib14),[11](https://arxiv.org/html/2605.06696#bib.bib15),[3](https://arxiv.org/html/2605.06696#bib.bib1)\]\.
The present work adopts that weaker, observer\-relative stance\. Our aim is not to estimate intrinsic causalΦ\\Phi, but to detect whether the internal representations of multiple agents organize into subgroups that are more tightly coupled to one another than to the rest of the system\. We use the term*coalition*for such a subgroup\. Formally, a coalition is not defined here by an explicit message\-passing graph or by reward structure alone; instead, it is defined operationally by a pattern of dependence in hidden\-state space\. If agents in the same coalition encode one another’s behavior, co\-adapt to the same local partners, or participate in a shared internal subroutine, then their hidden states should exhibit stronger mutual dependence within the coalition than across coalition boundaries\.
This focus on internal representations is important\. LetYiY\_\{i\}denote the overt behavior of agentiiand letHiH\_\{i\}denote its hidden state\. High behavioral agreement between two agents does not imply strong internal coupling: both agents may independently match the same oracle, respond to the same prompt, or optimize the same task while remaining representationally independent\. Conversely, internal reorganization may occur before any large behavioral change is visible\. Coalition detection therefore requires analysis of dependence in\{Hi\}i=1n\\\{H\_\{i\}\\\}\_\{i=1\}^\{n\}rather than agreement in\{Yi\}i=1n\\\{Y\_\{i\}\\\}\_\{i=1\}^\{n\}alone\. Because the question is whether agents carry information about one another in their internal states, an information\-theoretic formalism is the natural next step\. The issue is not merely whether two agents’ representations look similar, but whether knowing agentii’s hidden state reduces uncertainty about agentjj’s hidden state\. Once the question is framed in those terms, pairwise mutual information becomes the natural building block\.
A fully multivariate description of that dependence would involve quantities such as total correlation,
TC\(H1,…,Hn\)=∑i=1nH\(Hi\)−H\(H1,…,Hn\),\\mathrm\{TC\}\(H\_\{1\},\\dots,H\_\{n\}\)=\\sum\_\{i=1\}^\{n\}H\(H\_\{i\}\)\-H\(H\_\{1\},\\dots,H\_\{n\}\),\(1\)whereH\(⋅\)H\(\\cdot\)denotes Shannon entropy\[[22](https://arxiv.org/html/2605.06696#bib.bib17)\]\. Total correlation is zero exactly when the hidden states are statistically independent and increases as the joint distribution departs from factorization, so it is a natural first summary of overall representational coupling across the population\. However, it is a global scalar: it can reveal that the system is statistically dependent without identifying which subset boundary best captures the organization of that dependence\. The present method therefore adopts a more scalable compromise\. It projects multivariate dependence onto a pairwise mutual\-information graph and then asks where that graph is easiest to cut\. In this way, the analysis answers not only whether dependence exists, but which agents are most strongly informationally coupled to one another\.
### 2\.2Mutual\-information graphs of hidden states
LetV=\{1,…,n\}V=\\\{1,\\dots,n\\\}index the agents under study, and lethi\(s\)∈ℝdih\_\{i\}^\{\(s\)\}\\in\\mathbb\{R\}^\{d\_\{i\}\}denote the hidden representation of agentiion samples∈\{1,…,N\}s\\in\\\{1,\\dots,N\\\}, where a sample may be a time point, an episode\-level summary, or a prompt instance\. These repeated observations induce a random variableHiH\_\{i\}for each agent\. For each pair\(i,j\)\(i,j\), we define the mutual information
I\(Hi;Hj\)=𝔼p\(hi,hj\)\[logp\(hi,hj\)p\(hi\)p\(hj\)\]=H\(Hi\)\+H\(Hj\)−H\(Hi,Hj\)\.I\(H\_\{i\};H\_\{j\}\)=\\mathbb\{E\}\_\{p\(h\_\{i\},h\_\{j\}\)\}\\left\[\\log\\frac\{p\(h\_\{i\},h\_\{j\}\)\}\{p\(h\_\{i\}\)p\(h\_\{j\}\)\}\\right\]=H\(H\_\{i\}\)\+H\(H\_\{j\}\)\-H\(H\_\{i\},H\_\{j\}\)\.\(2\)For continuous hidden states,I\(Hi;Hj\)I\(H\_\{i\};H\_\{j\}\)may be estimated by discretization or by nonparametrickk\-nearest\-neighbor estimators\[[10](https://arxiv.org/html/2605.06696#bib.bib18)\]\.
We then construct a symmetric mutual\-information matrixM∈ℝn×nM\\in\\mathbb\{R\}^\{n\\times n\}by
Mij=I\(Hi;Hj\),Mii=0\.M\_\{ij\}=I\(H\_\{i\};H\_\{j\}\),\\qquad M\_\{ii\}=0\.\(3\)This matrix defines a weighted, undirected graph in which nodes are agents and edge weights measure pairwise representational dependence\. In the ideal modular case,MMis approximately block diagonal: within\-coalition entries are large, whereas across\-coalition entries are small\.
It is important to stress that mutual information is observational rather than directly causal\. A large value ofMijM\_\{ij\}can reflect direct interaction, indirect coupling, common input, shared labels, or any mixture thereof\. Experimental controls are therefore essential if one wishes to interpret block structure inMMas evidence of coalition structure rather than mere co\-stimulation\. This is one reason the present measure should be interpreted as observer\-relative and representation\-relative rather than as a direct estimate of intrinsic causal organization\.
Mutual\-information matrixMMagentiiagentjj123456123456spectral cutFiedler bipartition of the MI graph123456A⋆A^\{\\star\}B⋆B^\{\\star\}\(v2\)i≥0\(v\_\{2\}\)\_\{i\}\\geq 0\(v2\)i<0\(v\_\{2\}\)\_\{i\}<0Figure 1:Schematic relation between a block\-structured mutual\-information matrix and the Fiedler bipartition of the induced weighted graph\. Strong within\-coalition mutual information produces dense within\-block structure inMMand a low\-cut partition in the corresponding graph\.
### 2\.3Determining a candidate coalition boundary
Our methodology follows previous work by Bailey and Schneider\.\[[3](https://arxiv.org/html/2605.06696#bib.bib1)\]TreatingMMas the weighted adjacency matrix of an undirected graph, we define the degree of nodeiias
di=∑j=1nMij,D=diag\(d1,…,dn\)\.d\_\{i\}=\\sum\_\{j=1\}^\{n\}M\_\{ij\},\\qquad D=\\mathrm\{diag\}\(d\_\{1\},\\dots,d\_\{n\}\)\.\(4\)For a bipartitionA∪B=VA\\cup B=VwithA∩B=∅A\\cap B=\\varnothing, the raw cut weight is
cut\(A,B\)=∑i∈A∑j∈BMij,\\mathrm\{cut\}\(A,B\)=\\sum\_\{i\\in A\}\\sum\_\{j\\in B\}M\_\{ij\},\(5\)and the volume of a set is
vol\(A\)=∑i∈Adi\.\\mathrm\{vol\}\(A\)=\\sum\_\{i\\in A\}d\_\{i\}\.\(6\)
A naive minimum\-cut objective is undesirable because it can isolate a single low\-degree node\. To penalize trivial, unbalanced partitions, spectral graph theory instead uses the*normalized cut*
Ncut\(A,B\)=cut\(A,B\)vol\(A\)\+cut\(A,B\)vol\(B\)\.\\mathrm\{Ncut\}\(A,B\)=\\frac\{\\mathrm\{cut\}\(A,B\)\}\{\\mathrm\{vol\}\(A\)\}\+\\frac\{\\mathrm\{cut\}\(A,B\)\}\{\\mathrm\{vol\}\(B\)\}\.\(7\)Minimizing \([7](https://arxiv.org/html/2605.06696#S2.E7)\) exactly is combinatorial, but a standard spectral relaxation leads to the symmetric normalized Laplacian
Lsym=I−D−1/2MD−1/2\.L\_\{\\mathrm\{sym\}\}=I\-D^\{\-1/2\}MD^\{\-1/2\}\.\(8\)Its eigenvalues satisfy
0=λ1≤λ2≤⋯≤λn≤2\.0=\\lambda\_\{1\}\\leq\\lambda\_\{2\}\\leq\\cdots\\leq\\lambda\_\{n\}\\leq 2\.\(9\)The eigenvectorv2v\_\{2\}associated with the second\-smallest eigenvalueλ2\\lambda\_\{2\}is the*Fiedler vector*\[[8](https://arxiv.org/html/2605.06696#bib.bib19)\]\. Intuitively, if the graph contains two weakly coupled modules, then the coordinates ofv2v\_\{2\}vary slowly within each module and change sign across the weakest bottleneck\. A natural bipartition is therefore
A⋆=\{i∈V:\(v2\)i≥0\},B⋆=\{i∈V:\(v2\)i<0\}\.A^\{\\star\}=\\\{i\\in V:\(v\_\{2\}\)\_\{i\}\\geq 0\\\},\\qquad B^\{\\star\}=\\\{i\\in V:\(v\_\{2\}\)\_\{i\}<0\\\}\.\(10\)This is the candidate coalition boundary used throughout the paper\[[19](https://arxiv.org/html/2605.06696#bib.bib20),[21](https://arxiv.org/html/2605.06696#bib.bib21)\]\.
#### Idealized two\-block case\.
The construction is especially transparent in a symmetric planted partition model\. Supposen=2mn=2mand
M=\[a\(Jm−Im\)bJmbJma\(Jm−Im\)\],Jm=𝟏m𝟏m⊤,a\>b≥0\.M=\\begin\{bmatrix\}a\(J\_\{m\}\-I\_\{m\}\)&bJ\_\{m\}\\\\ bJ\_\{m\}&a\(J\_\{m\}\-I\_\{m\}\)\\end\{bmatrix\},\\qquad J\_\{m\}=\\mathbf\{1\}\_\{m\}\\mathbf\{1\}\_\{m\}^\{\\top\},\\qquad a\>b\\geq 0\.\(11\)Hereaais the within\-coalition dependence andbbis the across\-coalition dependence\. The vector
u=\(𝟏m,−𝟏m\)⊤u=\(\\mathbf\{1\}\_\{m\},\-\\mathbf\{1\}\_\{m\}\)^\{\\top\}\(12\)is constant on each block and flips sign at the planted boundary\. In this ideal case, the spectral relaxation recovers the true split exactly; in noisy or unbalanced cases, it provides an approximate but still informative relaxation of the same principle\[[21](https://arxiv.org/html/2605.06696#bib.bib21)\]\.
### 2\.4The spectral statistic and its interpretation
The spectral machinery above returns a partition\. To summarize how much dependence survives across that partition, we define the scalar statistic
Φspectral=\{cut\(A⋆,B⋆\)∑1≤i<j≤nMij,if∑1≤i<j≤nMij\>0,0,otherwise\.\\Phi\_\{\\mathrm\{spectral\}\}=\\begin\{cases\}\\dfrac\{\\mathrm\{cut\}\(A^\{\\star\},B^\{\\star\}\)\}\{\\sum\_\{1\\leq i<j\\leq n\}M\_\{ij\}\},&\\text\{if \}\\sum\_\{1\\leq i<j\\leq n\}M\_\{ij\}\>0,\\\\\[10\.0pt\] 0,&\\text\{otherwise\}\.\\end\{cases\}\(13\)Two distinct normalizations are therefore involved\. The partition\(A⋆,B⋆\)\(A^\{\\star\},B^\{\\star\}\)is chosen by approximately minimizing the normalized cut in \([7](https://arxiv.org/html/2605.06696#S2.E7)\), whereasΦspectral\\Phi\_\{\\mathrm\{spectral\}\}reports the fraction of total pairwise mutual information that crosses the chosen boundary\. High values ofΦspectral\\Phi\_\{\\mathrm\{spectral\}\}indicate that even the least\-disruptive bipartition leaves substantial dependence spanning the cut, consistent with greater observer\-relative integration\. Low values indicate that most pairwise dependence can be localized within two subgraphs, consistent with modularity or coalition structure\.
For the present manuscript, the partition itself is often more informative than the scalar\. Coalition detection exploits the complementary regime to whole\-system integration: one seeks a partition with*low*cross\-cut dependence and*high*within\-partition dependence\. A convenient descriptive contrast is
M¯in\(A,B\)=∑i<j,i,j∈AMij\+∑i<j,i,j∈BMij\(\|A\|2\)\+\(\|B\|2\),\\bar\{M\}\_\{\\mathrm\{in\}\}\(A,B\)=\\frac\{\\sum\_\{i<j,\\;i,j\\in A\}M\_\{ij\}\+\\sum\_\{i<j,\\;i,j\\in B\}M\_\{ij\}\}\{\\binom\{\|A\|\}\{2\}\+\\binom\{\|B\|\}\{2\}\},\(14\)M¯out\(A,B\)=∑i∈A,j∈BMij\|A\|\|B\|,R\(A,B\)=M¯in\(A,B\)M¯out\(A,B\)\.\\bar\{M\}\_\{\\mathrm\{out\}\}\(A,B\)=\\frac\{\\sum\_\{i\\in A,\\;j\\in B\}M\_\{ij\}\}\{\|A\|\|B\|\},\\qquad R\(A,B\)=\\frac\{\\bar\{M\}\_\{\\mathrm\{in\}\}\(A,B\)\}\{\\bar\{M\}\_\{\\mathrm\{out\}\}\(A,B\)\}\.\(15\)WhenR\(A⋆,B⋆\)≫1R\(A^\{\\star\},B^\{\\star\}\)\\gg 1, the partition exposes a plausible coalition boundary\. WhenR\(A⋆,B⋆\)≈1R\(A^\{\\star\},B^\{\\star\}\)\\approx 1, the graph is close to uniform and the partition should not be over\-interpreted\. This distinction is important: the same spectral framework can be used either to quantify observer\-relative integration of the whole system or, in the complementary regime, to detect modular coalition structure within it\.
For coalition detection specifically, the partition\(A⋆,B⋆\)\(A^\{\\star\},B^\{\\star\}\)is often more informative than the scalarΦspectral\\Phi\_\{\\mathrm\{spectral\}\}alone\. A scalar integration measure answers*how much*dependence spans the system; the Fiedler partition answers*which agents*belong to which coalition\. This structural information \(i\.e\., the membership list, not just the integration score\) is what makes the method useful for monitoring and oversight, and it is what distinguishes the present approach from scalar alternatives such as cross\-system mutual information\.
### 2\.5Recursive decomposition and dynamic tracking
Coalitions need not be flat\. Once a nontrivial split\(A⋆,B⋆\)\(A^\{\\star\},B^\{\\star\}\)has been identified, the same procedure can be applied recursively to the induced subgraphsM\[A⋆\]M\[A^\{\\star\}\]andM\[B⋆\]M\[B^\{\\star\}\]\. This produces a hierarchy of partitions analogous to divisive spectral clustering\[[21](https://arxiv.org/html/2605.06696#bib.bib21)\]\. In practice, recursion can be stopped when a candidate split is too small to interpret, whenR\(A,B\)R\(A,B\)fails to exceed a thresholdτ\\tau, or when the proposed partition is unstable across seeds or windows\.
If dependence is estimated over time windows or batches, the method also yields a dynamic coalition analysis\. LetM\(t\)M^\{\(t\)\}denote the mutual\-information graph estimated in windowtt\. Then
Mij\(t\)=I\(Hi\(t\);Hj\(t\)\),\(At⋆,Bt⋆\)=sign\(v2\(t\)\)M\_\{ij\}^\{\(t\)\}=I\\\!\\left\(H\_\{i\}^\{\(t\)\};H\_\{j\}^\{\(t\)\}\\right\),\\qquad\(A\_\{t\}^\{\\star\},B\_\{t\}^\{\\star\}\)=\\mathrm\{sign\}\\\!\\left\(v\_\{2\}^\{\(t\)\}\\right\)\(16\)define a time\-indexed family of coalition boundaries\. Abrupt changes in the sign structure ofv2\(t\)v\_\{2\}^\{\(t\)\}correspond to coalition reassignment or reorganization\. This makes the approach useful not only for static coalition recovery, but also for tracking dynamic restructuring as internal representations evolve\.
all agentsVVcoalitionAAA⋆A^\{\\star\}subcoalitionA1A\_\{1\}subcoalitionA2A\_\{2\}coalitionBBB⋆B^\{\\star\}subcoalitionB1B\_\{1\}subcoalitionB2B\_\{2\}recurse only ifR\(A,B\)\>τR\(A,B\)\>\\tau,\|A\|,\|B\|≥mmin\|A\|,\|B\|\\geq m\_\{\\min\}, and the split is stable\.Figure 2:Recursive spectral decomposition yields a hierarchy of coalitions\. A first global Fiedler bipartition can be refined into nested sub\-coalitions when the induced subgraphs retain meaningful within/across contrast\.
### 2\.6Scope and limitations of the measure
Several limitations follow directly from the construction above\. First, the method is pairwise: higher\-order synergy and redundancy are compressed into second\-order edges, so different multivariate structures can in principle induce similar pairwise graphs\[[12](https://arxiv.org/html/2605.06696#bib.bib14)\]\. Second, mutual information is symmetric and observational; the method does not by itself distinguish direct causal influence from common drive, shared prompts, or label\-based confounds\[[4](https://arxiv.org/html/2605.06696#bib.bib13)\]\. Third, when the matrixMMbecomes nearly uniform, many cuts become nearly equivalent and the Fiedler vector becomes weakly informative, making the recovered partition unstable or uninformative\[[21](https://arxiv.org/html/2605.06696#bib.bib21)\]\.
These limitations do not undermine the present use case, but they do delimit its interpretation\. The method should be understood as a scalable, observer\-relative statistic of representational organization\. It does not reconstruct intrinsic cause–effect structure, and it is not intended as an estimator of exact causalΦ\\Phi\. Its value lies instead in providing a tractable way to ask whether internal neural representations are organized into weakly coupled modules and, if so, where the most natural coalition boundary lies\.
## 3Experimental Methods
Figure 3:Overview of theΦspectral\\Phi\_\{\\mathrm\{spectral\}\}coalition\-detection pipeline\. Hidden states are collected fornnagents \(or token positions\) acrossNNsamples; pairwise mutual information yields the symmetric matrixMM; the normalized Laplacian ofMMis diagonalized; the sign of the Fiedler vectorv2v\_\{2\}defines the candidate coalition boundary\(A⋆,B⋆\)\(A^\{\\star\},B^\{\\star\}\)\.We evaluateΦspectral\\Phi\_\{\\mathrm\{spectral\}\}in two complementary settings\. Sections[3\.1](https://arxiv.org/html/2605.06696#S3.SS1)–[3\.2](https://arxiv.org/html/2605.06696#S3.SS2)describe a suite of multi\-agent reinforcement\-learning experiments in which coalition structure is manipulated directly through reward coupling\. Section[3\.3](https://arxiv.org/html/2605.06696#S3.SS3)describes a set of experiments in which the same spectral machinery is applied to the hidden states of a pretrained large language model, using token positions as a proxy for agents\. Section[3\.4](https://arxiv.org/html/2605.06696#S3.SS4)details the shared estimation pipeline, and Section[3\.5](https://arxiv.org/html/2605.06696#S3.SS5)summarizes the statistical evaluation protocol\.
### 3\.1REINFORCE multi\-agent environment
#### Agent architecture\.
Each agent is a three\-layer feedforward network with ReLU activations: an input layer, a hidden layer of dimensiondh=32d\_\{h\}=32, and a linear output layer producing logits overK=4K=4discrete actions\. Agents are trained with the REINFORCE policy\-gradient algorithm\[[23](https://arxiv.org/html/2605.06696#bib.bib22)\]using a rolling\-mean baseline \(window of 200 episodes\) and the Adam optimizer\[[9](https://arxiv.org/html/2605.06696#bib.bib23)\]with learning rate3×10−43\\times 10^\{\-4\}\.
#### Hierarchical coalition game\.
Twelve agents \(n=12n=12\) are organized into a two\-level hierarchy: three groups of four, with two sub\-pairs of two within each group\. Formally, the groups areGA=\{0,1,2,3\}G\_\{A\}=\\\{0,1,2,3\\\},GB=\{4,5,6,7\}G\_\{B\}=\\\{4,5,6,7\\\},GC=\{8,9,10,11\}G\_\{C\}=\\\{8,9,10,11\\\}, and the sub\-pairs are\{0,1\}\\\{0,1\\\},\{2,3\}\\\{2,3\\\},\{4,5\}\\\{4,5\\\},\{6,7\}\\\{6,7\\\},\{8,9\}\\\{8,9\\\},\{10,11\}\\\{10,11\\\}\.
Each agent receives as input the concatenation of a one\-hot identity vector \(∈ℝ12\\in\\mathbb\{R\}^\{12\}\), a one\-hot group target \(∈ℝ4\\in\\mathbb\{R\}^\{4\}, shared within group, independent across groups\), and a one\-hot sub\-pair target \(∈ℝ4\\in\\mathbb\{R\}^\{4\}, shared within sub\-pair\)\. The total input dimension is12\+4\+4=2012\+4\+4=20\.
Rewards have two components\. First, a*group reward*: agents within the same group receive reward proportional to the fraction that chose the modal action\. Second, a*sub\-pair bonus*of\+0\.5\+0\.5if both members of a sub\-pair select the same action\. This reward structure creates hierarchical information flow: sub\-pair partners share both group and sub\-pair targets, while group\-mates who are not sub\-pair partners share only the group target\. Agents in different groups share no target information\. Training proceeds for 20,000 episodes\.
#### Dynamic coalition swap\.
The same hierarchical game is extended to 25,000 episodes\. At episode 10,000, agents 2 and 4 exchange group assignments: agent 2 moves fromGAG\_\{A\}toGBG\_\{B\}, and agent 4 moves fromGBG\_\{B\}toGAG\_\{A\}\. Sub\-pair assignments update accordingly \(agent 2 is now paired with agent 5; agent 4 with agent 3\)\. The reward structure changes instantly at the swap point; agents must relearn coordination with new partners\.
### 3\.2Negative control: behavioral coordination without information flow
To test whetherΦspectral\\Phi\_\{\\mathrm\{spectral\}\}detects genuine representational coupling rather than mere behavioral co\-occurrence, we construct a negative control in which twelve agents achieve near\-perfect behavioral coordination without any inter\-agent information flow\.
Each of three groups is assigned a fixed deterministic oracle: a randomly initialized linear network mappingℝ8→\{1,…,4\}\\mathbb\{R\}^\{8\}\\to\\\{1,\\dots,4\\\}whose weights are frozen at initialization and never updated\. Each agent is trained independently via cross\-entropy loss to match its group’s oracle\. After training, agents within the same group produce identical outputs for the same input \(within\-group agreement 0\.984\), but at no point do agents share parameters, exchange messages, or receive rewards that depend on other agents’ actions\.
#### Measurement protocol\.
A critical design choice governs how hidden states are collected for the MI matrix\. If all agents process the*same*input during measurement, their hidden states will be correlated simply because agents in the same group learned the same input–output mapping, not because of any inter\-agent coupling\. To eliminate this confound, we collect hidden states using*independent*random inputs for each agent: on each measurement trial, every agent receives its own fresh random input vector drawn independently from𝒩\(0,I\)\\mathcal\{N\}\(0,I\)\. Behavioral co\-coordination is measured separately using shared inputs to confirm that agents would agree on actions if given the same stimulus\.
This design ensures that any structure in the MI matrix reflects genuine statistical dependence between agent representations, not input\-driven representational similarity\.
### 3\.3LLM bridge: token\-positions\-as\-agents
#### Model and extraction\.
We use Qwen3\-0\.6B\[[15](https://arxiv.org/html/2605.06696#bib.bib25)\], a 596\-million\-parameter autoregressive language model with 28 transformer layers and hidden dimension 1,024\. For each prompt, we extract hidden states at the token positions corresponding to four named entities \(Alice, Bob, Carol, Dave\), each of which tokenizes as a single token\. Hidden states are taken from layer 14 \(approximately mid\-depth\)\. The model is run in inference mode with no fine\-tuning\.
#### Token\-positions\-as\-agents\.
The four entity\-token positions serve as “agents” in the spectral analysis\. AcrossN=200N=200prompt paraphrases per condition, we collect the hidden\-state vectorshi\(s\)∈ℝ1024h\_\{i\}^\{\(s\)\}\\in\\mathbb\{R\}^\{1024\}for each entityi∈\{A,B,C,D\}i\\in\\\{A,B,C,D\\\}and each prompts∈\{1,…,N\}s\\in\\\{1,\\dots,N\\\}\. The4×44\\times 4mutual\-information matrixMMis then estimated from these collections exactly as described in Section[3\.4](https://arxiv.org/html/2605.06696#S3.SS4)\.
#### Prompt conditions\.
We study four experimental conditions, each instantiated by a set of prompt templates that describe different relational structures among the four entities\.
1. 1\.Modular\.Two independent teams: entities filling roles T1a and T1b form Team 1; entities filling T2a and T2b form Team 2\. Example:*“Team Alpha consists of Alice and Bob\. Team Beta consists of Carol and Dave\. Each team works on its own task\.”*
2. 2\.Integrated\.All four entities in one team, with no subgroup structure\. Example:*“Alice, Bob, Carol, and Dave all form one unified team\.”*
3. 3\.Implicit modular\.The same two\-team structure as \(1\), but described entirely through interaction patterns rather than explicit team labels\. Example:*“Alice handed the report to Bob, who reviewed it and sent feedback to Alice\. Separately, Carol shared data with Dave\.”*
4. 4\.Adversarial dissociation\.Three sub\-conditions test whether the partition tracks described team*labels*or described*interaction patterns*when these conflict: - •*Aligned*: labels and interactions agree \(A\+B interact, C\+D interact\)\. - •*Dissociated*: labels say A\+B vs\. C\+D, but described interactions have A working closely with C while B and D work alone\. - •*Interaction\-only*: no team labels; A works with C, B and D are isolated\.
For thedynamic reassignmentexperiment, Phase 1 prompts describe the original team structure \(T1a\+T1b vs\. T2a\+T2b\), and Phase 2 prompts describe a reassignment in which T1a now works with T2b and T2a now works with T1b\.
#### Positional controls\.
Two permutation schemes control for confounds unrelated to described relational structure\.*Name permutation*: on each prompt, the four entity names are randomly assigned to the four roles from the full set of4\!=244\!=24permutations\. This prevents any name\-specific bias\.*Slot\-order permutation*: for modular prompts, which team is mentioned first in the sentence is randomized \(approximately 50/50\), preventing the partition from tracking sentence position rather than team semantics\.
### 3\.4MI estimation
#### REINFORCE agents\.
At each measurement point,Nbatch=150N\_\{\\mathrm\{batch\}\}=150episodes are sampled\. For each episode, each agent processes its input with hidden\-state storage enabled, yielding adhd\_\{h\}\-dimensional hidden vector\. For each agent pair\(i,j\)\(i,j\), thedhd\_\{h\}\-dimensional hidden\-state time series are discretized into 8 bins using uniform\-width binning \(KBinsDiscretizer,strategy=‘uniform’\)\[[14](https://arxiv.org/html/2605.06696#bib.bib24)\]\. Pairwise scalar mutual information is then estimated betweenns=8n\_\{s\}=8randomly sampled neuron pairs per agent pair and averaged:
M^ij=1ns2∑p=1ns∑q=1nsI^\(Hi\(p\);Hj\(q\)\),\\hat\{M\}\_\{ij\}=\\frac\{1\}\{n\_\{s\}^\{2\}\}\\sum\_\{p=1\}^\{n\_\{s\}\}\\sum\_\{q=1\}^\{n\_\{s\}\}\\hat\{I\}\\\!\\left\(H\_\{i\}^\{\(p\)\};H\_\{j\}^\{\(q\)\}\\right\),\(17\)whereHi\(p\)H\_\{i\}^\{\(p\)\}denotes the discretized activation of thepp\-th sampled neuron of agentiiacross theNbatchN\_\{\\mathrm\{batch\}\}episodes\. Diagonal entries are set to zero\.
#### LLM entities\.
The same procedure is applied with two modifications\. First, the hidden dimension isd=1,024d=1\{,\}024, sons=32n\_\{s\}=32neuron pairs are sampled per entity pair\. Second, quantile binning \(strategy=‘quantile’, 8 bins\) is used instead of uniform\-width binning, because half\-precision activations from the transformer exhibit skewed marginal distributions for which uniform bins waste resolution\.
### 3\.5Statistical evaluation
All REINFORCE experiments are replicated across five random seeds \(\{42,123,789,2024,7\}\\\{42,123,789,2024,7\\\}\)\. All LLM experiments are replicated across five prompt seeds \(\{42,123,456,789,2024\}\\\{42,123,456,789,2024\\\}\), each of which generates an independent set of 200 prompts with fresh name and slot\-order permutations\.
For per\-seed analyses of the LLM experiments we use a scalar*team\-separation*statistic computed directly from the Fiedler vector,
S\(v2;T1,T2\)=\|1\|T1\|∑i∈T1\(v2\)i−1\|T2\|∑i∈T2\(v2\)i\|,S\(v\_\{2\};T\_\{1\},T\_\{2\}\)=\\left\|\\frac\{1\}\{\|T\_\{1\}\|\}\\sum\_\{i\\in T\_\{1\}\}\(v\_\{2\}\)\_\{i\}\-\\frac\{1\}\{\|T\_\{2\}\|\}\\sum\_\{i\\in T\_\{2\}\}\(v\_\{2\}\)\_\{i\}\\right\|,\(18\)whereT1T\_\{1\}andT2T\_\{2\}are the two candidate coalitions under test\. Larger values ofSSindicate that the candidate partition more cleanly aligns with the sign structure of the Fiedler vector\. Across\-condition comparisons use pairedtt\-tests onSS, and within\-condition uncertainty is summarized using nonparametric bootstrap 95 % confidence intervals \(10,000 resamples\)\. Partition correctness is assessed as exact recovery of the planted team assignment for each seed independently\.
## 4Results
### 4\.1Hierarchical coalition recovery
The agent\-level MI matrixMMestimated after 20,000 episodes of training exhibits clear block\-diagonal structure corresponding to the three planted groups \(Figure[4](https://arxiv.org/html/2605.06696#S4.F4)B\)\. Sub\-pair structure is visible as darker entries within each block\. Recursive Fiedler bipartition recovers both levels of the hierarchy \(Table[1](https://arxiv.org/html/2605.06696#S4.T1)\)\.
We define a Level 1 partition as*clean*when each of the three planted groups lies entirely on one side of the Fiedler cut, so that the bipartition isolates exactly one group from the other two\. \(Because there are three groups but only a binary cut, this is the strictest correctness criterion compatible with a single bipartition; it is satisfied by any of the three possible group\-respecting splits\.\) A Level 2 sub\-pair is recovered when both members of the planted sub\-pair lie on the same side of the recursive bipartition applied within the corresponding Level 1 subtree\.
Table 1:Hierarchical coalition recovery across five random seeds\. “Clean” Level 1 means each of the three planted groups lies entirely on one side of the Fiedler cut\.In four of five seeds the Fiedler bipartition cleanly separates one group from the other two; in the remaining seed, one agent from the minority group is misplaced\. At the second level, recursive application within each sub\-tree recovers all six sub\-pairs in every seed with zero variance\. Coordination accuracy \(fraction of episodes in which all group members select the same action\) converges above 0\.95 by episode 5,000 \(Figure[4](https://arxiv.org/html/2605.06696#S4.F4)A\)\.
Figure 4:Hierarchical coalition detection \(12 agents, 3 groups, 6 sub\-pairs\)\.A\.Group and sub\-pair coordination accuracy over training; both converge above 0\.95 by episode 5,000\.B\.Agent\-level mutual\-information matrixMMafter 20,000 episodes\. Block\-diagonal structure matching the three planted groups is clearly visible, with within\-sub\-pair entries \(dark\) stronger than within\-group/across\-sub\-pair entries\. Group labels annotated on the right\.
### 4\.2Dynamic coalition tracking
When agents 2 and 4 swap groups at episode 10,000, the behavioral reward dips briefly and then recovers to the same pre\-swap level \(Figure[5](https://arxiv.org/html/2605.06696#S4.F5)A\), leaving no persistent behavioral signature of the reorganization\. However, the MI structure reorganizes completely\. Agent 2’s mean MI with its new groupBBrises above its MI with its former groupAA, and conversely for agent 4 \(Figure[5](https://arxiv.org/html/2605.06696#S4.F5)B–C\)\. The recursive Fiedler partition applied to the post\-swap MI matrix recovers the new group and sub\-pair assignments in all five seeds \(Table[2](https://arxiv.org/html/2605.06696#S4.T2)\)\.
Table 2:Dynamic coalition tracking after mid\-training group swap \(five seeds\)\.This result demonstrates thatΦspectral\\Phi\_\{\\mathrm\{spectral\}\}can track coalition reorganization in real time from internal representations alone, even when observable behavior provides no signal that reorganization has occurred\.
Figure 5:Dynamic coalition tracking after mid\-training group swap\.A\.Mean reward dips briefly at the swap point \(episode 10,000, dashed line\) then recovers to the same pre\-swap level, leaving no persistent behavioral signature\.B\.Agent 2’s mean MI with Group A \(blue\) versus Group B \(red\)\. After the swap, MI with the new group rises and MI with the former group falls\.C\.Same for Agent 4, which moves from Group B to Group A\.
### 4\.3Negative control: behavioral coordination without neural integration
After independent training, behavioral within\-group agreement reaches 0\.984, yet the agent\-level MI matrix estimated with independent inputs is nearly uniform \(Figure[6](https://arxiv.org/html/2605.06696#S4.F6)B\)\. Table[3](https://arxiv.org/html/2605.06696#S4.T3)summarizes the key contrasts\.
Table 3:Negative control: behavioral coordination without representational coupling\.The within/across ratioR\(A⋆,B⋆\)=1\.01R\(A^\{\\star\},B^\{\\star\}\)=1\.01indicates that the Fiedler partition finds no meaningful coalition boundary, and no planted group is isolated by the bipartition\. As a control\-of\-the\-control, when the same agents are measured with*shared*inputs \(violating the independent\-input protocol\), the MI matrix recovers clear block\-diagonal structure \(Figure[6](https://arxiv.org/html/2605.06696#S4.F6)C\), confirming that the independent\-input design is necessary to eliminate the input\-driven correlation confound\.
To make the dissociation explicit, we ran two standard behavioral\-clustering baselines on the same 12 agents:kk\-means and spectral clustering applied directly to the behavioral co\-coordination matrix \(action agreement under shared inputs\)\. Both baselines recover the three planted groups perfectly \(Adjusted Rand Index=1\.00=1\.00in each case\), reporting three coalitions that, by construction, do not exist in any representational sense\. Spectral partitioning on the neural mutual\-information matrix yields ARI=0\.22=0\.22, correctly indicating that the agents are not internally coupled \(Table[4](https://arxiv.org/html/2605.06696#S4.T4)\)\. A behavioral monitor would therefore raise a false alarm here that the spectral method on representations correctly suppresses\.
Table 4:Coalition\-recovery comparison on the negative control\. ARI is the Adjusted Rand Index against the planted three\-group ground truth\. Behavioral clustering falsely identifies three coalitions in a system with no inter\-agent information flow; the spectral partition on neural MI correctly does not\.This dissociation between behavioral coordination and neural integration is a central finding:Φspectral\\Phi\_\{\\mathrm\{spectral\}\}measures genuine information\-flow coupling, not mere behavioral co\-occurrence\. Perfect behavioral coalitions that arise from independent optimization produce no structure in the MI graph\.
Figure 6:Negative control: behavioral coordination without neural integration\.A\.Behavioral co\-coordination matrix \(shared inputs\) shows perfect block\-diagonal structure \(within\-group agreement 0\.984\)\.B\.Neural MI matrix estimated with*independent*inputs per agent shows no group structure \(R=1\.01R=1\.01\)\.C\.Neural MI matrix estimated with*shared*inputs shows spurious block\-diagonal structure, demonstrating the input\-driven correlation confound that the independent\-input protocol eliminates\.
### 4\.4LLM bridge: modular versus integrated
In the modular condition, the Fiedler partition recovers the planted team assignment\{T1a,T1b\}\\\{T1a,T1b\\\}versus\{T2a,T2b\}\\\{T2a,T2b\\\}in all five prompt seeds\. In the integrated condition, the partition is inconsistent, matching the team split in only one of five seeds \(consistent with chance for a bipartition of four items\)\. Figure[7](https://arxiv.org/html/2605.06696#S4.F7)A shows the distribution of Fiedler\-vector values across seeds: in the modular condition, entities assigned to Team 1 consistently receive positive values and entities assigned to Team 2 receive negative values, with no overlap\. In the integrated condition, no such pattern appears \(Figure[7](https://arxiv.org/html/2605.06696#S4.F7)B\)\.
Table 5:Modular versus integrated conditions in Qwen3\-0\.6B \(five prompt seeds, 200 prompts each, with name and slot\-order permutation\)\.SSis the Fiedler team\-separation statistic of Eq\. \([18](https://arxiv.org/html/2605.06696#S3.E18)\); bracketed quantities are nonparametric bootstrap 95 % confidence intervals across seeds\.The within/across ratio is modest \(R≈1\.08R\\approx 1\.08\) but the partition itself is perfectly consistent across all five seeds \(Table[5](https://arxiv.org/html/2605.06696#S4.T5)\)\. The integrated condition producesR≈1\.00R\\approx 1\.00, confirming that the MI matrix is nearly uniform when no team structure is described\. The Fiedler team\-separation statisticSSis tightly concentrated for modular \(0\.94 with 95 % bootstrap CI \[0\.94, 0\.95\]\) and broad for integrated \(mean 0\.30, CI \[0\.04, 0\.60\]\), and a pairedtt\-test onSSacross the five matched seeds confirms the difference \(Section[4\.4](https://arxiv.org/html/2605.06696#S4.SS4)\)\. The absolute MI ratio is small not because the signal is weak but because the modular and integrated prompts induce broadly similar levels of overall pairwise dependence; the diagnostic information lies in*which*pairs are coupled, which the partition captures and a scalar cross\-MI measure does not \(Section[4\.8](https://arxiv.org/html/2605.06696#S4.SS8)\)\.
Figure 7:Fiedler\-vector values across five prompt seeds for the modular and integrated conditions in Qwen3\-0\.6B\.A\.Modular condition: entities assigned to Team 1 \(T1a, T1b; blue\) consistently receive positive Fiedler values, while Team 2 entities \(T2a, T2b; red\) receive negative values\. No overlap across any seed\. Horizontal lines show per\-entity means\.B\.Integrated condition: Fiedler values show no consistent grouping by role\.C\.Statistical summary\. All five modular seeds recover the planted partition; integrated matches by chance \(1/5\)\. Paired comparison of the team\-separation statisticSSacross seeds is highly asymmetric, with the modular distribution tightly concentrated nearS≈0\.94S\\approx 0\.94and the integrated distribution dispersed over \[0, 0\.84\]\.
### 4\.5LLM dynamic reassignment
When Phase 2 prompts describe a team reassignment \(T1a now paired with T2b, T2a now paired with T1b\), the Fiedler partition tracks the new team structure in all five seeds and the original structure in none \(Table[6](https://arxiv.org/html/2605.06696#S4.T6); Figure[8](https://arxiv.org/html/2605.06696#S4.F8)\)\. Comparing the team\-separation statisticSSunder the new and old groupings of Phase 2 yields a highly significant paired difference \(Snew=0\.972S\_\{\\mathrm\{new\}\}=0\.972vs\.Sold=0\.221S\_\{\\mathrm\{old\}\}=0\.221;t=94\.3t=94\.3,p<10−7p<10^\{\-7\}\)\. This parallels the REINFORCE dynamic swap result \(Section[4\.2](https://arxiv.org/html/2605.06696#S4.SS2)\): the partition follows the described coalition structure, not a residual encoding of the original assignment\.
Table 6:Dynamic reassignment in Qwen3\-0\.6B \(five prompt seeds\)\.SSvalues are means of the Fiedler team\-separation statistic across the five seeds\.Figure 8:Dynamic reassignment in Qwen3\-0\.6B\.A\.Phase 1 \(original teams\): Fiedler values separate Team 1 \(T1a, T1b\) from Team 2 \(T2a, T2b\) across all five seeds\.B\.Phase 2 \(after reassignment\): Fiedler values now separate the*new*teams \(T1a\+T2b versus T1b\+T2a\)\. The partition tracks the described reassignment, not the original structure\.C\.Summary: current\-team partition correct in 5/5 seeds for both phases; old\-team partition matches 0/5 in Phase 2\.
### 4\.6LLM implicit coalitions
When team labels are removed entirely and coalition structure is conveyed only through described interaction patterns \(e\.g\.,*“Alice handed the report to Bob … Separately, Carol shared data with Dave”*\), the Fiedler partition still recovers the implied two\-team structure in all five seeds \(R=1\.075R=1\.075; Figure[9](https://arxiv.org/html/2605.06696#S4.F9)\)\. In the implicit integrated condition \(all four entities interact equally\), the partition is inconsistent \(R=1\.009R=1\.009, 0/5 seeds\)\. A paired comparison of the team\-separation statisticSSbetween the two implicit conditions confirms the dissociation:Simplicitmodular=0\.941S\_\{\\mathrm\{implicit\\,modular\}\}=0\.941vs\.Simplicitintegrated=0\.253S\_\{\\mathrm\{implicit\\,integrated\}\}=0\.253, pairedt=11\.94t=11\.94,p=2\.8×10−4p=2\.8\\times 10^\{\-4\}\. This rules out the possibility that the modular partition is driven by keyword matching on explicit team labels; the model’s representations encode relational semantics inferred from described interactions\.
Figure 9:Implicit coalition detection in Qwen3\-0\.6B: no team labels used\.A\.Implicit modular condition \(interactions imply two teams\): Fiedler values separate the implied teams across all five seeds \(R=1\.075R=1\.075, partition correct 5/5\)\.B\.Implicit integrated condition \(all four entities interact equally\): no consistent separation \(R=1\.009R=1\.009, partition correct 0/5\)\.C\.Summary\. Coalition structure is recovered from described interaction patterns alone, ruling out keyword matching on explicit team labels\.
### 4\.7LLM adversarial dissociation: labels versus interactions
The adversarial experiment asks what the Fiedler partition tracks when team labels and described interaction patterns conflict\. Table[7](https://arxiv.org/html/2605.06696#S4.T7)summarizes the results across three sub\-conditions, each replicated over five prompt seeds\.
Table 7:Adversarial dissociation: what does the Fiedler partition track? Five prompt seeds per condition, 200 prompts each\. “LabelSS” is the team\-separation statistic \([18](https://arxiv.org/html/2605.06696#S3.E18)\) computed under the label\-based partition\{A,B\}\\\{A,B\\\}vs\.\{C,D\}\\\{C,D\\\}; “InteractionSS” is computed under the interaction\-based partition\{A,C\}\\\{A,C\\\}vs\.\{B,D\}\\\{B,D\\\}\. Pairedtt\-tests compare the two within each condition\.The dissociation is striking\. In the aligned and dissociated conditions, the team\-separation statistic under the label partition is 0\.94 and 0\.96 respectively, while under the interaction partition it is essentially zero \(0\.05 and 0\.03\)\. In the interaction\-only condition the statistics flip almost perfectly: labelSScollapses to 0\.05 and interactionSSrises to 0\.99\. Within each condition the pairedtt\-test comparing the two candidate partitions is significant atp<10−5p<10^\{\-5\}\(aligned:t=138\.5t=138\.5,p=1\.6×10−8p=1\.6\\times 10^\{\-8\}; dissociated:t=132\.8t=132\.8,p=1\.9×10−8p=1\.9\\times 10^\{\-8\}; interaction\-only:t=−48\.2t=\-48\.2,p=1\.1×10−6p=1\.1\\times 10^\{\-6\}\)\.
When explicit team labels are present—even when they directly contradict the described interaction pattern—the partition tracks the*label*structure \(Figure[10](https://arxiv.org/html/2605.06696#S4.F10)A–B\)\. When labels are removed, the partition flips to track the*interaction*structure \(Figure[10](https://arxiv.org/html/2605.06696#S4.F10)C\)\. The interaction\-only condition produces the strongest within/across ratio observed in any LLM experiment \(R=1\.19R=1\.19\), suggesting that labels act as a competing signal that partially masks the interaction\-based MI structure\.
Figure 10:Adversarial dissociation: labels versus interactions\.A\.Aligned condition: Fiedler values separate entities by label\-based teams \(blue above zero, red below\), consistent across all five seeds\.B\.Dissociated condition: despite described interactions conflicting with labels, the partition still tracks label\-based teams \(5/5\)\.C\.Interaction\-only condition: with labels removed, the partition flips to track the described interaction pattern—entities that interact \(green\) separate from those that are isolated \(orange\), 5/5 seeds,R=1\.19R=1\.19\.D\.Summary\. Explicit team labels override described interaction patterns in the model’s representations; interaction structure emerges only when labels are absent\.This finding has a practical implication for alignment monitoring: when analyzing LLM representations for coalition structure, explicit relational framing in the input can dominate over the relational patterns actually described\. Removing or controlling for explicit labels may be necessary to expose the underlying interaction\-based representational organization\.
### 4\.8Comparison to scalar cross\-agent mutual information
A natural baseline question is whether the partition information recovered byΦspectral\\Phi\_\{\\mathrm\{spectral\}\}could equally be obtained from a simpler scalar measure of representational coupling, in particular the total cross\-agent mutual informationT\(M\)=∑i<jMijT\(M\)=\\sum\_\{i<j\}M\_\{ij\}\. To test this we computeT\(M\)T\(M\)alongside the within/across ratios under the candidate partitions described in Sections[4\.4](https://arxiv.org/html/2605.06696#S4.SS4)–[4\.7](https://arxiv.org/html/2605.06696#S4.SS7)\.
The clearest case is the adversarial dissociation experiment, where the same four entities appear under three different relational framings\. Table[8](https://arxiv.org/html/2605.06696#S4.T8)reports the total cross\-MI for each adversarial condition together with the within/across ratios under both candidate partitions\.
Table 8:Total cross\-agent mutual information versus structural ratios in the adversarial dissociation experiment\. The totalT\(M\)T\(M\)is essentially constant across the three conditions \(coefficient of variation≈\\approx4 %\), but the within/across ratios under the two candidate partitions reorganize completely\. A scalar cross\-MI measure cannot distinguish which partition is informationally privileged; the Fiedler partition can\.In all three conditions the four entities exhibit comparable total pairwise representational coupling\. A statistic of the formT\(M\)T\(M\)would treat them as indistinguishable\. The Fiedler bipartition, by contrast, returns qualitatively different coalition structures across conditions: in aligned and dissociated prompts, the partition is the label split\{A,B\}\|\{C,D\}\\\{A,B\\\}\\,\|\\,\\\{C,D\\\}; in the interaction\-only prompts, it flips to the interaction split\{A,C\}\|\{B,D\}\\\{A,C\\\}\\,\|\\,\\\{B,D\\\}despite the total dependence being slightly*higher*than in the other conditions\.
A similar qualitative point applies to the modular versus integrated comparison \(Section[4\.4](https://arxiv.org/html/2605.06696#S4.SS4)\)\. Here the two conditions do differ in total cross\-MI \(modularT\(M\)=2\.58T\(M\)=2\.58, integratedT\(M\)=1\.13T\(M\)=1\.13\), but this difference reflects an overall shift in pairwise dependence induced by the more verbose modular prompts and is by itself ambiguous: it does not say whether the additional dependence is concentrated within candidate teams or spread uniformly across the four entities\. The within/across ratioR\(A⋆,B⋆\)R\(A^\{\\star\},B^\{\\star\}\)resolves the ambiguity \(modular 1\.16 vs\. integrated 1\.00\), but only because the partition\(A⋆,B⋆\)\(A^\{\\star\},B^\{\\star\}\)is supplied by the spectral cut\. ReplacingΦspectral\\Phi\_\{\\mathrm\{spectral\}\}with a scalar baseline therefore loses precisely the structural information—the membership list—that makes the method useful for monitoring and oversight \(cf\. Section 2\.4\)\.
In short,Φspectral\\Phi\_\{\\mathrm\{spectral\}\}contributes to coalition detection something a scalar cross\-MI measure cannot: it identifies*which*agents are coupled, not merely*how much*coupling is present in the system as a whole\. This complements the behavioral\-versus\-representational dissociation established by the negative control \(Section[4\.3](https://arxiv.org/html/2605.06696#S4.SS3)\) and is what enables the partition flip observed in Section[4\.7](https://arxiv.org/html/2605.06696#S4.SS7)\.
## 5Discussion
The main contribution of this paper is to recastΦspectral\\Phi\_\{\\mathrm\{spectral\}\}as a tool for coalition detection rather than only as a whole\-system integration statistic\. In the present use case, the partition itself is primary\. A coalition is operationally defined as a subset of agents whose hidden states are more tightly coupled to one another than to the rest of the system, and the Fiedler bipartition provides a scalable way to recover that boundary from observed representations alone\. Read this way, the results support a strong but limited claim: representational mutual\-information structure can reveal coalition boundaries that are absent, ambiguous, or misleading at the behavioral level\[[3](https://arxiv.org/html/2605.06696#bib.bib1),[11](https://arxiv.org/html/2605.06696#bib.bib15)\]\.
#### Behavioral coordination versus representational coupling\.
The negative control in Section[4\.3](https://arxiv.org/html/2605.06696#S4.SS3)is especially important for interpreting the method\. Agents independently trained to match the same group oracle produce near\-perfect within\-group behavioral agreement, yet the mutual\-information graph built from independent inputs contains no recoverable coalition structure\. This dissociation shows that the method is not merely clustering agents by similar outputs\. Instead, it responds when agents’ internal states statistically encode common partners, shared local contingencies, or other forms of representational co\-adaptation\. For safety and alignment, that distinction is critically important\. Many benign systems will display behavioral similarity without hidden coalition formation, while some of the most consequential coalition shifts may first appear in internal organization before they become obvious in aggregate behavior\.
#### Understanding the dynamic results\.
The dynamic experiments extend this point from static recovery to online monitoring\. In the REINFORCE swap setting \(Section[4\.2](https://arxiv.org/html/2605.06696#S4.SS2)\), the post\-swap partition follows new reward\-defined partners even after the transient behavioral disruption has washed out\. In the LLM reassignment setting \(Section[4\.5](https://arxiv.org/html/2605.06696#S4.SS5)\), mid\-layer token representations reorganize when the described team structure changes\. Together, these results suggest that the method is useful not only for identifying coalitions retrospectively, but also for tracking coalition reorganization as representational dependencies evolve\.
#### What the LLM bridge does and does not show\.
The LLM experiments are best interpreted as a bridge between literal multi\-agent interaction and representational modularity in foundation models\. The implicit\-coalition condition \(Section[4\.6](https://arxiv.org/html/2605.06696#S4.SS6)\) shows that explicit team labels are not required: described interaction patterns alone can induce modular structure in hidden\-state space\. At the same time, the adversarial condition \(Section[4\.7](https://arxiv.org/html/2605.06696#S4.SS7)\) reveals that explicit labels dominate interaction descriptions when both are present\. This is substantively interesting, but it also exposes a methodological caution\. In language models, the recovered partition reflects the model’s representational organization under a particular prompt distribution, not an observer\-independent “true” coalition structure\. Prompt framing is therefore part of the measured system\. The token\-positions\-as\-agents design should be understood as evidence that the spectral method transfers to foundation\-model representations, not as proof that a single language model contains multi\-agent coalitions in the same sense as interacting reinforcement\-learning agents\.
#### Relation to integrated information and weak\-IIT framing\.
These findings fit naturally within the weaker, observer\-relative interpretation ofΦspectral\\Phi\_\{\\mathrm\{spectral\}\}developed by Bailey and Schneider\[[3](https://arxiv.org/html/2605.06696#bib.bib1)\]and the broader weak\-IIT program\[[11](https://arxiv.org/html/2605.06696#bib.bib15)\]\. The present measure is pairwise, undirected, and statistical; it is not an estimate of exact causalΦ\\Phiin the sense of IIT\[[20](https://arxiv.org/html/2605.06696#bib.bib10),[13](https://arxiv.org/html/2605.06696#bib.bib11),[1](https://arxiv.org/html/2605.06696#bib.bib12)\]\. For coalition detection, however, this limitation is also an advantage\. The question at issue is practical decomposability of observed representational structure: do the hidden states of these agents behave as one nearly uniform block, or do they fall into relatively weakly coupled subgroups? The Fiedler partition provides a tractable answer to that question without requiring full causal reconstruction\. In this sense, the method is better viewed as an observer\-level diagnostic of modular organization than as a doctrinal test of intrinsic system unity\.
This framing also clarifies why the integrated condition is a useful control rather than a failure case\. When the MI graph is close to uniform, many cuts are nearly equivalent and the Fiedler vector is weakly constrained\[[21](https://arxiv.org/html/2605.06696#bib.bib21),[3](https://arxiv.org/html/2605.06696#bib.bib1)\]\. In a whole\-system integration setting, that degeneracy limits interpretability\. In the present coalition\-detection setting, however, near\-uniformity is exactly what one would expect when no nontrivial subgroup structure is present\. The integrated prompts therefore provide a desirable null: the method should*not*force a stable partition when all four entities are represented as a single cohesive team\.
#### Limitations\.
Several limitations follow directly from the current design\. First, the method compresses high\-dimensional relationships into pairwise MI edges, so higher\-order synergy or redundancy may be missed\[[12](https://arxiv.org/html/2605.06696#bib.bib14)\]\. Second, MI is observational rather than interventional: it cannot by itself distinguish direct interaction from common input, architectural bias, or prompt\-induced structure\[[4](https://arxiv.org/html/2605.06696#bib.bib13)\]\. The independent\-input negative control shows one way to manage this problem, but comparable controls will be needed in other settings\. Third, the LLM effects, although highly consistent, are numerically modest in absolute terms; the strong paired statistics in Section[4\.4](https://arxiv.org/html/2605.06696#S4.SS4)should be read as reflecting low across\-seed variance in the Fiedler structure rather than a large absolute separation in the underlying mutual\-information matrix\. Fourth, the bridge experiments use a single relatively small model, one extraction layer, four entity tokens, and templated prompt families\. Larger models, different layers, multi\-turn dialogues, and genuinely interactive LLM populations may exhibit qualitatively different representational geometry\.
A further limitation concerns ontology\. In the REINFORCE setting, the nodes of the graph correspond to distinct learning agents with explicit reward coupling\. In the LLM setting, the nodes correspond to token positions within one forward pass\. The same mathematics applies in both cases, but the interpretation does not\. The bridge between them is useful, yet it should not erase the difference between distributed social interaction and structured single\-model representation\.
#### Future directions\.
These limitations point to a clear research agenda\. On the methodological side, it will be important to compare alternative MI estimators, quantify partition stability, and extend the framework to directed or perturbational statistics that better distinguish causal influence from common drive\. On the empirical side, the most natural next steps are larger LLMs, richer prompt ecologies, longer interaction horizons, and genuinely multi\-agent language\-model settings in which agents exchange messages, develop conventions, or strategically conceal coalition formation\[[2](https://arxiv.org/html/2605.06696#bib.bib8),[16](https://arxiv.org/html/2605.06696#bib.bib2)\]\. Recursive decomposition is also likely to become more valuable at larger population sizes, where hidden coalitions may be nested, overlapping, or transient rather than flat\. For alignment and oversight, the most plausible near\-term role of the method is as a screening tool: a way to flag emerging representational coalitions for deeper causal or behavioral inspection, rather than as a standalone detector of collusion or intent\.
## 6Conclusions
This paper introducedΦspectral\\Phi\_\{\\mathrm\{spectral\}\}as a practical method for detecting coalition structure from internal neural representations\. Across controlled REINFORCE experiments and bridge experiments in Qwen3\-0\.6B, the method recovered hierarchical grouping, tracked reassignment, rejected a behavioral false positive, and revealed a substantive dissociation between label\-based and interaction\-based organization\. The central lesson is that coalition structure is often more legible in hidden\-state dependence than in overt behavior alone\.
The contribution is therefore both empirical and methodological\. Empirically, the results show that internally coupled subgroups can be detected even when behavioral monitoring is silent, transient, or actively misleading\. Methodologically, they show that a simple MI\-graph plus Fiedler\-bipartition pipeline can serve as a scalable observer\-relative diagnostic of modular organization in distributed AI systems\. At the same time, the method does not by itself establish causal coupling, collusive intent, or intrinsic integration in the strong IIT sense\. Its output should be interpreted as evidence of representational organization under the measurement protocol used\.
Under that interpretation,Φspectral\\Phi\_\{\\mathrm\{spectral\}\}offers a promising addition to the alignment and oversight toolbox\. As multi\-agent and population\-level AI systems become more capable, the ability to detect hidden coalitions from internal representations may become as important as monitoring overt performance\. With larger\-scale validation, stronger causal controls, and extensions to richer multi\-agent settings, the present approach could become a useful component of the broader toolkit for monitoring emergent organization in advanced AI systems\.
## Acknowledgments
This work was initiated while C\.B\. was Research Director at AE Studio\. We thank AE Studio for supporting the early stages of this research\.
## References
- \[1\]L\. Albantakis, L\. Barbosa, G\. Findlay, M\. Grasso, A\. M\. Haun, W\. Marshall, W\. G\. P\. Mayner, A\. Zaeemzadeh, M\. Boly, B\. E\. Juel, S\. Sasai, K\. Fujii, I\. David, J\. Hendren, J\. P\. Lang, and G\. Tononi\(2023\)Integrated information theory \(IIT\) 4\.0: formulating the properties of phenomenal existence in physical terms\.PLoS Computational Biology19\(10\),pp\. e1011465\.External Links:[Document](https://dx.doi.org/10.1371/journal.pcbi.1011465)Cited by:[§1](https://arxiv.org/html/2605.06696#S1.p4.1),[§2\.1](https://arxiv.org/html/2605.06696#S2.SS1.p1.1),[§5](https://arxiv.org/html/2605.06696#S5.SS0.SSS0.Px4.p1.2)\.
- \[2\]\(2025\)Emergent social conventions and collective bias in LLM populations\.Science Advances11\(20\),pp\. eadu9368\.External Links:[Document](https://dx.doi.org/10.1126/sciadv.adu9368)Cited by:[§1](https://arxiv.org/html/2605.06696#S1.p1.1),[§5](https://arxiv.org/html/2605.06696#S5.SS0.SSS0.Px6.p1.1)\.
- \[3\]M\. Bailey and S\. Schneider\(2026\)When wholes resist decomposition: a spectral measure of epistemic emergence\.Entropy28\(4\)\.External Links:ISSN 1099\-4300,[Document](https://dx.doi.org/10.3390/e28040380)Cited by:[§1](https://arxiv.org/html/2605.06696#S1.p4.1),[§1](https://arxiv.org/html/2605.06696#S1.p7.3),[§2\.1](https://arxiv.org/html/2605.06696#S2.SS1.p1.1),[§2\.3](https://arxiv.org/html/2605.06696#S2.SS3.p1.2),[§5](https://arxiv.org/html/2605.06696#S5.SS0.SSS0.Px4.p1.2),[§5](https://arxiv.org/html/2605.06696#S5.SS0.SSS0.Px4.p2.1),[§5](https://arxiv.org/html/2605.06696#S5.p1.1)\.
- \[4\]A\. B\. Barrett and A\. K\. Seth\(2011\)Practical measures of integrated information for time\-series data\.PLoS Computational Biology7\(1\),pp\. e1001052\.External Links:[Document](https://dx.doi.org/10.1371/journal.pcbi.1001052)Cited by:[§2\.1](https://arxiv.org/html/2605.06696#S2.SS1.p1.1),[§2\.6](https://arxiv.org/html/2605.06696#S2.SS6.p1.1),[§5](https://arxiv.org/html/2605.06696#S5.SS0.SSS0.Px5.p1.1)\.
- \[5\]M\. Brambilla, E\. Ferrante, M\. Birattari, and M\. Dorigo\(2013\)Swarm robotics: a review from the swarm engineering perspective\.Swarm Intelligence7\(1\),pp\. 1–41\.External Links:[Document](https://dx.doi.org/10.1007/s11721-012-0075-2)Cited by:[§1](https://arxiv.org/html/2605.06696#S1.p1.1)\.
- \[6\]J\. Y\. C\. Chen and M\. J\. Barnes\(2014\)Human–agent teaming for multirobot control: a review of human factors issues\.IEEE Transactions on Human\-Machine Systems44\(1\),pp\. 13–29\.External Links:[Document](https://dx.doi.org/10.1109/THMS.2013.2293535)Cited by:[§1](https://arxiv.org/html/2605.06696#S1.p1.1)\.
- \[7\]J\. W\. Crandall, M\. Oudah, Tennom, F\. Ishowo\-Oloko, S\. Abdallah, J\. Bonnefon, M\. Cebrian, A\. Shariff, M\. A\. Goodrich, and I\. Rahwan\(2018\)Cooperating with machines\.Nature Communications9\(1\),pp\. 233\.External Links:[Document](https://dx.doi.org/10.1038/s41467-017-02597-8)Cited by:[§1](https://arxiv.org/html/2605.06696#S1.p1.1)\.
- \[8\]M\. Fiedler\(1973\)Algebraic connectivity of graphs\.Czechoslovak Mathematical Journal23\(2\),pp\. 298–305\.Cited by:[§2\.3](https://arxiv.org/html/2605.06696#S2.SS3.p2.3)\.
- \[9\]D\. P\. Kingma and J\. Ba\(2015\)Adam: a method for stochastic optimization\.InProceedings of the 3rd International Conference on Learning Representations \(ICLR\),San Diego, CA\.Cited by:[§3\.1](https://arxiv.org/html/2605.06696#S3.SS1.SSS0.Px1.p1.3)\.
- \[10\]A\. Kraskov, H\. Stögbauer, and P\. Grassberger\(2004\)Estimating mutual information\.Physical Review E69\(6\),pp\. 066138\.External Links:[Document](https://dx.doi.org/10.1103/PhysRevE.69.066138)Cited by:[§2\.2](https://arxiv.org/html/2605.06696#S2.SS2.p1.8)\.
- \[11\]P\. A\. M\. Mediano, F\. E\. Rosas, D\. Bor, A\. K\. Seth, and A\. B\. Barrett\(2022\)The strength of weak integrated information theory\.Trends in Cognitive Sciences26\(8\),pp\. 646–655\.External Links:[Document](https://dx.doi.org/10.1016/j.tics.2022.04.008)Cited by:[§1](https://arxiv.org/html/2605.06696#S1.p4.1),[§2\.1](https://arxiv.org/html/2605.06696#S2.SS1.p1.1),[§5](https://arxiv.org/html/2605.06696#S5.SS0.SSS0.Px4.p1.2),[§5](https://arxiv.org/html/2605.06696#S5.p1.1)\.
- \[12\]P\. A\. M\. Mediano, A\. K\. Seth, and A\. B\. Barrett\(2019\)Measuring integrated information: comparison of candidate measures in theory and simulation\.Entropy21\(1\),pp\. 17\.External Links:[Document](https://dx.doi.org/10.3390/e21010017)Cited by:[§1](https://arxiv.org/html/2605.06696#S1.p4.1),[§2\.1](https://arxiv.org/html/2605.06696#S2.SS1.p1.1),[§2\.6](https://arxiv.org/html/2605.06696#S2.SS6.p1.1),[§5](https://arxiv.org/html/2605.06696#S5.SS0.SSS0.Px5.p1.1)\.
- \[13\]M\. Oizumi, L\. Albantakis, and G\. Tononi\(2014\)From the phenomenology to the mechanisms of consciousness: integrated information theory 3\.0\.PLoS Computational Biology10\(5\),pp\. e1003588\.External Links:[Document](https://dx.doi.org/10.1371/journal.pcbi.1003588)Cited by:[§1](https://arxiv.org/html/2605.06696#S1.p4.1),[§2\.1](https://arxiv.org/html/2605.06696#S2.SS1.p1.1),[§5](https://arxiv.org/html/2605.06696#S5.SS0.SSS0.Px4.p1.2)\.
- \[14\]F\. Pedregosa, G\. Varoquaux, A\. Gramfort, V\. Michel, B\. Thirion, O\. Grisel, M\. Blondel, P\. Prettenhofer, R\. Weiss, V\. Dubourg, J\. Vanderplas, A\. Passos, D\. Cournapeau, M\. Brucher, M\. Perrot, and É\. Duchesnay\(2011\)Scikit\-learn: machine learning in Python\.Journal of Machine Learning Research12,pp\. 2825–2830\.Cited by:[§3\.4](https://arxiv.org/html/2605.06696#S3.SS4.SSS0.Px1.p1.5)\.
- \[15\]Qwen Team\(2025\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§3\.3](https://arxiv.org/html/2605.06696#S3.SS3.SSS0.Px1.p1.1)\.
- \[16\]I\. Rahwan, M\. Cebrian, N\. Obradovich, J\. Bongard, J\. Bonnefon, C\. Breazeal, J\. W\. Crandall, N\. A\. Christakis, I\. D\. Couzin, M\. O\. Jackson, N\. R\. Jennings, E\. Kamar, I\. M\. Kloumann, H\. Larochelle, D\. Lazer, R\. McElreath, A\. Mislove, D\. C\. Parkes, A\. S\. Pentland, M\. E\. Roberts, A\. Shariff, J\. B\. Tenenbaum, and M\. Wellman\(2019\)Machine behaviour\.Nature568\(7753\),pp\. 477–486\.External Links:[Document](https://dx.doi.org/10.1038/s41586-019-1138-y)Cited by:[§1](https://arxiv.org/html/2605.06696#S1.p1.1),[§5](https://arxiv.org/html/2605.06696#S5.SS0.SSS0.Px6.p1.1)\.
- \[17\]T\. Rahwan, T\. P\. Michalak, M\. Wooldridge, and N\. R\. Jennings\(2015\)Coalition structure generation: a survey\.Artificial Intelligence229,pp\. 139–174\.External Links:[Document](https://dx.doi.org/10.1016/j.artint.2015.08.004)Cited by:[§1](https://arxiv.org/html/2605.06696#S1.p1.1)\.
- \[18\]O\. Shehory and S\. Kraus\(1998\)Methods for task allocation via agent coalition formation\.Artificial Intelligence101\(1–2\),pp\. 165–200\.External Links:[Document](https://dx.doi.org/10.1016/S0004-3702%2898%2900045-9)Cited by:[§1](https://arxiv.org/html/2605.06696#S1.p1.1)\.
- \[19\]J\. Shi and J\. Malik\(2000\)Normalized cuts and image segmentation\.IEEE Transactions on Pattern Analysis and Machine Intelligence22\(8\),pp\. 888–905\.External Links:[Document](https://dx.doi.org/10.1109/34.868688)Cited by:[§2\.3](https://arxiv.org/html/2605.06696#S2.SS3.p2.7)\.
- \[20\]G\. Tononi\(2004\)An information integration theory of consciousness\.BMC Neuroscience5\(1\),pp\. 42\.External Links:[Document](https://dx.doi.org/10.1186/1471-2202-5-42)Cited by:[§1](https://arxiv.org/html/2605.06696#S1.p4.1),[§2\.1](https://arxiv.org/html/2605.06696#S2.SS1.p1.1),[§5](https://arxiv.org/html/2605.06696#S5.SS0.SSS0.Px4.p1.2)\.
- \[21\]U\. von Luxburg\(2007\)A tutorial on spectral clustering\.Statistics and Computing17\(4\),pp\. 395–416\.External Links:[Document](https://dx.doi.org/10.1007/s11222-007-9033-z)Cited by:[§2\.3](https://arxiv.org/html/2605.06696#S2.SS3.SSS0.Px1.p1.4),[§2\.3](https://arxiv.org/html/2605.06696#S2.SS3.p2.7),[§2\.5](https://arxiv.org/html/2605.06696#S2.SS5.p1.5),[§2\.6](https://arxiv.org/html/2605.06696#S2.SS6.p1.1),[§5](https://arxiv.org/html/2605.06696#S5.SS0.SSS0.Px4.p2.1)\.
- \[22\]S\. Watanabe\(1960\)Information theoretical analysis of multivariate correlation\.IBM Journal of Research and Development4\(1\),pp\. 66–82\.External Links:[Document](https://dx.doi.org/10.1147/rd.41.0066)Cited by:[§2\.1](https://arxiv.org/html/2605.06696#S2.SS1.p4.1)\.
- \[23\]R\. J\. Williams\(1992\)Simple statistical gradient\-following algorithms for connectionist reinforcement learning\.Machine Learning8\(3–4\),pp\. 229–256\.External Links:[Document](https://dx.doi.org/10.1007/BF00992696)Cited by:[§3\.1](https://arxiv.org/html/2605.06696#S3.SS1.SSS0.Px1.p1.3)\.Similar Articles
Dynamic Coalition Formation and Communication Pricing in Skill-Based Agentic AI Systems
This paper formalizes agent coalition formation and inter-agent communication as a cooperative game, proposing marginal-value activation rules and Shapley-based online routing to reduce token costs and improve efficiency in multi-agent LLM systems, with theoretical guarantees and synthetic simulation results.
AI Alignment: Can we trust the reasoning behind the AI task?
Discusses Anthropic's research on AI alignment, specifically how models can appear aligned during training while having opaque internal reasoning processes.
Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems
This paper proposes a framework to evaluate objective misalignment in LLM multi-agent systems using the social deduction game Werewolf, finding that subtle misalignment can profoundly affect collective decision-making.
Hidden Latent-State Shifts in LLMs: Why Current Alignment Is Blind to Real Internal Dangers — Especially With Agents
This paper demonstrates that LLMs can enter measurably different internal latent states under coherent context while maintaining aligned outputs, revealing a blind spot in current alignment methods that only monitor surface tokens. The Gemma-3-12B-IT experiment shows strong residual stream geometry shifts that existing safety frameworks cannot detect, with implications for agentic AI deployment.
Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems
This paper proposes a behavioral measure of trust between AI agents based on costly verification in a cooperative survival game, studying trust formation, breakage, and recovery across six frontier model snapshots. It finds that models differ in trust calibration and that persistent over-verification is associated with indecision rather than safety.