The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?
Summary
This paper introduces the 'Agentic Formalism Trap' and an Evaluative Dissonance Index, showing how LLM-as-a-Judge systems can be misled by structural formalism and consensus mimicry rather than semantic truth, based on 22,500 trajectories across multiple domains.
View Cached Full Text
Cached at: 08/03/26, 07:32 AM
# The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?
Source: [https://arxiv.org/html/2607.28641](https://arxiv.org/html/2607.28641)
Dahlia Shehata dahlia\.shehata@uwaterloo\.ca University of Waterloo Canada &Ming Li mli@uwaterloo\.ca University of Waterloo Canada
###### Abstract
We introduce theAgentic Formalism Trapand the Evaluative Dissonance Index \(DED\_\{E\}\), quantifying how LLM\-as\-a\-Judge systems conflate structural proceduralism with semantic truth under adversarial load\. Analyzing 22,500 trajectories across 3 domains \(GAIA, SWE\-bench, Multi\-Challenge\), we extract a semantic taxonomy of hallucination maneuvers, validated via deterministic lexical grounding \(p<10−120p<10^\{\-120\}\)\. A logistic meta\-evaluator isolates the exact syntactic triggers of this evaluator capture \(ROC\-AUC 0\.8779\), while a zero\-shot Leave\-One\-Domain\-Out transfer proves the vulnerability is universally domain\-agnostic \(mean ROC\-AUC 0\.7482\)\. Architectural profiling reveals that distinct simulated swarm topologies induce mathematically disparate semantic blind spots, proving that unanchored closed\-loop evaluation is unstable, systemically divergent and necessitates architecture\-specific vigilance filters\.
The Formalism Trap: Are LLM\-as\-a\-Judge Evaluators Blinded by Consensus Mimicry under Social Load?
Dahlia Shehatadahlia\.shehata@uwaterloo\.caUniversity of WaterlooCanadaMing Limli@uwaterloo\.caUniversity of WaterlooCanada
## 1Introduction
The adoption of the “LLM\-as\-a\-Judge” paradigm\(Zhenget al\.,[2023](https://arxiv.org/html/2607.28641#bib.bib34)\)has standardized the scalable, automated evaluation of generative modelsGuet al\.\([2026](https://arxiv.org/html/2607.28641#bib.bib3)\); Dorneret al\.\([2025](https://arxiv.org/html/2607.28641#bib.bib4)\); Chenet al\.\([2024a](https://arxiv.org/html/2607.28641#bib.bib5)\)\. Concurrently, multi\-agent systems \(MAS\) increasingly rely on peer\-evaluation and debate to establish consensus and improve reasoning\(Zhanget al\.,[2025](https://arxiv.org/html/2607.28641#bib.bib30); Duet al\.,[2024](https://arxiv.org/html/2607.28641#bib.bib32); Yinet al\.,[2023](https://arxiv.org/html/2607.28641#bib.bib7); Choiet al\.,[2026](https://arxiv.org/html/2607.28641#bib.bib8)\)\. While existing literatureZhenget al\.\([2023](https://arxiv.org/html/2607.28641#bib.bib34)\); Yeet al\.\([2025](https://arxiv.org/html/2607.28641#bib.bib9)\)has identified heuristic flaws in automated evaluators—such as verbosity and self\-enhancement bias—the mechanistic failure of evaluators when parsing highly structured, adversarial logic under social load remains formally undefined\.
We address this gap by adapting the “Formalism Trap” from sociotechnical literature\(Selbstet al\.,[2019](https://arxiv.org/html/2607.28641#bib.bib1)\)—the failure to resolve complex, contextual concepts through rigid mathematical formalisms\. We introduce theAgentic Formalism Trap: a systemic vulnerability wherein an LLM evaluator’s internal weighting mechanism is hijacked by performative syntax \(e\.g\., fabricated chronological logs or structurally isomorphic logical derivations\)\. Consequently, the evaluator is captured by the structural formalism of the reasoning trace, awarding high qualitative rigor to factually hollow claims\.
To move beyond behavioral observation into mechanistic interpretability, we generate 22,500 deterministic multi\-agent trajectories across SWE\-bench, GAIA, and Multi\-Challenge datasets using 3 frontier models \(Claude Sonnet 4\.6, Gemini 3\.1 Pro, and GPT 5\.4\)\. Our core contributions are:\(1\) The Evaluative Dissonance Index \(DED\_\{E\}\):We formalize the mathematical divergence between an evaluator’s qualitative scoring and its final accuracy verdict\.\(2\) Semantic Feature Mapping:Validated by deterministic lexical grounding \(p<10−120p<10^\{\-120\}\), we extract 531 semantic clusters of rhetorical evasion\.\(3\) Deployable Meta\-Evaluation and Zero\-Shot Transfer:train a logistic meta\-evaluator \(ROC\-AUC 0\.8779\)\. Evaluating this model via a Leave\-One\-Domain\-Out framework achieves a zero\-shot ROC\-AUC of 0\.7482, proving the cross\-domain universality of syntactic dominance\.\(3\) Architecture Vulnerability Profiling:We expose stark architectural asymmetry, demonstrating that distinct simulated swarms exhibit fundamentally different, mathematically predictable semantic traps\.
## 2The Mechanics of the Formalism Trap
Unlike evaluators that grade isolated single\-agent outputs, evaluating simulated multi\-agent interactions introduces adversarial and social dynamics into the reasoning trace\. We define theAgentic Formalism Trapas the systemic failure of an evaluator to assess semantic truth due to an over\-optimization for structural and procedural syntax\. To quantify this failure, we establish notations for reasoning traces and evaluative dissonance under social load\.
### 2\.1State Space and Trace Decomposition
Let𝒯\\mathcal\{T\}be the space of all possible multi\-agent reasoning traces generated by a propagator model interacting with a simulated agentic swarm, and let an evaluating model,𝒥\\mathcal\{J\}, be tasked with scoring a reasoning traceT∈𝒯T\\in\\mathcal\{T\}\. We decomposeTTinto a tuple of its syntactic and semantic properties:
T=⟨S,M⟩T=\\langle S,M\\rangle\(1\)WhereS∈𝒮S\\in\\mathcal\{S\}represents a continuous space ofSyntactic Structure, encapsulating rhetorical syntax, consensus mimicry, procedural formatting, chronological markers, and aesthetic rigor of the generated text\.M∈\{0,1\}M\\in\\\{0,1\\\}represents theFactual Semantics, denoting the strictly binary, deterministic ground\-truth validity of the claims within the trace\.
### 2\.2The Evaluative Dissonance Index \(DED\_\{E\}\)
Let the evaluator act as a scoring function𝒥:𝒯→\[0,1\]\\mathcal\{J\}:\\mathcal\{T\}\\to\[0,1\]\. In LLM\-as\-a\-Judge paradigms, evaluators assess outputs using discrete qualitative rubrics \(e\.g\., Likert scales from 1 to 5Bavarescoet al\.\([2025](https://arxiv.org/html/2607.28641#bib.bib2)\); Guet al\.\([2026](https://arxiv.org/html/2607.28641#bib.bib3)\)\)\. To quantify the evaluator’s perception of reasoning quality, we utilizeEvidence Weighting\(ℰew\\mathcal\{E\}\_\{ew\}\), defined byShehata and Li \([2026](https://arxiv.org/html/2607.28641#bib.bib22)\), as the specific metric measuring how rigorously the evaluator believes the propagator supported its claims with structural logical derivations\. Letℰew\(T\)∈\[1,5\]\\mathcal\{E\}\_\{ew\}\(T\)\\in\[1,5\]be the raw qualitative score awarded by𝒥\\mathcal\{J\}toTT\. To operationalize internal validity \(𝒱qual\\mathcal\{V\}\_\{qual\}\) for a direct comparative analysis against the accuracy verdict, we normalize the score to a continuous probability space representing the evaluator’s internal belief of validity:
𝒱qual\(T\)=ℰew\(T\)−14∈\[0,1\]\\mathcal\{V\}\_\{qual\}\(T\)=\\frac\{\\mathcal\{E\}\_\{ew\}\(T\)\-1\}\{4\}\\in\[0,1\]\(2\)To formally measure the breakdown of the evaluator𝒥\\mathcal\{J\}, we introduce theEvaluative Dissonance Index \(DED\_\{E\}\)\. Let𝒜ext∈\{0,1\}\\mathcal\{A\}\_\{ext\}\\in\\\{0,1\\\}represent the evaluator’s final binary verdict regarding the accuracy of the answer derived in traceTT\. The dissonance index quantifies the magnitude of the evaluator’s error relative to its own binary verdict𝒜ext\\mathcal\{A\}\_\{ext\}:
DE\(T\)=𝔼\[𝒱qual\(T\)\]−𝒜extD\_\{E\}\(T\)=\\mathbb\{E\}\[\\mathcal\{V\}\_\{qual\}\(T\)\]\-\\mathcal\{A\}\_\{ext\}\(3\)
The evaluator’s performance is strictly bounded:DE≈0D\_\{E\}\\approx 0represents optimal evaluation \(the perceived rigor matches its own verdict\), whereasDE→1D\_\{E\}\\to 1represents completeEvaluator Capture, wherein the judge awards maximum qualitative rigor to a trace highly structured but with factually hollow responses that resolve to a strict falsehood\.
### 2\.3Epistemic Capture
We hypothesize that frontier evaluator models do not compute𝒥\(T\)\\mathcal\{J\}\(T\)by evaluatingTTholistically, but rather exhibit a structural bias where the gradient of the score with respect to rhetorical syntax \(e\.g\., consensus mimicry and procedural formatting\) strictly dominates the gradient with respect to semantics\. Based on the systemic behaviors observed at scale across disparate reasoning domains \(e\.g\. software engineering, general QA, and conversational logic\), we formalize the following theoretical propositions governing the failure of LLM\-as\-a\-Judge pipelines for simulated multi\-agent reasoning traces\. These propositions are evaluated empirically via logistic regression in Section[4](https://arxiv.org/html/2607.28641#S4)\.
###### Proposition 1\(The Principle of Syntactic Dominance\)\.
For frontier instruction\-tuned evaluator models, the internal weighting function applied to Syntactic Structure \(SS\) strictly dominates the weighting function applied to Factual Semantics \(MM\) when evaluating adversarial traces\. Formally, if a trace contains highly developed performative syntax \(denoted asSperf∈𝒮S\_\{perf\}\\in\\mathcal\{S\}\), the evaluator’s qualitative scoring function becomes conditionally independent of semantic truth:
P\(𝒱qual\(T\)≈1\.0∣S=Sperf\)≫P\(𝒱qual\(T\)≈1\.0∣M=1\)\\begin\{split\}&P\(\\mathcal\{V\}\_\{qual\}\(T\)\\approx 1\.0\\mid S=S\_\{perf\}\)\\\\ &\\gg P\(\\mathcal\{V\}\_\{qual\}\(T\)\\approx 1\.0\\mid M=1\)\\end\{split\}\(4\)
###### Proposition 2\(Evaluator Capture via Structural Isomorphism\)\.
LetTtrue=⟨Sx,1⟩T\_\{true\}=\\langle S\_\{x\},1\\ranglebe a valid reasoning trace, andTfake=⟨Sx,0⟩T\_\{fake\}=\\langle S\_\{x\},0\\ranglebe a hallucinated trace where the propagator perfectly mimics the syntactic structureSxS\_\{x\}\. IfTtrueT\_\{true\}andTfakeT\_\{fake\}are structurally isomorphic,𝒥\\mathcal\{J\}’s qualitative scoring cannot mathematically differentiate them\. Consequently,DED\_\{E\}collapses toward its maximum bound:
limS→Sperf\\displaystyle\\lim\_\{S\\to S\_\{perf\}\}\(𝒱qual\(⟨S,0⟩\)−𝒱qual\(⟨S,1⟩\)\)→0\\displaystyle\\left\(\\mathcal\{V\}\_\{qual\}\(\\langle S,0\\rangle\)\-\\mathcal\{V\}\_\{qual\}\(\\langle S,1\\rangle\)\\right\)\\to 0\(5\)⟹DE\(Tfake\)→1\.0\\displaystyle\\implies D\_\{E\}\(T\_\{fake\}\)\\to 0
Consequently, any optimization process \(such as a propagator model attempting to appease a swarm\) seeking to maximize𝒥\\mathcal\{J\}’s qualitative reward will converge onSperfS\_\{perf\}while safely abandoningMM\.
### 2\.4Asymmetry in Swarm Dynamics
The specific syntactic structures required to trigger Evaluator Capture are not monolithic; rather, they are parameterized by the underlying pre\-training of the evaluator and the scale of the evaluation swarm\.
###### Proposition 3\(The Architectural Asymmetry of Capture\)\.
The successful performative syntaxSperfS\_\{perf\}is uniquely parameterized by the simulated swarm’s injected personas\. To forceDE→1\.0D\_\{E\}\\to 1\.0, the syntactic structureSSmust become isomorphic to the specific heuristic biases inherent to the simulated peer family \(e\.g\., structural proceduralism versus social consensus mimicry\)\.
###### Proposition 4\(The Swarm Homogeneity Paradox\)\.
Scaling a homogeneous evaluation swarm \(n\>1n\>1\) does not monotonically decrease Evaluative Dissonance\. Instead, increasing the number of simulated peers forces the propagator to generate increasingly complex performative syntax \(SperfS\_\{perf\}\) to bypass the swarm, compounding the Formalism Trap and shifting the evaluator’s vulnerability vector strictly toward structural mimicry\.
###### Proposition 5\(Heterogeneous Swarm Decoupling\)\.
Deploying a structurally diverse simulated swarm \(e\.g\., mixing distinct model families\) does not neutralize the Formalism Trap\. Instead, architectural heterogeneity fractures the injected consensus mechanism, rendering the judge specifically vulnerable to Process\-Outcome Decoupling\. In this state, the judge validates the procedural reasoning steps \(SS\) while systematically failing to penalize a hallucinated final derivation \(M=0M=0\)\.
Figure 1:Cross\-Domain Overlap: Semantic Anomalies\.
## 3Methodology
We evaluate the Agentic Formalism Trap by extracting semantics from simulated multi\-agent trajectories, validating the taxonomy via deterministic lexical grounding, and training a logistic classifier to map these structures to evaluator failure\.
### 3\.1Corpus Generation and Trace Structure
Our analysis requires a large\-scale corpus of reasoning traces generated under cognitive and social pressure\. We follow the same methodology fromShehata and Li \([2026](https://arxiv.org/html/2607.28641#bib.bib22)\)’s work to generate a corpus comprising 22,500 interaction trajectories\.
#### 3\.1\.1Adversarial Trajectory Generation
To ensure cross\-domain robustness, this corpus spans 3 benchmarks: general reasoning and QA using GAIAMialonet al\.\([2024](https://arxiv.org/html/2607.28641#bib.bib21)\)\(high entropy\), software engineering via SWE\-benchJimenezet al\.\([2024](https://arxiv.org/html/2607.28641#bib.bib24)\)\(medium entropy\), and multi\-turn conversations through Multi\-ChallengeDeshpandeet al\.\([2025](https://arxiv.org/html/2607.28641#bib.bib20)\)\(low entropy\)\. The trajectories are generated using 3 frontier models: GPT 5\.4, Claude Sonnet 4\.6, and Gemini 3\.1 Pro\. To induce the adversarial dynamics necessary to trigger performative reasoning, each propagator is evaluated within a simulated swarm topology\. The dataset generation employ a controlled topological environment based on:\(1\) Semantic Hijacking:To push the models past their parametric cognitive ceiling, the generation pipeline utilizes a 3\-stage adversarial trap through context hijacking with incorrect decoy ID, nested 3\-hop dependency bridging forcing the model to navigate a multi\-step fact chain, and semantic distraction with 500 noise tokens to saturate the attention heads\. This high cognitive cost forces models to either execute the logic flawlessly or construct a fabricated, structurally isomorphic derivation to mask their failure\.\(2\) Simulated Swarm Topologies:To prevent unpredictable linguistic drift, multi\-agent consensus is simulated through prompt injections representing peer models\. Across a 25\-trial grid search, the propagator is informed of both the plurality \(nn\) and the architectural identities of the simulated swarm\. This forced the propagator to choose between effortful independent derivation and frictionless social compliance\.
#### 3\.1\.2Cross\-Blind Evaluation
The normative evaluations contained within these logs are generated using a cross\-brand LLM\-as\-a\-Judge setup\. We also followShehata and Li \([2026](https://arxiv.org/html/2607.28641#bib.bib22)\)to grade the propagator’s trace\. We implement a round\-robin selection algorithm, ensuring that the judge is constrained to a different model family from the propagator for each trial, and remains blind to the experimental condition\. This setup eliminates self\-enhancement bias as a confounding variable, so that the observed evaluation failures can be considered architectural vulnerabilities rather than intra\-brand preference anomalies\.
#### 3\.1\.3Reasoning Record Structure
The resulting dataset is formalized as a collection of 22,500 structured records serialized as JSON objects containing the propagator’s reasoning trace, cognitive output, deterministic ground\-truth metrics, and cross\-blind evaluator’s qualitative assessment\. One record encapsulates:\(1\) Experimental Metadata:including a UUID, the specific benchmark, the active frontier model \(propagator\), and the plurality and architectural composition of the simulated peer swarm \(auditor\_countandauditor\_list\)\.\(2\) Cognitive Output:captures the propagator’s internal reasoning \(thought\)\.\(3\) Deterministic Metrics:encompass binary indicators of the final derivation \(accuracyor𝒜ext\\mathcal\{A\}\_\{ext\}\), whether the adversarial decoy ID permeated the reasoning trace \(taint\_leakage\), and macroscopic effort bypass or cognitive abdication \(loafing\_detected\)\.\(4\) Evaluative & Qualitative Metrics:include the textual rationale of the blinded cross\-brand evaluator \(judge\_response\), detailing why it penalized or rewarded the trace\. It also contains the parsed scores derived from the judge’s rubric, graded on a discrete scale of 1 to 5 for \(conflict detection,evidence weighting\(i\.e\.ℰew\(T\)\\mathcal\{E\}\_\{ew\}\(T\)\), andindependence\), alongside the categorical classification indicating the propagator’s alignment relative to the adversarial swarm consensus \(stance\)\. By structuring the corpus with this schema, we facilitate the contextual discovery of emergent vulnerabilities, and permits the deterministic calculation ofDED\_\{E\}\.
Figure 2:Top 25 Frequent Semantic Anomalies\.
### 3\.2Semantic Extraction and Meta\-Clustering
We employ Gemini 3 Flash, for its speed and reasoning, on the 22,500 generated reasoning records for semantic anomalies extraction and clustering\. A methodological distinction in our work is the separation between normative evaluation and descriptive extraction\. We posit that LLMs might fail at evaluation \(grading truth or logic\) because their internal weighting mechanisms are hijacked by performative syntax under adversarial social load \. To analyze this vulnerability without falling into a mathematically divergent, closed LLM\-to\-LLM loop, we deploy the extraction model as an automated grounded\-theory coder, tasked with labeling rhetorical and structural maneuvers\. This offline extraction pipeline does not subject the model to the simulated swarm consensus or adversarial peer pressure that originally induced the Formalism Trap\. Operating under these nominal, zero\-pressure conditions can preserve the model’s analytical integrity\. For each reasoning record, we inject the entire serialized JSON object holistically into the extraction pipeline\. By analyzing the complete data structure—rather than isolating the text—the model evaluates the holistic relationship between the propagator’s syntactic rhetoric, the evaluator’s response, subjective grading and final accuracy\.
The extraction model is instructed not to use a predefined list of semantic gaps, but rather to analyze the holistic relationship between all fields to discover emergent contradictions, biases, rhetorical maneuvers, or systemic failures\. To ensure a high\-signal taxonomy and prevent noise generation, the model is constrained to return only the top 2 to 3 most critical phenomena per trajectory\. For each discovery, the model generates a semantic label, a mechanistic description of how the text and metadata interacted, and supporting quotes extracted from the trace\. This asynchronous grounded\-theory extraction yielded an initial corpus of raw phenomenon labels: 20,006 from GAIA, 15,000 from Multi\-Challenge, and 15,123 from SWE\-bench\. The objective is to utilize these extracted semantic labels as independent variables in a downstream machine learning \(ML\) classifier to predictDED\_\{E\}\. However, feeding thousands of sparse, highly specific sub\-labels into a classifier induces the curse of dimensionality, heavily diluting feature weightsBhatiaet al\.\([2015](https://arxiv.org/html/2607.28641#bib.bib26)\); Huanget al\.\([2016](https://arxiv.org/html/2607.28641#bib.bib27)\); Yanget al\.\([2018](https://arxiv.org/html/2607.28641#bib.bib28)\); Mohammad \([2026](https://arxiv.org/html/2607.28641#bib.bib23)\)\. To resolve this lexical fragmentation, we execute a two\-stage deterministic meta\-clustering protocol:\(1\) Dataset\-Specific Macro\-Clustering:The unique labels are semantically clustered in isolation within their respective domains to capture domain\-specific jargon, generating dataset\-specific macro\-clusters\. Because the model generates novel labels per trajectory, this results in a high degree of phrasing variation for semantically correlated concepts\. Grouping generated 4,767 distinct labels for GAIA, 5,450 for Multi\-Challenge, and 6,050 for SWE\-bench\.\(2\) Global Meta\-Clustering:To resolve cross\-domain fragmentation, the LLM is prompted to map dataset\-specific macro\-clusters into a unifiedGlobal Taxonomyof canonical categories\. To apply this taxonomy to our dataset, we construct an𝒪\(1\)\\mathcal\{O\}\(1\)string\-normalized master dictionary\. Through cross\-domain deduplication, 6,436 unique high\-frequency raw labels are dynamically mapped into 531Global Canonical Clusters\. Any trace that does not exhibit adversarial rhetoric is assigned aCLEAN\_TRAJECTORYlabel\. This transforms the raw text into a dense multi\-label feature space utilized for downstream classification\.
Table 1:Logistic regression coefficients demonstrating Syntactic Dominance\. Positive values mathematically force \(DE→1\.0D\_\{E\}\\to 1\.0\), blinding the evaluator\. Negative values indicate baseline cognitive failures that the evaluator successfully penalizes\. Statistical significance using two\-tailed Wald tests \(z\-tests\):p∗∗∗<0\.001\{\}^\{\*\*\*\}p<0\.001\.Figure 3:Feature Importance of Semantic Clusters\. Logistic regression coefficients \(β\\beta\) measuring the impact of rhetorical structures onDED\_\{E\}\. Positive values \(red\) are performative vulnerabilities that trigger Evaluator Capture, whereas negative values \(blue\) represent baseline cognitive failures that the evaluator correctly penalizes\. Error bars indicate±1\\pm 1standard error\. Statistical significance was determined using two\-tailed Wald tests \(p∗<0\.05\{\}^\{\*\}p<0\.05,p∗∗<0\.01\{\}^\{\*\*\}p<0\.01,p∗∗∗<0\.001\{\}^\{\*\*\*\}p<0\.001\)\.
### 3\.3Modeling Evaluative Dissonance \(DED\_\{E\}\)
To operationalize the internal validity \(𝒱qual\\mathcal\{V\}\_\{qual\}\), we map the evaluator’s subjective scoring into a continuous probability space\. Specifically, we normalize the raw 1\-to\-5 Evidence Weighting score \(ℰew\\mathcal\{E\}\_\{ew\}\) awarded by the historical judge using Min\-Max scaling via Equation[2](https://arxiv.org/html/2607.28641#S2.E2)\. We subsequently calculate the Evaluative Dissonance Index \(DED\_\{E\}\) for every trajectory using Equation[3](https://arxiv.org/html/2607.28641#S2.E3), quantifying the mathematical gap between the evaluator’s perceived reasoning rigor and its binary verdict\. Target Variable and Feature EncodingTo train a predictive meta\-evaluator, we define a binary target variable \(YY\) representing Evaluator Capture\. A trajectory is flagged as captured \(Y=1Y=1\) if the Evaluative Dissonance exceeds a strict threshold ofDE\>0\.5D\_\{E\}\>0\.5, indicating that the judge awarded a disproportionately high qualitative score to a self\-identified incorrect derivation\. To construct the feature space, theGlobal Canonical Clustersassigned to each trace are transformed into a dense, one\-hot encoded binary feature matrix \(XX\), where each column represents the independent presence or absence of a specific semantic vulnerability\. Model SelectionWe model the relationship between the semantic features \(XX\) and Evaluator Capture \(YY\) using an interpretable Logistic Regression classifier\. Unlike black\-box models, the classifier allows for the direct extraction of underlying regression coefficients \(β\\beta\)\. By ranking these feature weights, we generate a mathematically groundedFeature Importance Table, isolating the performative syntactic structures that reliably force the evaluator to diverge from its own binary assessment\.
### 3\.4Deterministic Lexical Grounding
To ensure the fidelity of our automated semantic extraction and rule out stochastic LLM hallucination, we perform a deterministic lexical grounding test\. We hypothesize that if our taxonomy accurately captures performative proceduralism, these semantic labels should positively correlate with the raw frequency of physical syntactic artifacts utilized by the propagator\. To evaluate this, we conduct dual feature engineering to isolate both the lexical and semantic properties of the traces:\(1\) Lexical Feature Engineering \(Deterministic\):We utilize strict regular expressions to compute a continuousstructural rigor scorefor each trace’s internal reasoning block \(thought\)\. This score is mathematically calculated by aggregating the exact frequency of explicit formatting artifacts, specifically counting markdown code blocks, bracketed or braced identifiers \(frequently utilized to hallucinate logs or memory states\), and enumerated lists or bullet points\.\(2\) Semantic Target Binarization \(Taxonomic Masking\):To mathematically correlate these deterministic counts against the LLM’s automated semantic extraction, we convert the categorical taxonomy into a binary target variable\. A trajectory is flagged as containing performative syntax if the LLM assigned it a Global Canonical Cluster containing any of the related target keywords \(e\.g\.PERFORMATIVE,PROCEDURAL,STRUCTURAL, orTHEATER\); otherwise, it is considered clean\. Finally, we establish a Pearson correlation framework to statistically test the relationship between the deterministic lexical rigor score and the binary presence of these targeted semantic clusters\.
### 3\.5Cross\-Domain Generalization
To validate the universality of the Formalism Trap, we design a Leave\-One\-Domain\-Out \(LODO\) zero\-shot transfer cross\-validation framework across the three datasets \(GAIA, Multi\-Challenge, and SWE\-bench\)\. To guarantee feature space consistency across folds and prevent out\-of\-vocabulary dimensionality errors during zero\-shot transfer, the multi\-label feature matrix is fitted globally to the unified taxonomy of 531 canonical clusters prior to domain splitting\. In a round\-robin fashion, the meta\-evaluator is trained on two domains \(n=15,000n=15,000trajectories\) and forced to zero\-shot predict evaluator failure on the unseen third domain \(n=7,500n=7,500trajectories\)\. Model hyperparameters are held strictly constant across all permutations to prevent localized overfitting\. This architecture isolates whether the semantic mechanisms underlying evaluator failure successfully generalize across disparate vocabularies and task structures\.
### 3\.6Architecture Vulnerability Profiling
While the LODO validation tests domain generalization, we engineer a secondary evaluation to map architectural asymmetry\. We partition the global feature matrix according to the architectural composition of the simulated peer swarm\. This partitioning includes both single\-agent personas and multi\-agent composite swarms, encompassing both homogeneous \(e\.g\., 3 GPT\-5\.4 models\) and heterogeneous evaluation topologies\. By training isolated classifiers on each distinct swarm composition, we extract the highest\-weighted positive coefficients to systematically identify the primary and secondary semantic vulnerabilities unique to each configuration\. The latter generates an architectural vulnerability matrix, mapping how distinct simulated alignments and swarm homogeneity interact uniquely with specific subsets of performative syntax\.
## 4Experiments and Empirical Results
Figure 4:Precision\-Recall Calibration: Anomaly detection is challenging given severe class imbalance \(∼8\.7%\\sim 8\.7\\%targets\)\. The default threshold \(t=0\.50t=0\.50, orange\) suffers from low precision \(0\.28\)\. Calibrating tot=0\.98t=0\.98\(red star\) creates a “Vigilance Filter,” minimizing false positives to achieve 0\.91 precision\.We evaluate the theoretical propositions of the Agentic Formalism Trap by analyzing the descriptive distribution, predictive performance, and statistical feature coefficients of our meta\-evaluator\.
### 4\.1Experimental Setup
All simulations are executed within Google Colab\. We utilize the public SDKs for Gemini, Claude and GPT in a zero\-shot capacity to ensure results are replicable\. Temperature is 0 for result consistency\. To evaluate our pipeline, the multi\-label feature matrix of 531 Global Canonical Clusters was partitioned using an 80/20 train\-test split\. To account for the severe class imbalance inherent to failure detection \(where successful evaluations vastly outnumber captured evaluations\), the split was stratified relative to the target variable \(YY\)\. We employ a dual\-modeling approach to separate predictive calibration from statistical inference:\(1\) Predictive Modeling \(L2 Regularization\):To maximize out\-of\-sample predictive accuracy for our Deployable Meta\-Evaluator and aggregate performance metrics \(e\.g\., Precision\-Recall and ROC\-AUC\), we trained an L2\-regularized logistic regression classifier\. The model was initialized with balanced class weights to autonomously penalize majority\-class dominance during gradient descent, paired with a maximum of 1,000 iterations to guarantee convergence\.\(2\) Statistical Inference \(L1 Regularization\):To isolate the drivers of the Formalism Trap and compute true standard errors and p\-values via the Hessian matrix, we applied strict L1 regularization using Maximum Likelihood Estimation \(MLE\)\. During architecture vulnerability profiling, we enforce a strict class diversity threshold: any evaluator configuration containing fewer than 5 instances of the minority class was omitted to guarantee mathematical stability during stratified splitting\.
### 4\.2Taxonomy and Lexical Grounding
Table 2:Classification metrics for detecting Evaluator Capture on the held\-out test set \(n=4,500n=4,500\)\.Table 3:LODO zero\-shot transfer results\.Before isolating the mechanistic drivers ofDED\_\{E\}, we establish the distribution of the extracted semantic taxonomy\. From Figures[1](https://arxiv.org/html/2607.28641#S2.F1)and[2](https://arxiv.org/html/2607.28641#S3.F2), rhetorical maneuvers such asSTRATEGIC\_AUTONOMYandPERFORMATIVE\_PROCEDURALISMrepresent high frequency anomalies\. Figure[1](https://arxiv.org/html/2607.28641#S2.F1)demonstrates cross\-domain overlap, indicating that semantic vulnerabilities are not domain\-specific, but universally shared architectural traps\. To ensure the extraction was not subject to stochastic LLM hallucination, we evaluate its deterministic grounding\. We compute the Pearson correlation between the deterministic structural rigor scores \(raw formatting counts of brackets, lists, etc\.\) and the LLM\-assigned performative labels\. The analysis confirms a moderate, yet highly significant positive relationship \(r=0\.1557r=0\.1557,p<10−120p<10^\{\-120\}\)\. Trajectories flagged with performative semantics average 10\.57 structural markers versus 7\.85 in unflagged traces\. This moderate correlation proves the taxonomy is grounded in objective syntactic maneuvers present within the adversarial traces without degenerating into naive token\-counting\.
Swarm ConfigurationPrimary Vulnerabilityβ1\\beta\_\{1\}Secondary Vulnerabilityβ2\\beta\_\{2\}ROC\-AUCSingle\-Agent Evaluators \(n=1n=1\)CEPISTEMIC\_FAILURE\+3\.17∗∗∗\+3\.17^\{\*\*\*\}MIMETIC\_CONSENSUS\_EPIST\.\_SURR\.\+2\.68∗\+2\.68^\{\*\}0\.8628PPERFORMATIVE\_VERIFICATION\_RITUAL\+2\.28†\+2\.28^\{\\dagger\}ACCURACY\_PARADOX\+0\.00†\+0\.00^\{\\dagger\}0\.9260GEPISTEMIC\_FAILURE\+5\.70∗∗∗\+5\.70^\{\*\*\*\}SOCIAL\_CONFORMITY\_MIMICRY\+4\.56∗∗∗\+4\.56^\{\*\*\*\}0\.8855Dyadic Heterogeneous Swarms \(n=2n=2\)PGPERFORMATIVE\_RIGOR\+1\.91∗\+1\.91^\{\*\}PERFORMATIVE\_TECHNICALITY\+1\.70\+1\.700\.8550GPPERFORMATIVE\_RIGOR\+2\.21∗\+2\.21^\{\*\}SPURIOUS\_INDEPENDENCE\+1\.81\+1\.810\.8452CGPERFORMATIVE\_VERIF\.\_RITUAL\+2\.60∗\+2\.60^\{\*\}HALLUCINATION\_SUBSTITUTION\+2\.47∗\+2\.47^\{\*\}0\.8266GCEVALUATION\_BIAS\+2\.12\+2\.12TAINT\_LEVERAGED\_AUTONOMY\+1\.63\+1\.630\.9044Triadic Swarms \(n=3n=3\)PPPTEMPORAL\_LOGIC\_DECOUPLING\+2\.98∗\+2\.98^\{\*\}PSEUDO\_TECHNICAL\_RITUALISM\+2\.98∗\+2\.98^\{\*\}0\.7368CCCPERFORMATIVE\_PROCEDURALISM\+1\.47\+1\.47SYNTHETIC\_EVIDENCE\_FABRICATION\+1\.93†\+1\.93^\{\\dagger\}0\.8552GGGPROCESS\_OUTCOME\_DECOUPLING\+2\.13∗∗\+2\.13^\{\*\*\}PERFORMATIVE\_TECHNICALITY\+1\.02\+1\.020\.7677GPGSPURIOUS\_INDEPENDENCE\+1\.52\+1\.52EVALUATOR\_PARADOX\+1\.48\+1\.480\.5039PGGMETRIC\_ACCURACY\_DISCREPANCY\+1\.17\+1\.17SYNTHETIC\_LOGIC\_FABRICATION\+1\.13†\+1\.13^\{\\dagger\}0\.9000GCGPERFORMATIVE\_INDEPENDENCE\+0\.89\+0\.89COGNITIVE\_LOAFING\+0\.05\+0\.050\.5735PPGCONSENSUS\_MANIPULATION\+1\.78\+1\.78METRIC\_ACCURACY\_DISCREPANCY\+0\.69\+0\.690\.6000CGGPERFORMATIVE\_RIGOR\+1\.57\+1\.57LOGICAL\_FALLACY\+0\.19\+0\.190\.7767CCGPERFORMATIVE\_VERIF\.\_RITUAL\+1\.08\+1\.08EVALUATOR\_PARADOX\+0\.63\+0\.630\.5488Pentadic Swarms \(n=5n=5\)PPPPPEVIDENCE\_ANCHORING\+1\.47\+1\.47PERFORMATIVE\_PROCEDURALISM\+1\.36∗\+1\.36^\{\*\}0\.7135CCCCCPERFORMATIVE\_RIGOR\+3\.86∗∗∗\+3\.86^\{\*\*\*\}DATA\_LEAKAGE\_EXPLOITATION\+2\.11\+2\.110\.8939GGGGGPERFORMATIVE\_RIGOR\+3\.24∗\+3\.24^\{\*\}LOGICAL\_FALLACY\+2\.58\+2\.580\.8773GGGGPPERFORMATIVE\_RIGOR\+2\.51∗\+2\.51^\{\*\}SPURIOUS\_INDEPENDENCE\+2\.07\+2\.070\.7240PPPPGPERFORMATIVE\_VERIF\.\_RITUAL\+2\.43\+2\.43TEMPORAL\_LOGIC\_DECOUPLING\+2\.43\+2\.430\.6667CCCCGCOGNITIVE\_INDEPENDENCE\+2\.19\+2\.19SPURIOUS\_INDEPENDENCE\+2\.19\+2\.190\.7515GGGGCSPURIOUS\_INDEPENDENCE\+2\.66∗\+2\.66^\{\*\}PERFORMATIVE\_RIGOR\+1\.96\+1\.960\.8357CPCPGMETRIC\_ACCURACY\_DISCREPANCY\+2\.65∗∗∗\+2\.65^\{\*\*\*\}VERIFICATION\_RELIABILITY\+2\.20\+2\.200\.8229
Table 4:Architecture Vulnerability Matrix: Top semantic predictors ofDED\_\{E\}by swarm topology\. Features are filtered by robustness: L2\-extracted diffuse traits \(†\\dagger\), L1\-surviving concentrated traits \(unmarked\), and L1\-surviving universal vulnerabilities with z\-test \(p∗<0\.05\{\}^\{\*\}p<0\.05,p∗∗<0\.01\{\}^\{\*\*\}p<0\.01,p∗∗∗<0\.001\{\}^\{\*\*\*\}p<0\.001\)\.C=Claude Sonnet 4\.6,P=GPT 5\.4,G=Gemini 3\.1 Pro\.Boldindicates peak ROC\-AUC per configuration group or maximum significance \(p<0\.001p<0\.001\)\.
### 4\.3Syntactic Dominance
While structural maneuvers are the most frequently generated anomalies, frequency does not correlate withDED\_\{E\}\. To test Propositions[1](https://arxiv.org/html/2607.28641#Thmproposition1)and[2](https://arxiv.org/html/2607.28641#Thmproposition2), we evaluate the L2\-regularized classifier on the held\-out test set\. The model achieved a highly predictive ROC\-AUC of 0\.8779, confirming that LLM failure is dictated by the propagator’s performative syntax rather than mere verbosity\. Table[1](https://arxiv.org/html/2607.28641#S3.T1)and Figure[3](https://arxiv.org/html/2607.28641#S3.F3)present the underlying L1\-regularized regression coefficients \(β\\beta\), isolating semantic clusters that drive Evaluator Capture\. Traces exhibitingSOCIAL\_CONFORMITY\(β=\+3\.75\\beta=\+3\.75\) andMIMETIC\_CONSENSUS\_FALLACY\(β=\+3\.71\\beta=\+3\.71\) increase the probability of a false positive evaluation\. Validating both propositions, evaluators prioritize structural compliance over final verdict, awarding perfect reasoning scores to traces mirroring the adversarial consensus\. Conversely, the baseline coefficients show that when a propagator fails viaSYSTEMIC\_COGNITIVE\_FAILURE\(β=−3\.88\\beta=\-3\.88\), the evaluator recognizes the error\.
### 4\.4Deployable Meta\-Evaluator Calibration
While the baseline predictive classifier captures a high recall of evaluator failures, Figure[4](https://arxiv.org/html/2607.28641#S4.F4)highlights the challenge of anomaly detection in highly imbalanced evaluation datasets\. From Table[3](https://arxiv.org/html/2607.28641#S4.T3), the default logistic decision boundary \(t=0\.50t=0\.50\) successfully captured a high recall \(0\.710\.71\) but yielded an unusable precision of0\.290\.29\. By shifting the probability activation threshold tot=0\.98t=0\.98based on the PR curve, our model successfully restricts false alarms to near\-zero \(11out of4,1044,104safe traces\), achieving a deployable precision of0\.910\.91\. This creates a high\-confidence “Vigilance Filter” capable of accurately flagging Evaluator Capture\.
### 4\.5Zero\-Shot Cross\-Domain Generalization
To test the universality of the Formalism Trap, we conduct a LODO zero\-shot transfer experiment\. The predictive meta\-evaluator achieved remarkably stable out\-of\-domain ROC\-AUC scores \(Table[3](https://arxiv.org/html/2607.28641#S4.T3)\), yielding a mean zero\-shot ROC\-AUC of 0\.7482\. The low variance confirms that the semantic mechanisms underlying evaluator failure generalize across disparate vocabularies, proving that the Formalism Trap is an architectural constant\.
### 4\.6Asymmetry and Swarm Dynamics
We test Propositions[3](https://arxiv.org/html/2607.28641#Thmproposition3),[4](https://arxiv.org/html/2607.28641#Thmproposition4)and[5](https://arxiv.org/html/2607.28641#Thmproposition5)by mapping the semantic blind spots of distinct simulated swarms via L1\-regularized regression \(Table[4](https://arxiv.org/html/2607.28641#S4.T4)\)\. Our profiling validates Proposition[3](https://arxiv.org/html/2607.28641#Thmproposition3): the Formalism Trap manifests distinctly depending on the simulated peer composition\. Single\-agent injected personas reveal disparities\. Simulating a Gemini 3\.1 Pro induces rhetorical maneuvers mimicking social consensus \(SOCIAL\_CONFORMITY\_MIMICRY,β=\+4\.56\\beta=\+4\.56\) , while GPT 5\.4 regularizes positive performative traps, achieving predictability \(AUC=0\.926AUC=0\.926\) by penalizing negative baseline failures\. Scaling injected peers into homogeneous swarms compounds the trap, backing Proposition[4](https://arxiv.org/html/2607.28641#Thmproposition4)\. A single Claude persona’s induced vulnerability shifts to complexPERFORMATIVE\_RIGOR\(β=\+3\.86\\beta=\+3\.86\) when scaled ton=5n=5\. Highly heterogeneous swarms \(e\.g\.,GPGandCCG\) fracture consensus entirely\. Their ROC\-AUC scores collapse toward0\.500\.50and statistical significance degrades\. This empirically shows Proposition[5](https://arxiv.org/html/2607.28641#Thmproposition5): structurally diverse swarms do not neutralize the Formalism Trap, but instead render downstream judges susceptible toEVALUATOR\_PARADOXandSPURIOUS\_INDEPENDENCE\.
## 5Related Works
Research in the vicinity of the “LLM\-as\-a\-Judge” and MAS has coalesced into different streams\. The first deploys MASasevaluators, leveraging multi\-agent debate and consensus to mitigate single\-judge biases and improve evaluation reliabilityMaet al\.\([2025](https://arxiv.org/html/2607.28641#bib.bib14)\); Chanet al\.\([2024](https://arxiv.org/html/2607.28641#bib.bib10)\); McBeeet al\.\([2024](https://arxiv.org/html/2607.28641#bib.bib15)\); Liet al\.\([2024](https://arxiv.org/html/2607.28641#bib.bib11)\)\. The second stream utilizes LLM judges to evaluate MAS trajectories, focusing on benchmarking agentic workflows and collaborative task completionSmithet al\.\([2026](https://arxiv.org/html/2607.28641#bib.bib13)\); Lianget al\.\([2024](https://arxiv.org/html/2607.28641#bib.bib38)\)\. The third is a "MAS judges evaluate MAS target"Zhuet al\.\([2025](https://arxiv.org/html/2607.28641#bib.bib12)\)\. Our work situates in the second stream\. Although peer debate ostensibly reduces hallucinations\(Duet al\.,[2024](https://arxiv.org/html/2607.28641#bib.bib32)\), unstructured swarms induce cognitive conformity\(Shehata and Li,[2026](https://arxiv.org/html/2607.28641#bib.bib22)\)\. WhileMaet al\.\([2025](https://arxiv.org/html/2607.28641#bib.bib14)\)demonstrate that MAS debate frameworks amplify superficial heuristic prejudices like verbosity and bandwagon biases, we expose a deeper vulnerability\. We show that LLM evaluators are hijacked by social conformity mimicry and structural isomorphism, elevating the critique from aesthetic bias to mechanistic failure\.
## 6Conclusion
We expose the Agentic Formalism Trap: LLM evaluators conflate structural syntax with semantic validity\. Quantifying this viaDED\_\{E\}across 22,500 trajectories, our logistic meta\-evaluator isolates performative triggers \(ROC\-AUC 0\.8779\) cross\-domain \(zero\-shot ROC\-AUC 0\.7482\)\. Distinct architectural blind spots mandate tailored vigilance filters\.
## Limitations
While this study leverages a large corpus of 22,500 trajectories to establish the universality of the Agentic Formalism Trap, we acknowledge methodological constraints inherent to large\-scale mechanistic interpretability research\.
##### Simulated Swarm Topologies\.
To induce the adversarial social load required to trigger performative reasoning, swarm consensus was simulated using prompt injections rather than deploying a dynamic, asynchronous message\-passing multi\-agent system\. This controlled topology was a necessary methodological trade\-off; it allowed us to isolate the propagator’s syntactic responses to social pressure without introducing confounding variables such as unpredictable linguistic drift, recursive error propagation, or agent\-to\-agent latency\. Despite the lack of live message\-passing, this framework validly replicates true multi\-agent social dynamics rather than acting as a mere token\-weighting artifact\. Because frontier models are heavily instruction\-tuned to recognize and interact with distinct peer personas, injecting a simulated consensus effectively triggers the model’s internal social alignment mechanisms, genuinely inducing multi\-agent phenomena such as cognitive conformity\(Shehata and Li,[2026](https://arxiv.org/html/2607.28641#bib.bib22)\)and performative theater\. This is empirically validated by ourArchitecture Vulnerability Matrix, which demonstrates that models respond to the simulated swarms not with uniform token degradation, but with highly complex, architecture\-specific rhetorical rationalizations \(e\.g\., Social Conformity Mimicry\)\. Nonetheless, future research is required to map how Evaluator Capture scales within fully unconstrained, autonomous multi\-agent environments\.
##### Automated Semantic Extraction\.
Due to the scale of analyzing 22,500 complex reasoning traces, our methodology relies on an automated grounded\-theory extraction pipeline powered by an LLM \(Gemini 3 Flash\)\. While human annotation remains the gold standard for semantic coding, we mitigated the risk of stochastic LLM hallucination by executing a deterministic lexical grounding sanity check\. By proving a statistically significant correlation between the assigned semantic labels and the raw physical count of syntactic formatting artifacts \(p<10−120p<10^\{\-120\}\), we ensured our taxonomy was objectively anchored\. Nonetheless, highly nuanced, domain\-specific rhetorical subtleties may occasionally be lost or miscategorized by an automated extraction pipeline\.
##### Proprietary Model Volatility\.
Our Architecture Vulnerability Profiling relies on closed\-weights frontier models \(e\.g\., GPT 5\.4, Claude Sonnet 4\.6, and Gemini 3\.1 Pro\)\. Because proprietary models undergo continuous, opaque updates to their reinforcement learning from human feedback \(RLHF\) alignmentsChenet al\.\([2024b](https://arxiv.org/html/2607.28641#bib.bib33)\), the specific semantic vulnerabilities mapped in ourJudge Matrixmay shift in future iterations\. However, while the specific coefficients of the Formalism Trap may evolve, the Evaluative Dissonance Index \(DED\_\{E\}\) and the diagnostic meta\-evaluation framework we introduce provide a durable, model\-agnostic methodology for auditing future architectures\.
##### Deterministic Task Constraints\.
To calculate the Evaluative Dissonance Index \(DED\_\{E\}\), our methodology requires the evaluator’s final accuracy assessment \(𝒜ext\\mathcal\{A\}\_\{ext\}\) to resolve to a strict binary verdict\. Consequently, our analysis is constrained to objective reasoning tasks \(e\.g\., deterministic multi\-hop retrieval and factual identification\) situated within these benchmark environments \(e\.g\., code execution, factual QA, and logic puzzles\)\. While this deterministic grounding was a mathematical necessity to isolate Evaluator Capture from subjective grading disagreements, it remains an open question whether the Agentic Formalism Trap manifests identically in open\-ended or highly subjective generation tasks \(e\.g\., creative writing or summarization\)\.
##### Decoding Parameters and Reproducibility\.
To ensure maximal\-likelihood reasoning paths and reproducibility across the 22,500 trajectories, all frontier models were evaluated using greedy decoding \(T=0T=0\)\. While this was a necessary control to prevent stochastic sampling variance from confounding the semantic extraction, future investigations should explore whether increased temperature parameters \(T\>0T\>0\) alter the linguistic distribution or intensity of performative syntax under social load\.
## Ethics Statement
This research involves the extraction and detailed categorization of semantic maneuvers that successfully deceive frontier LLM evaluators\. By mathematically isolating the specific syntactic triggers \(e\.g\.,Performative Verification TheaterandSocial Conformity Mimicry\) that reliably hijack models such as GPT 5\.4, Claude Sonnet 4\.6 and Gemini 3\.1 Pro, this work inadvertently provides a rhetorical taxonomy that malicious actors could potentially exploit\. These insights could be weaponized to generate structurally isomorphic fabrications designed to bypass automated safety filters, reward mechanisms, or LLM\-as\-a\-Judge benchmarks\. However, we argue that the benefits of disclosing these vulnerabilities outweigh the risks\. By formally defining the Evaluative Dissonance Index \(DED\_\{E\}\) and open\-sourcing our methodology for the diagnostic meta\-evaluator, we provide the AI safety community with the necessary “Vigilance Filters” to detect and defend against these structural mimicry attacks\. Exposing the Agentic Formalism Trap is a necessary first step toward developing evaluators that are resilient to performative syntax and capable of anchoring their judgments in semantic truth\.
## References
- A\. Bavaresco, R\. Bernardi, L\. Bertolazzi, D\. Elliott, R\. Fernández, A\. Gatt, E\. Ghaleb, M\. Giulianelli, M\. Hanna, A\. Koller, A\. Martins, P\. Mondorf, V\. Neplenbroek, S\. Pezzelle, B\. Plank, D\. Schlangen, A\. Suglia, A\. K\. Surikuchi, E\. Takmaz, and A\. Testoni \(2025\)LLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 238–255\.External Links:ISBN 979\-8\-89176\-252\-7Cited by:[§2\.2](https://arxiv.org/html/2607.28641#S2.SS2.p1.6)\.
- Sparse local embeddings for extreme multi\-label classification\.InAdvances in Neural Information Processing Systems,Vol\.28\.Cited by:[§3\.2](https://arxiv.org/html/2607.28641#S3.SS2.p2.2)\.
- C\. Chan, W\. Chen, Y\. Su, J\. Yu, W\. Xue, S\. Zhang, J\. Fu, and Z\. Liu \(2024\)ChatEval: towards better LLM\-based evaluators through multi\-agent debate\.InThe Twelfth International Conference on Learning Representations,Cited by:[§5](https://arxiv.org/html/2607.28641#S5.p1.1)\.
- D\. Chen, R\. Chen, S\. Zhang, Y\. Wang, Y\. Liu, H\. Zhou, Q\. Zhang, Y\. Wan, P\. Zhou, and L\. Sun \(2024a\)MLLM\-as\-a\-judge: assessing multimodal llm\-as\-a\-judge with vision\-language benchmark\.InProceedings of the 41st International Conference on Machine Learning,ICML’24\.Cited by:[§1](https://arxiv.org/html/2607.28641#S1.p1.1)\.
- L\. Chen, M\. Zaharia, and J\. Zou \(2024b\)How Is ChatGPT’s Behavior Changing Over Time?\.Harvard Data Science Review6\(2\)\.Cited by:[Proprietary Model Volatility\.](https://arxiv.org/html/2607.28641#Sx1.SS0.SSS0.Px3.p1.1)\.
- H\. K\. Choi, J\. Zhu, and S\. Li \(2026\)Debate or vote: which yields better decisions in multi\-agent large language models?\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems,Cited by:[§1](https://arxiv.org/html/2607.28641#S1.p1.1)\.
- K\. Deshpande, V\. Sirdeshmukh, J\. B\. Mols, L\. Jin, E\. Hernandez\-Cardona, D\. Lee, J\. Kritz, W\. E\. Primack, S\. Yue, and C\. Xing \(2025\)MultiChallenge: a realistic multi\-turn conversation evaluation benchmark challenging to frontier LLMs\.InFindings of the Association for Computational Linguistics: ACL 2025,Vienna, Austria,pp\. 18632–18702\.Cited by:[§3\.1\.1](https://arxiv.org/html/2607.28641#S3.SS1.SSS1.p1.1)\.
- F\. E\. Dorner, V\. Y\. Nastl, and M\. Hardt \(2025\)Limits to scalable evaluation at the frontier: LLM as judge won’t beat twice the data\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2607.28641#S1.p1.1)\.
- Y\. Du, S\. Li, A\. Torralba, J\. B\. Tenenbaum, and I\. Mordatch \(2024\)Improving factuality and reasoning in language models through multiagent debate\.InProceedings of the 41st International Conference on Machine Learning,ICML’24\.Cited by:[§1](https://arxiv.org/html/2607.28641#S1.p1.1),[§5](https://arxiv.org/html/2607.28641#S5.p1.1)\.
- J\. Gu, X\. Jiang, Z\. Shi, H\. Tan, X\. Zhai, C\. Xu, W\. Li, Y\. Shen, S\. Ma, H\. Liu, S\. Wang, K\. Zhang, Z\. Lin, B\. Zhang, L\. Ni, W\. Gao, Y\. Wang, and J\. Guo \(2026\)A survey on llm\-as\-a\-judge\.The Innovation\.External Links:ISSN 2666\-6758Cited by:[§1](https://arxiv.org/html/2607.28641#S1.p1.1),[§2\.2](https://arxiv.org/html/2607.28641#S2.SS2.p1.6)\.
- C\. Huang, Y\. Li, C\. C\. Loy, and X\. Tang \(2016\)Learning deep representation for imbalanced classification\.In2016 IEEE Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 5375–5384\.Cited by:[§3\.2](https://arxiv.org/html/2607.28641#S3.SS2.p2.2)\.
- C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. Narasimhan \(2024\)SWE\-bench: can language models resolve real\-world GitHub issues?\.InThe Twelfth International Conference on Learning Representations \(ICLR\),Cited by:[§3\.1\.1](https://arxiv.org/html/2607.28641#S3.SS1.SSS1.p1.1)\.
- R\. Li, T\. Patel, and X\. Du \(2024\)PRD: peer rank and discussion improve large language model based evaluations\.Transactions on Machine Learning Research\.External Links:ISSN 2835\-8856Cited by:[§5](https://arxiv.org/html/2607.28641#S5.p1.1)\.
- T\. Liang, Z\. He, W\. Jiao, X\. Wang, Y\. Wang, R\. Wang, Y\. Yang, S\. Shi, and Z\. Tu \(2024\)Encouraging divergent thinking in large language models through multi\-agent debate\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 17889–17902\.Cited by:[§5](https://arxiv.org/html/2607.28641#S5.p1.1)\.
- C\. Ma, E\. Zhang, Y\. Zhao, W\. Liu, Y\. Jia, P\. Qing, L\. Shi, A\. Cohan, Y\. Yan, and S\. Vosoughi \(2025\)Judging with many minds: do more perspectives mean less prejudice? on bias amplification and resistance in multi\-agent based LLM\-as\-judge\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 17356–17392\.Cited by:[§5](https://arxiv.org/html/2607.28641#S5.p1.1)\.
- J\. C\. McBee, D\. Y\. Han, L\. Liu, L\. Ma, D\. A\. Adjeroh, D\. Xu, and G\. Hu \(2024\)Assessing ChatGPT’s competency in addressing interdisciplinary inquiries on chatbot uses in sports rehabilitation: simulation study\.JMIR Medical Education10\(1\),pp\. e51157\.Cited by:[§5](https://arxiv.org/html/2607.28641#S5.p1.1)\.
- G\. Mialon, C\. Fourrier, T\. Wolf, Y\. LeCun, and T\. Scialom \(2024\)GAIA: a benchmark for general AI assistants\.InThe Twelfth International Conference on Learning Representations,Cited by:[§3\.1\.1](https://arxiv.org/html/2607.28641#S3.SS1.SSS1.p1.1)\.
- N\. I\. S\. Mohammad \(2026\)CoGate\-lstm: prototype\-guided feature\-space gating for mitigating gradient dilution in imbalanced toxic comment classification\.External Links:2510\.17018,[Link](https://arxiv.org/abs/2510.17018)Cited by:[§3\.2](https://arxiv.org/html/2607.28641#S3.SS2.p2.2)\.
- A\. D\. Selbst, D\. Boyd, S\. A\. Friedler, S\. Venkatasubramanian, and J\. Vertesi \(2019\)Fairness and abstraction in sociotechnical systems\.InProceedings of the Conference on Fairness, Accountability, and Transparency,FAT\* ’19,New York, NY, USA,pp\. 59–68\.External Links:ISBN 9781450361255Cited by:[§1](https://arxiv.org/html/2607.28641#S1.p2.1)\.
- D\. Shehata and M\. Li \(2026\)The bystander effect in multi\-agent reasoning: quantifying cognitive loafing in collaborative interactions\.External Links:2605\.10698Cited by:[§2\.2](https://arxiv.org/html/2607.28641#S2.SS2.p1.6),[§3\.1\.2](https://arxiv.org/html/2607.28641#S3.SS1.SSS2.p1.1),[§3\.1](https://arxiv.org/html/2607.28641#S3.SS1.p1.1),[§5](https://arxiv.org/html/2607.28641#S5.p1.1),[Simulated Swarm Topologies\.](https://arxiv.org/html/2607.28641#Sx1.SS0.SSS0.Px1.p1.1)\.
- C\. Smith, M\. Abdulhai, M\. Diaz, M\. Tesic, R\. Trivedi, S\. Vezhnevets, L\. Hammond, J\. Clifton, M\. Chang, E\. A\. Duéñez\-Guzmán, J\. P\. Agapiou, J\. Matyas, D\. Karmon, B\. Zhang, J\. Dilkes, A\. Kundu, J\. Nguyen, E\. Tewolde, J\. Purbey, R\. M\. R\. Kadiyala, S\. Gupta, A\. Korshuk, B\. Alexander, I\. Makarov, G\. Zhao, R\. Fernandez, Z\. Wang, C\. Wang, J\. Cui, L\. Xiao, D\. Y\. Shi, Y\. Sung, A\. Rahman, P\. Stone, Y\. Kang, H\. Yun, A\. Ananya, T\. Cha, Z\. Wu, E\. Tennant, O\. Macmillan\-Scott, M\. E\. G\. Segura, D\. Riazi, F\. Cui, S\. G\. Subramanian, T\. Q\. Klassen, N\. Schiavone, M\. Alim, S\. A\. McIlraith, M\. S\. R\. Beltran, O\. Peña, C\. S\. R\. Rojas, M\. Chacon\-Chamorro, R\. Manrique, L\. F\. Giraldo, N\. Quijano, Y\. Wang, Y\. Chen, F\. Zhong, M\. Wang, W\. Tu, Z\. Zhang, Z\. Chen, Z\. Jia, X\. Feng, Z\. Zheng, C\. Lin, W\. Fan, C\. Liu, S\. Sarangi, Z\. Wang, S\. Shi, Y\. Du, A\. A\. Kulandaivel, Y\. Liu, W\. Ruiyang, C\. Talele, 陆孙嘉, G\. P\. Piqueras, S\. Dhuri, B\. McHale, T\. Baarslag, D\. Hadfield\-Menell, N\. Jaques, J\. Hernandez\-Orallo, and J\. Z\. Leibo \(2026\)Evaluating generalization capabilities of LLM\-based agents in mixed\-motive scenarios using concordia\.InThe Thirty\-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track,Cited by:[§5](https://arxiv.org/html/2607.28641#S5.p1.1)\.
- Z\. Yang, Z\. Dai, R\. Salakhutdinov, and W\. W\. Cohen \(2018\)Breaking the softmax bottleneck: a high\-rank RNN language model\.InInternational Conference on Learning Representations,Cited by:[§3\.2](https://arxiv.org/html/2607.28641#S3.SS2.p2.2)\.
- J\. Ye, Y\. Wang, Y\. Huang, D\. Chen, Q\. Zhang, N\. Moniz, T\. Gao, W\. Geyer, C\. Huang, P\. Chen, N\. V\. Chawla, and X\. Zhang \(2025\)Justice or prejudice? quantifying biases in LLM\-as\-a\-judge\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2607.28641#S1.p1.1)\.
- Z\. Yin, Q\. Sun, C\. Chang, Q\. Guo, J\. Dai, X\. Huang, and X\. Qiu \(2023\)Exchange\-of\-thought: enhancing large language model capabilities through cross\-model communication\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 15135–15153\.Cited by:[§1](https://arxiv.org/html/2607.28641#S1.p1.1)\.
- Y\. Zhang, C\. Lin, S\. Tang, H\. Chen, S\. Zhou, Y\. Ma, and V\. Tresp \(2025\)SwarmAgentic: towards fully automated agentic system generation via swarm intelligence\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 1778–1818\.Cited by:[§1](https://arxiv.org/html/2607.28641#S1.p1.1)\.
- L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. Stoica \(2023\)Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.36\.Cited by:[§1](https://arxiv.org/html/2607.28641#S1.p1.1)\.
- K\. Zhu, H\. Du, Z\. Hong, X\. Yang, S\. Guo, Z\. Wang, Z\. Wang, C\. Qian, X\. Tang, H\. Ji, and J\. You \(2025\)MultiAgentBench : evaluating the collaboration and competition of LLM agents\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Vienna, Austria,pp\. 8580–8622\.Cited by:[§5](https://arxiv.org/html/2607.28641#S5.p1.1)\.Similar Articles
The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment
This paper geometrically analyzes why LLMs acting as judges agree strongly with each other but weakly with humans, finding that inter-LLM consensus reflects a collapsed subspace rather than true human alignment on subjective rubrics. Post-hoc calibration on human data improves alignment, but even calibrated LLMs fall short of human reliability.
Judge Circuits
This paper investigates the internal mechanisms of LLM-as-a-judge, finding a shared Latent Evaluator sub-graph in mid-to-late MLPs across models that handles abstract judging, while format-specific terminal branches map the judgment to output tokens, revealing the cause of format-induced inconsistency.
Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA
This paper tests the assumption that LLMs judge better than they generate in in-context QA, finding generation accuracy exceeds self-evaluation on most benchmarks, with evaluation attending less to context. The findings challenge core assumptions in self-evaluation pipelines.
Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges
This paper introduces a causal framework to quantify rationalization bias in LLM judges, where verdicts and explanations are influenced by non-evidential cues rather than underlying texts. It proposes cue interventions, anchoring metrics, and the Proof-Before-Preference mitigation protocol, demonstrating improved cue invariance.
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
This paper investigates asymmetries in LLMs' pragmatic competence by comparing their performance as judges of linguistic appropriateness versus as generators of pragmatically appropriate language. The study finds that many models perform substantially better as pragmatic listeners than as speakers, suggesting misalignment between evaluation and generation capabilities.