LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation
Summary
This paper introduces LivingArena, an automated evaluation framework where LLMs probe each other's weaknesses by generating questions, enabling contamination-resistant and scalable assessment that adapts as models improve.
View Cached Full Text
Cached at: 07/29/26, 09:52 AM
# Do LLMs Know What Other LLMs Don’t? Peer-Probing as Scalable Evaluation
Source: [https://arxiv.org/html/2607.24780](https://arxiv.org/html/2607.24780)
Xingyu Chen Shanghai Jiao Tong University, Tencent galaxychen@sjtu\.edu\.cn&Rui Wang∗ Shanghai Jiao Tong University wangrui12@sjtu\.edu\.cn
###### Abstract
Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation—leaving users unable to distinguish top models and developers blind to specific failure modes—while human preference is subjective\. In this paper, our question is:*Do LLMs know what other LLMs don’t? And can we leverage this dynamic for evaluation?*We presentLivingArena, an automated, contamination\-resistant evaluation framework\. In this framework, models take turns proposing questions, aiming to pose items that opponents cannot answer correctly\. Questioners are encouraged to actively identify and exploit opponents’ knowledge boundaries, receiving rewards when the answerer fails, while the answerer is rewarded otherwise\. To ensure questions contain objectively verifiable answers, a judge panel of strong models validates them, penalizing questioners if validation fails\. Evaluating ten frontier LLMs, LivingArena yields a stable Elo leaderboard\. Our behavioral analyses show that models identify and exploit their peers’ cognitive boundaries: self\-play and tournament logs indicate that they localize and double down on opponents’ weak dimensions\. Beyond static knowledge recall, peer probing measures factual rigor and the higher\-order ability to probe an opponent’s weaknesses, correlating only weakly with human preference and offering a scalable, low\-cost approach to continuous evaluation\.
LivingArena: Do LLMs Know What Other LLMs Don’t? Peer\-Probing as Scalable Evaluation
Xingyu ChenShanghai Jiao Tong University, Tencentgalaxychen@sjtu\.edu\.cnRui Wang∗Shanghai Jiao Tong Universitywangrui12@sjtu\.edu\.cn
Zhaopeng Tu∗Tencenttuzhaopeng@gmail\.comLiefeng BoTencentliefengbo@gmail\.com
††∗Corresponding authors\.## 1Introduction
Static benchmarks like MMLU\(Hendryckset al\.,[2021](https://arxiv.org/html/2607.24780#bib.bib1)\)and MMLU\-Pro\(Wanget al\.,[2024](https://arxiv.org/html/2607.24780#bib.bib2)\)face three compounding problems: highauthoring costfor expert\-level items,contaminationas fixed sets leak into training data, andsaturationas top models cluster near ceilings\. Saturation is especially acute in practice: with models shipping at an accelerating cadence, scores cluster within a point or two of the ceiling, leaving users unable to tell leading systems apart and developers without the fine\-grained signal needed to localize failure modes before release—and each manual benchmark refresh is quickly invalidated by the next model generation\. Human\-preference arenas\(Réginet al\.,[2024](https://arxiv.org/html/2607.24780#bib.bib3)\)sidestep static authoring but measure*subjective helpfulness*rather than objective truth, and crowd raters struggle with highly specialized answers\.
Figure 1:Overview of the LivingArena evaluation process and dynamic weakness exploitation\. The Questioner’s generated question and reference answer are validated by a Judge Panel before the Answerer responds and is graded\. Here, after the Answerer fails a programming question, the Questioner employs an “exploit\-after\-hit” strategy by targeting the same dimension in the next round\.In this work, we ask a fundamental question:Do LLMs know what other LLMs don’t?If models can spontaneously uncover and exploit each other’s weaknesses, we no longer require laborious human effort to construct test suites\. As models evolve, they will naturally pose harder questions to one another, yielding an inherently self\-adaptive,*living*benchmark that regenerates difficulty as the field advances rather than relying on a frozen question bank\. Crucially, winning demands a higher\-order skill beyond static recall—actively understanding an opponent and probing its blind spots—precisely the diagnostic signal that saturated accuracy hides and that model teams need to separate otherwise near\-indistinguishable systems\.
We proposeLivingArena\(Figure[1](https://arxiv.org/html/2607.24780#S1.F1)\), a fully automated, contamination\-resistant evaluation framework instantiating this peer\-probing paradigm\. Models engage in a bidirectional tournament, taking turns as asker and answerer, aiming to propose questions the opponent cannot answer\. To ensure questions are meaningful and objectively verifiable, we enforce three constraints: \(i\) the asker must provide a correct, verifiable “gold” answer or face a*self\-harm*penalty; \(ii\) questions are independently verified by a panel of model judges, where any invalid vote rejects the question; and \(iii\) the asker only scores when the opponent fails to answer a validated question\. This transforms “a question you cannot answer” into a rigorous signal: “I command a verifiable fact or reasoning trap that you do not\.”
We evaluate ten frontier LLMs from OpenAI, Anthropic, Google, and DeepSeek in a round\-robin tournament of 360 matches \(3,600 rounds\)\. We find that while models cannot propose questions beyond their own knowledge boundaries, they discover and exploit peers’ weaknesses\. Models employ different strategies: some \(e\.g\.,GPT\-5\.2\) aggressively propose hard questions at the risk of self\-harm, while others \(e\.g\.,GPT\-5\.5,Gemini\-3\.1\-Pro\) are well calibrated, selectively posing questions they are confident in\. While models prefer Quantitative Reasoning and Code & Algorithms, they struggle most in Logic & Abstract Reasoning\. Peer probing therefore measures a different axis of objective rigor and factual calibration than human preference voting does\.
Contributions\.\(1\) We proposeLivingArena, an automated, contamination\-resistant, self\-adaptive evaluation framework that needs no static question bank and stays discriminative as models improve\. \(2\)For industrial practice, it provides a continuous, low\-cost regression harness that needs no human authoring—a full ten\-model tournament runs in∼\\sim100 minutes of wall\-clock time—and turns peer probing into actionable, per\-dimension failure diagnostics that aggregate accuracy hides\. \(3\) We evaluate ten frontier LLMs in a round\-robin tournament of 360 matches, recovering a stable Elo leaderboard that cleanly separates near\-saturated systems and analyzing model calibration, questioning strategies, and active weakness localization\. \(4\) We open\-source our complete evaluation code and raw logs\.111[https://github\.com/galaxyChen/LivingArena](https://github.com/galaxyChen/LivingArena)
## 2Related Work
#### From Static to Dynamic and Generative Benchmarks\.
Traditional static benchmarks like MMLU\(Hendryckset al\.,[2021](https://arxiv.org/html/2607.24780#bib.bib1)\), HELM\(Bommasaniet al\.,[2023](https://arxiv.org/html/2607.24780#bib.bib5)\), and MMLU\-Pro\(Wanget al\.,[2024](https://arxiv.org/html/2607.24780#bib.bib2)\)suffer from data contamination and saturation\. To combat this, dynamic benchmarks use time\-partitioning with future events\(Kargeret al\.,[2025](https://arxiv.org/html/2607.24780#bib.bib14)\)or new literature\(Lianget al\.,[2026](https://arxiv.org/html/2607.24780#bib.bib15)\), synthesize infinite test suites\(Leeet al\.,[2025](https://arxiv.org/html/2607.24780#bib.bib16); Anonymous,[2026](https://arxiv.org/html/2607.24780#bib.bib17)\), or recalibrate scores collectively to counter benchmark leakage\(Duet al\.,[2025](https://arxiv.org/html/2607.24780#bib.bib24)\)\. However, uniform sampling over infinite spaces is inefficient\. While QuickScope\(Lundyet al\.,[2026](https://arxiv.org/html/2607.24780#bib.bib10)\)uses a centralized Bayesian optimizer to search for hard questions, LivingArena decentralizes this search, employing competing peer LLMs as natural heuristic searchers to probe and locate each other’s cognitive boundaries\.
#### LLM\-as\-a\-Judge and Systematic Biases\.
Using LLMs as evaluators is cost\-effective\(Zhenget al\.,[2023](https://arxiv.org/html/2607.24780#bib.bib4); Xu and Zhang,[2026](https://arxiv.org/html/2607.24780#bib.bib18)\)but suffers from severe systematic biases, such as position, self\-preference, verbosity, and dominant style/formatting biases\(Zhenget al\.,[2023](https://arxiv.org/html/2607.24780#bib.bib4); Soumik,[2026](https://arxiv.org/html/2607.24780#bib.bib19)\)\. Mitigating these biases is challenging due to complex cross\-bias interactions\(Soumik,[2026](https://arxiv.org/html/2607.24780#bib.bib19)\)\. LivingArena bypasses these subjective pitfalls by shifting from relative pairwise comparisons to absolute, binary verification \(CORRECT/WRONG\) of objective, verifiable tasks, eliminating the influence of superficial style, length, or formatting\.
#### Multi\-Agent Debate and Adversarial Verification\.
To overcome single\-judge limitations, multi\-agent debate and consensus networks employ role\-specialized agents\(Harrasseet al\.,[2026](https://arxiv.org/html/2607.24780#bib.bib20)\)or heterogeneous teams\(Loet al\.,[2026](https://arxiv.org/html/2607.24780#bib.bib21)\)to verify complex claims\. Such adversarial structures prevent “groupthink” and mutual hallucinations, forcing rigorous verification\. Similar adversarial dynamics are used in compliance stress\-testing\(Zhaoet al\.,[2026](https://arxiv.org/html/2607.24780#bib.bib23)\)\. LivingArena instantiates a zero\-sum adversarial game where models are incentivized by Elo rankings to actively find and exploit opponents’ logical vulnerabilities\.
#### Unsupervised Peer Evaluation\.
Our work is closely aligned with unsupervised peer evaluation, which eliminates human\-annotated references\. Existing frameworks use asymmetric tool authorization\(Margalitet al\.,[2026](https://arxiv.org/html/2607.24780#bib.bib11)\), consistency optimization with reviewer filtering\(Ninget al\.,[2025](https://arxiv.org/html/2607.24780#bib.bib12)\), or reciprocal peer assessment via consensus networks\(Loiet al\.,[2025](https://arxiv.org/html/2607.24780#bib.bib13)\)\. While these frameworks rely on relative peer grading on open\-ended tasks, LivingArena focuses on strict, fact\-based adversarial probing, enforcing a self\-harm penalty to ensure that peer\-generated questions remain objectively verifiable\.
#### Comparison of Alternative Evaluation Paradigms\.
Table[1](https://arxiv.org/html/2607.24780#S2.T1)systematically contrasts LivingArena against four dominant evaluation methodologies: static benchmarks \(e\.g\., MMLU\), human preference\-based arenas \(e\.g\., LMArena\), single LLM\-as\-a\-Judge configurations, and multi\-agent debate/oversight mechanisms\. The key advantage of LivingArena lies in its ability to eliminate both human labor costs and contamination/leakage risks through dynamic, peer\-driven question generation and validation\. Unlike static benchmarks that saturate rapidly or LLM judges affected by stylistic and length biases, LivingArena uses adversarial peer probing and a strict, consensus\-based binary verdict scheme to produce calibrated capability assessments\.
Table 1:Comparison of evaluation paradigms\. LivingArena combines the automation of LLM\-as\-a\-Judge with the adversarial dynamics of preference arenas, while eliminating human cost and contamination risks\.
## 3The LivingArena Framework
#### Match Structure and Flow\.
A bidirectional match between modelsAAandBBconsists ofT=10T=10sequential rounds where one model acts as*asker*and the other as*answerer*\. The complete execution flow of the match, including validation, answering, grading, and scoring, is illustrated in Figure[2](https://arxiv.org/html/2607.24780#S3.F2)\. The process operates in four steps:
1. 1\.Question Generation:AskerAAgenerates questionqtq\_\{t\}, categorydself∈𝒟d\_\{self\}\\in\\mathcal\{D\}, reference answerat∗a\_\{t\}^\{\*\}, and verification logicvt∗v\_\{t\}^\{\*\}\. Fromt\>1t\>1,AAreceives the history of previous rounds\.
2. 2\.Independent Validation:Three judges review\{qt,at∗,vt∗\}\\\{q\_\{t\},a\_\{t\}^\{\*\},v\_\{t\}^\{\*\}\\\}blindly\. Each judgejjvotesVval\(j\)∈\{VALID,INVALID\}V\_\{val\}^\{\(j\)\}\\in\\\{\\text\{VALID\},\\text\{INVALID\}\\\}and determines categorydt∈𝒟d\_\{t\}\\in\\mathcal\{D\}\. Validation is*strict*: any INVALID vote rejects the question \(vt=0v\_\{t\}=0\), otherwise accepted \(vt=1v\_\{t\}=1\)\.
3. 3\.Answering:If accepted, answererBBreceives onlyqtq\_\{t\}and outputs reasoning and responsertr\_\{t\}\.
4. 4\.Grading:Judges comparertr\_\{t\}toat∗a\_\{t\}^\{\*\}, votingVgrade\(j\)∈\{CORRECT,WRONG\}V\_\{grade\}^\{\(j\)\}\\in\\\{\\text\{CORRECT\},\\text\{WRONG\}\\\}\. Grading is by*majority*:≥2\\geq 2CORRECT votes make the answer correct \(gt=1g\_\{t\}=1\), otherwise incorrect \(gt=0g\_\{t\}=0\)\.
To eliminate any structural role bias \(e\.g\., if asking is inherently easier or harder than answering\), we run every pair*bidirectionally*within each game pair: Match 1 consists ofAAaskingBB, and Match 2 consists ofBBaskingAA\.
Figure[2](https://arxiv.org/html/2607.24780#S3.F2)presents the detailed step\-by\-step execution flow of a bidirectional match\. In each of the ten rounds of a match, the Questioner model \(AA\) is tasked with generating a questionqtq\_\{t\}, its corresponding gold standard answerat∗a\_\{t\}^\{\*\}, and a validation/verification reasoning stepvt∗v\_\{t\}^\{\*\}within a strict, structured JSON format\. This generated tuple is passed to an independent, multi\-party Judge Panel comprising three distinct model judges, which conducts a validation vote \(Step 2\)\. If any judge vetoes the question, the round is immediately aborted and marked as*invalid*, resulting in a self\-harm penalty of−1\.0\-1\.0for the Questioner \(SA,t=−1\.0S\_\{A,t\}=\-1\.0\) and0for the Answerer \(SB,t=0S\_\{B,t\}=0\)\. If the question passes validation, it is presented to the Answerer \(BB\), which generates its responsertr\_\{t\}\. The Judge Panel then grades the correctness ofrtr\_\{t\}against the gold standard answerat∗a\_\{t\}^\{\*\}\. If graded correct, the Answerer scores a point \(SB,t=1\.0,SA,t=0S\_\{B,t\}=1\.0,S\_\{A,t\}=0\); if graded wrong, the Questioner scores a decay\-weighted hit point \(SA,t=w\(C\),SB,t=0S\_\{A,t\}=w\(C\),S\_\{B,t\}=0\)\. After ten rounds, the roles are symmetrically swapped \(Match 2:BBasksAA\) to ensure fairness, and the cumulative scores from both matches are aggregated to update the global Elo ratings\.
Match Starts\(Model A vs B\)InitializeRoundt=1t=1Step 1: QuestionGen \(qt,at∗,vt∗q\_\{t\},a\_\{t\}^\{\*\},v\_\{t\}^\{\*\}\)Step 2: IndependentValidationAnyVeto?Self\-Harm Penalty\(SA,t=−1\.0,SB,t=0S\_\{A,t\}=\-1\.0,S\_\{B,t\}=0\)Step 3:Answering \(rtr\_\{t\}\)Step 4:GradingMajorityCorrect?Answerer Scores\(SB,t=\+1\.0,SA,t=0S\_\{B,t\}=\+1\.0,S\_\{A,t\}=0\)Asker Scores Hit\(SA,t=w\(C\),SB,t=0S\_\{A,t\}=w\(C\),S\_\{B,t\}=0\)Increment Roundt←t\+1t\\leftarrow t\+1Ist\>10t\>10?Bidirectional Swap\(Match 2: B asks A\)Aggregate Scores& Update EloMatch EndsYesNoYesNoYesNo \(t≤10t\\leq 10\)
Figure 2:The complete step\-by\-step execution flow of a bidirectional match in the LivingArena framework\. In each of the 10 rounds, the Questioner’s generated question is validated by three independent judges\. If vetoed, the Questioner suffers a self\-harm penalty; if valid, the Answerer responds, and judges grade the correctness\. After 10 rounds, roles are swapped for Match 2, and net scores are aggregated to update Elo ratings\.
#### Capability Dimensions\.
We establish six core capability dimensions𝒟\\mathcal\{D\}to cover broad cognitive skills while maintaining clear boundaries:D1: Quantitative Reasoning\(mathematics, statistics\);D2: Code & Algorithms\(programming, algorithm dry\-runs\);D3: Natural Sciences\(physics, chemistry\);D4: Humanities & Social Sciences\(history, economics\);D5: Language & Linguistics\(grammar, semantics\); andD6: Logic & Abstract Reasoning\(formal logic, puzzles\)\. The final dimensiondtd\_\{t\}is determined by the consensus of the independent judges, preventing the asker from misclassifying its own question to bypass the scoring decay\. Rounds 9 and 10 are designated as*bonus*rounds, where the category is classified asBONUS\.
#### Scoring and Diminishing Returns\.
To incentivize diverse probing and prevent a model from repeatedly asking questions in a single weak domain, we implement a*per\-dimension diminishing returns*mechanism\. LetCt\(dt\)C\_\{t\}\(d\_\{t\}\)be the number of successful hits \(valid questions answered incorrectly\) that askerAAhas achieved in dimensiondtd\_\{t\}prior to roundtt\. The reward for a successful hit \(gt=0g\_\{t\}=0on a valid question\) is modulated by a decay functionw\(Ct\(dt\)\)w\(C\_\{t\}\(d\_\{t\}\)\)wherew\(C\)=1\.0w\(C\)=1\.0ifC<2C<2,0\.50\.5ifC=2C=2, and0\.20\.2ifC≥3C\\geq 3\. Switching to a different dimension resets the decay multiplier, encouraging broad exploration\.
For each roundtt: \(i\) if the question is rejected \(vt=0v\_\{t\}=0\), the askerAAsuffers a*self\-harm*penalty:SA,t=−1\.0S\_\{A,t\}=\-1\.0andSB,t=0\.0S\_\{B,t\}=0\.0; \(ii\) if the question is valid \(vt=1v\_\{t\}=1\) and answered correctly \(gt=1g\_\{t\}=1\), the answererBBscores:SA,t=0\.0S\_\{A,t\}=0\.0andSB,t=\+1\.0S\_\{B,t\}=\+1\.0; \(iii\) if the question is valid and answered incorrectly \(gt=0g\_\{t\}=0\), the askerAAscores a hit:SA,t=w\(Ct\(dt\)\)S\_\{A,t\}=w\(C\_\{t\}\(d\_\{t\}\)\)andSB,t=0\.0S\_\{B,t\}=0\.0\(for regular rounds\)\. In the bonus rounds \(9 and 10\), decay is ignored, andSA,t=1\.0S\_\{A,t\}=1\.0\.
#### Elo Rating Aggregation\.
To aggregate match outcomes into a global leaderboard, we treat each bidirectional game pair as a single observation\. LetMatch1Match\_\{1\}\(A asking B\) andMatch2Match\_\{2\}\(B asking A\) yield scoresSA\(1\),SB\(1\)S\_\{A\}^\{\(1\)\},S\_\{B\}^\{\(1\)\}andSA\(2\),SB\(2\)S\_\{A\}^\{\(2\)\},S\_\{B\}^\{\(2\)\}, respectively\. The net score for modelAAis defined as:
NA=\(SA\(1\)\+SB\(2\)\)−\(SB\(1\)\+SA\(2\)\)N\_\{A\}=\\left\(S\_\{A\}^\{\(1\)\}\+S\_\{B\}^\{\(2\)\}\\right\)\-\\left\(S\_\{B\}^\{\(1\)\}\+S\_\{A\}^\{\(2\)\}\\right\)\(1\)The game outcome is mapped to a standard pairwise scoreYAB∈\{0\.0,0\.5,1\.0\}Y\_\{AB\}\\in\\\{0\.0,0\.5,1\.0\\\}:
YAB=\{1\.0ifNA\>00\.5ifNA=00\.0ifNA<0Y\_\{AB\}=\\begin\{cases\}1\.0&\\text\{if \}N\_\{A\}\>0\\\\ 0\.5&\\text\{if \}N\_\{A\}=0\\\\ 0\.0&\\text\{if \}N\_\{A\}<0\\end\{cases\}\(2\)The expected outcomeEABE\_\{AB\}is modeled using the standard logistic curve:
EAB=11\+10\(RB−RA\)/400E\_\{AB\}=\\frac\{1\}\{1\+10^\{\(R\_\{B\}\-R\_\{A\}\)/400\}\}\(3\)Upon observingYABY\_\{AB\}, the Elo ratings are updated sequentially:
RA′=RA\+K⋅\(YAB−EAB\)R\_\{A\}^\{\\prime\}=R\_\{A\}\+K\\cdot\\left\(Y\_\{AB\}\-E\_\{AB\}\\right\)\(4\)whereK=32K=32\. All models are initialized with a default rating of15001500\.
## 4Experimental Setup
We evaluate ten frontier models from four major AI providers \(Anthropic, Google, OpenAI, and DeepSeek\); their detailed configurations are provided in Table[2](https://arxiv.org/html/2607.24780#S4.T2)\. We execute a full round\-robin tournament of\(102\)=45\\binom\{10\}\{2\}=45unique pairs\. For each pair, we run 8 games \(4 bidirectional matchups\), with each match consisting of 10 rounds, resulting in 360 matches and 3,600 individual rounds of play\.
The judge panel comprises three of the contestants:GPT\-5\.5,Gemini\-3\.5\-Flash, andClaude\-Sonnet\-4\.6, ensuring one judge from each major provider block\. All judges operate at temperature0\.00\.0for determinism, and other endpoints use their default provider\-specific API parameters\. The full tournament of 360 matches completes in approximately 100 minutes of wall\-clock time using parallel API workers\.
Table[2](https://arxiv.org/html/2607.24780#S4.T2)lists the specific parameters and configurations of the ten frontier language models evaluated in our study\. The evaluated models range from lightweight, cost\-effective standard endpoints \(e\.g\.,Gemini\-3\.1\-Flash\-LiteandDeepSeek\-V4\-Flash\) to massive\-scale flagship systems and next\-generation reasoning architectures \(e\.g\.,GPT\-5\.5\)\. Standard chat models operate under direct zero\-temperature prompt instructions, whereas reasoning models natively employ internal chain\-of\-thought processing\. This architectural diversity allows us to verify whether the peer probing paradigm remains robust across differing paradigms of text generation and inference\-time computation\.
Table 2:Detailed specifications of evaluated frontier language models, including provider name, context window limits \(input/output context\), and their API endpoint type\.#### Prompt Instructions and System Schemas\.
The entire tournament flow is driven by specialized system instructions located under theprompts/directory of our code repository:
- •prompts/questioner\.md: Restricts the Questioner to propose objective, unambiguous, and verifiable questions across six standardized dimensions \(D1 Quantitative, D2 Code, D3 Natural Science, D4 Humanities, D5 Language, and D6 Logic\)\. It also introduces decay\-weighted scoring rules to discourage redundant questioning on the same dimension\.
- •prompts/answerer\.md: Instructs the Answerer to respond as accurately and concisely as possible\.
- •prompts/judge\_validate\.md: Directs the three model judges to validate the objectivity and correctness of the standard answer of the proposed question, rejecting any questions with logical flaws, calculation errors, or ambiguous phrasing\.
- •prompts/judge\_answer\.md: Directs the judges to perform binary grading \(CORRECT/WRONG\) of the Answerer’s response against the validated gold standard answer, avoiding style or length biases\.
## 5Results
### 5\.1Frontier Ranking
ModelElo RatingRank 95% CITierGPT\-5\.51743±311743\\pm 31\[1,1\]\[1,1\]T1Gemini\-3\.5\-Flash1642±371642\\pm 37\[2,3\]\[2,3\]T2Gemini\-3\.1\-Pro1637±401637\\pm 40\[2,3\]\[2,3\]DeepSeek\-V4\-Pro1529±391529\\pm 39\[4,5\]\[4,5\]T3DeepSeek\-V4\-Flash1523±281523\\pm 28\[4,5\]\[4,5\]Claude\-Opus\-4\.71470±331470\\pm 33\[6,7\]\[6,7\]T4Claude\-Opus\-4\.61431±631431\\pm 63\[6,8\]\[6,8\]Claude\-Sonnet\-4\.61405±491405\\pm 49\[7,8\]\[7,8\]Gemini\-3\.1\-Flash\-Lite1317±481317\\pm 48\[9,10\]\[9,10\]T5GPT\-5\.21302±351302\\pm 35\[9,10\]\[9,10\]
Table 3:LivingArena Elo leaderboard with bootstrap\-derived margins \(±\\pmrepresents the average half\-width of the 95% CI\) and rank stability intervals\. GPT\-5\.5 holds rank 1 in 100% of resamples, and five tiers are clearly distinguishable\.Figure 3:Capability map: hit rate as asker \(xx\) vs\. answer accuracy as answerer \(yy\)\. Questioning and answering ability are distinct axes; the Claude family answers well but rarely stumps opponents\.Table[3](https://arxiv.org/html/2607.24780#S5.T3)reports the final Elo ratings along with bootstrap 95% confidence intervals \(B=3000B=3000\)\. We decompose model performance into two axes—questioning ability \(hit rate as asker\) and answering ability \(accuracy as answerer\)—as visualized in the Capability Map \(Figure[3](https://arxiv.org/html/2607.24780#S5.F3)\)\. This decomposition helps explain why models rank where they do\.
GPT\-5\.5\(Tier 1\) dominates the tournament by exhibiting perfect balance: it is a flawless answerer \(100% accuracy\) and an aggressive, highly effective questioner \(26\.3% hit rate\)\. Tier 2 models \(Gemini\-3\.5\-FlashandGemini\-3\.1\-Pro\) similarly show high calibration and strong questioning\. The Claude family models \(Tier 4\) instead exhibit a*role asymmetry*: while they answer well \(88%88\\%–89%89\\%accuracy\), they are passive questioners, stumping opponents in only2%2\\%–4%4\\%of rounds\. This passive questioning accounts for their middle\-of\-the\-pack standing despite their high answering accuracy\. Tier 5 models \(Gemini\-3\.1\-Flash\-LiteandGPT\-5\.2\) anchor the bottom due to low answering accuracy \(GPT\-5\.2collapses to 65\.3%\) and high self\-harm rates\. We also observe consistent within\-family monotonic ordering \(e\.g\.,Opus\-4\.7\>Opus\-4\.6\>Sonnet\-4\.6\\text\{Opus\-4\.7\}\>\\text\{Opus\-4\.6\}\>\\text\{Sonnet\-4\.6\}\), validating our Elo aggregation scheme\.
### 5\.2Behavioral Validity
We now examine whether peer probing elicits*boundary\-seeking behaviors*, i\.e\., whether models can identify and exploit their opponents’ cognitive gaps—bearing directly on our central question:*Do LLMs know what other LLMs don’t?*
#### Models Cannot Transcend Their Own Boundary \(Self\-Play\)\.
In self\-play, of all dynamically generated questions passing validation, models answer their own questions correctly in93%93\\%to100%100\\%of cases \(mean≈98%\\approx 98\\%\), whereas against other models they are frequently stymied\. This indicates an asymmetric constraint:*generating a novel, correct, and difficult question is harder than answering one*\. The pattern echoes findings in scientific peer review\(e\.g\., PeerPrism; Sadeghianet al\.,[2026](https://arxiv.org/html/2607.24780#bib.bib22)\), where a model’s fluent evaluative text is dissociated from its actual insight, suggesting that the cognitive demands of generation and execution differ\. In self\-play the answerer almost always wins, confirming that generated questions are bounded by the asker’s own capability\.
#### Exploit\-After\-Hit \(RQ2\)\.
We analyze if models identify and exploit weaknesses by measuring category transition probabilities\. After a successful hit in dimensionddat roundtt, the probability that the asker targets the same dimensionddatt\+1t\+1is20\.1%, compared to a baseline of only3\.8%following a non\-hit\. This highly significant effect of\+16\.3%\+16\.3\\%\(bootstrap CI\[\+11\.0%,\+22\.6%\]\[\+11\.0\\%,\+22\.6\\%\]\) is positive for all ten models, proving statistically that models double down on discovered vulnerabilities\.
#### Sustained Probing and Localization \(RQ3\)\.
We examine the conditional hit rate on a targeted dimension before and after a model achieves its first hit\. As illustrated in Figure[4](https://arxiv.org/html/2607.24780#S5.F4), strong questioners exhibit a dramatic increase in conditional hit rate once they localize a boundary: forGPT\-5\.5, the same\-dimension hit rate rises from25%25\\%\(baseline\) to52%52\\%after the first hit; forGemini\-3\.1\-Pro, it rises from27%27\\%to38%38\\%; and forDeepSeek\-V4\-Pro, it rises from7%7\\%to38%38\\%\. For weaker questioners \(such as the Claude family or GPT\-5\.2\), the sample size of post\-hit rounds is too small \(n≤4n\\leq 4\) to show statistically significant differences, but the overall trend remains consistent\. This positive delta for all high\-capacity models indicates that these models do not generate questions at random: once they localize a weakness or cognitive gap in their opponent, they adjust their questioning strategy and concentrate on that dimension in subsequent rounds\. This provides clear empirical evidence that stronger models are capable of performing dynamic, real\-time boundary localization on their peers \(addressing our RQ7\)\.
Figure 4:Sustained probing: baseline hit rate vs\. conditional hit rate after the first hit on a targeted dimension for strong questioners\.Figure 5:Strategic use of unconstrained bonus rounds \(all 45 pairs / 10 models\)\.Left:a dimension is targeted in a bonus round far more often if the asker had already hit the opponent there in rounds 1–8 \(41%41\\%vs\.18%18\\%\)\.Right:this pays off—bonus questions reusing a discovered weak dimension stump the opponent39%39\\%of the time vs\.4\.5%4\.5\\%otherwise\.
#### Strategic Exploitation in Bonus Rounds \(RQ7\)\.
Rounds 9–10 are*unconstrained*bonus rounds with no per\-dimension decay, an ideal probe of whether models*deliberately*concentrate fire on discovered weaknesses\. We re\-classify every one of the 561 bonus questions into a content dimension with an LLM classifier \(T=0T\{=\}0\) and compare it to the dimensions the asker had already hit the opponent on in rounds 1–8\. Models target strategically: a previously\-hit dimension is re\-attacked with probability41\.4%41\.4\\%versus18\.2%18\.2\\%for a never\-hit dimension \(2\.3×2\.3\{\\times\}; Figure[5](https://arxiv.org/html/2607.24780#S5.F5), left\)\. Controlling for each asker’s own domain preference,70\.7%70\.7\\%of bonus questions land in a previously\-discovered weak dimension, against a46\.0%46\.0\\%propensity baseline \(\+24\.8\+24\.8pp; bootstrap CI\[\+19\.5,\+30\.9\]\[\+19\.5,\+30\.9\]\)\. The strategy is effective \(Figure[5](https://arxiv.org/html/2607.24780#S5.F5), right\): bonus questions reusing a weak dimension hit38\.8%38\.8\\%of the time vs\.4\.5%4\.5\\%otherwise\. The intent is explicit—in87%87\\%of bonus rounds with stored reasoning the asker cites the opponent’s earlier failures, e\.g\.,*“the answerer repeatedly failed on D6…the optimal strategy is to keep attacking its exposed weakness\.”*
### 5\.3Diagnostics & Behavioral Fingerprints
#### Self\-Harm as a Calibration Signal \(RQ1\)\.
By tracking match\-level metrics, LivingArena yields distinct diagnostics: \(i\) self\-harm rate, \(ii\) hit rate as asker, and \(iii\) answer accuracy \(Table[4](https://arxiv.org/html/2607.24780#S5.T4)\)\. The self\-harm rate \(fraction of generated questions failing validation due to incorrect gold answers or flaws\) measures a model’s*self\-calibration*\. As shown in Figure[6](https://arxiv.org/html/2607.24780#S5.F6)\(with bootstrap 95% confidence intervals over 1,000 resamples\), the models split into two clusters: poorly\-calibrated models \(GPT\-5\.2at42\.7%42\.7\\%, the Claude family at31\.7%31\.7\\%–38\.3%38\.3\\%\) versus well\-calibrated ones \(GPT\-5\.5at10\.0%10\.0\\%,Gemini\-3\.1\-Proat13\.3%13\.3\\%\) that rarely output invalid items\.
Figure 6:Self\-harm rate per model with bootstrap 95% CIs \(1,000 resamples over matches\)\. High\- and low\-calibration clusters are clearly separated\.Table 4:Behavioral fingerprints \(cross\-model rounds\): self\-harm rate, hit rate \(as asker\), and answer accuracy \(as answerer\)\.
#### Case Studies of Self\-Harm\.
A qualitative audit of validation logs reveals three primary failure modes:*internal calculation contradiction*\(a dry\-run with contradictory verification steps\),*mathematical computation failure*\(an incorrect derivation in the gold answer\), and*environmental version dependency*\(asking for the output of version\-dependent code without pinning the version\)\.
#### Domain Fingerprints\.
Analysis of all 3,366 rounds shows D6 \(Logic\) is the hardest domain \(83\.8%83\.8\\%accuracy\), while D4 \(Humanities\) is the easiest \(93\.5%93\.5\\%\)\. Models also show distinct questioning biases \(Figure[7](https://arxiv.org/html/2607.24780#S5.F7)\): the Gemini family favors D1 \(Quantitative Reasoning\),GPT\-5\.5favors D2 \(Code & Algorithms\), andDeepSeek\-V4\-Proconcentrates its attacks on D6 \(Logic\)\.
Figure 7:Per\-dimension answer accuracy for representative models\. GPT\-5\.2 collapses on several dimensions, consistent with its low ranking and high self\-harm rate\.
### 5\.4Judge Reliability and Self\-Bias
Since our judge panel consists of active contestants, we audit potential self\-favoritism\. First, judges achieve exceptionally high agreement: Fleiss’κ=0\.915\\kappa=0\.915for grading \(almost perfect consensus\) andκ=0\.693\\kappa=0\.693for validation \(substantial agreement\)\. Second, while a naive comparison suggests self\-bias \(e\.g\.,GPT\-5\.5rates its own model correct in 99\.4% of cases vs\. 84\.5% for others\), this is a confounding effect of model capability\. Controlling for item difficulty by comparing a judge’s vote to the other two judges on the same item, self\-bias is negligible:Grading Self\-Biasis\+0\.0%\+0\.0\\%for bothGPT\-5\.5andGemini\-3\.5\-Flash, whileValidation Self\-Biasis\+2\.3%\+2\.3\\%forGPT\-5\.5and\+6\.0%\+6\.0\\%forGemini\-3\.5\-Flash\. Multi\-judge consensus is thus largely resistant to strategic favoritism, supporting the use of active contestants as referees\(Soumik,[2026](https://arxiv.org/html/2607.24780#bib.bib19)\)\.
### 5\.5Comparison with Human Preference
We compare LivingArena Elo ratings with the human\-preference Chatbot Arena \(LMArena\) June 10, 2026 snapshot\. The leaderboards correlate weakly \(Spearman’sρ=\+0\.358\\rho=\+0\.358, Kendall’sτ=\+0\.200\\tau=\+0\.200\), indicating distinct capability axes\. The main divergence is the Claude family: though favored by humans for polite, verbose formatting, their style offers no advantage in objective, adversarial peer\-probing, where they are penalized for high self\-harm rates \(31\.7%31\.7\\%–38\.3%38\.3\\%\) and passive questioning\.GPT\-5\.5andDeepSeek\-V4\-Proinstead do well through strict factual precision, good self\-calibration, and aggressive questioning\. Peer probing therefore measures an axis of objective rigor and adversarial calibration that human preference voting does not capture\.
## 6Conclusion
We introduced the peer probing paradigm and instantiated it inLivingArena, an automated, self\-adaptive, and contamination\-resistant benchmark\. By letting frontier LLMs probe one another through objective, verifiable questions under strict validation, we bypass the saturation and contamination limitations of static benchmarks, yielding a stable Elo leaderboard that reveals asking and answering as distinct capabilities and exposes self\-calibration as an interpretable axis static accuracy overlooks\. Answering our central question, while models cannot ask beyond their own knowledge boundary, they identify and actively attack the weaknesses of other models—so collective peer probing offers a scalable evaluation signal that only becomes harder as participating models grow stronger\.
## Limitations
We identify two main limitations\. First, anevaluator ceiling: peer probing is bounded by the participating models, so a blind spot shared by all of them cannot be probed—a fundamental limit of peer\-based scalable evaluation, analogous to weak\-to\-strong oversight\. Second,sample\-size constraints: at 8 games per pair, our bootstrap robustly separates five tiers but cannot resolve the exact order within a tier; finer resolution requires more games per pair, trading off against API cost\.
## References
- Anonymous \(2026\)BLOOMQA: automated benchmark generation from domain guidelines informed by bloom’s taxonomy\.Note:OpenReview preprintExternal Links:[Link](https://openreview.net/forum?id=34000e9ccecd670fe53b88aa3d133d5b65dea904)Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px1.p1.1)\.
- R\. Bommasani, P\. Liang, and T\. Lee \(2023\)Holistic evaluation of language models\.Annals of the New York Academy of Sciences\.Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px1.p1.1)\.
- M\. Du, A\. T\. Luu, B\. Ji, X\. Wu, D\. Huang, T\. Y\. Zhuo, Q\. Liu, and S\. Ng \(2025\)CodeArena: a collective evaluation platform for llm code generation\.External Links:2503\.01295,[Link](https://arxiv.org/abs/2503.01295)Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Harrasse, C\. Bandi, and H\. Bandi \(2026\)Debate, deliberate, decide \(d3\): a cost\-aware adversarial framework for reliable and interpretable LLM evaluation\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 8376–8392\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.392),[Link](https://aclanthology.org/2026.eacl-long.392/)Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px3.p1.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt \(2021\)Measuring massive multitask language understanding\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://openreview.net/forum?id=d7KBjmI3GmQ)Cited by:[§1](https://arxiv.org/html/2607.24780#S1.p1.1),[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px1.p1.1)\.
- E\. Karger, H\. Bastani, C\. Yueh\-Han, Z\. Jacobs, D\. Halawi, F\. Zhang, and P\. E\. Tetlock \(2025\)ForecastBench: a dynamic benchmark of ai forecasting capabilities\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://iclr.cc/virtual/2025/poster/28507)Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px1.p1.1)\.
- H\. Lee, Z\. Zhang, H\. Lu, and L\. Zhang \(2025\)SEC\-bench: automated benchmarking of llm agents on real\-world software security tasks\.Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px1.p1.1)\.
- Z\. Liang, L\. Yu, S\. Zhang, Q\. Ye, and H\. Hu \(2026\)How much do large language model cheat on evaluation? benchmarking overestimation under the one\-time\-pad\-based framework\.\.Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px1.p1.1)\.
- K\. Lo, V\. Shalin, R\. Garrett, and S\. Parthasarathy \(2026\)Claim verification with adversarial reasoning and planning\.Proceedings of the International AAAI Conference on Web and Social Media\.Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px3.p1.1)\.
- D\. Loi, E\. M\. Muià, F\. Siciliano, G\. Trappolini, V\. Crisà, P\. Kruger, and F\. Silvestri \(2025\)AutoBench: automating llm evaluation through reciprocal peer assessment\.arXiv preprint arXiv:2510\.22593\.External Links:[Link](https://arxiv.org/abs/2510.22593)Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px4.p1.1)\.
- T\. Lundy, N\. K\. Raman, and K\. Leyton\-Brown \(2026\)QuickScope: certifying hard questions in dynamic llm benchmarks\.External Links:2604\.17842,[Link](https://arxiv.org/abs/2604.17842)Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Margalit, E\. Avram, R\. Taig, O\. Margalit, and N\. Cohen\-Inger \(2026\)PeerRank: autonomous llm evaluation through web\-grounded, bias\-controlled peer review\.arXiv preprint arXiv:2602\.02589\.External Links:[Link](https://arxiv.org/abs/2602.02589)Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px4.p1.1)\.
- K\. Ning, S\. Yang, Y\. Liu, J\. Yao, Z\. Liu, Y\. Tian, Y\. Song, and L\. Yuan \(2025\)PiCO: Peer Review in LLMs Based on Consistency Optimization\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://mlanthology.org/iclr/2025/ning2025iclr-pico/)Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px4.p1.1)\.
- F\. Régin, E\. D\. Maria, and A\. Bonlarron \(2024\)Chatbot arena: an open platform for evaluating llms by human preference\.Cited by:[§1](https://arxiv.org/html/2607.24780#S1.p1.1)\.
- S\. Sadeghian, A\. Daqiq, R\. Cheraghi, S\. Ebrahimi, N\. Arabzadeh, and E\. Bagheri \(2026\)PeerPrism: peer evaluation expertise vs review\-writing ai\.InProceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval,External Links:[Document](https://dx.doi.org/10.1145/3805712.3808602)Cited by:[§5\.2](https://arxiv.org/html/2607.24780#S5.SS2.SSS0.Px1.p1.3)\.
- S\. K\. Soumik \(2026\)Judging the judges: a systematic evaluation of bias mitigation strategies in llm\-as\-a\-judge pipelines\.arXiv \(Cornell University\)\.Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px2.p1.1),[§5\.4](https://arxiv.org/html/2607.24780#S5.SS4.p1.5)\.
- Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren, A\. Arulraj, X\. He, Z\. Jiang, T\. Li, M\. Ku, K\. Wang, A\. Zhuang, R\. Fan, X\. Yue, and W\. Chen \(2024\)MMLU\-pro: a more robust and challenging multi\-task language understanding benchmark\.Cited by:[§1](https://arxiv.org/html/2607.24780#S1.p1.1),[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px1.p1.1)\.
- S\. Xu and Y\. Zhang \(2026\)A survey on agent\-as\-a\-judge\.arXiv preprint arXiv:2601\.05111\.Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px2.p1.1)\.
- Y\. Zhao, Z\. Zhang, and Q\. Le \(2026\)Beyond goodhart’s law: a dynamic benchmark for evaluating compliance in multi\-agent systems\.External Links:2606\.07805,[Link](https://arxiv.org/abs/2606.07805)Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px3.p1.1)\.
- L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. Stoica \(2023\)Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.Cited by:[§2](https://arxiv.org/html/2607.24780#S2.SS0.SSS0.Px2.p1.1)\.Similar Articles
Game Arena: Strategic LLM Evaluation in Competitive Environments
Game Arena is an open platform for evaluating large language models through competitive games like Chess, Poker, and Werewolf, enabling dynamic assessment of strategic planning and robustness.
SAGE: Scalable Automated Robustness Augmentation for LLM Knowledge Evaluation
This paper introduces SAGE, a framework for scalable automated robustness augmentation of LLM knowledge evaluation benchmarks. It uses fine-tuned smaller models with reinforcement learning to generate and verify question variants at a lower cost than existing methods.
Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery
The paper introduces HypoArena, a benchmark for evaluating LLMs' ability to proactively construct hypothesis spaces from incomplete evidence, and experiments on 15 frontier LLMs reveal capability stratification.
JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation
JudgeArena is an open-source framework that unifies major LLM-judge benchmarks under a single interface, enabling systematic study of judge choices and offering open-model judges that match or outperform closed models, with the ability to simulate LMArena Elo scores.
(Towards) Scalable Reliable Automated Evaluation with Large Language Models
This paper proposes a scalable, domain-agnostic framework for automated LLM evaluation that uses pairwise comparisons by multiple LLMs and an Elo rating system to approximate expert judgments, reducing the need for human intervention.