Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
Summary
Co-RL is a cooperative multi-agent framework that enables unsupervised reasoning improvement in language and vision-language models by using peer-derived rewards, reducing reliance on ground-truth labels and mitigating training collapse through increased cohort diversity.
View Cached Full Text
Cached at: 08/19/26, 10:25 AM
# Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
Source: [https://arxiv.org/html/2608.17253](https://arxiv.org/html/2608.17253)
Yuexin Bian\*Yunjie TianAffiliation:University of Exeter ByteDance\[3pt\]\{yijiangli, nuno\}@ucsd\.eduDi FuAffiliation:University of Exeter ByteDance\[3pt\]\{yijiangli, nuno\}@ucsd\.eduTianjin HuangYuanyuan ShiZiang XiaoNuno VasconcelosYijiang Li\[4pt\] Johns Hopkins University UC San Diego
###### Abstract
Reinforcement learning \(RL\) has emerged as a powerful approach for improving reasoning in language and vision\-language models, yet its strongest successes still depend heavily on ground\-truth supervision \(e\.g\., verifiable reward\)\. Such annotations are costly to obtain and become increasingly scarce as reasoning capabilities advance beyond what humans can reliably evaluate\. Self\-rewarding RL reduces this dependence by enabling models to derive reward signals from their own completions\. However, training solely on self\-generated feedback can reinforce existing biases and suboptimal behaviors, reduce response diversity, and ultimately lead to homogenized responses and training collapse\. In this work, we show that unsupervised reasoning can emerge through cooperative multi\-agent training\. We introduceCo\-RL, a framework in which multiple decoupled models, sharing no parameters, are simultaneously optimized through RL using rewards derived from their peers\. We further show that increasing cohort diversity, through heterogeneous model families, sizes, and rephrased training samples, reduces the correlated errors that drive self\-reinforcing feedback loops\. This diversity consistently improves reasoning performance, maintains behavioral diversity, and mitigates training collapse\. Across text\-only and multimodal domains,Co\-RLconsistently outperforms the base models and prior label\-free approaches, while matching or surpassing supervised methods,without access to any ground\-truth labels\. Concretely,Co\-RLyields average gains of 3\.0–8\.6% across seven text\-only benchmarks for LLMs and 2\.3–7\.2% across four multimodal benchmarks for VLMs\. Code is available at[https://github\.com/DrStranded/Co\-RL](https://github.com/DrStranded/Co-RL)\.
††footnotetext:∗Equal contribution\.†Project lead\.††footnotetext:Corresponding to Yijiang Li: yijiangli@ucsd\.edu, Nuno Vasconcelos: nuno@ucsd\.edu## 1Introduction
Reinforcement learning with verifiable rewards \(RLVR\) has emerged as a powerful approach for improving reasoning in large language models\([Lightman et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib40);[DeepSeek\-AI 2025](https://arxiv.org/html/2608.17253#bib.bib17)\), yet its strongest successes still depend heavily on ground\-truth supervision\. Such supervision is costly to obtain and becomes increasingly scarce as target reasoning capabilities approach or surpass what humans can reliably evaluate\([Yue et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib76)\)\. Self\-rewarding RL reduces this dependence by deriving rewards from the model’s own completions, incorporating signals such as agreement with its majority\-vote prediction\([Zuo et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib88)\), self\-certainty\([Zhao et al\. 2026](https://arxiv.org/html/2608.17253#bib.bib87)\), predictive entropy\([Prabhudesai et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib51)\), or consistency across paraphrased inputs or moving\-average policies\([Zhang et al\. 2026b](https://arxiv.org/html/2608.17253#bib.bib84)\)\. However, these signals remain within a single model’s own predictions\. Without an external reference, such self\-reinforcement can amplify existing biases and suboptimal behaviors, reduce response diversity, and ultimately lead to increasingly homogeneous outputs or even training collapse\.
This raises a fundamental question:*how can a model obtain a sufficiently independent learning signal to improve without any ground\-truth supervision?*In this work, we show that such a signal can emerge from independently trained models\. Since their errors are not perfectly correlated, each model can provide corrective feedback that the other cannot derive from its own generations\. We therefore take the learning signal from a separate model, giving each agent*decorrelated supervision*: a target produced by independently updated weights that is less likely to echo its own biases\([Blum and Mitchell 1998](https://arxiv.org/html/2608.17253#bib.bib3);[Li et al\. 2023](https://arxiv.org/html/2608.17253#bib.bib37)\)\.
Building on this insight, we introduceCo\-RL, a cooperative multi\-agent label\-free RL method in which multiple decoupled models, sharing no parameters, are optimized simultaneously using rewards derived from their cohorts\. Given an unlabeled prompt, each agent samples multiple completions and aggregates them into a pseudo\-answer through majority voting\([Wang et al\. 2023](https://arxiv.org/html/2608.17253#bib.bib63)\)\. The completions of one agent are then rewarded against another agent’s pseudo\-answer, and the resulting rewards drive policy optimization with GRPO\([Shao et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib55)\)or REINFORCE\+\+\([Hu et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib28)\)\. Unlike self\-rewarding methods, where the policy and the reward come from the same model,Co\-RLdraws supervision from an independently updated partner, which breaks the feedback loop that amplifies a model’s own bias\.
What a cohort teaches depends on how different its mistakes are\. Highly similar models tend to make correlated errors and may reinforce the same incorrect answers\. Therefore, cohort diversity is crucial to the effectiveness ofCo\-RL\. Thus, we push it as far as the models allow: different families of architecture and pretrained weights, model size, and input formulation \(e\.g\., rephrased prompt\([Zhang et al\. 2026b](https://arxiv.org/html/2608.17253#bib.bib84)\)\)\. These differences expose agents to distinct inductive biases and decision boundaries, reducing correlated errors and strengthening the corrective signal available to their cohorts\.
By incorporating multiple agents in the training loop,Co\-RLnaturally constitutes a multi\-agent RL framework\. However, unlike prior multi\-agent methods built around debate or iterative communication\([Du et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib20);[Liang et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib38)\),Co\-RLrequires no interaction between agents beyond the reward stage and eliminates the need for an external LLM judge or learned reward model\([Xue et al\. 2026](https://arxiv.org/html/2608.17253#bib.bib69);[Park et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib49);[Zhang et al\. 2026a](https://arxiv.org/html/2608.17253#bib.bib77)\)\. The resulting framework is lightweight and symmetric: each agent serves simultaneously as a learner and a source of supervision, allowing all agents to improve within a single training run\.
Across text\-only and multimodal domains, diverse model families, and training settings,Co\-RLconsistently improves upon the base models and prior label\-free approaches, while matching or surpassing supervised training in several settings,without access to ground\-truth labels\. Across seven text\-only benchmarks spanning mathematical reasoning, coding, and knowledge\-intensive reasoning,Co\-RLimproves four LLMs by 3\.0–8\.6% on average across seven text\-only benchmarks, outperforming the strongest self\-rewarding baselines by 0\.8–2\.0%\. The gains extend consistently to multimodal reasoning, whereCo\-RLimproves five VLMs ranging from 2B to 12B parameters by 2\.3–7\.2% across MathVision, MathVerse, MathVista, and We\-Math\. Under the controlled evaluation setting of CoMAS\([Xue et al\. 2026](https://arxiv.org/html/2608.17253#bib.bib69)\),Co\-RLfurther outperforms prior multi\-agent RL methods by 4\.0% on average while using only half as many agents\. Together, these results demonstrate that cross\-agent supervision provides an effective and general learning signal across modalities, model families, and reasoning domains without relying on labeled supervision or external judges\.
## 2Related Work
Self\-rewarding RL\.While RLVR effectively improves LLM reasoning\([Shao et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib55);[DeepSeek\-AI 2025](https://arxiv.org/html/2608.17253#bib.bib17);[Yu et al\. 2025b](https://arxiv.org/html/2608.17253#bib.bib74)\), its reliance on high\-quality ground\-truth labels remains a major bottleneck\([Yue et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib76)\)\. Recent work therefore explores self\-rewarding mechanisms that learn from unlabeled data\. One prominent direction derives reward signals solely from a model’s own behavior, such as majority voting\([Wang et al\. 2023](https://arxiv.org/html/2608.17253#bib.bib63)\), self\-consistency\([Zuo et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib88)\), self\-certainty\([Zhao et al\. 2026](https://arxiv.org/html/2608.17253#bib.bib87);[Li et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib36)\), or predictive entropy\([Prabhudesai et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib51);[Zhang et al\. 2025a](https://arxiv.org/html/2608.17253#bib.bib78)\)\. This line extends earlier work on self\-rewarding models\([Yuan et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib75)\)and self\-play supervision\([Chen et al\. 2024c](https://arxiv.org/html/2608.17253#bib.bib13)\), and now includes unsupervised self\-training\([Xu et al\. 2025a](https://arxiv.org/html/2608.17253#bib.bib67);[Fang et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib21)\), self\-correction based on a model’s own judgments\([Xiong et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib66)\), and zero\-data self\-evolution, where one or more models generate their own curricula\([Zhao et al\. 2025a](https://arxiv.org/html/2608.17253#bib.bib85);[Huang et al\. 2026a](https://arxiv.org/html/2608.17253#bib.bib29);[Liu et al\. 2026](https://arxiv.org/html/2608.17253#bib.bib41)\)\. However, because these methods rely exclusively on a single model’s own view, they can reinforce and amplify existing biases and errors without providing an external corrective signal, often leading to substantial training collapse\.
Co\-training, cross\-view supervision, and its origins\.Using a second view to supervise a learner is a longstanding idea\. Co\-training shows that two conditionally independent views can teach each other\([Blum and Mitchell 1998](https://arxiv.org/html/2608.17253#bib.bib3)\), while subsequent analysis demonstrates that highly similar views tend to reinforce the same errors: the more alike the two views are, the more they simply agree on the same mistakes\([Li et al\. 2023](https://arxiv.org/html/2608.17253#bib.bib37)\)\. Related principles underlie self\-supervised representation learning, where augmentations, momentum targets, or stop\-gradient operations prevent collapse\([Chen et al\. 2020](https://arxiv.org/html/2608.17253#bib.bib9);[Grill et al\. 2020](https://arxiv.org/html/2608.17253#bib.bib25);[Caron et al\. 2021](https://arxiv.org/html/2608.17253#bib.bib4)\), and deep mutual learning, where peer networks provide reciprocal supervision\([Zhang et al\. 2018](https://arxiv.org/html/2608.17253#bib.bib82)\)\. Co\-rewarding\([Zhang et al\. 2026b](https://arxiv.org/html/2608.17253#bib.bib84)\)brings this insight to label\-free RL by deriving rewards from either paraphrased questions or a slowly updated model copy\. However, because both views originate from the same model, their errors remain strongly correlated, only partially satisfying the independence required for effective two\-view learning\.
Multi\-agent RL\.Several methods improve reasoning by letting multiple models interact\. At inference time, multi\-agent debate and round\-table consensus have models critique and revise one another\([Du et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib20);[Liang et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib38);[Chen et al\. 2024a](https://arxiv.org/html/2608.17253#bib.bib6);[Sun et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib57)\), and general orchestration frameworks compose such agents into pipelines\([Wu et al\. 2023](https://arxiv.org/html/2608.17253#bib.bib65);[Zhang et al\. 2024b](https://arxiv.org/html/2608.17253#bib.bib83);[Kim et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib32)\)\. A more recent line trains the agents: CoMAS turns interactions scored by an LLM judge into rewards\([Xue et al\. 2026](https://arxiv.org/html/2608.17253#bib.bib69)\), MAPoRL and MARFT co\-train agents against a learned or shared reward\([Park et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib49);[Liao et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib39)\), and other systems reinforce or bootstrap multi\-agent cooperation directly\([Motwani et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib48);[Chen et al\. 2025c](https://arxiv.org/html/2608.17253#bib.bib10);[Zhao et al\. 2025b](https://arxiv.org/html/2608.17253#bib.bib86);[Zhang et al\. 2026a](https://arxiv.org/html/2608.17253#bib.bib77);[Chen et al\. 2025d](https://arxiv.org/html/2608.17253#bib.bib11)\)\.
RLVR for multimodal reasoning\.RLVR has rapidly expanded to vision\-language models, with R1\-style rule\-based rewards underpinning a broad family of methods\([Huang et al\. 2026b](https://arxiv.org/html/2608.17253#bib.bib30);[Shen et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib56);[Feng et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib22);[Liu et al\. 2025b](https://arxiv.org/html/2608.17253#bib.bib43);[Meng et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib47);[Peng et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib50);[Yang et al\. 2025b](https://arxiv.org/html/2608.17253#bib.bib72);[Wang et al\. 2025b](https://arxiv.org/html/2608.17253#bib.bib60)\)\. Existing work mainly improves reward design and training stability through curriculum or perception\-aware rewards\([Deng et al\. 2025a](https://arxiv.org/html/2608.17253#bib.bib18);[Yu et al\. 2025a](https://arxiv.org/html/2608.17253#bib.bib73);[Wang et al\. 2025a](https://arxiv.org/html/2608.17253#bib.bib58)\), data augmentation and selection\([Liu et al\. 2025a](https://arxiv.org/html/2608.17253#bib.bib42);[Wang et al\. 2025c](https://arxiv.org/html/2608.17253#bib.bib62)\), and staged supervised\-to\-RL training\([Deng et al\. 2025b](https://arxiv.org/html/2608.17253#bib.bib19);[Chen et al\. 2025a](https://arxiv.org/html/2608.17253#bib.bib5)\), but still largely assumes verifiable labels\. Label\-free multimodal RL remains underexplored, and existing methods again rely on a single model’s own signals\([Wei et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib64)\)\. Cross\-view supervision is especially promising here, since VLMs rarely share a vision encoder, pairing InternViT\([Chen et al\. 2024b](https://arxiv.org/html/2608.17253#bib.bib12)\), SigLIP\([Gemma Team 2025](https://arxiv.org/html/2608.17253#bib.bib23)\), or a native\-resolution encoder\([Bai et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib2);[Zhang et al\. 2025c](https://arxiv.org/html/2608.17253#bib.bib81);[Zhang et al\. 2025b](https://arxiv.org/html/2608.17253#bib.bib80)\)with different LLMs, yielding distinctive families of models\.
## 3Preliminary
### 3\.1Reinforcement Learning for Reasoning
Given a promptx∼𝒟x\\sim\\mathcal\{D\}, an autoregressive language modelπθ\\pi\_\{\\theta\}generates a responsey=\(y1,…,yT\)y=\(y\_\{1\},\\ldots,y\_\{T\}\)according toπθ\(y∣x\)=∏t=1Tπθ\(yt∣x,y<t\)\\pi\_\{\\theta\}\(y\\mid x\)=\\prod\_\{t=1\}^\{T\}\\pi\_\{\\theta\}\(y\_\{t\}\\mid x,y\_\{<t\}\)\. Reinforcement learning optimizes the policy using a scalar rewardr\(x,y\)r\(x,y\)that evaluates the quality of the generated response\. The corresponding objective is
𝒥\(θ\)=𝔼x∼𝒟,y∼πθ\(⋅∣x\)\[r\(x,y\)\]\.\\mathcal\{J\}\(\\theta\)=\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\},\\,y\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x\)\}\\left\[r\(x,y\)\\right\]\.\(1\)For reasoning tasks, the responseyytypically contains both an intermediate reasoning trajectory and a final answer, while the reward is commonly assigned at the response level based on answer correctness or another outcome\-based criterion\.
Among policy optimization algorithms, Group Relative Policy Optimization \(GRPO\) has been widely adopted as it improves reasoning performance without requiring a separately trained critic model\([Shao et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib55);[DeepSeek\-AI 2025](https://arxiv.org/html/2608.17253#bib.bib17)\)\. Given a promptxx, the old policyπθold\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}samples a group ofKKresponses\{y1,…,yK\}\\\{y^\{1\},\\ldots,y^\{K\}\\\}and assigns each response a rewardrk=r\(x,yk\)r^\{k\}=r\(x,y^\{k\}\)\. GRPO estimates the advantage of each response by normalizing rewards within the group:A^k=rk−mean\(\{rj\}j=1K\)std\(\{rj\}j=1K\)\+ϵ\\hat\{A\}^\{k\}=\\frac\{r^\{k\}\-\\operatorname\{mean\}\(\\\{r^\{j\}\\\}\_\{j=1\}^\{K\}\)\}\{\\operatorname\{std\}\(\\\{r^\{j\}\\\}\_\{j=1\}^\{K\}\)\+\\epsilon\}, whereϵ\\epsilonis a small constant for numerical stability\. Letρk,t\(θ\)=πθ\(ytk∣x,y<tk\)πθold\(ytk∣x,y<tk\)\\rho\_\{k,t\}\(\\theta\)=\\frac\{\\pi\_\{\\theta\}\(y\_\{t\}^\{k\}\\mid x,y\_\{<t\}^\{k\}\)\}\{\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(y\_\{t\}^\{k\}\\mid x,y\_\{<t\}^\{k\}\)\}denote the token\-level importance ratio\. GRPO optimizes the clipped surrogate objective
𝒥GRPO\(θ\)=𝔼\[1K∑k=1K1\|yk\|∑t=1\|yk\|\[min\(ρk,t\(θ\)A^k,clip\(ρk,t\(θ\),1−δ,1\+δ\)A^k\)−β𝒟KLk,t\]\],\\mathcal\{J\}\_\{\\mathrm\{GRPO\}\}\(\\theta\)=\\mathbb\{E\}\\\!\\left\[\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\frac\{1\}\{\|y^\{k\}\|\}\\sum\_\{t=1\}^\{\|y^\{k\}\|\}\\left\[\\min\\\!\\left\(\\rho\_\{k,t\}\(\\theta\)\\hat\{A\}^\{k\},\\operatorname\{clip\}\\\!\\left\(\\rho\_\{k,t\}\(\\theta\),1\-\\delta,1\+\\delta\\right\)\\hat\{A\}^\{k\}\\right\)\-\\beta\\mathcal\{D\}\_\{\\mathrm\{KL\}\}^\{k,t\}\\right\]\\right\],\(2\)whereδ\\deltais the clipping threshold,β\\betacontrols the strength of KL regularization, and𝒟KLk,t\\mathcal\{D\}\_\{\\mathrm\{KL\}\}^\{k,t\}penalizes deviation from a reference policyπref\\pi\_\{\\mathrm\{ref\}\}\. GRPO typically uses verifiable rewards from ground\-truth answers or external verifiers\. However, such supervision can be costly or difficult to obtain at scale\. The following subsection considers model\-generated rewards as an alternative, enabling reinforcement learning on unlabeled prompts\.
### 3\.2Self\-rewarding RL
In the absence of external verifiers, reward functions can be constructed directly from a model’s own outputs, thereby enabling RL on unlabeled prompts\. Given an unlabeled promptxx, the policyπθ\\pi\_\{\\theta\}samples responsesy1,…,yK\{y^\{1\},\\ldots,y^\{K\}\}, with answersak=g\(yk\)a^\{k\}=g\(y^\{k\}\)extracted from each response\. TTRL\([Zuo et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib88)\)defines the reward for responseyky^\{k\}asrmajk=𝟏\[ak=a^θ\(x\)\]r\_\{\\mathrm\{maj\}\}^\{k\}=\\mathbf\{1\}\[a^\{k\}=\\hat\{a\}\_\{\\theta\}\(x\)\], wherea^θ\(x\)=argmaxb∑j=1K𝟏\[aj=b\]\\hat\{a\}\_\{\\theta\}\(x\)=\\arg\\max\_\{b\}\\sum\_\{j=1\}^\{K\}\\mathbf\{1\}\[a^\{j\}=b\]is the policy’s majority\-vote answer\. Intuitor\([Zhao et al\. 2026](https://arxiv.org/html/2608.17253#bib.bib87)\)instead measures token\-level confidence usingrconfk=1\|yk\|∑t=1\|yk\|DKL\(𝐔∥πθ\(⋅∣x,y<tk\)\)r\_\{\\mathrm\{conf\}\}^\{k\}=\\frac\{1\}\{\|y^\{k\}\|\}\\sum\_\{t=1\}^\{\|y^\{k\}\|\}D\_\{\\mathrm\{KL\}\}\(\\mathbf\{U\}\\\|\\pi\_\{\\theta\}\(\\cdot\\mid x,y\_\{<t\}^\{k\}\)\), where𝐔\\mathbf\{U\}denotes the uniform distribution over the vocabulary\. RENT\([Prabhudesai et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib51)\)operates on the predictive distribution, assigningrentk=−1\|yk\|∑t=1\|yk\|ℋ\(πθ\(⋅∣x,y<tk\)\)r\_\{\\mathrm\{ent\}\}^\{k\}=\-\\frac\{1\}\{\|y^\{k\}\|\}\\sum\_\{t=1\}^\{\|y^\{k\}\|\}\\mathcal\{H\}\(\\pi\_\{\\theta\}\(\\cdot\\mid x,y\_\{<t\}^\{k\}\)\), whereℋ\(⋅\)\\mathcal\{H\}\(\\cdot\)denotes entropy\. Thus, TTRL\([Zuo et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib88)\)rewards responses that agree with the policy’s majority\-vote prediction, whereas Intuitor\([Zhao et al\. 2026](https://arxiv.org/html/2608.17253#bib.bib87)\)and RENT\([Prabhudesai et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib51)\)derive rewards from the model’s token\-level predictive distributions\. Despite their different reward constructions, these methods ultimately produce scalar rewards that can be used for policy optimization\. Under GRPO, the self\-generated rewardsrkk=1K\{r^\{k\}\}\_\{k=1\}^\{K\}are normalized into group\-relative advantagesA^k\\hat\{A\}^\{k\}and used directly in Eq\. \([2](https://arxiv.org/html/2608.17253#S3.E2)\)\.
Despite their effectiveness, these signals originates from the same policy being optimized: without an external reference, training may reinforce existing biases and suboptimal behaviors, reduce response diversity, and eventually lead to homogenized responses or training collapse\. We show this in \(b\) and \(c\) of Figure[2](https://arxiv.org/html/2608.17253#S4.F2), where prolonged training causes TTRL to degenerate and lead to training collapse\.
Figure 1:Comparison ofCo\-RLwith prior label\-free RL methods\. Both TTRL and Co\-rewarding derive rewards from self\-generated agreement, and CoMAS scores multi\-turn interactions with one of its own agents acting as judge\.Co\-RLinstead derives rewards directly from peer votes\. Beyond two agents, the votes pass along a directed ring \(N=3N\{=\}3shown\)\.
## 4Method
### 4\.1Co\-Reinforcement Learning \(Co\-RL\)
This motivates a fundamental question:*how can a model obtain a sufficiently independent learning signal without any ground\-truth supervision?*Our key insight is that such a signal can emerge from independently trained models, since their errors are decorrelated; one model can provide corrective feedback that another cannot derive from its own generations\([Blum and Mitchell 1998](https://arxiv.org/html/2608.17253#bib.bib3);[Li et al\. 2023](https://arxiv.org/html/2608.17253#bib.bib37)\)\. We provide an overview of our method in the right of Figure[1](https://arxiv.org/html/2608.17253#S3.F1), in comparison with prior self\-rewarding RL paradigms\.
Building on this principle, we introduce Co\-Reinforcement Learning \(Co\-RL\), a label\-free multi\-agent reinforcement learning method in which multiple decoupled agents independently generate completions to unlabeled problems and supervise one another through majority\-voted pseudo reward from its cohort\. The agents share neither parameters nor gradients; their optimization is coupled solely through the rewards they provide to one another\.
Formally, letx∼𝒟x\\sim\\mathcal\{D\}denote an unlabeled reasoning problem\. We consider a cohort ofNNagents with independently parameterized policies\{πθn\}n=1N\\\{\\pi\_\{\\theta\_\{n\}\}\\\}\_\{n=1\}^\{N\}\. The agents may be initialized from the same pretrained model or from models of different families and sizes\. For each unlabeled reasoning problemxx, each agentn∈\{1,…,N\}n\\in\\\{1,\\ldots,N\\\}independently rollout a group ofKKcompletionsynk∼i\.i\.d\.πθn\(⋅∣x\),k=1,…,Ky\_\{n\}^\{k\}\\overset\{\\mathrm\{i\.i\.d\.\}\}\{\\sim\}\\pi\_\{\\theta\_\{n\}\}\(\\cdot\\mid x\),k=1,\\ldots,Kand extracts the corresponding final answers asank=g\(ynk\)a\_\{n\}^\{k\}=g\(y\_\{n\}^\{k\}\), withg\(⋅\)g\(\\cdot\)denoting an answer\-extraction function\.
For each agentnn,Co\-RLconstructs a supervision target exclusively from the answers generated by one designated peer\. Specifically, the pseudo\-label is constructed as
a^−n\(x\)∈argmaxb∑j=1K\[an−1j=b\],\\hat\{a\}\_\{\-n\}\(x\)\\in\\arg\\max\_\{b\}\\sum\_\{j=1\}^\{K\}\\mathbf\{1\}\\\!\\left\[a\_\{n\-1\}^\{j\}=b\\right\],\(3\)wherebbranges over the answers produced by agentn−1n\-1, with the index taken cyclically so that agent11is supervised by agentNN\. Thus,a^−n\(x\)\\hat\{a\}\_\{\-n\}\(x\)represents the majority\-vote answer of the peer that supervises agentnn\. The reward assigned to thekk\-th response of agentnnis thenrnk=\[ank=a^−n\(x\)\]r\_\{n\}^\{k\}=\\mathbf\{1\}\\\!\\left\[a\_\{n\}^\{k\}=\\hat\{a\}\_\{\-n\}\(x\)\\right\]\. A response therefore receives reward11when its extracted answer agrees with the constructed pseudo\-label, and reward00otherwise\. In the two\-agent setting, the two agents supervise each other\. Crucially, each agent does not contribute to its own supervision targeta^−n\(x\)\\hat\{a\}\_\{\-n\}\(x\)\.
Policy Optimization\.Without loss of generality, we use GRPO\([Shao et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib55)\)to train our agents\. For each agent, the rewards\{rnk\}k=1K\\\{r\_\{n\}^\{k\}\\\}\_\{k=1\}^\{K\}are normalized within its own rollout group to obtain group\-relative advantages, which are then used in the GRPO objective in Eq\.[4](https://arxiv.org/html/2608.17253#S4.E4)to optimize the agent\. Let𝜽=\(θ1,…,θN\)\\boldsymbol\{\\theta\}=\(\\theta\_\{1\},\\ldots,\\theta\_\{N\}\)\. The overall training objective can be written as
max𝜽𝒥Co\-RL\(𝜽\)=1N∑n=1N𝒥GRPO\(θn,\{ynk,rnk\}k=1K\),\\max\_\{\\boldsymbol\{\\theta\}\}\\;\\mathcal\{J\}\_\{\\mathrm\{Co\\text\{\-\}RL\}\}\(\\boldsymbol\{\\theta\}\)=\\frac\{1\}\{N\}\\sum\_\{n=1\}^\{N\}\\mathcal\{J\}\_\{\\mathrm\{GRPO\}\}\\left\(\\theta\_\{n\};\\\{y\_\{n\}^\{k\},r\_\{n\}^\{k\}\\\}\_\{k=1\}^\{K\}\\right\),\(4\)where
𝒥GRPO\(θn\)=𝔼\[1K∑k=1K1\|ynk\|∑t=1\|ynk\|min\(ρn,tkA^nk,clip\(ρn,tk,1−ϵ,1\+ϵ\)A^nk\)\]−βDKL\(πθn∥πref,n\)\.\\displaystyle\\mathcal\{J\}\_\{\\mathrm\{GRPO\}\}\(\\theta\_\{n\}\)=\\mathbb\{E\}\\Bigg\[\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\frac\{1\}\{\|y\_\{n\}^\{k\}\|\}\\sum\_\{t=1\}^\{\|y\_\{n\}^\{k\}\|\}\\min\\Big\(\\rho\_\{n,t\}^\{k\}\\hat\{A\}\_\{n\}^\{k\},\\,\\operatorname\{clip\}\(\\rho\_\{n,t\}^\{k\},1\-\\epsilon,1\+\\epsilon\)\\hat\{A\}\_\{n\}^\{k\}\\Big\)\\Bigg\]\-\\beta D\_\{\\mathrm\{KL\}\}\\left\(\\pi\_\{\\theta\_\{n\}\}\\,\\\|\\,\\pi\_\{\\mathrm\{ref\},n\}\\right\)\.\(5\)A^nk=rnk−mean\(\{rnj\}j=1K\)std\(\{rnj\}j=1K\),ρn,tk=πθn\(yn,tk∣xn,yn,<tk\)πθnold\(yn,tk∣xn,yn,<tk\)\.\\hat\{A\}\_\{n\}^\{k\}=\\frac\{r\_\{n\}^\{k\}\-\\operatorname\{mean\}\\\!\\left\(\\\{r\_\{n\}^\{j\}\\\}\_\{j=1\}^\{K\}\\right\)\}\{\\operatorname\{std\}\\\!\\left\(\\\{r\_\{n\}^\{j\}\\\}\_\{j=1\}^\{K\}\\right\)\},\\hskip 18\.49988pt\\rho\_\{n,t\}^\{k\}=\\frac\{\\pi\_\{\\theta\_\{n\}\}\\left\(y\_\{n,t\}^\{k\}\\mid x\_\{n\},y\_\{n,<t\}^\{k\}\\right\)\}\{\\pi\_\{\\theta\_\{n\}^\{\\mathrm\{old\}\}\}\\left\(y\_\{n,t\}^\{k\}\\mid x\_\{n\},y\_\{n,<t\}^\{k\}\\right\)\}\.\(6\)Figure[3](https://arxiv.org/html/2608.17253#S4.F3)illustrates the two\-agent setting\. At each training step, all rollouts and majority\-vote pseudo\-labels are computed before any policy update\. Given the resulting rewards, each policy is then updated independently\.
Algorithm[1](https://arxiv.org/html/2608.17253#alg1)summarizes the training procedure, in which all agents are updated at each optimization step\. Our framework is also compatible with other policy optimization methods that support sequence\-level rewards\.
Algorithm 1Co\-Reinforcement Learning \(Co\-RL\)1:Agents
\{πθn\}n=1N\\\{\\pi\_\{\\theta\_\{n\}\}\\\}\_\{n=1\}^\{N\}initialized from different models
2:foreach training stepdo
3:Sample a batch of unlabeled prompts
ℬ⊂𝒟\\mathcal\{B\}\\subset\\mathcal\{D\}
4:for
x∈ℬx\\in\\mathcal\{B\}do
5:for
n=1,…,Nn=1,\\ldots,Ndo⊳\\trianglerightGenerate responses
6:Sample
ynk∼i\.i\.d\.πθn\(⋅∣x\)y\_\{n\}^\{k\}\\overset\{\\mathrm\{i\.i\.d\.\}\}\{\\sim\}\\pi\_\{\\theta\_\{n\}\}\(\\cdot\\mid x\)and extract answers
ank←g\(ynk\)a\_\{n\}^\{k\}\\leftarrow g\(y\_\{n\}^\{k\}\)for
k=1,…,Kk=1,\\ldots,K
7:endfor
8:for
n=1,…,Nn=1,\\ldots,Ndo⊳\\trianglerightConstruct cross\-agent rewards
9:Compute
a^−n\(x\)\\hat\{a\}\_\{\-n\}\(x\)using Eq\. \([3](https://arxiv.org/html/2608.17253#S4.E3)\)
10:for
k=1,…,Kk=1,\\ldots,Kdo
11:Assign
rnk←\[ank=a^−n\(x\)\]r\_\{n\}^\{k\}\\leftarrow\\mathbf\{1\}\\\!\\left\[a\_\{n\}^\{k\}=\\hat\{a\}\_\{\-n\}\(x\)\\right\]
12:endfor
13:endfor
14:endfor
15:Update
\{θn\}n=1N\\\{\\theta\_\{n\}\\\}\_\{n=1\}^\{N\}with one GRPO step using the objective in \([4](https://arxiv.org/html/2608.17253#S4.E4)\)
16:endfor
### 4\.2Diverse Cohorts Enable Unsupervised Reasoner
What a cohort teaches depends on how different its mistakes are\. Highly similar models tend to make correlated errors and may reinforce the same incorrect answers\. Therefore, cohort diversity plays a crucial role inCo\-RL\. We push this diversity as far as the framework allows: policy optimization \(Sec[4\.1](https://arxiv.org/html/2608.17253#S4.SS1)\), model families and sizes, and input formation\. Each source of diversity induces distinct inductive biases and reasoning behaviors, reducing correlated errors and strengthening the corrective signal available through peer supervision\.
Decoupled policy optimization\.As discussed in Sec\.[4\.1](https://arxiv.org/html/2608.17253#S4.SS1),Co\-RLrealizes decoupled policy optimization by training two policies independently, with interaction occurring only through the reward in Eq\. \([4](https://arxiv.org/html/2608.17253#S4.E4)\)\. Each policy maintains its own parameters and optimizer state with no gradient propagated between policies\. Thus, this decoupled policy optimization design also serves as a source of diversity: independent updates prevent the two policies from being directly coupled, maintaining less\-correlated predictions throughout the training process\. We show in Figure[2](https://arxiv.org/html/2608.17253#S4.F2)\(b\) and \(c\) that decoupled policy optimization stabilizes training and yields consistent improvements over self\-rewarding methods\.
Diversity across model families and sizes\.Our primary source of diversity comes from independently pretrained model families\. Distinct families differ throughout the model\-development pipeline, including architecture, tokenization, pretraining data, and post\-training, each of which brings inductive biases that can benefit co\-learning\. For example, Qwen2\.5 and Llama 3 differ in their tokenizers, vocabulary sizes, architectural choices, and pretraining corpora\([Yang et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib70);[Grattafiori et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib24)\), while Gemma introduces a 256K\-token vocabulary and interleaved local–global attention\([Gemma Team 2025](https://arxiv.org/html/2608.17253#bib.bib23)\)\. Such differences are even more pronounced for VLMs, whose families employ distinct vision encoders: Qwen2\.5\-VL uses a natively trained dynamic\-resolution ViT\([Bai et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib2)\), InternVL adopts InternViT\([Chen et al\. 2024b](https://arxiv.org/html/2608.17253#bib.bib12)\), and Gemma 3 uses SigLIP\([Gemma Team 2025](https://arxiv.org/html/2608.17253#bib.bib23)\)\. Appendix[A](https://arxiv.org/html/2608.17253#A1)further quantifies this diversity through the error\-overlap ratios across model families\. As shown in Figure[2](https://arxiv.org/html/2608.17253#S4.F2)\(d\), models from different families exhibit substantially less overlap in their errors\. This complementary error structure is reflected by lower inter\-model agreement during pseudo\-labeling \(Figure[2](https://arxiv.org/html/2608.17253#S4.F2)\(a\)\) and, in turn, yields more accurate pseudo\-labels \(Figure[2](https://arxiv.org/html/2608.17253#S4.F2)\(b\)\) and higher performance \(Figure[2](https://arxiv.org/html/2608.17253#S4.F2)\(c\)\) throughout training\.
Figure 2:\(a\) Agreement between the two models, \(b\) pseudo\-label accuracy, and \(c\) evaluation performance; \(d\) Error overlap before RL for two pairs each from a different family, the same family, and the same model under a different seed\.Model size provides an additional, orthogonal source of diversity\. Models with different capacities exhibit distinct reasoning and prediction behaviors\. Pairing models of different sizes therefore yields different error profiles, allowing each policy to provide informative reward signals on examples that the other does not solve correctly\. The three\-agent run in Section[6\.2](https://arxiv.org/html/2608.17253#S6.SS2)trains models of different sizes together\.
Diversity across input formation\.Despite model\- and optimization\-level decoupling, both agents are still trained on identical prompts\. We further diversify their training data by rewriting each MATH problem with DeepSeek\-V3\([DeepSeek\-AI 2024](https://arxiv.org/html/2608.17253#bib.bib16)\), training one agent on the original prompt and the other on its rewrite\. The rewrite preserves the answer and sample order while typically recasting the problem into a different concrete scenario rather than performing superficial lexical substitutions\. The two agents therefore solve semantically equivalent problems expressed in different forms, reducing correlated errors induced by prompt\-specific phrasing\. Appendix[C](https://arxiv.org/html/2608.17253#A3)provides examples\.
We provide an illustration of diversifiedCo\-RLin Figure[3](https://arxiv.org/html/2608.17253#S4.F3)\.
Figure 3:Overview ofCo\-RLwith two agents\. Each agent samplesKKresponses to the same unlabeled question and generates a pseudo label with majority vote\. Each rollout is then rewarded by agreement with the cohort’s pseudo label, and updates its own policy, with no sharing parameters and no gradient exchange except the cross\-reward process\.
## 5Theoretical Analysis
In this section, we theoretically characterize the learning dynamics ofCo\-RLand compare them with self\-rewarding GRPO, where each agent uses its own majority vote as the pseudo\-label\. A key limitation of self\-rewarding is that an agent learns from a pseudo\-label derived from its own predictions\. As a result, when the agent is systematically wrong, its supervision signal is likely to reinforce the same error rather than correct it\. In contrast,Co\-RLderives supervision from other agents, allowing one agent’s correct prediction to provide a corrective signal for another agent’s mistake\. Specifically, we ignore clipping and the KL term for simplicity\. Our analysis shows thatCo\-RLcan exploit complementary strengths across agents to correct errors that self\-rewarding would otherwise reinforce, thereby expanding the set of prompts that converge to the correct answer\.
To obtain a tractable characterization, we consider a fixed promptxxand reduce the induced answer distribution to two outcomes: the correct answera⋆a^\{\\star\}and an aggregate incorrect answer\. Letpn=Prπθn\(a=a⋆∣x\)p\_\{n\}=\\Pr\_\{\\pi\_\{\\theta\_\{n\}\}\}\(a=a^\{\\star\}\\mid x\)denote theprobability that agentnnassigns to the correct answer; in the single\-agent setting, we simply writepp\. We assume an odd numberKKof independently sampled rollouts, so that the majority vote has no ties, and letη\\etadenote the GRPO update rate\. Our experiments use an evenKK, where a tie is resolved deterministically rather than discarded\.
### 5\.1Comparison of training dynamics under self\-rewarding and cross\-agent supervision\.
We study how the probability of generating the correct answerppevolves during training\.
###### Proposition 1\(Training dynamics under self\-rewarding and cross\-agent supervision\)\.
Under the above definitions, the probability dynamics take the following forms\.
\[1\] Self\-rewarding\.LetC∼Bin\(K,p\)C\\sim\\operatorname\{Bin\}\(K,p\)denote the number of correct responses among theKKrollouts\. The probability dynamics satisfy
p˙=ηp\(1−p\)𝔼C∼Bin\(K,p\)\[sign\(C−K2\)C\(K−C\)K\]\.\\dot\{p\}=\\eta\\,p\(1\-p\)\\mathbb\{E\}\_\{C\\sim\\operatorname\{Bin\}\(K,p\)\}\\left\[\\operatorname\{sign\}\\\!\\left\(C\-\\frac\{K\}\{2\}\\right\)\\frac\{\\sqrt\{C\(K\-C\)\}\}\{K\}\\right\]\.\(7\)
\[2\]Co\-RL: Cross\-agent supervision\.For agentnn, letCn∼Bin\(K,pn\)C\_\{n\}\\sim\\operatorname\{Bin\}\(K,p\_\{n\}\)denote the number of correct responses among itsKKrollouts, and letZ−n=𝟏\[a^−n\(x\)=a⋆\]Z\_\{\-n\}=\\mathbf\{1\}\[\\hat\{a\}\_\{\-n\}\(x\)=a^\{\\star\}\]indicate whether the pseudo\-label constructed from the other agents is correct\. The probability dynamics satisfy
p˙n=ηnpn\(1−pn\)𝔼\[sign\(Z−n−12\)\]𝔼Cn∼Bin\(K,pn\)\[Cn\(K−Cn\)K\]\.\\dot\{p\}\_\{n\}=\\eta\_\{n\}p\_\{n\}\(1\-p\_\{n\}\)\\mathbb\{E\}\\left\[\\operatorname\{sign\}\\\!\\left\(Z\_\{\-n\}\-\\frac\{1\}\{2\}\\right\)\\right\]\\mathbb\{E\}\_\{C\_\{n\}\\sim\\operatorname\{Bin\}\(K,p\_\{n\}\)\}\\left\[\\frac\{\\sqrt\{C\_\{n\}\(K\-C\_\{n\}\)\}\}\{K\}\\right\]\.\(8\)
In the two\-agent setting, defineϕK\(p\)=2VK\(p\)−1\\phi\_\{K\}\(p\)=2V\_\{K\}\(p\)\-1as the signed majority\-vote direction andqK\(p\)=ηp\(1−p\)𝔼C∼Bin\(K,p\)\[C\(K−C\)K\]\>0q\_\{K\}\(p\)=\\eta p\(1\-p\)\\mathbb\{E\}\_\{C\\sim\\operatorname\{Bin\}\(K,p\)\}\\left\[\\frac\{\\sqrt\{C\(K\-C\)\}\}\{K\}\\right\]\>0as the positive update magnitude\. Then Eq\. \([8](https://arxiv.org/html/2608.17253#S5.E8)\) reduces to
p˙A=qK\(pA\)ϕK\(pB\),p˙B=qK\(pB\)ϕK\(pA\)\.\\dot\{p\}\_\{A\}=q\_\{K\}\(p\_\{A\}\)\\phi\_\{K\}\(p\_\{B\}\),\\qquad\\dot\{p\}\_\{B\}=q\_\{K\}\(p\_\{B\}\)\\phi\_\{K\}\(p\_\{A\}\)\.\(9\)
###### Proof\.
A complete proof is provided in Appendix[B\.1](https://arxiv.org/html/2608.17253#A2.SS1)\. ∎
Proposition[1](https://arxiv.org/html/2608.17253#Thmproposition1)highlights the structural difference between the two training mechanisms\. Under self\-rewarding, the pseudo\-label is constructed from the same rollout group being optimized, making the update*self\-confirming*: so the agent reinforces whichever answer is currently more likely, whether correct or not\. In contrast,Co\-RLdecouples supervision from the optimized agent: sinceqK\(pA\)\>0q\_\{K\}\(p\_\{A\}\)\>0, the update direction of agentAAis determined entirely byϕK\(pB\)\\phi\_\{K\}\(p\_\{B\}\), the supervision signal provided by agentBB\.
### 5\.2Cross\-agent supervision enlarges the basin of correct convergence\.
Following Proposition[1](https://arxiv.org/html/2608.17253#Thmproposition1), Proposition[2](https://arxiv.org/html/2608.17253#Thmproposition2)shows that self\-rewarding is*self\-confirming*: the expected GRPO update amplifies the currently favored answer, regardless of its correctness\. Thus, whenp<1/2p<1/2, self\-rewarding further suppresses the correct answer instead of correcting the error\.
###### Proposition 2\(Self\-confirming dynamics\)\.
For oddKK,sign\(GKself\(p\)\)=sign\(p−12\)\\operatorname\{sign\}\(G\_\{K\}^\{\\mathrm\{self\}\}\(p\)\)=\\operatorname\{sign\}\(p\-\\frac\{1\}\{2\}\)\. Consequently,p\(0\)<12⇒p\(t\)→0p\(0\)<\\frac\{1\}\{2\}\\Rightarrow p\(t\)\\rightarrow 0andp\(0\)\>12⇒p\(t\)→1p\(0\)\>\\frac\{1\}\{2\}\\Rightarrow p\(t\)\\rightarrow 1\.
###### Proof\.
A complete proof is provided in Appendix[B\.2](https://arxiv.org/html/2608.17253#A2.SS2)\. ∎
We next state our main result in Theorem[1](https://arxiv.org/html/2608.17253#Thmtheorem1), which characterizes the basin of attraction under cross\-agent supervision\.
###### Theorem 1\(Co\-RL enlarges the basin of correct convergence\)\.
Consider the symmetric two\-agent dynamics in Eq\. \([9](https://arxiv.org/html/2608.17253#S5.E9)\) with interior initialization\(pA\(0\),pB\(0\)\)∈\(0,1\)2\(p\_\{A\}\(0\),p\_\{B\}\(0\)\)\\in\(0,1\)^\{2\}\. The correct and incorrect consensus states\(1,1\)\(1,1\)and\(0,0\)\(0,0\)are asymptotically stable, while\(1/2,1/2\)\(1/2,1/2\)is a saddle point whose interior separatrix ispA\+pB=1p\_\{A\}\+p\_\{B\}=1\. Consequently,
pA\(0\)\+pB\(0\)\>1⟹\(pA\(t\),pB\(t\)\)→\(1,1\),p\_\{A\}\(0\)\+p\_\{B\}\(0\)\>1\\quad\\Longrightarrow\\quad\(p\_\{A\}\(t\),p\_\{B\}\(t\)\)\\rightarrow\(1,1\),whereas
pA\(0\)\+pB\(0\)<1⟹\(pA\(t\),pB\(t\)\)→\(0,0\)\.p\_\{A\}\(0\)\+p\_\{B\}\(0\)<1\\quad\\Longrightarrow\\quad\(p\_\{A\}\(t\),p\_\{B\}\(t\)\)\\rightarrow\(0,0\)\.
###### Proof\.
A complete proof is provided in Appendix[B\.3](https://arxiv.org/html/2608.17253#A2.SS3)\. ∎
Takeaway:Co\-RLleverages complementary strengths to expand correct convergence\.Consider an extreme example where\(pA,pB\)=\(0\.9,0\.2\)\(p\_\{A\},p\_\{B\}\)=\(0\.9,0\.2\)on half of the prompts and\(pA,pB\)=\(0\.2,0\.9\)\(p\_\{A\},p\_\{B\}\)=\(0\.2,0\.9\)on the other half\. Each agent has average accuracy0\.550\.55, but their strengths are perfectly complementary\. Under self\-rewarding, each agent succeeds only on the half of prompts wherep\>1/2p\>1/2, yielding a final accuracy of0\.50\.5\. In contrast, underCo\-RL,pA\+pB=1\.1\>1p\_\{A\}\+p\_\{B\}=1\.1\>1on every prompt, so Theorem[1](https://arxiv.org/html/2608.17253#Thmtheorem1)predicts that both agents converge to the correct answer on all prompts, achieving accuracy1\.01\.0\. Thus, cross\-agent supervision can exploit complementary expertise to correct errors that self\-rewarding would otherwise reinforce\.
## 6Experiments
### 6\.1Experimental Setup
Datasets\.For the main experiments, language models are trained on the level 3 to 5 split of MATH\([Hendrycks et al\. 2021b](https://arxiv.org/html/2608.17253#bib.bib27)\), following MARTI\([Zhang et al\. 2026a](https://arxiv.org/html/2608.17253#bib.bib77)\)\. Vision\-language models are trained on two multimodal math datasets\. We use MMR1\-Math\([Leng et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib35)\)to match the training data of our baseline, MM\-UPT\([Wei et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib64)\), and additionally use multimodal\-open\-r1\([LMMs\-Lab 2025](https://arxiv.org/html/2608.17253#bib.bib44)\)to verify that the observed gains are not specific to a particular training dataset\.
Models\.For language models, we consider Qwen3\-1\.7B\([Yang et al\. 2025a](https://arxiv.org/html/2608.17253#bib.bib71)\), Qwen2\.5\-3B paired with Llama\-3\.2\-3B\-Instruct, and Qwen2\.5\-7B paired with Llama\-3\.1\-8B\-Instruct\([Yang et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib70);[Grattafiori et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib24)\)\. For vision\-language models, we use Qwen2\.5\-VL \(3B, 7B\), InternVL3\.5 \(2B, 8B\), and Gemma\-3 \(4B, 12B\)\([Bai et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib2);[Chen et al\. 2024b](https://arxiv.org/html/2608.17253#bib.bib12);[Gemma Team 2025](https://arxiv.org/html/2608.17253#bib.bib23)\), with models paired at comparable scales\.
Baselines\.On language models, we compare our method against state\-of\-the\-art label\-free self\-rewarding methods, including TTRL\([Zuo et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib88)\), Intuitor\([Zhao et al\. 2026](https://arxiv.org/html/2608.17253#bib.bib87)\), RENT\([Prabhudesai et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib51)\)and Co\-rewarding\-II\([Zhang et al\. 2026b](https://arxiv.org/html/2608.17253#bib.bib84)\)\. On vision\-language models, we compare with the recent self\-rewarding method MM\-UPT\([Wei et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib64)\), which adopts the TTRL reward formulation\. We also include GRPO with ground\-truth rewards \(GT\-Reward\) as a supervised reference\([Shao et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib55)\)\. For multi\-agent RL baselines, we compare against MAPoRL\([Park et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib49)\)and CoMAS\([Xue et al\. 2026](https://arxiv.org/html/2608.17253#bib.bib69)\)\.
Training details\.All runs use the AdamW optimizer and sample rollouts at a temperature of 1\.0\. Following Co\-rewarding\([Zhang et al\. 2026b](https://arxiv.org/html/2608.17253#bib.bib84)\), language models use a learning rate of3×10−63\\times 10^\{\-6\}, an effective batch of 128 prompts per agent and a 3072\-token cap, and train withK=12K=12responses per prompt for 2 epochs\. Following R1\-V\([Chen et al\. 2025b](https://arxiv.org/html/2608.17253#bib.bib7)\), vision\-language models use a learning rate of1×10−61\\times 10^\{\-6\}, a 1024\-token cap andK=8K=8responses per prompt for 1 epoch\. Because the rollout and policy distributions drift apart on Gemma\-3, the vision\-language runs additionally apply a token\-level importance\-sampling correction\. All experiments use a single node of eight H100 GPUs, four per agent\.
Evaluation\.Language models are scored on seven reasoning benchmarks\. For mathematics, we use GSM8K\([Cobbe et al\. 2021](https://arxiv.org/html/2608.17253#bib.bib14)\), MATH\-500\([Lightman et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib40)\)and AMC\([Project\-Numina 2024](https://arxiv.org/html/2608.17253#bib.bib52)\)\. For code generation, we use HumanEval\([Chen et al\. 2021](https://arxiv.org/html/2608.17253#bib.bib8)\), MBPP\([Austin et al\. 2021](https://arxiv.org/html/2608.17253#bib.bib1)\)and LiveCodeBench\([Jain et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib31)\)\. For science, we use GPQA\([Rein et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib54)\)\. Vision\-language models are scored on four multimodal math benchmarks, MathVision\([Wang et al\. 2024a](https://arxiv.org/html/2608.17253#bib.bib59)\), MathVerse\([Zhang et al\. 2024a](https://arxiv.org/html/2608.17253#bib.bib79)\), MathVista\([Lu et al\. 2024](https://arxiv.org/html/2608.17253#bib.bib45)\)and We\-Math\([Qiao et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib53)\)\. Detailed experimental settings and the complete results are provided in Appendix[D](https://arxiv.org/html/2608.17253#A4)\.
### 6\.2Main Results
Co\-RLoutperforms all self\-rewarding methods on language models\.We first compareCo\-RLagainst single\-agent self\-rewarding baselines\. Table[1](https://arxiv.org/html/2608.17253#S6.T1)reports results against TTRL, RENT, Intuitor, and Co\-Rewarding\-II, with GRPO using ground\-truth rewards included as a supervised reference\. Appendix[D\.1](https://arxiv.org/html/2608.17253#A4.SS1)extends the same comparison to 7B and 8B models\. We evaluate three variants of our framework: Same family jointly trains two independent agents initialized from the same base model; Different family pairs one agent from each of the two model families; Different family\+ further applies data decoupling\. Notably, the readily accessible Same family setting already yields substantial gains over the base models, improving the average performance by 8\.0% and 4\.0% for Qwen2\.5\-3B and Llama\-3\.2\-3B\-Instruct, respectively\. Introducing cross\-family diversity provides further benefits overall, and, when combined with data decoupling \(Different family\+\), achieves the strongest label\-free average performance for both model families\.
Table 1:Full performance across seven benchmarks for 3B models \(%\)\. For each benchmark, the best label\-free result is shown inboldand the second best isunderlined, with ties sharing the marking\. Base and GT\-Reward serve as references and are excluded from the ranking\.Co\-RL\(Same family\) trains two agents initialized from the same base model\.Co\-RL\(Different family\) pairs one agent from each of the two families\.Co\-RL\(Different family\+\) further decouples the training data\. Appendix[D\.1](https://arxiv.org/html/2608.17253#A4.SS1)extends the comparison to 7B and 8B models\.Co\-RLoutperforms Multi\-Agent RL\.We next compareCo\-RLwith existing multi\-agent RL methods\. While these approaches also train multiple models jointly, existing state\-of\-the\-art methods typically rely on an additional judging mechanism to construct rewards, such as an LLM judge in CoMAS or a learned reward model in MAPoRL\. Following CoMAS, we adopt the same experimental setup, official implementation, and evaluation benchmarks to ensure a fair comparison\. We report the CoMAS results as presented in the original paper\. As shown in Table[2](https://arxiv.org/html/2608.17253#S6.T2),Co\-RLachieves the best average performance and leads on five of the seven benchmarks, outperforming CoMAS by 4\.0% while using only half as many agents and requiring no additional judging mechanism\.
ScalingCo\-RLto Three Agents\.We further extendCo\-RLbeyond the two\-agent setting by jointly training Qwen2\.5\-3B, Llama\-3\.2\-3B\-Instruct, and Qwen3\-1\.7B in a single run\. For each agent, we compareCo\-RLagainst the same model trained independently using ground\-truth rewards \(GT\-Reward\) or self\-generated majority\-vote rewards \(TTRL\)\. Results are reported in Table[3](https://arxiv.org/html/2608.17253#S6.T3)\.Co\-RLconsistently improves all three base models, with average gains of 7\.8%, 6\.0%, and 8\.2%, respectively\. Despite using no ground\-truth supervision, three\-agentCo\-RLmatches or outperforms GT\-Reward in average performance for all three models, while also outperforming TTRL for Qwen2\.5\-3B and Llama\-3\.2\-3B\-Instruct\. These results suggest thatCo\-RLnaturally extends beyond pairwise training, allowing multiple heterogeneous agents to benefit from cross\-agent supervision within a shared training run\.
Table 2:Comparison under the CoMAS multi\-agent RL setting \(%\)\. All methods train Qwen2\.5\-3B\-Instruct on the same prompt mixture and are evaluated following the CoMAS protocol\. Results for prior methods are reported from[Xue et al\. 2026](https://arxiv.org/html/2608.17253#bib.bib69)\.Co\-RLTransfers to Vision\-Language Models\.Vision\-language model families differ in both their visual encoders and language backbones, providing a stronger test ofCo\-RLunder heterogeneous architectures\. We therefore extendCo\-RLto multimodal mathematical reasoning\. Table[4](https://arxiv.org/html/2608.17253#S6.T4)reports results for Qwen2\.5\-VL\-3B and InternVL3\.5\-2B, while Appendix[D\.2](https://arxiv.org/html/2608.17253#A4.SS2)extends the evaluation to three model families from 7B to 12B\.Co\-RLachieves the best average performance in three of the four 2B\-3B settings and remains competitive with ground\-truth supervision\. The gains persist at larger scales, whereCo\-RLconsistently outperforms TTRL and even surpasses GT\-Reward for Gemma\-3\-12B, demonstrating thatCo\-RLgeneralizes beyond text\-only models\.
Table 3:Three\-agentCo\-RLwith heterogeneous model families \(%\)\. Qwen2\.5\-3B, Llama\-3\.2\-3B\-Instruct, and Qwen3\-1\.7B are jointly trained in a singleCo\-RLrun\. For each model, we compare against the base model, training with ground\-truth rewards \(GT\-Reward\), and self\-rewarding with majority\-vote pseudo\-labels \(TTRL\)\.Table 4:Vision\-language results for the small pair, Qwen2\.5\-VL\-3B with InternVL3\.5\-2B, trained separately on open\-r1 and MMR1 \(%\)\. Base is graded once with the corrected multiple\-choice grader and is therefore identical across the two training sets\. Base and GT\-Reward serve as references and are excluded from the ranking\.Figure 4:Training dynamics at four scales, one column per backbone \(Qwen2\.5\-3B, Llama\-3\.2\-3B, Qwen2\.5\-7B, Llama\-3\.1\-8B\)\. \(a\) MATH\-500 validation accuracy, \(b\) standard deviation of the reward within a rollout group, normalized to its value at the first step, and \(c\) mean completion length\. Runs marked diverged leave the plotted range\.
## 7Ablation Study
#### Training dynamics and stability\.
We further examine the training dynamics ofCo\-RLin Appendix[D\.3](https://arxiv.org/html/2608.17253#A4.SS3)\. As shown in Figure[4](https://arxiv.org/html/2608.17253#S6.F4), across text models,Co\-RLmaintains stable reward variation and completion lengths, whereas self\-rewarding baselines can exhibit reward collapse, length degeneration, or divergence\. We observe a similar pattern for VLMs: the agents retain partial agreement while the accuracy of their exchanged pseudo\-labels improves throughout training\. These results are consistent with our theory: by decoupling an agent’s update from its own predictions, cross\-agent supervision avoids self\-reinforcing errors and preserves an informative learning signal throughout training\.
#### Controlling for training and inference budgets\.
To match the two\-agent training budget ofCo\-RL, we construct a self\-rewarding baseline with the same two base models\. Each model is trained independently with TTRL\. At inference, bothCo\-RLand TTRL ensemble the two models by pooling four rollouts from each for majority voting, thereby matching both training and test\-time budgets\. As shown in Appendix[D\.4](https://arxiv.org/html/2608.17253#A4.SS4),Co\-RLconsistently achieves the best average score across text and multimodal settings, showing that the gains come from cross\-agent supervision rather than additional compute or ensembling alone\.
## 8Conclusion
In this work, we introducedCo\-RL, a label\-free multi\-agent RL framework for reasoning tasks\. In our framework, multiple agents learn from rewards constructed from their peers’ predictions rather than ground\-truth labels or external judges\. Across text\-only and multimodal reasoning benchmarks,Co\-RLconsistently improves diverse LLMs and VLMs, outperforming prior self\-rewarding and multi\-agent RL approaches and, in many settings, matching or surpassing training with ground\-truth rewards\. Our theoretical analysis shows that cross\-agent supervision expands the set of initial conditions that converge to the correct solution, allowingCo\-RLto correct errors that self\-rewarding RL would otherwise reinforce\. An important direction for future work is to understand how the number, diversity, and interaction topology of agents shape cross\-agent learning, and to develop adaptive supervision mechanisms that more effectively exploit complementary expertise\.
## References
- Austin et al\. \[2021\]Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton\.Program synthesis with large language models\.*arXiv preprint arXiv:2108\.07732*, 2021\.
- Bai et al\. \[2025\]Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al\.Qwen2\.5\-VL technical report\.*arXiv preprint arXiv:2502\.13923*, 2025\.
- Blum and Mitchell \[1998\]Avrim Blum and Tom Mitchell\.Combining labeled and unlabeled data with co\-training\.In*Conference on Computational Learning Theory \(COLT\)*, pages 92–100, 1998\.
- Caron et al\. \[2021\]Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin\.Emerging properties in self\-supervised vision transformers\.In*IEEE/CVF International Conference on Computer Vision \(ICCV\)*, pages 9630–9640, 2021\.
- Chen et al\. \[2025a\]Hardy Chen, Haoqin Tu, Fali Wang, Hui Liu, Xianfeng Tang, Xinya Du, Yuyin Zhou, and Cihang Xie\.SFT or RL? an early investigation into training R1\-like reasoning large vision\-language models\.*arXiv preprint arXiv:2504\.11468*, 2025a\.
- Chen et al\. \[2024a\]Justin Chih\-Yao Chen, Swarnadeep Saha, and Mohit Bansal\.ReConcile: Round\-table conference improves reasoning via consensus among diverse LLMs\.In*Annual Meeting of the Association for Computational Linguistics \(ACL\)*, pages 7066–7085, 2024a\.
- Chen et al\. \[2025b\]Liang Chen, Lei Li, Haozhe Zhao, Yifan Song, and Vinci\.R1\-V: Reinforcing super generalization ability in vision\-language models with less than $3\.[https://github\.com/Deep\-Agent/R1\-V](https://github.com/Deep-Agent/R1-V), 2025b\.Accessed: 2026\-08\-17\.
- Chen et al\. \[2021\]Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al\.Evaluating large language models trained on code\.*arXiv preprint arXiv:2107\.03374*, 2021\.
- Chen et al\. \[2020\]Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton\.A simple framework for contrastive learning of visual representations\.In*International Conference on Machine Learning \(ICML\)*, pages 1597–1607, 2020\.
- Chen et al\. \[2025c\]Weize Chen, Jiarui Yuan, Chen Qian, Cheng Yang, Zhiyuan Liu, and Maosong Sun\.Optima: Optimizing effectiveness and efficiency for LLM\-based multi\-agent system\.In*Findings of the Association for Computational Linguistics: ACL 2025*, pages 11534–11557, 2025c\.
- Chen et al\. \[2025d\]Yixing Chen, Yiding Wang, Siqi Zhu, Haofei Yu, Tao Feng, Muhan Zhang, Mostofa Patwary, and Jiaxuan You\.Multi\-agent evolve: LLM self\-improve through co\-evolution\.*arXiv preprint arXiv:2510\.23595*, 2025d\.
- Chen et al\. \[2024b\]Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al\.InternVL: Scaling up vision foundation models and aligning for generic visual\-linguistic tasks\.In*IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\)*, pages 24185–24198, 2024b\.
- Chen et al\. \[2024c\]Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu\.Self\-play fine\-tuning converts weak language models to strong language models\.In*International Conference on Machine Learning \(ICML\)*, 2024c\.
- Cobbe et al\. \[2021\]Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman\.Training verifiers to solve math word problems\.*arXiv preprint arXiv:2110\.14168*, 2021\.
- Cohen \[1960\]Jacob Cohen\.A coefficient of agreement for nominal scales\.*Educational and Psychological Measurement*, 20\(1\):37–46, 1960\.
- DeepSeek\-AI \[2024\]DeepSeek\-AI\.DeepSeek\-V3 technical report\.*arXiv preprint arXiv:2412\.19437*, 2024\.
- DeepSeek\-AI \[2025\]DeepSeek\-AI\.DeepSeek\-R1: Incentivizing reasoning capability in LLMs via reinforcement learning\.*arXiv preprint arXiv:2501\.12948*, 2025\.
- Deng et al\. \[2025a\]Huilin Deng, Ding Zou, Rui Ma, Hongchen Luo, Yang Cao, and Yu Kang\.Boosting the generalization and reasoning of vision language models with curriculum reinforcement learning\.*arXiv preprint arXiv:2503\.07065*, 2025a\.
- Deng et al\. \[2025b\]Yihe Deng, Hritik Bansal, Fan Yin, Nanyun Peng, Wei Wang, and Kai\-Wei Chang\.OpenVLThinker: Complex vision\-language reasoning via iterative SFT\-RL cycles\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2025b\.
- Du et al\. \[2024\]Yilun Du, Shuang Li, Antonio Torralba, Joshua B\. Tenenbaum, and Igor Mordatch\.Improving factuality and reasoning in language models through multiagent debate\.In*International Conference on Machine Learning \(ICML\)*, 2024\.
- Fang et al\. \[2025\]Wenkai Fang, Shunyu Liu, Yang Zhou, Kongcheng Zhang, Tongya Zheng, Kaixuan Chen, Mingli Song, and Dacheng Tao\.SeRL: Self\-play reinforcement learning for large language models with limited data\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2025\.
- Feng et al\. \[2025\]Yicheng Feng, Yijiang Li, Wanpeng Zhang, Sipeng Zheng, Hao Luo, Zihao Yue, and Zongqing Lu\.VideoOrion: Tokenizing object dynamics in videos\.In*IEEE/CVF International Conference on Computer Vision \(ICCV\)*, pages 20401–20412, 2025\.
- Gemma Team \[2025\]Gemma Team\.Gemma 3 technical report\.*arXiv preprint arXiv:2503\.19786*, 2025\.
- Grattafiori et al\. \[2024\]Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al\-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al\.The Llama 3 herd of models\.*arXiv preprint arXiv:2407\.21783*, 2024\.
- Grill et al\. \[2020\]Jean\-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al\.Bootstrap your own latent: A new approach to self\-supervised learning\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, pages 21271–21284, 2020\.
- Hendrycks et al\. \[2021a\]Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt\.Measuring massive multitask language understanding\.In*International Conference on Learning Representations \(ICLR\)*, 2021a\.
- Hendrycks et al\. \[2021b\]Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt\.Measuring mathematical problem solving with the MATH dataset\.In*Advances in Neural Information Processing Systems \(NeurIPS\), Datasets and Benchmarks Track*, 2021b\.
- Hu et al\. \[2025\]Jian Hu, Jason Klein Liu, Haotian Xu, and Wei Shen\.REINFORCE\+\+: Stabilizing critic\-free policy optimization with global advantage normalization\.*arXiv preprint arXiv:2501\.03262*, 2025\.
- Huang et al\. \[2026a\]Chengsong Huang, Wenhao Yu, Xiaoyang Wang, Hongming Zhang, Zongxia Li, Ruosen Li, Jiaxin Huang, Haitao Mi, and Dong Yu\.R\-Zero: Self\-evolving reasoning LLM from zero data\.In*International Conference on Learning Representations \(ICLR\)*, 2026a\.
- Huang et al\. \[2026b\]Wenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye, Zhe Xu, Yao Hu, Shaohui Lin, et al\.Vision\-R1: Incentivizing reasoning capability in multimodal large language models\.In*International Conference on Learning Representations \(ICLR\)*, 2026b\.
- Jain et al\. \[2025\]Naman Jain, King Han, Alex Gu, Wen\-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar\-Lezama, Koushik Sen, and Ion Stoica\.LiveCodeBench: Holistic and contamination free evaluation of large language models for code\.In*International Conference on Learning Representations \(ICLR\)*, 2025\.
- Kim et al\. \[2024\]Yubin Kim, Chanwoo Park, Hyewon Jeong, Yik S\. Chan, Xuhai Xu, Daniel McDuff, Hyeonhoon Lee, Marzyeh Ghassemi, Cynthia Breazeal, and Hae W\. Park\.MDAgents: An adaptive collaboration of LLMs for medical decision\-making\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2024\.
- Krogh and Vedelsby \[1994\]Anders Krogh and Jesper Vedelsby\.Neural network ensembles, cross validation, and active learning\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, volume 7, 1994\.
- Kuncheva and Whitaker \[2003\]Ludmila I\. Kuncheva and Christopher J\. Whitaker\.Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy\.*Machine Learning*, 51\(2\):181–207, 2003\.
- Leng et al\. \[2025\]Sicong Leng, Jing Wang, Jiaxi Li, Hao Zhang, Zhiqiang Hu, Boqiang Zhang, Yuming Jiang, Hang Zhang, Xin Li, Lidong Bing, et al\.MMR1: Enhancing multimodal reasoning with variance\-aware sampling and open resources\.*arXiv preprint arXiv:2509\.21268*, 2025\.
- Li et al\. \[2025\]Pengyi Li, Matvey Skripkin, Alexander Zubrey, Andrey Kuznetsov, and Ivan Oseledets\.Confidence is all you need: Few\-shot RL fine\-tuning of language models\.*arXiv preprint arXiv:2506\.06395*, 2025\.
- Li et al\. \[2023\]Yijiang Li, Xinjiang Wang, Lihe Yang, Litong Feng, Wayne Zhang, and Ying Gao\.Diverse cotraining makes strong semi\-supervised segmentor\.In*IEEE/CVF International Conference on Computer Vision \(ICCV\)*, 2023\.
- Liang et al\. \[2024\]Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, and Zhaopeng Tu\.Encouraging divergent thinking in large language models through multi\-agent debate\.In*Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*, 2024\.
- Liao et al\. \[2025\]Junwei Liao, Muning Wen, Jun Wang, and Weinan Zhang\.MARFT: Multi\-agent reinforcement fine\-tuning\.*arXiv preprint arXiv:2504\.16129*, 2025\.
- Lightman et al\. \[2024\]Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe\.Let’s verify step by step\.In*International Conference on Learning Representations \(ICLR\)*, 2024\.
- Liu et al\. \[2026\]Bo Liu, Simon Yu, Zichen Liu, Leon Guertler, Penghui Qi, Daniel Balcells, Mickel Liu, Cheston Tan, Weiyan Shi, Min Lin, et al\.SPIRAL: Self\-play on zero\-sum games incentivizes reasoning via multi\-agent multi\-turn reinforcement learning\.In*International Conference on Learning Representations \(ICLR\)*, 2026\.
- Liu et al\. \[2025a\]Xiangyan Liu, Jinjie Ni, Zijian Wu, Chao Du, Longxu Dou, Haonan Wang, Tianyu Pang, and Michael Shieh\.NoisyRollout: Reinforcing visual reasoning with data augmentation\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2025a\.
- Liu et al\. \[2025b\]Ziyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, and Jiaqi Wang\.Visual\-RFT: Visual reinforcement fine\-tuning\.In*IEEE/CVF International Conference on Computer Vision \(ICCV\)*, pages 2034–2044, 2025b\.
- LMMs\-Lab \[2025\]LMMs\-Lab\.Multimodal open R1\.[https://huggingface\.co/datasets/lmms\-lab/multimodal\-open\-r1\-8k\-verified](https://huggingface.co/datasets/lmms-lab/multimodal-open-r1-8k-verified), 2025\.Accessed: 2026\-08\-17\.
- Lu et al\. \[2024\]Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai\-Wei Chang, Michel Galley, and Jianfeng Gao\.MathVista: Evaluating mathematical reasoning of foundation models in visual contexts\.In*International Conference on Learning Representations \(ICLR\)*, 2024\.
- Ma et al\. \[2025\]Xueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang, Zejun Ma, and Wenhu Chen\.General\-Reasoner: Advancing LLM reasoning across all domains\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2025\.
- Meng et al\. \[2025\]Fanqing Meng, Lingxiao Du, Zongkai Liu, Zhixiang Zhou, Quanfeng Lu, Daocheng Fu, Tiancheng Han, Botian Shi, Wenhai Wang, Junjun He, et al\.MM\-Eureka: Exploring the frontiers of multimodal reasoning with rule\-based reinforcement learning\.*arXiv preprint arXiv:2503\.07365*, 2025\.
- Motwani et al\. \[2024\]Sumeet Ramesh Motwani, Chandler Smith, Rocktim Jyoti Das, Rafael Rafailov, Ivan Laptev, Philip H\. S\. Torr, Fabio Pizzati, Ronald Clark, and Christian Schroeder de Witt\.MALT: Improving reasoning with multi\-agent LLM training\.*arXiv preprint arXiv:2412\.01928*, 2024\.
- Park et al\. \[2025\]Chanwoo Park, Seungju Han, Xingzhi Guo, Asuman E\. Ozdaglar, Kaiqing Zhang, and Joo\-Kyung Kim\.MAPoRL: Multi\-agent post\-co\-training for collaborative large language models with reinforcement learning\.In*Annual Meeting of the Association for Computational Linguistics \(ACL\)*, 2025\.
- Peng et al\. \[2025\]Yingzhe Peng, Gongrui Zhang, Miaosen Zhang, Zhiyuan You, Jie Liu, Qipeng Zhu, Kai Yang, Xingzhong Xu, Xin Geng, and Xu Yang\.LMM\-R1: Empowering 3B LMMs with strong reasoning abilities through two\-stage rule\-based RL\.*arXiv preprint arXiv:2503\.07536*, 2025\.
- Prabhudesai et al\. \[2025\]Mihir Prabhudesai, Lili Chen, Alex Ippoliti, Katerina Fragkiadaki, Hao Liu, and Deepak Pathak\.Maximizing confidence alone improves reasoning\.*arXiv preprint arXiv:2505\.22660*, 2025\.
- Project\-Numina \[2024\]Project\-Numina\.AIMO validation AMC\.[https://huggingface\.co/datasets/AI\-MO/aimo\-validation\-amc](https://huggingface.co/datasets/AI-MO/aimo-validation-amc), 2024\.Accessed: 2026\-08\-17\.
- Qiao et al\. \[2025\]Runqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu, Chong Sun, Xiaoshuai Song, Jiapeng Wang, Zhuoma Gongque, Shanglin Lei, Yifan Zhang, et al\.We\-Math: Does your large multimodal model achieve human\-like mathematical reasoning?In*Annual Meeting of the Association for Computational Linguistics \(ACL\)*, 2025\.
- Rein et al\. \[2024\]David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R\. Bowman\.GPQA: A graduate\-level google\-proof Q&A benchmark\.In*Conference on Language Modeling \(COLM\)*, 2024\.
- Shao et al\. \[2024\]Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y\. K\. Li, Yang Wu, et al\.DeepSeekMath: Pushing the limits of mathematical reasoning in open language models\.*arXiv preprint arXiv:2402\.03300*, 2024\.
- Shen et al\. \[2025\]Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqian Zhang, et al\.VLM\-R1: A stable and generalizable R1\-style large vision\-language model\.*arXiv preprint arXiv:2504\.07615*, 2025\.
- Sun et al\. \[2024\]Qiushi Sun, Zhangyue Yin, Xiang Li, Zhiyong Wu, Xipeng Qiu, and Lingpeng Kong\.Corex: Pushing the boundaries of complex reasoning through multi\-model collaboration\.In*Conference on Language Modeling \(COLM\)*, 2024\.
- Wang et al\. \[2025a\]Haozhe Wang, Chao Qu, Zuming Huang, Wei Chu, Fangzhen Lin, and Wenhu Chen\.VL\-Rethinker: Incentivizing self\-reflection of vision\-language models with reinforcement learning\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2025a\.
- Wang et al\. \[2024a\]Ke Wang, Junting Pan, Weikang Shi, Zimu Lu, Houxing Ren, Aojun Zhou, Mingjie Zhan, and Hongsheng Li\.Measuring multimodal mathematical reasoning with MATH\-Vision dataset\.In*Advances in Neural Information Processing Systems \(NeurIPS\), Datasets and Benchmarks Track*, 2024a\.
- Wang et al\. \[2025b\]Peiyu Wang, Yichen Wei, Yi Peng, Xiaokun Wang, Weijie Qiu, Wei Shen, Tianyidan Xie, Jiangbo Pei, Jianhao Zhang, Yunzhuo Hao, et al\.Skywork R1V2: Multimodal hybrid reinforcement learning for reasoning\.*arXiv preprint arXiv:2504\.16656*, 2025b\.
- Wang et al\. \[2024b\]Xiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu, Jieyu Zhang, Satyen Subramaniam, Arjun R\. Loomba, Shichang Zhang, Yizhou Sun, and Wei Wang\.SciBench: Evaluating college\-level scientific problem\-solving abilities of large language models\.In*International Conference on Machine Learning \(ICML\)*, 2024b\.
- Wang et al\. \[2025c\]Xiyao Wang, Zhengyuan Yang, Chao Feng, Hongjin Lu, Linjie Li, Chung\-Ching Lin, Kevin Lin, Furong Huang, and Lijuan Wang\.SoTA with less: MCTS\-guided sample selection for data\-efficient visual reasoning self\-improvement\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2025c\.
- Wang et al\. \[2023\]Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou\.Self\-consistency improves chain of thought reasoning in language models\.In*International Conference on Learning Representations \(ICLR\)*, 2023\.
- Wei et al\. \[2025\]Lai Wei, Yuting Li, Chen Wang, Yue Wang, Linghe Kong, Weiran Huang, and Lichao Sun\.Unsupervised post\-training for multi\-modal LLM reasoning via GRPO\.*arXiv preprint arXiv:2505\.22453*, 2025\.
- Wu et al\. \[2023\]Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, et al\.AutoGen: Enabling next\-gen LLM applications via multi\-agent conversation\.*arXiv preprint arXiv:2308\.08155*, 2023\.
- Xiong et al\. \[2025\]Wei Xiong, Hanning Zhang, Chenlu Ye, Lichang Chen, Nan Jiang, and Tong Zhang\.Self\-rewarding correction for mathematical reasoning\.*arXiv preprint arXiv:2502\.19613*, 2025\.
- Xu et al\. \[2025a\]Fangzhi Xu, Hang Yan, Chang Ma, Haiteng Zhao, Qiushi Sun, Kanzhi Cheng, Junxian He, Jun Liu, and Zhiyong Wu\.Genius: A generalizable and purely unsupervised self\-training framework for advanced reasoning\.In*Annual Meeting of the Association for Computational Linguistics \(ACL\)*, pages 13153–13167, 2025a\.
- Xu et al\. \[2025b\]Zhangchen Xu, Yang Liu, Yueqin Yin, Mingyuan Zhou, and Radha Poovendran\.KodCode: A diverse, challenging, and verifiable synthetic dataset for coding\.In*Findings of the Association for Computational Linguistics: ACL 2025*, pages 6980–7008, 2025b\.
- Xue et al\. \[2026\]Xiangyuan Xue, Yifan Zhou, Guibin Zhang, Zaibin Zhang, Yijiang Li, Chen Zhang, Zhenfei Yin, Philip Torr, Wanli Ouyang, and Lei Bai\.CoMAS: Co\-evolving multi\-agent systems via interaction rewards\.In*International Conference on Learning Representations \(ICLR\)*, 2026\.
- Yang et al\. \[2024\]An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al\.Qwen2\.5 technical report\.*arXiv preprint arXiv:2412\.15115*, 2024\.
- Yang et al\. \[2025a\]An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al\.Qwen3 technical report\.*arXiv preprint arXiv:2505\.09388*, 2025a\.
- Yang et al\. \[2025b\]Yi Yang, Xiaoxuan He, Hongkun Pan, Xiyan Jiang, Yan Deng, Xingtao Yang, Haoyu Lu, Dacheng Yin, Fengyun Rao, Minfeng Zhu, et al\.R1\-OneVision: Advancing generalized multimodal reasoning through cross\-modal formalization\.In*IEEE/CVF International Conference on Computer Vision \(ICCV\)*, pages 2376–2385, 2025b\.
- Yu et al\. \[2025a\]En Yu, Kangheng Lin, Liang Zhao, Yana Wei, Yuang Peng, Haoran Wei, Jianjian Sun, Chunrui Han, Zheng Ge, Xiangyu Zhang, et al\.Perception\-R1: Pioneering perception policy with reinforcement learning\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2025a\.
- Yu et al\. \[2025b\]Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, et al\.DAPO: An open\-source LLM reinforcement learning system at scale\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2025b\.
- Yuan et al\. \[2024\]Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston\.Self\-rewarding language models\.*arXiv preprint arXiv:2401\.10020*, 2024\.
- Yue et al\. \[2025\]Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Yang Yue, Shiji Song, and Gao Huang\.Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model?In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2025\.
- Zhang et al\. \[2026a\]Kaiyan Zhang, Kai Tian, Runze Liu, Sihang Zeng, Xuekai Zhu, Guoli Jia, Yuchen Fan, Xingtai Lv, Yuxin Zuo, Che Jiang, et al\.MARTI: A framework for multi\-agent LLM systems reinforced training and inference\.In*International Conference on Learning Representations \(ICLR\)*, 2026a\.
- Zhang et al\. \[2025a\]Qingyang Zhang, Haitao Wu, Changqing Zhang, Peilin Zhao, and Yatao Bian\.Right question is already half the answer: Fully unsupervised LLM reasoning incentivization\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2025a\.
- Zhang et al\. \[2024a\]Renrui Zhang, Dongzhi Jiang, Yichi Zhang, Haokun Lin, Ziyu Guo, Pengshuo Qiu, Aojun Zhou, Pan Lu, Kai\-Wei Chang, Yu Qiao, et al\.MathVerse: Does your multi\-modal LLM truly see the diagrams in visual math problems?In*European Conference on Computer Vision \(ECCV\)*, pages 169–186, 2024a\.
- Zhang et al\. \[2025b\]Wanpeng Zhang, Yicheng Feng, Hao Luo, Yijiang Li, Zihao Yue, Sipeng Zheng, and Zongqing Lu\.Unified multimodal understanding via byte\-pair visual encoding\.In*IEEE/CVF International Conference on Computer Vision \(ICCV\)*, pages 12976–12986, 2025b\.
- Zhang et al\. \[2025c\]Wanpeng Zhang, Zilong Xie, Yicheng Feng, Yijiang Li, Xingrun Xing, Sipeng Zheng, and Zongqing Lu\.From pixels to tokens: Byte\-pair encoding on quantized visual modalities\.In*International Conference on Learning Representations \(ICLR\)*, 2025c\.
- Zhang et al\. \[2018\]Ying Zhang, Tao Xiang, Timothy M\. Hospedales, and Huchuan Lu\.Deep mutual learning\.In*IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\)*, pages 4320–4328, 2018\.
- Zhang et al\. \[2024b\]Yusen Zhang, Ruoxi Sun, Yanfei Chen, Tomas Pfister, Rui Zhang, and Sercan Ö\. Arık\.Chain of agents: Large language models collaborating on long\-context tasks\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2024b\.
- Zhang et al\. \[2026b\]Zizhuo Zhang, Jianing Zhu, Xinmu Ge, Zihua Zhao, Xuan Li, Xiao Feng, Jiangchao Yao, and Bo Han\.Co\-rewarding: Stable self\-supervised RL for eliciting reasoning in large language models\.In*International Conference on Learning Representations \(ICLR\)*, 2026b\.
- Zhao et al\. \[2025a\]Andrew Zhao, Yiran Wu, Tong Wu, Quentin Xu, Yang Yue, Matthieu Lin, Shenzhi Wang, Qingyun Wu, Zilong Zheng, and Gao Huang\.Absolute zero: Reinforced self\-play reasoning with zero data\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2025a\.
- Zhao et al\. \[2025b\]Wanjia Zhao, Mert Yuksekgonul, Shirley Wu, and James Zou\.SiriuS: Self\-improving multi\-agent systems via bootstrapped reasoning\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2025b\.
- Zhao et al\. \[2026\]Xuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine, and Dawn Song\.Learning to reason without external rewards\.In*International Conference on Learning Representations \(ICLR\)*, 2026\.
- Zuo et al\. \[2025\]Yuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu, Ganqu Cui, Xuekai Zhu, Haozhan Li, Xinwei Long, Ermo Hua, Biqing Qi, et al\.TTRL: Test\-time reinforcement learning\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2025\.
## Appendix APre\-training Error Decoupling
Section[4\.2](https://arxiv.org/html/2608.17253#S4.SS2)argues that a peer helps when its errors do not overlap with the agent’s own and Figure[2](https://arxiv.org/html/2608.17253#S4.F2)\(d\) previews this overlap for three kinds of model pairs\. This appendix reports the full measurement\. All numbers are computed on base checkpoints before any RL, so they describe the pretrained models themselves rather than anything our training produces\.
#### Setup and metrics\.
We score base checkpoints on 500 MATH problems at levels 3 to 5, zero\-shot and single\-sample atT=0\.8T\{=\}0\.8, with rule\-based extraction and equivalence checking\. For a pair of models, every problem lands in one of the four cells of Table[5](https://arxiv.org/html/2608.17253#A1.T5), and the four diversity measures\[[Kuncheva and Whitaker 2003](https://arxiv.org/html/2608.17253#bib.bib34)\]are counts over these cells\. Complementarityccis the share of problems where exactly one model is correct, so one model can correct the other\. Oracle accuracyuuis the share of problems outside the both\-wrong cell, which is the accuracy a perfect selector would reach\[[Krogh and Vedelsby 1994](https://arxiv.org/html/2608.17253#bib.bib33)\]\. Wrong\-agreementwwis the part of the both\-wrong cell where the two models also return the same answer, which is the case a majority vote cannot detect\. Cohen’sκ\\kappa\[[Cohen 1960](https://arxiv.org/html/2608.17253#bib.bib15)\]measures how strongly the two models land in the same cells beyond chance, so lower values mean more decoupled errors\. All four numbers come from the same four cells\.κ\\kappaandccare therefore two readings of one measurement rather than independent evidence, and across the twelve pairs below they correlate atr=−0\.98r=\-0\.98\.
Table 5:The four outcomes for a pair of models A and B\. Every problem falls into exactly one cell, and all four diversity measures are counts over these cells\.
#### Three levels of decoupling\.
Table[6](https://arxiv.org/html/2608.17253#A1.T6)groups pairs by what the two models differ in\.*Seed only*pairs a checkpoint with itself under a different sampling seed, so the two views share every weight and differ in generation noise alone\.*Same family*pairs models of one lineage across sizes or generations\.*Different family*pairs models with separate architectures and pretraining data\. The first level is whatCo\-RLhas by construction, and the third is what a different\-family cohort adds\.
#### The three levels separate without overlap\.
Every different\-family pair reachesκ≤0\.42\\kappa\\leq 0\.42andc≥29\.4c\\geq 29\.4, and every same\-family and seed\-only pair sits atκ≥0\.51\\kappa\\geq 0\.51andc≤24\.6c\\leq 24\.6\. No pair falls between the two groups, at either scale\. Averaged within a group, crossing families lowersκ\\kappafrom0\.530\.53to0\.380\.38and raisesccfrom23\.223\.2to30\.830\.8\. The same\-family and seed\-only groups are not distinguishable from one another, so changing the size, the generation or the sampling seed within a lineage leaves the error structure where it was\.
#### Holding the anchor fixed gives the same ordering\.
Group averages mix models of different strength, so we also fix one model and vary only its partner \(Table[7](https://arxiv.org/html/2608.17253#A1.T7)\)\. Both anchors give a monotone ladder\. Relative to pairing a model with itself, a same\-family partner lowersκ\\kappaby0\.040\.04and0\.070\.07, and a different\-family partner lowers it by0\.180\.18and0\.160\.16while raisingccby 9\.2% and 10\.4%\. Because the seed\-only row is the anchor paired with itself, capability is identical along each ladder and the source of the partner is the only variable\. Wrong\-agreement follows the same direction, and is lowest for the different\-family partner under both anchors, though it does not order the middle rows\.
#### What this means forCo\-RL\.
Two models trained by different groups on different data fail on different problems, and no amount of resampling or rescaling within one lineage reproduces that\. A same\-family cohort still gives each agent a target it did not produce itself, which is enough for stable training, but the target repeats what the agent would have answered anyway on three quarters of the problems\. A different\-family cohort is the cheapest way to buy the remaining headroom\.
DecouplingPairκ↓\\kappa\\ \\downarrowc↑c\\ \\uparrow\(%\)w↓w\\ \\downarrow\(%\)u↑u\\ \\uparrow\(%\)*3B tier*different familyLlama\-3\.2\-3B×\\timesPhi\-3\.5\-mini0\.3132\.83\.053\.0different familyQwen2\.5\-3B×\\timesLlama\-3\.2\-3B0\.3831\.22\.463\.0different familyQwen2\.5\-3B×\\timesPhi\-3\.5\-mini0\.3831\.24\.055\.4different familyQwen2\.5\-3B×\\timesMiniCPM3\-4B0\.4129\.44\.460\.4same familyQwen2\.5\-3B×\\timesQwen3\-1\.7B\-Base0\.5224\.24\.263\.2seed onlyQwen3\-1\.7B\-Base×\\timesitself0\.5224\.05\.066\.4seed onlyQwen2\.5\-3B×\\timesitself0\.5622\.04\.462\.6*7B tier*different familyQwen2\.5\-7B×\\timesLlama\-3\.1\-8B0\.4229\.41\.871\.4same familyQwen2\.5\-7B×\\timesQwen2\.5\-3B0\.5124\.23\.869\.8same familyQwen2\.5\-7B×\\timesQwen3\-1\.7B\-Base0\.5124\.44\.070\.4seed onlyLlama\-3\.1\-8B×\\timesitself0\.5124\.63\.062\.0seed onlyQwen2\.5\-7B×\\timesitself0\.5819\.05\.274\.6Table 6:Error decoupling before RL, by what the two models differ in, sorted byκ\\kappawithin each block\.Table 7:One model held fixed, partner varied\. Capability is identical to the seed\-only row along each ladder, so the source of the partner is the only variable\.
## Appendix BComplete Proof
### B\.1Proof of Proposition[1](https://arxiv.org/html/2608.17253#Thmproposition1)
In this section, we derive the reward\-induced GRPO dynamics in Proposition[1](https://arxiv.org/html/2608.17253#Thmproposition1)\. We first derive a common update expression under a fixed pseudo\-label and then specialize it to self\-rewarding and cross\-agent supervision\.
For a fixed promptxx, letXk=𝟏\[ak=a⋆\]X^\{k\}=\\mathbf\{1\}\[a^\{k\}=a^\{\\star\}\]indicate whether thekk\-th rollout is correct\. Under the binary reduction,Xk∼i\.i\.d\.Bernoulli\(p\)X^\{k\}\\overset\{\\mathrm\{i\.i\.d\.\}\}\{\\sim\}\\operatorname\{Bernoulli\}\(p\), and
C=∑k=1KXk∼Bin\(K,p\)C=\\sum\_\{k=1\}^\{K\}X^\{k\}\\sim\\operatorname\{Bin\}\(K,p\)denotes the number of correct responses among theKKrollouts\. To analyze the induced binary dynamics, we use the log\-odds
ℓ=logp1−p\\ell=\\log\\frac\{p\}\{1\-p\}as a one\-dimensional coordinate\. Equivalently,p=σ\(ℓ\)=1/\(1\+e−ℓ\)p=\\sigma\(\\ell\)=1/\(1\+e^\{\-\\ell\}\)\. We analyze the reward\-induced component of GRPO in the infinitesimal\-update limit, where clipping is locally inactive\.
GRPO update under a fixed pseudo\-label\.LetZ∈\{0,1\}Z\\in\\\{0,1\\\}indicate whether a fixed pseudo\-label is correct, whereZ=1Z=1corresponds toa⋆a^\{\\star\}andZ=0Z=0to the aggregate incorrect answer\. The binary reward for rolloutkkis
rk=𝟏\[Xk=Z\]=\(1−Z\)\+\(2Z−1\)Xk\.r^\{k\}=\\mathbf\{1\}\[X^\{k\}=Z\]=\(1\-Z\)\+\(2Z\-1\)X^\{k\}\.SinceC=∑k=1KXkC=\\sum\_\{k=1\}^\{K\}X^\{k\}, the group\-mean reward is
r¯=\(1−Z\)\+\(2Z−1\)CK,\\bar\{r\}=\(1\-Z\)\+\(2Z\-1\)\\frac\{C\}\{K\},and therefore
rk−r¯=\(2Z−1\)\(Xk−CK\)\.r^\{k\}\-\\bar\{r\}=\(2Z\-1\)\\left\(X^\{k\}\-\\frac\{C\}\{K\}\\right\)\.\(10\)
The within\-group standard deviation of the binary rewards is
sK\(C\)≜1K∑j=1K\(rj−r¯\)2=C\(K−C\)K\.s\_\{K\}\(C\)\\triangleq\\sqrt\{\\frac\{1\}\{K\}\\sum\_\{j=1\}^\{K\}\(r^\{j\}\-\\bar\{r\}\)^\{2\}\}=\\frac\{\\sqrt\{C\(K\-C\)\}\}\{K\}\.For0<C<K0<C<K, the normalized group\-relative advantage is therefore
A^k=rk−r¯sK\(C\)=2Z−1sK\(C\)\(Xk−CK\)\.\\widehat\{A\}^\{k\}=\\frac\{r^\{k\}\-\\bar\{r\}\}\{s\_\{K\}\(C\)\}=\\frac\{2Z\-1\}\{s\_\{K\}\(C\)\}\\left\(X^\{k\}\-\\frac\{C\}\{K\}\\right\)\.ForC∈\{0,K\}C\\in\\\{0,K\\\}, all rewards are identical and the centered update is zero\.
Treating the advantages as fixed during the policy update, the reward\-induced GRPO gradient in the log\-odds coordinate is
gℓ\(C,Z\)\\displaystyle g\_\{\\ell\}\(C,Z\)≜∂∂ℓ\[1K∑k=1KA^klogPrℓ\(Xk\)\]\\displaystyle\\triangleq\\frac\{\\partial\}\{\\partial\\ell\}\\left\[\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\widehat\{A\}^\{k\}\\log\\Pr\_\{\\ell\}\(X^\{k\}\)\\right\]=1K∑k=1KA^k∂∂ℓ\[Xklogp\+\(1−Xk\)log\(1−p\)\]\\displaystyle=\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\widehat\{A\}^\{k\}\\frac\{\\partial\}\{\\partial\\ell\}\\left\[X^\{k\}\\log p\+\(1\-X^\{k\}\)\\log\(1\-p\)\\right\]=1K∑k=1KA^k\(Xk−p\)\\displaystyle=\\frac\{1\}\{K\}\\sum\_\{k=1\}^\{K\}\\widehat\{A\}^\{k\}\(X^\{k\}\-p\)=2Z−1KsK\(C\)∑k=1K\(Xk−CK\)\(Xk−p\)\.\\displaystyle=\\frac\{2Z\-1\}\{Ks\_\{K\}\(C\)\}\\sum\_\{k=1\}^\{K\}\\left\(X^\{k\}\-\\frac\{C\}\{K\}\\right\)\(X^\{k\}\-p\)\.Since∑k=1K\(Xk−C/K\)=0\\sum\_\{k=1\}^\{K\}\(X^\{k\}\-C/K\)=0, the terms involvingppcancel\. Using\(Xk\)2=Xk\(X^\{k\}\)^\{2\}=X^\{k\}and∑k=1KXk=C\\sum\_\{k=1\}^\{K\}X^\{k\}=Cgives
∑k=1K\(Xk−CK\)Xk=C−C2K=C\(K−C\)K\.\\sum\_\{k=1\}^\{K\}\\left\(X^\{k\}\-\\frac\{C\}\{K\}\\right\)X^\{k\}=C\-\\frac\{C^\{2\}\}\{K\}=\\frac\{C\(K\-C\)\}\{K\}\.Hence,
gℓ\(C,Z\)=\(2Z−1\)C\(K−C\)K\.g\_\{\\ell\}\(C,Z\)=\(2Z\-1\)\\frac\{\\sqrt\{C\(K\-C\)\}\}\{K\}\.\(11\)
\[1\] Self\-rewarding\.Under majority\-vote self\-rewarding, the pseudo\-label is determined by the same rollout group:
Zself=\[C\>K2\]\.Z\_\{\\mathrm\{self\}\}=\\mathbf\{1\}\\\!\\left\[C\>\\frac\{K\}\{2\}\\right\]\.SinceKKis odd,
2Zself−1=sign\(C−K2\)\.2Z\_\{\\mathrm\{self\}\}\-1=\\operatorname\{sign\}\\\!\\left\(C\-\\frac\{K\}\{2\}\\right\)\.Substituting into Eq\. \([11](https://arxiv.org/html/2608.17253#A2.E11)\) gives
gℓ\(C,Zself\)=sign\(C−K2\)C\(K−C\)K\.g\_\{\\ell\}\(C,Z\_\{\\mathrm\{self\}\}\)=\\operatorname\{sign\}\\\!\\left\(C\-\\frac\{K\}\{2\}\\right\)\\frac\{\\sqrt\{C\(K\-C\)\}\}\{K\}\.For a GRPO update with learning rateη\\eta,
ℓ\+−ℓ=ηgℓ\(C,Zself\)\.\\ell^\{\+\}\-\\ell=\\eta\\,g\_\{\\ell\}\(C,Z\_\{\\mathrm\{self\}\}\)\.SinceC∼Bin\(K,p\)C\\sim\\operatorname\{Bin\}\(K,p\)is random, the conditional expected one\-step update is
𝔼\[ℓ\+−ℓ∣p\]=η𝔼C∼Bin\(K,p\)\[sign\(C−K2\)C\(K−C\)K\]\.\\mathbb\{E\}\[\\ell^\{\+\}\-\\ell\\mid p\]=\\eta\\,\\mathbb\{E\}\_\{C\\sim\\operatorname\{Bin\}\(K,p\)\}\\left\[\\operatorname\{sign\}\\\!\\left\(C\-\\frac\{K\}\{2\}\\right\)\\frac\{\\sqrt\{C\(K\-C\)\}\}\{K\}\\right\]\.In the infinitesimal\-update limit, the corresponding mean log\-odds dynamics are
ℓ˙=η𝔼C∼Bin\(K,p\)\[sign\(C−K2\)C\(K−C\)K\]\.\\dot\{\\ell\}=\\eta\\,\\mathbb\{E\}\_\{C\\sim\\operatorname\{Bin\}\(K,p\)\}\\left\[\\operatorname\{sign\}\\\!\\left\(C\-\\frac\{K\}\{2\}\\right\)\\frac\{\\sqrt\{C\(K\-C\)\}\}\{K\}\\right\]\.Sincep=σ\(ℓ\)p=\\sigma\(\\ell\),
p˙\\displaystyle\\dot\{p\}=∂p∂ℓℓ˙\\displaystyle=\\frac\{\\partial p\}\{\\partial\\ell\}\\dot\{\\ell\}=ηp\(1−p\)𝔼C∼Bin\(K,p\)\[sign\(C−K2\)C\(K−C\)K\],\\displaystyle=\\eta\\,p\(1\-p\)\\mathbb\{E\}\_\{C\\sim\\operatorname\{Bin\}\(K,p\)\}\\left\[\\operatorname\{sign\}\\\!\\left\(C\-\\frac\{K\}\{2\}\\right\)\\frac\{\\sqrt\{C\(K\-C\)\}\}\{K\}\\right\],which recovers Eq\. \([7](https://arxiv.org/html/2608.17253#S5.E7)\)\.
\[2\]Co\-RL: Cross\-agent supervision\.For agentnn, letXnk=𝟏\[ank=a⋆\]X\_\{n\}^\{k\}=\\mathbf\{1\}\[a\_\{n\}^\{k\}=a^\{\\star\}\]and
Cn=∑k=1KXnk∼Bin\(K,pn\)\.C\_\{n\}=\\sum\_\{k=1\}^\{K\}X\_\{n\}^\{k\}\\sim\\operatorname\{Bin\}\(K,p\_\{n\}\)\.Let
Z−n=𝟏\[a^−n\(x\)=a⋆\]Z\_\{\-n\}=\\mathbf\{1\}\[\\hat\{a\}\_\{\-n\}\(x\)=a^\{\\star\}\]indicate whether the pseudo\-label constructed from the other agents is correct\. Conditional on the prompt and current policies,Z−nZ\_\{\-n\}is independent of agentnn’s rollout group, and hence
Z−n⟂Cn\.Z\_\{\-n\}\\perp C\_\{n\}\.
Applying Eq\. \([11](https://arxiv.org/html/2608.17253#A2.E11)\) to agentnngives
gℓn\(Cn,Z−n\)=sign\(Z−n−12\)Cn\(K−Cn\)K,g\_\{\\ell\_\{n\}\}\(C\_\{n\},Z\_\{\-n\}\)=\\operatorname\{sign\}\\\!\\left\(Z\_\{\-n\}\-\\frac\{1\}\{2\}\\right\)\\frac\{\\sqrt\{C\_\{n\}\(K\-C\_\{n\}\)\}\}\{K\},whereℓn=logpn1−pn\\ell\_\{n\}=\\log\\frac\{p\_\{n\}\}\{1\-p\_\{n\}\}\. Therefore,
𝔼\[ℓn\+−ℓn∣pn,p−n\]\\displaystyle\\mathbb\{E\}\[\\ell\_\{n\}^\{\+\}\-\\ell\_\{n\}\\mid p\_\{n\},p\_\{\-n\}\]=ηn𝔼\[sign\(Z−n−12\)Cn\(K−Cn\)K\]\\displaystyle=\\eta\_\{n\}\\,\\mathbb\{E\}\\left\[\\operatorname\{sign\}\\\!\\left\(Z\_\{\-n\}\-\\frac\{1\}\{2\}\\right\)\\frac\{\\sqrt\{C\_\{n\}\(K\-C\_\{n\}\)\}\}\{K\}\\right\]=ηn𝔼\[sign\(Z−n−12\)\]𝔼Cn∼Bin\(K,pn\)\[Cn\(K−Cn\)K\],\\displaystyle=\\eta\_\{n\}\\,\\mathbb\{E\}\\left\[\\operatorname\{sign\}\\\!\\left\(Z\_\{\-n\}\-\\frac\{1\}\{2\}\\right\)\\right\]\\mathbb\{E\}\_\{C\_\{n\}\\sim\\operatorname\{Bin\}\(K,p\_\{n\}\)\}\\left\[\\frac\{\\sqrt\{C\_\{n\}\(K\-C\_\{n\}\)\}\}\{K\}\\right\],where the second equality follows fromZ−n⟂CnZ\_\{\-n\}\\perp C\_\{n\}\.
Taking the infinitesimal\-update limit and using∂pn/∂ℓn=pn\(1−pn\)\\partial p\_\{n\}/\\partial\\ell\_\{n\}=p\_\{n\}\(1\-p\_\{n\}\)yields
p˙n\\displaystyle\\dot\{p\}\_\{n\}=∂pn∂ℓnℓ˙n\\displaystyle=\\frac\{\\partial p\_\{n\}\}\{\\partial\\ell\_\{n\}\}\\dot\{\\ell\}\_\{n\}=ηnpn\(1−pn\)𝔼\[sign\(Z−n−12\)\]𝔼Cn∼Bin\(K,pn\)\[Cn\(K−Cn\)K\],\\displaystyle=\\eta\_\{n\}p\_\{n\}\(1\-p\_\{n\}\)\\mathbb\{E\}\\left\[\\operatorname\{sign\}\\\!\\left\(Z\_\{\-n\}\-\\frac\{1\}\{2\}\\right\)\\right\]\\mathbb\{E\}\_\{C\_\{n\}\\sim\\operatorname\{Bin\}\(K,p\_\{n\}\)\}\\left\[\\frac\{\\sqrt\{C\_\{n\}\(K\-C\_\{n\}\)\}\}\{K\}\\right\],which recovers Eq\. \([8](https://arxiv.org/html/2608.17253#S5.E8)\)\.
Symmetric two\-agent dynamics\.We finally specialize the result to two symmetric agentsAAandBB\. For agentAA, its pseudo\-label is the majority vote of agentBB’s rollouts\. Hence, ifCB∼Bin\(K,pB\)C\_\{B\}\\sim\\operatorname\{Bin\}\(K,p\_\{B\}\),
Z−A=\[CB\>K2\],Z\_\{\-A\}=\\mathbf\{1\}\\\!\\left\[C\_\{B\}\>\\frac\{K\}\{2\}\\right\],and therefore
𝔼\[sign\(Z−A−12\)\]\\displaystyle\\mathbb\{E\}\\left\[\\operatorname\{sign\}\\\!\\left\(Z\_\{\-A\}\-\\frac\{1\}\{2\}\\right\)\\right\]=𝔼CB∼Bin\(K,pB\)\[sign\(CB−K2\)\]\\displaystyle=\\mathbb\{E\}\_\{C\_\{B\}\\sim\\operatorname\{Bin\}\(K,p\_\{B\}\)\}\\left\[\\operatorname\{sign\}\\\!\\left\(C\_\{B\}\-\\frac\{K\}\{2\}\\right\)\\right\]=2VK\(pB\)−1=ϕK\(pB\)\.\\displaystyle=2V\_\{K\}\(p\_\{B\}\)\-1=\\phi\_\{K\}\(p\_\{B\}\)\.Similarly,
𝔼\[sign\(Z−B−12\)\]=ϕK\(pA\)\.\\mathbb\{E\}\\left\[\\operatorname\{sign\}\\\!\\left\(Z\_\{\-B\}\-\\frac\{1\}\{2\}\\right\)\\right\]=\\phi\_\{K\}\(p\_\{A\}\)\.
Assuming the same learning rateη\\etafor both agents and defining
qK\(p\)=ηp\(1−p\)𝔼C∼Bin\(K,p\)\[C\(K−C\)K\]\>0,q\_\{K\}\(p\)=\\eta p\(1\-p\)\\mathbb\{E\}\_\{C\\sim\\operatorname\{Bin\}\(K,p\)\}\\left\[\\frac\{\\sqrt\{C\(K\-C\)\}\}\{K\}\\right\]\>0,we obtain
p˙A=qK\(pA\)ϕK\(pB\),p˙B=qK\(pB\)ϕK\(pA\),\\dot\{p\}\_\{A\}=q\_\{K\}\(p\_\{A\}\)\\phi\_\{K\}\(p\_\{B\}\),\\qquad\\dot\{p\}\_\{B\}=q\_\{K\}\(p\_\{B\}\)\\phi\_\{K\}\(p\_\{A\}\),which recovers Eq\. \([9](https://arxiv.org/html/2608.17253#S5.E9)\)\.
### B\.2Proof of Proposition[2](https://arxiv.org/html/2608.17253#Thmproposition2)
We prove that majority\-vote self\-rewarding reinforces the currently favored answer and therefore induces self\-confirming dynamics\. For convenience, denote
GKself\(p\)=𝔼C∼Bin\(K,p\)\[sign\(C−K2\)C\(K−C\)K\]\.G\_\{K\}^\{\\mathrm\{self\}\}\(p\)=\\mathbb\{E\}\_\{C\\sim\\operatorname\{Bin\}\(K,p\)\}\\left\[\\operatorname\{sign\}\\\!\\left\(C\-\\frac\{K\}\{2\}\\right\)\\frac\{\\sqrt\{C\(K\-C\)\}\}\{K\}\\right\]\.Expanding the expectation gives
GKself\(p\)=∑c=0Ksign\(c−K2\)c\(K−c\)KPrp\(C=c\)\.G\_\{K\}^\{\\mathrm\{self\}\}\(p\)=\\sum\_\{c=0\}^\{K\}\\operatorname\{sign\}\\\!\\left\(c\-\\frac\{K\}\{2\}\\right\)\\frac\{\\sqrt\{c\(K\-c\)\}\}\{K\}\\Pr\_\{p\}\(C=c\)\.Since the update magnitude satisfies
c\(K−c\)K=\(K−c\)cK,\\frac\{\\sqrt\{c\(K\-c\)\}\}\{K\}=\\frac\{\\sqrt\{\(K\-c\)c\}\}\{K\},each eventC=c\>K/2C=c\>K/2can be paired with its symmetric eventC=K−c<K/2C=K\-c<K/2\. Moreover, the update magnitude vanishes atc=0c=0andc=Kc=K\. Hence,
GKself\(p\)=∑c=\(K\+1\)/2K−1c\(K−c\)K\[Prp\(C=c\)−Prp\(C=K−c\)\]\.G\_\{K\}^\{\\mathrm\{self\}\}\(p\)=\\sum\_\{c=\(K\+1\)/2\}^\{K\-1\}\\frac\{\\sqrt\{c\(K\-c\)\}\}\{K\}\\left\[\\Pr\_\{p\}\(C=c\)\-\\Pr\_\{p\}\(C=K\-c\)\\right\]\.\(12\)
Forp∈\(0,1\)p\\in\(0,1\)andc\>K/2c\>K/2,
Prp\(C=c\)Prp\(C=K−c\)\\displaystyle\\frac\{\\Pr\_\{p\}\(C=c\)\}\{\\Pr\_\{p\}\(C=K\-c\)\}=\(Kc\)pc\(1−p\)K−c\(KK−c\)pK−c\(1−p\)c\\displaystyle=\\frac\{\\binom\{K\}\{c\}p^\{c\}\(1\-p\)^\{K\-c\}\}\{\\binom\{K\}\{K\-c\}p^\{K\-c\}\(1\-p\)^\{c\}\}=\(p1−p\)2c−K,\\displaystyle=\\left\(\\frac\{p\}\{1\-p\}\\right\)^\{2c\-K\},where we use\(Kc\)=\(KK−c\)\\binom\{K\}\{c\}=\\binom\{K\}\{K\-c\}\. Since2c−K\>02c\-K\>0,
Prp\(C=c\)−Prp\(C=K−c\)\{\>0,p\>12,=0,p=12,<0,p<12\.\\Pr\_\{p\}\(C=c\)\-\\Pr\_\{p\}\(C=K\-c\)\\begin\{cases\}\>0,&p\>\\frac\{1\}\{2\},\\\\ =0,&p=\\frac\{1\}\{2\},\\\\ <0,&p<\\frac\{1\}\{2\}\.\\end\{cases\}The factorc\(K−c\)/K\\sqrt\{c\(K\-c\)\}/Kin Eq\. \([12](https://arxiv.org/html/2608.17253#A2.E12)\) is strictly positive for0<c<K0<c<K\. Therefore, every term in the sum has the same sign, yielding
sign\(GKself\(p\)\)=sign\(p−12\)\.\\operatorname\{sign\}\\\!\\left\(G\_\{K\}^\{\\mathrm\{self\}\}\(p\)\\right\)=\\operatorname\{sign\}\\\!\\left\(p\-\\frac\{1\}\{2\}\\right\)\.\(13\)
We next characterize the limiting behavior\. From Eq\. \([7](https://arxiv.org/html/2608.17253#S5.E7)\),
p˙=ηp\(1−p\)GKself\(p\)\.\\dot\{p\}=\\eta p\(1\-p\)G\_\{K\}^\{\\mathrm\{self\}\}\(p\)\.If0<p\(0\)<1/20<p\(0\)<1/2, Eq\. \([13](https://arxiv.org/html/2608.17253#A2.E13)\) impliesp˙<0\\dot\{p\}<0wheneverp∈\(0,1/2\)p\\in\(0,1/2\)\. Thus,p\(t\)p\(t\)is monotonically decreasing and bounded below by zero, and therefore converges to somep∞∈\[0,1/2\)p\_\{\\infty\}\\in\[0,1/2\)\.
Supposep∞\>0p\_\{\\infty\}\>0\. By Eq\. \([13](https://arxiv.org/html/2608.17253#A2.E13)\),
p∞\(1−p∞\)GKself\(p∞\)<0\.p\_\{\\infty\}\(1\-p\_\{\\infty\}\)G\_\{K\}^\{\\mathrm\{self\}\}\(p\_\{\\infty\}\)<0\.Since the right\-hand side of Eq\. \([7](https://arxiv.org/html/2608.17253#S5.E7)\) is continuous,p˙\\dot\{p\}remains strictly negative in a neighborhood ofp∞p\_\{\\infty\}, which contradicts convergence to an interior limit\. Hence,
p\(0\)<12⟹p\(t\)→0\.p\(0\)<\\frac\{1\}\{2\}\\quad\\Longrightarrow\\quad p\(t\)\\rightarrow 0\.
Similarly, if1/2<p\(0\)<11/2<p\(0\)<1, thenp˙\>0\\dot\{p\}\>0\. Thus,p\(t\)p\(t\)is monotonically increasing and bounded above by one\. The same argument rules out any interior limit, giving
p\(0\)\>12⟹p\(t\)→1\.p\(0\)\>\\frac\{1\}\{2\}\\quad\\Longrightarrow\\quad p\(t\)\\rightarrow 1\.Finally, atp=1/2p=1/2, symmetry givesGKself\(1/2\)=0G\_\{K\}^\{\\mathrm\{self\}\}\(1/2\)=0, sop=1/2p=1/2is an unstable equilibrium\. This completes the proof\.
### B\.3Proof of Theorem[1](https://arxiv.org/html/2608.17253#Thmtheorem1)
We prove Theorem[1](https://arxiv.org/html/2608.17253#Thmtheorem1)for the symmetric two\-agent dynamics from Eq\. \([9](https://arxiv.org/html/2608.17253#S5.E9)\),
p˙A=qK\(pA\)ϕK\(pB\),p˙B=qK\(pB\)ϕK\(pA\),\\dot\{p\}\_\{A\}=q\_\{K\}\(p\_\{A\}\)\\phi\_\{K\}\(p\_\{B\}\),\\qquad\\dot\{p\}\_\{B\}=q\_\{K\}\(p\_\{B\}\)\\phi\_\{K\}\(p\_\{A\}\),\(14\)where
qK\(p\)=ηp\(1−p\)𝔼C∼Bin\(K,p\)\[C\(K−C\)K\]\>0q\_\{K\}\(p\)=\\eta p\(1\-p\)\\mathbb\{E\}\_\{C\\sim\\operatorname\{Bin\}\(K,p\)\}\\left\[\\frac\{\\sqrt\{C\(K\-C\)\}\}\{K\}\\right\]\>0forp∈\(0,1\)p\\in\(0,1\), andϕK\(p\)=2VK\(p\)−1\\phi\_\{K\}\(p\)=2V\_\{K\}\(p\)\-1\.
###### Proof\.
Symmetry of the dynamics\.We first establish two useful symmetries\. Since the distribution ofK−CK\-CunderC∼Bin\(K,p\)C\\sim\\operatorname\{Bin\}\(K,p\)isBin\(K,1−p\)\\operatorname\{Bin\}\(K,1\-p\)andC\(K−C\)\\sqrt\{C\(K\-C\)\}is invariant underC↦K−CC\\mapsto K\-C, we have
qK\(1−p\)=qK\(p\)\.q\_\{K\}\(1\-p\)=q\_\{K\}\(p\)\.Moreover, for oddKK, complementing every rollout reverses the majority outcome, so
VK\(1−p\)=1−VK\(p\)\.V\_\{K\}\(1\-p\)=1\-V\_\{K\}\(p\)\.Therefore,
qK\(1−p\)=qK\(p\),ϕK\(1−p\)=−ϕK\(p\)\.q\_\{K\}\(1\-p\)=q\_\{K\}\(p\),\\qquad\\phi\_\{K\}\(1\-p\)=\-\\phi\_\{K\}\(p\)\.\(15\)SinceVK\(p\)V\_\{K\}\(p\)is strictly increasing andVK\(1/2\)=1/2V\_\{K\}\(1/2\)=1/2,
sign\(ϕK\(p\)\)=sign\(p−12\)\.\\operatorname\{sign\}\(\\phi\_\{K\}\(p\)\)=\\operatorname\{sign\}\\\!\\left\(p\-\\frac\{1\}\{2\}\\right\)\.
A conserved quantity\.Define
FK\(p\)≜∫1/2pϕK\(u\)qK\(u\)𝑑u\.F\_\{K\}\(p\)\\triangleq\\int\_\{1/2\}^\{p\}\\frac\{\\phi\_\{K\}\(u\)\}\{q\_\{K\}\(u\)\}\\,du\.\(16\)Along any interior trajectory of Eq\. \([14](https://arxiv.org/html/2608.17253#A2.E14)\),
ddt\[FK\(pA\)−FK\(pB\)\]\\displaystyle\\frac\{d\}\{dt\}\\left\[F\_\{K\}\(p\_\{A\}\)\-F\_\{K\}\(p\_\{B\}\)\\right\]=ϕK\(pA\)qK\(pA\)p˙A−ϕK\(pB\)qK\(pB\)p˙B\\displaystyle=\\frac\{\\phi\_\{K\}\(p\_\{A\}\)\}\{q\_\{K\}\(p\_\{A\}\)\}\\dot\{p\}\_\{A\}\-\\frac\{\\phi\_\{K\}\(p\_\{B\}\)\}\{q\_\{K\}\(p\_\{B\}\)\}\\dot\{p\}\_\{B\}=ϕK\(pA\)ϕK\(pB\)−ϕK\(pB\)ϕK\(pA\)\\displaystyle=\\phi\_\{K\}\(p\_\{A\}\)\\phi\_\{K\}\(p\_\{B\}\)\-\\phi\_\{K\}\(p\_\{B\}\)\\phi\_\{K\}\(p\_\{A\}\)=0\.\\displaystyle=0\.Hence,
FK\(pA\)−FK\(pB\)=constantF\_\{K\}\(p\_\{A\}\)\-F\_\{K\}\(p\_\{B\}\)=\\text\{constant\}\(17\)along every trajectory\.
From Eq\. \([15](https://arxiv.org/html/2608.17253#A2.E15)\),
FK\(1−p\)=FK\(p\)\.F\_\{K\}\(1\-p\)=F\_\{K\}\(p\)\.Furthermore,FK′\(p\)=ϕK\(p\)/qK\(p\)F\_\{K\}^\{\\prime\}\(p\)=\\phi\_\{K\}\(p\)/q\_\{K\}\(p\), soFKF\_\{K\}is strictly decreasing on\(0,1/2\)\(0,1/2\)and strictly increasing on\(1/2,1\)\(1/2,1\), withFK\(1/2\)=0F\_\{K\}\(1/2\)=0\.
Separatrix and basins of attraction\.Consider first the disagreement regionpA<1/2<pBp\_\{A\}<1/2<p\_\{B\}\. In this region,
p˙A\>0,p˙B<0,\\dot\{p\}\_\{A\}\>0,\\qquad\\dot\{p\}\_\{B\}<0,so the two agents move toward one another\.
SupposepA\+pB\>1p\_\{A\}\+p\_\{B\}\>1\. ThenpB\>1−pA\>1/2p\_\{B\}\>1\-p\_\{A\}\>1/2\. SinceFKF\_\{K\}is strictly increasing on\(1/2,1\)\(1/2,1\)andFK\(1−pA\)=FK\(pA\)F\_\{K\}\(1\-p\_\{A\}\)=F\_\{K\}\(p\_\{A\}\),
FK\(pB\)\>FK\(1−pA\)=FK\(pA\),F\_\{K\}\(p\_\{B\}\)\>F\_\{K\}\(1\-p\_\{A\}\)=F\_\{K\}\(p\_\{A\}\),and therefore
FK\(pA\)−FK\(pB\)<0\.F\_\{K\}\(p\_\{A\}\)\-F\_\{K\}\(p\_\{B\}\)<0\.\(18\)Because this quantity is conserved, agentBBcannot reach1/21/2before agentAA: ifpB=1/2p\_\{B\}=1/2, thenFK\(pA\)−FK\(pB\)=FK\(pA\)≥0F\_\{K\}\(p\_\{A\}\)\-F\_\{K\}\(p\_\{B\}\)=F\_\{K\}\(p\_\{A\}\)\\geq 0, contradicting Eq\. \([18](https://arxiv.org/html/2608.17253#A2.E18)\)\. Thus, agentAAcrosses the decision boundary first\.
OncepA,pB\>1/2p\_\{A\},p\_\{B\}\>1/2, Eq\. \([14](https://arxiv.org/html/2608.17253#A2.E14)\) gives
p˙A\>0,p˙B\>0\.\\dot\{p\}\_\{A\}\>0,\\qquad\\dot\{p\}\_\{B\}\>0\.Both probabilities therefore increase monotonically and are bounded above by one\. No interior point of\(1/2,1\)2\(1/2,1\)^\{2\}is an equilibrium becauseqK\(p\)\>0q\_\{K\}\(p\)\>0andϕK\(p\)\>0\\phi\_\{K\}\(p\)\>0there\. Hence,
\(pA\(t\),pB\(t\)\)⟶\(1,1\)\.\(p\_\{A\}\(t\),p\_\{B\}\(t\)\)\\longrightarrow\(1,1\)\.
Conversely, ifpA\+pB<1p\_\{A\}\+p\_\{B\}<1, thenpB<1−pAp\_\{B\}<1\-p\_\{A\}, and the same symmetry gives
FK\(pA\)−FK\(pB\)\>0\.F\_\{K\}\(p\_\{A\}\)\-F\_\{K\}\(p\_\{B\}\)\>0\.The conserved quantity now prevents agentAAfrom reaching1/21/2before agentBB\. Thus,BBcrosses below1/21/2first, after which both agents satisfy
p˙A<0,p˙B<0,\\dot\{p\}\_\{A\}<0,\\qquad\\dot\{p\}\_\{B\}<0,and consequently
\(pA\(t\),pB\(t\)\)⟶\(0,0\)\.\(p\_\{A\}\(t\),p\_\{B\}\(t\)\)\\longrightarrow\(0,0\)\.The casepB<1/2<pAp\_\{B\}<1/2<p\_\{A\}follows symmetrically\.
Now considerpA\+pB=1p\_\{A\}\+p\_\{B\}=1\. SettingpB=1−pAp\_\{B\}=1\-p\_\{A\}and using Eq\. \([15](https://arxiv.org/html/2608.17253#A2.E15)\),
p˙A\+p˙B\\displaystyle\\dot\{p\}\_\{A\}\+\\dot\{p\}\_\{B\}=qK\(pA\)ϕK\(1−pA\)\+qK\(1−pA\)ϕK\(pA\)\\displaystyle=q\_\{K\}\(p\_\{A\}\)\\phi\_\{K\}\(1\-p\_\{A\}\)\+q\_\{K\}\(1\-p\_\{A\}\)\\phi\_\{K\}\(p\_\{A\}\)=0\.\\displaystyle=0\.Hence, the line
pA\+pB=1p\_\{A\}\+p\_\{B\}=1\(19\)is invariant\. Along this line, ifpA<1/2<pBp\_\{A\}<1/2<p\_\{B\}, thenp˙A\>0\\dot\{p\}\_\{A\}\>0andp˙B<0\\dot\{p\}\_\{B\}<0; the reverse holds whenpB<1/2<pAp\_\{B\}<1/2<p\_\{A\}\. Thus, every interior trajectory on this line converges to\(1/2,1/2\)\(1/2,1/2\)\.
The cases where both agents initially lie on the same side of1/21/2follow directly from Eq\. \([14](https://arxiv.org/html/2608.17253#A2.E14)\)\. Combining all cases,
pA\(0\)\+pB\(0\)\>1⟹\(pA\(t\),pB\(t\)\)→\(1,1\),p\_\{A\}\(0\)\+p\_\{B\}\(0\)\>1\\quad\\Longrightarrow\\quad\(p\_\{A\}\(t\),p\_\{B\}\(t\)\)\\rightarrow\(1,1\),whereas
pA\(0\)\+pB\(0\)<1⟹\(pA\(t\),pB\(t\)\)→\(0,0\)\.p\_\{A\}\(0\)\+p\_\{B\}\(0\)<1\\quad\\Longrightarrow\\quad\(p\_\{A\}\(t\),p\_\{B\}\(t\)\)\\rightarrow\(0,0\)\.Therefore,pA\+pB=1p\_\{A\}\+p\_\{B\}=1is the interior separatrix between the two consensus basins\.
Stability of the equilibria\.In a neighborhood of\(1,1\)\(1,1\), both agents satisfypA,pB\>1/2p\_\{A\},p\_\{B\}\>1/2and hence both coordinates increase monotonically toward one\. Thus,\(1,1\)\(1,1\)is asymptotically stable\. By symmetry,\(0,0\)\(0,0\)is also asymptotically stable\.
Finally, consider\(1/2,1/2\)\(1/2,1/2\)\. SinceϕK\(1/2\)=0\\phi\_\{K\}\(1/2\)=0, letpA=1/2\+δAp\_\{A\}=1/2\+\\delta\_\{A\}andpB=1/2\+δBp\_\{B\}=1/2\+\\delta\_\{B\}\. Linearizing Eq\. \([14](https://arxiv.org/html/2608.17253#A2.E14)\) gives
\[δ˙Aδ˙B\]=qK\(1/2\)ϕK′\(1/2\)\[0110\]\[δAδB\]\.\\begin\{bmatrix\}\\dot\{\\delta\}\_\{A\}\\\\ \\dot\{\\delta\}\_\{B\}\\end\{bmatrix\}=q\_\{K\}\(1/2\)\\phi\_\{K\}^\{\\prime\}\(1/2\)\\begin\{bmatrix\}0&1\\\\ 1&0\\end\{bmatrix\}\\begin\{bmatrix\}\\delta\_\{A\}\\\\ \\delta\_\{B\}\\end\{bmatrix\}\.For oddKK,
ϕK′\(1/2\)=2K\(K−1\(K−1\)/2\)\(12\)K−1\>0\.\\phi\_\{K\}^\{\\prime\}\(1/2\)=2K\\binom\{K\-1\}\{\(K\-1\)/2\}\\left\(\\frac\{1\}\{2\}\\right\)^\{K\-1\}\>0\.The Jacobian therefore has eigenvalues
λ±=±qK\(1/2\)ϕK′\(1/2\),\\lambda\_\{\\pm\}=\\pm q\_\{K\}\(1/2\)\\phi\_\{K\}^\{\\prime\}\(1/2\),with corresponding eigenvectors\(1,1\)\(1,1\)and\(1,−1\)\(1,\-1\)\. Thus,\(1/2,1/2\)\(1/2,1/2\)has one unstable and one stable direction and is therefore a saddle point\. Its stable manifold is exactly the invariant linepA\+pB=1p\_\{A\}\+p\_\{B\}=1established above\.
Finally, Proposition[2](https://arxiv.org/html/2608.17253#Thmproposition2)gives the correct\-convergence basin under independent self\-rewarding as
ℬself\+=\{\(pA,pB\):pA\>12,pB\>12\}\.\\mathcal\{B\}\_\{\\mathrm\{self\}\}^\{\+\}=\\left\\\{\(p\_\{A\},p\_\{B\}\):p\_\{A\}\>\\frac\{1\}\{2\},\\;p\_\{B\}\>\\frac\{1\}\{2\}\\right\\\}\.ForCo\-RL, the result above gives
ℬCo\-RL\+=\{\(pA,pB\):pA\+pB\>1\}\.\\mathcal\{B\}\_\{\\mathrm\{\\textsc\{Co\-RL\}\}\}^\{\+\}=\\left\\\{\(p\_\{A\},p\_\{B\}\):p\_\{A\}\+p\_\{B\}\>1\\right\\\}\.Therefore,
ℬself\+⊊ℬCo\-RL\+,\\mathcal\{B\}\_\{\\mathrm\{self\}\}^\{\+\}\\subsetneq\\mathcal\{B\}\_\{\\mathrm\{\\textsc\{Co\-RL\}\}\}^\{\+\},which establishes the strictly larger basin of correct convergence\. ∎
## Appendix CRephrased Training Questions
The data\-decoupled runs train one agent on the original MATH questions and the other on a rephrased copy produced by DeepSeek\-V3\. The two copies are aligned by row and the answer is never changed\. The rewrites go beyond word substitution: most of them place the problem in a concrete scenario, roughly doubling the question length\. In a sample of 300 pairs, every rewrite preserved the answer and the row alignment\. Two representative pairs follow\.
Example 1*Answer:*2ORIGINALHow many vertical asymptotes does the graph ofy=2x2\+x−6y=\\frac\{2\}\{x^\{2\}\+x\-6\}have?REPHRASEDThe functionf\(t\)=2t2\+t−6f\(t\)=\\frac\{2\}\{t^\{2\}\+t\-6\}describes the temperature of a chemical reaction over timett\. How many vertical asymptotes appear on the graph of this function?
Example 2*Answer:*333\\sqrt\{3\}ORIGINALIn triangleABCABC,AB=AC=14AB=AC=14andBC=26BC=26\. What is the length of the shortest angle bisector inABCABC? Express your answer in simplest radical form\.REPHRASEDA triangular park has two equal sides of 14 meters and a third side of 26 meters\. The city plans a path from each corner that bisects its angle, and will build only the shortest one\. How long is that path? Express your answer in simplest radical form\.
## Appendix DComplete Experiment Results
### D\.1Results on 7B and 8B language models
We extend the comparison in Section[6\.2](https://arxiv.org/html/2608.17253#S6.SS2)to larger models, pairing Qwen2\.5\-7B with Llama\-3\.1\-8B\-Instruct\. Table[8](https://arxiv.org/html/2608.17253#A4.T8)reports the full results across the same seven benchmarks\. The trends observed for 3B models continue to hold at this scale\. AllCo\-RLvariants improve over their respective base models on average, and using agents from different model families generally provides stronger gains than using two agents initialized from the same family\. Further decoupling the training data yields the strongest overall variant, withCo\-RL\(Different family\+\) improving Qwen2\.5\-7B from 49\.0 to 53\.6 and Llama\-3\.1\-8B\-Instruct from 44\.7 to 47\.7 on average\.
Compared with prior label\-free methods,Co\-RL\(Different family\+\) achieves the best average performance for both models, outperforming the strongest self\-rewarding baseline by 0\.8% on Qwen2\.5\-7B and 1\.1% on Llama\-3\.1\-8B\-Instruct\. Notably, on Llama\-3\.1\-8B\-Instruct, it also surpasses GT\-Reward \(47\.7 vs\. 47\.1\) despite using no ground\-truth labels\. These results show that the benefits of cross\-agent supervision persist as model scale increases, and further support the importance of diversity across agents for effective label\-free learning\.
### D\.2Results on 7B–12B Vision\-Language Models
We further evaluateCo\-RLon larger vision\-language models from three different families: Qwen2\.5\-VL\-7B, InternVL3\.5\-8B, and Gemma\-3\-12B, with InternVL3\.5\-8B serving as the shared training partner\. Table[9](https://arxiv.org/html/2608.17253#A4.T9)reports results across the same four multimodal reasoning benchmarks\. The gains observed at smaller scales persist consistently:Co\-RLimproves the corresponding base models by 7\.2%, 6\.3%, and 5\.8% on average for Qwen2\.5\-VL\-7B, InternVL3\.5\-8B, and Gemma\-3\-12B, respectively, and outperforms TTRL for all three model families\.
Despite using no ground\-truth supervision,Co\-RLapproaches GT\-Reward on Qwen2\.5\-VL\-7B and InternVL3\.5\-8B, while surpassing it on Gemma\-3\-12B \(47\.56% vs\. 45\.17%\)\. In particular, the improvement remains consistent across models with substantially different visual encoders and language backbones\. Together with the 2B–3B results in Table[4](https://arxiv.org/html/2608.17253#S6.T4), these results show that the benefits of cross\-agent supervision extend across model scales and heterogeneous vision\-language architectures\.
Table 8:Results at 7B and 8B on the seven\-benchmark suite \(%\)\. Base and GT\-Reward serve as references and are excluded from the ranking\.Table 9:Vision\-language results at 7B to 12B on open\-r1, with InternVL3\.5\-8B as the shared partner\.
### D\.3Training Dynamics and Stability
#### Text models\.
Figure[4](https://arxiv.org/html/2608.17253#S6.F4)compares validation accuracy, reward standard deviation, and mean completion length throughout training\. Across all four model families,Co\-RLmaintains a non\-degenerate reward standard deviation and relatively stable completion lengths while steadily improving validation accuracy\. In contrast, several self\-rewarding methods become unstable during training\. RENT rapidly drives the reward standard deviation toward zero and often exhibits a sharp increase in completion length, accompanied by degraded accuracy or divergence\. Intuitor similarly shows substantial reductions in reward variation and, for some models, degenerate completion lengths\. TTRL is more stable, but its reward variation generally decreases and its validation performance remains belowCo\-RL\.
These observations align with the dynamics analyzed in Section[5](https://arxiv.org/html/2608.17253#S5)\. Under self\-rewarding, the supervision signal is determined by the optimized agent itself\. Consequently, incorrect predictions can be reinforced rather than corrected\. As the model increasingly agrees with its own pseudo\-labels, rollout rewards can also become homogeneous, reducing the group\-relative learning signal\. InCo\-RL, supervision is instead provided by another agent: since the update direction of one agent is determined by its peer’s supervision signal, errors that would be self\-reinforced can be corrected when the agents have complementary strengths\. The stable reward variation observed in Figure[4](https://arxiv.org/html/2608.17253#S6.F4)is consistent with this cross\-agent signal remaining informative throughout training\.
#### Vision\-language models\.
We observe the same qualitative behavior in the multimodal setting\. As shown in Figure[5](https://arxiv.org/html/2608.17253#A4.F5),Co\-RLcontinues to improve evaluation accuracy while maintaining stable completion lengths, whereas TTRL eventually degrades in both accuracy and response length\. More importantly, the agreement between the two agents remains well below full agreement throughout training, while the accuracy of the exchanged pseudo\-labels steadily increases\. Thus, the agents do not simply converge to identical behaviors; instead, they preserve meaningful differences while providing increasingly reliable supervision to one another\. This provides further empirical support for the mechanism predicted by our theory: cross\-agent learning benefits from complementary supervision rather than requiring the agents to collapse to the same predictions\.
Figure 5:Training dynamics for Qwen2\.5\-VL\-7B trained with InternVL3\.5\-8B on open\-r1\. \(a\) Evaluation accuracy, \(b\) mean completion length, and \(c\) the accuracy of the exchanged pseudo\-labels together with the agreement between the two agents\.Co\-RLkeeps improving and holds its completion length, while TTRL peaks and then degrades in both\.
### D\.4Controlling for the Two\-Agent Training Budget
Unlike single\-agent self\-rewarding baselines,Co\-RLjointly trains two agents\. To control for this additional training budget, we construct a matched self\-rewarding baseline that trains the*same two base models*independently with TTRL on the same prompts\. At inference time, we ensemble the two independently trained models: each model generates four responses, and the resulting eight responses are pooled for majority voting\. We apply the same ensemble protocol to the two agents trained withCo\-RL, such that the compared methods use both the same training\-model budget and the same test\-time sampling budget\. We additionally report each trained agent individually to separate the effect of training from that of ensembling\.
Tables[10](https://arxiv.org/html/2608.17253#A4.T10)and[11](https://arxiv.org/html/2608.17253#A4.T11)report the results for text and multimodal reasoning, respectively\. Simply training two self\-rewarding agents and ensembling their predictions provides only limited gains over the stronger individual model\. In contrast, ensembling the agents trained withCo\-RLconsistently gives the best macro\-average on the text benchmarks and on the two multimodal settings\. These results indicate that the advantage ofCo\-RLcannot be explained merely by training two models\. Rather, cross\-agent supervision produces agents whose predictions combine more effectively under the same training and inference budgets\.
Table 10:Matched\-budget comparison between TTRL andCo\-RLon text reasoning benchmarks\. Both settings train the same two base models, Qwen2\.5\-3B and Llama\-3\.2\-3B\-Instruct\. The ensemble rows pool four rollouts from each of the two models for majority voting \(maj@8,T=0\.6T=0\.6\)\.Avgis the macro\-average over the three benchmarks\. For each benchmark, the best result is inboldand the second best isunderlined, with ties sharing the marking\.Table 11:Matched\-budget comparison between TTRL andCo\-RLon multimodal reasoning benchmarks\. Both settings train the same two base models, Qwen2\.5\-VL\-3B and InternVL3\.5\-2B\. The ensemble rows pool four rollouts from each of the two models for majority voting \(maj@8,T=0\.6T=0\.6, top\-pp0\.95\)\. Rows are grouped by training set\. MMR1 runs use the corrected multiple\-choice grader and open\-r1 runs the legacy grader, so the two blocks are not compared against each other\.
### D\.5Evaluation Details
Language model evaluation\.All language benchmarks except LiveCodeBench run through lm\-evaluation\-harness\. Sampling uses a temperature of 0\.6 with top\-pp0\.95 and a 3072\-token generation budget\. Trained checkpoints are evaluated with their tokenizer’s chat template, which matches the prompt format used during training\. Base models never saw a template and are evaluated without one\. For AMC we sample eight responses per problem and report avg@8, averaged over three evaluation seeds\. All other benchmarks use a single sample\. GPQA uses the Diamond subset with a boxed\-answer prompt\. Answers are scored by rule\-based graders with math\-aware normalization, and multiple\-choice questions accept both the option letter and the option value\. LiveCodeBench \(release v6\) runs through its official harness at its default temperature of 0\.2 with one sample per problem\. The harness hard\-codes a gated Llama\-3 tokenizer for prompt construction, and we substitute the openly available Llama\-3\.2\-3B\-Instruct tokenizer, which carries the same chat template\.
CoMAS comparison\.The comparison with CoMAS uses its benchmark suite \(GSM8K, MATH\-500, HumanEval, MBPP, SciBench\[[Wang et al\. 2024b](https://arxiv.org/html/2608.17253#bib.bib61)\], GPQA and MMLU\[[Hendrycks et al\. 2021a](https://arxiv.org/html/2608.17253#bib.bib26)\]\), its driver and its graders, with training prompts drawn from their blended 2,000 problems from MATH\[[Hendrycks et al\. 2021b](https://arxiv.org/html/2608.17253#bib.bib27)\], KodCode\[[Xu et al\. 2025b](https://arxiv.org/html/2608.17253#bib.bib68)\]and WebInstruct\-verified\[[Ma et al\. 2025](https://arxiv.org/html/2608.17253#bib.bib46)\]\. Their protocol draws five samples per question at temperature 0\.7 and then issues a sixth call that reasons over the five drafts, and only the sixth response is graded\. On coding benchmarks this aggregation admits a loophole\. A response that quotes several candidate solutions has every code block executed, and grading stops at the first block that passes, so a model that quotes many candidates is effectively scored at pass@5 while a decisive one is scored at pass@1\. In our measurement this is worth 7\.3% to the untrained baseline and 2\.4% to our model\. On coding benchmarks we therefore keep the five\-sample budget but replace the aggregation with majority voting over candidates clustered by their execution behavior on the public example inputs\.
Three\-agent runs\.The three\-agent runs use the same training configuration as the two\-agent language runs and are evaluated with the protocol above\.
Vision\-language evaluation\.We follow the benchmark splits of MM\-UPT, the MathVision test set \(3,040 problems\), the MathVerse testmini split \(3,940, all five versions\), the MathVista testmini split \(1,000\) and the We\-Math testmini split \(1,740\)\. Decoding is greedy with a 16384\-token generation budget\. Images are resized so that the long side does not exceed 1024 pixels, matching the training\-time preprocessing\. Trained checkpoints are prompted in the format they were trained on, with reasoning in think tags and the final answer in answer tags\. Base models receive a standard boxed\-answer prompt\. Scoring is two\-stage\. A rule\-based pass extracts the final answer and grades it with math\-aware matching, and a response that never commits to an answer in a recognized format counts as incorrect\. Responses that follow the format but fail the rule match are passed to an LLM judge, Qwen2\.5\-32B\-Instruct at temperature 0, which accepts only semantically equivalent answers and rejects responses that are cut off\. The judge can only recover rule\-grading false negatives and never overturns a rule\-credited answer\. MMR1 runs use the corrected multiple\-choice grader and open\-r1 runs the legacy grader, so results across the two training sets are not compared\.
Engineering notes\.All fixes below ship with the released code\. None of them changes the training or evaluation semantics\. They repair crashes or a wrong backend choice in the underlying libraries\.
*Gemma\-3, embedding initialization under ZeRO\-3\.*At startup, the weight initializer zeroes the embedding row atpadding\_idx\. Under DeepSpeed ZeRO\-3, most ranks hold empty parameter shards, so this write fails before training begins\. The branch is reached only when a model setspadding\_idx, which Gemma\-3 does and Qwen2\.5\-VL does not\. We guard the initializer to skip embeddings whose local shard is empty\.
*Gemma\-3, batched prompt tokenization\.*The Gemma\-3 processor buildstoken\_type\_idsby stacking the unpadded prompts of a batch into one array\. Prompts of unequal length make this stacking fail at the first training step\. We wrap the processor to tokenize with padding and strip the padding through the attention mask immediately after\. The wrapper is a no\-op for every other processor\.
*Gemma\-3, log\-probability drift\.*Between the vLLM rollout engine and the training forward pass, Gemma\-3 shows a systematic per\-token log\-probability drift of about 0\.13\. This is an architectural discrepancy rather than a removable bug\. Following the reference recipes for this model, Gemma\-3 runs, and only Gemma\-3 runs, train with token\-level truncation of the importance\-sampling ratio\.
*Qwen2\.5\-VL, vision\-tower attention backend\.*In vLLM 0\.11\.2, the helper that selects the vision tower’s attention backend silently promotes xFormers to the bundled FlashAttention build\. That build supports head dimensions that are multiples of 32 only, and the Qwen2\.5\-VL vision tower has head dimension 80, so the model crashes at load\. We patch the helper to keep the original xFormers choice, which has no head\-dimension restriction\. Gemma\-3 and InternVL vision towers have head dimensions that are multiples of 32 and are unaffected, and later vLLM releases fix the bug\.
*InternVL3\.5, processor and tiling\.*We use the transformers\-native HF variants, whose checkpoints load throughAutoProcessorwithout the legacy remote\-code path\. Dynamic patch tiling is disabled at both training and evaluation, so the image\-token count per sample is identical in the two settings\. Model families are detected from each checkpoint’s configuration file rather than from directory names, sinceCo\-RLrun directories contain both partners’ names\.Similar Articles
DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning
This paper identifies that failures in visual reasoning often stem from breakdowns in dynamic cross-modal coordination between visual and textual evidence during chain-of-thought generation. It introduces DyCo-RL, a reinforcement learning framework that rewards effective cross-modal coordination, leading to improved reasoning performance.
CORA: Analyzing and bridging thinking-answer gap in Multimodal RLVR via Consistency-Oriented Reasoning Alignment
This paper analyzes the thinking-answer inconsistency in multimodal reinforcement learning with verifiable rewards (RLVR) for large vision-language models and proposes CORA, a method that introduces a consistency reward model and hybrid reward advantage splitting to improve faithfulness and task performance.
IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents
This paper introduces Isolated Bilateral Reinforcement Learning (IB-RL), a method where two dialogue roles co-evolve through joint rollouts while optimizing their own rewards independently. It addresses the static-counterpart mismatch in RL for strategic dialogue, showing improved generalization to unseen counterparts in Vehicle TeleSales and Deal-or-No-Deal benchmarks.
SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design
The paper introduces SocialRL, a multi-turn reinforcement learning framework that enhances the social intelligence of large language models through delayed reward propagation and fine-grained process rewards, achieving notable improvements in goal completion for dialogue systems.
CoRA: Confidence-Rationale Alignment for Reliable Chain-of-Thought Reasoning
This paper introduces CoRA, a GRPO-based reinforcement learning framework that aligns LLM confidence with generated rationales to improve the reliability of chain-of-thought reasoning, achieving up to 26.51% reduction in misalignment error across multiple benchmarks.