通过门控与衰减的在策略蒸馏提升 OCR 保真度

arXiv cs.AI 论文

摘要

论文提出了 GAD-RL 方法,在视觉-语言模型的联合后训练阶段对在策略蒸馏的监督信号进行自适应门控与衰减,以提升 OCR 转录的保真度。在 Qwen3.5-2B 上,该方法在 CHAOS-Bench 上取得了 59.92% 的 Micro Recall,超过 GRPO 8.45 个百分点。

arXiv:2609.38282v1 Announce Type: new Abstract: Vision-language models may rewrite anomalous text in images into linguistically plausible expressions, compromising OCR transcription faithfulness. Sequence-level task rewards and local teacher guidance are complementary, but guidance from the same teacher may not remain equally effective as the student improves. Offline analysis shows that supervision from a fixed teacher becomes progressively less favorable as the student improves, both across training checkpoints and across response groups with different task rewards. Motivated by this observation, we introduce GAD-RL, which adaptively regulates teacher supervision during joint post-training according to the student's current task performance and local distributions. A frozen teacher conditions on reference transcriptions and student-generated prefixes. GAD-RL disables distillation for response groups containing an output with task reward at least 0.95 and continuously attenuates distillation strength as group-mean reward increases. It also weights forward KL by the student's probability of the teacher's Top-1 token, moderating local auxiliary updates when student support for that candidate is low. On Qwen3.5-2B, GAD-RL achieves 59.92% Micro Recall on CHAOS-Bench, surpassing GRPO and GRPO+OPD (fixed-weight) by 8.45 and 4.43 percentage points, respectively, while achieving an Overall score of 91.18 on OmniDocBench v1.6.
查看原文
查看缓存全文

缓存时间: 2026/10/01 09:40

# Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation
Source: [https://arxiv.org/html/2609.38282](https://arxiv.org/html/2609.38282)
Zuming Huang11footnotemark:1Kexuan RenJun HuangWei Chu

###### Abstract

Vision\-language models may rewrite anomalous text in images into linguistically plausible expressions, compromising OCR transcription faithfulness\. Sequence\-level task rewards and local teacher guidance are complementary, but guidance from the same teacher may not remain equally effective as the student improves\. Offline analysis shows that supervision from a fixed teacher becomes progressively less favorable as the student improves, both across training checkpoints and across response groups with different task rewards\. Motivated by this observation, we introduceGAD\-RL, which adaptively regulates teacher supervision during joint post\-training according to the student’s current task performance and local distributions\. A frozen teacher conditions on reference transcriptions and student\-generated prefixes\. GAD\-RL disables distillation for response groups containing an output with task reward at least 0\.95 and continuously attenuates distillation strength as group\-mean reward increases\. It also weights forward KL by the student’s probability of the teacher’s Top\-1 token, moderating local auxiliary updates when student support for that candidate is low\. On Qwen3\.5\-2B, GAD\-RL achieves 59\.92% Micro Recall on CHAOS\-Bench, surpassing GRPO and GRPO\+OPD \(fixed\-weight\) by 8\.45 and 4\.43 percentage points, respectively, while achieving an Overall score of 91\.18 on OmniDocBench v1\.6\.

## 1Introduction

Document parsing aims to faithfully convert the text and structure in page images into machine\-readable representations\. Vision\-language models \(VLMs\) integrate text recognition, layout understanding, and structured generation, advancing complex document parsing\([Bai et al\., 2025](https://arxiv.org/html/2609.38282#bib.bib2);[Poznanski et al\., 2025](https://arxiv.org/html/2609.38282#bib.bib3)\)\. Yet strong parsing ability does not guarantee faithful transcription: misspellings, visually confusable characters, or unexpected words may be replaced with linguistically plausible expressions that contradict the image\. This linguistic\-prior hallucination turns reading into implicit rewriting\([Yao et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib1);[Lee et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib8)\)\. We focus on OCR transcription faithfulness: when visual evidence conflicts with linguistic priors, models should preserve the content actually visible in the image, even when the source itself contains errors, while maintaining reliable parsing of ordinary text and document structure\.

Figure 1:Student performance improves while directional supervisory SNR declines\.The same frozen teacher scores GRPO checkpoints offline on a separate analysis set of 1,000 synthetic pages\. \(a\) Signal composition on perturbed\-word\-associated tokens: relative to student probabilities at the same prefixes, teacher signals are beneficial when favoring correct emitted tokens or disfavoring errors, and potentially harmful in the reverse direction; small differences are neutral \(Appendix[B\.4](https://arxiv.org/html/2609.38282#A2.SS4)\)\. \(b\) Beneficial\-to\-harmful token\-count ratio on a log scale; the dotted line marks equal counts\. \(c\) Student perturbed\-word Micro Recall\.To optimize anomalous\-content retention and transcription quality, we adopt Group Relative Policy Optimization\([Shao et al\., 2024](https://arxiv.org/html/2609.38282#bib.bib10), GRPO;\), which has been applied to document parser post\-training\([Poznanski et al\., 2025](https://arxiv.org/html/2609.38282#bib.bib3);[Wang et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib4)\)\. Unlike supervised fine\-tuning \(SFT\), which learns under reference prefixes, GRPO uses relative rewards within groups of student\-generated responses\. However, identical group rewards yield zero group\-relative advantages\([Yu et al\., 2025](https://arxiv.org/html/2609.38282#bib.bib11)\), while rarely sampled faithful outputs receive limited direct reinforcement\. We therefore complement GRPO with on\-policy distillation \(OPD\), which provides token\-level supervision at student\-generated prefixes\([Agarwal et al\., 2024](https://arxiv.org/html/2609.38282#bib.bib12)\)\.

Recent work has explored regulating on\-policy teacher supervision through student\-aware target reformulation\([Jang et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib28)\)or trajectory\-level selection\([Akhondzadeh et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib29)\)\. Our focus is complementary: we study how the utility of a fixed teacher changes with the student’s current task performance, and use this observation to regulate when and how strongly the teacher intervenes during joint GRPO post\-training\. Offline analysis of GRPO\-only checkpoints shows rising perturbed\-word Micro Recall \(7\.18% to 98\.06%\) and declining directional supervisory signal\-to\-noise ratio \(SNR; 12\.57 to 0\.114\) on a fixed analysis set \(Figure[1](https://arxiv.org/html/2609.38282#S1.F1)\)\.

Motivated by these observations, we propose GAD\-RL, a policy\-state\-aware framework that regulates teacher intervention according to the student’s current task performance\. Group\-level gating determines when teacher supervision remains active, while reward\-aware attenuation continuously reduces its strength as group performance improves\. Separately, to avoid overly strong local FKL updates under substantial teacher–student disagreement, we use a simple student\-probability weight to moderate the distillation gradient without altering the teacher target\.

We evaluate GAD\-RL on GlitchText\([Yao et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib1)\)and CHAOS\-Bench\([Li et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib27)\), covering controlled text anomalies and character perturbations in realistic document layouts\. On CHAOS\-Bench, GAD\-RL improves perturbed\-word Micro Recall from 51\.47% for GRPO to 59\.92%, a gain of8\.45 percentage points, while maintaining comparable general document\-parsing performance \(Table[1](https://arxiv.org/html/2609.38282#S4.T1)\)\. Comparisons with fixed\-weight distillation, step\-based decay, and low\-score response selection examine state\-aware regulation against simpler controls \(Section[5\.2](https://arxiv.org/html/2609.38282#S5.SS2)\)\.

Our main contributions are:

- •Revealing changes in teacher supervision with student state\.Through token\-level analysis, we identify a decline in the same frozen teacher’s directional supervisory signal\-to\-noise ratio as the student improves, providing empirical motivation for dynamically regulating teacher intervention\.
- •Introducing state\-aware teacher intervention for OCR faithfulness\.Based on the observed reward\-dependent variation in supervision quality, GAD\-RL combines group\-level mastery gating with continuous reward\-aware attenuation to regulate when and how strongly a frozen teacher intervenes during joint GRPO post\-training\. We additionally use student\-probability\-weighted FKL to moderate large local distillation updates under teacher–student disagreement\.
- •Improving transcription faithfulness across student backbones\.GAD\-RL improves CHAOS\-Bench Recall over GRPO by 8\.45 and 3\.82 percentage points on Qwen3\.5\-2B and Qwen3\-VL\-2B, respectively, while maintaining comparable OmniDocBench performance\. Ablations and controlled comparisons further support the effectiveness of regulating teacher intervention according to the student’s current task performance\.

## 2Related Work

Document parsing\.LayoutLM and LayoutLMv3 integrate text, image, and layout features\([Xu et al\., 2020](https://arxiv.org/html/2609.38282#bib.bib16)\); Donut, Pix2Struct, and Nougat recover structured content from images\([Kim et al\., 2022](https://arxiv.org/html/2609.38282#bib.bib17);[Lee et al\., 2023](https://arxiv.org/html/2609.38282#bib.bib18);[Blecher et al\., 2024](https://arxiv.org/html/2609.38282#bib.bib19)\)\. RL\-based parsing uses verifiable rewards in olmOCR 2\([Poznanski et al\., 2025](https://arxiv.org/html/2609.38282#bib.bib3)\)and layout\-aware rewards in Infinity\-Parser\([Wang et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib4)\)\. OvisOCR2 combines real\-document annotations with synthetic HTML\-derived pages, then uses SFT, RL on a larger branch, OPD into a compact parser, and model fusion\([Lu et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib34)\)\. OmniDocBench measures general parsing quality\([Ouyang et al\., 2025](https://arxiv.org/html/2609.38282#bib.bib20)\)\. However, current VLM\-based parsers remain vulnerable to text perturbations: FaithC4 reveals unfaithful rewriting and error amplification on unperturbed content\([Lee et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib8)\), while CHAOS\-Bench, introduced with HunyuanOCR\-1\.5, reports low perturbed\-word recall across evaluated OCR models\([Li et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib27)\)\.

Multimodal hallucination mitigation\.VCD, OPERA, and VISTA modify decoding or visual representations\([Leng et al\., 2024](https://arxiv.org/html/2609.38282#bib.bib21);[Huang et al\., 2024](https://arxiv.org/html/2609.38282#bib.bib22);[Li et al\., 2025b](https://arxiv.org/html/2609.38282#bib.bib9)\)\. PAR targets OCR overcorrection through positional perturbation and attention recycling at inference time, without additional training\([Yao et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib1)\)\. Training methods use factual feedback\([Sun et al\., 2024](https://arxiv.org/html/2609.38282#bib.bib23)\), preference optimization\([Zhao et al\., 2023](https://arxiv.org/html/2609.38282#bib.bib24);[Yu et al\., 2024](https://arxiv.org/html/2609.38282#bib.bib25)\), or phrase\-level alignment\([Sarkar et al\., 2025](https://arxiv.org/html/2609.38282#bib.bib26)\)\. Our work focuses on OCR transcription faithfulness: mitigating overcorrection driven by linguistic priors while preserving the text actually present in document images\.

On\-policy learning and distillation\.GKD supports on\-policy distillation jointly with RL\([Agarwal et al\., 2024](https://arxiv.org/html/2609.38282#bib.bib12)\)\. KDRL combines GRPO and reverse KL, with linear coefficient decay and reward\-guided response/group masks; group distillation stops when any response succeeds\([Xu et al\., 2025](https://arxiv.org/html/2609.38282#bib.bib32)\)\. OPSD uses privileged\-context self\-distillation\([Zhao et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib13)\); TGPO supplies teacher\-preferred next tokens at student prefixes under large policy divergence\([Liu et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib14)\)\. Veto constructs student\-dependent geometric targets\([Jang et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib28)\)\.

Teacher supervision reliability and control\.RG\-OPD filters trajectories by agreement between verifier advantages and teacher–student log\-likelihood gaps\([Akhondzadeh et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib29)\)\. RSTG \(*Distill Where You Fail*\) combines all\-incorrect zero\-variance group selection, teacher\-confidence weighting, token selection, and SFT on correct teacher trajectories\([Han et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib15)\)\. I\-SDPO routes all\-incorrect groups to privileged self\-distillation and any\-success groups to GRPO\([Zhang et al\., 2026a](https://arxiv.org/html/2609.38282#bib.bib31)\)\. PACED favors intermediate student pass rates based on cross\-problem gradient SNR\([Xu et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib30)\)\. Analyzing teacher noise,[Ding and Zhang \(2026\)](https://arxiv.org/html/2609.38282#bib.bib33)find that fixed negative advantages on low\-probability sampled tokens can match OPD in their reverse\-KL reasoning\-distillation settings\. We track a fixed teacher’s correctness\-aware beneficial\-to\-harmful token\-count ratio across OCR student checkpoints as transcription improves\. GAD\-RL combines group gating and continuous group\-mean\-reward attenuation with GRPO; student\-probability weighting separately moderates local FKL updates\.

## 3Methods

GAD\-RL combines GRPO with state\-aware on\-policy distillation from a frozen teacher\. Group\-level gating and reward\-aware attenuation weaken or disable teacher guidance as task performance improves, while student\-adaptive forward KL scales token\-level updates\. GRPO remains active throughout training\.

Let\(x,z\)∼𝒟\(x,z\)\\sim\\mathcal\{D\}be a document example, wherexxcontains the page image and parsing instruction, andzzis the ground truth \(GT\)\. The student isπθ\\pi\_\{\\theta\}\. The frozen teacherπθT\\pi\_\{\\theta\_\{\\mathrm\{T\}\}\}receivesxTx\_\{\\mathrm\{T\}\}, a text prompt combiningzzwith a content\-preserving rewrite instruction, without the page image\. For each input, the rollout policyπθold\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}generates a group𝒴=\{y\(i\)\}i=1G\\mathcal\{Y\}=\\\{y^\{\(i\)\}\\\}\_\{i=1\}^\{G\}, withy\(i\)∼πθold\(⋅∣x\)y^\{\(i\)\}\\sim\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(\\cdot\\mid x\)\. Here,iiindexes responses,ttindexes tokens, andTi=\|y\(i\)\|T\_\{i\}=\|y^\{\(i\)\}\|is the valid response length\. Student and teacher score tokens under the same generated prefixy<t\(i\)y\_\{<t\}^\{\(i\)\}\. We write𝔼roll\\mathbb\{E\}\_\{\\mathrm\{roll\}\}for expectation over\(x,z\)∼𝒟\(x,z\)\\sim\\mathcal\{D\}and these sampled groups\.

#### GRPO objective\.

LetRi=R⁡\(y\(i\),z\)R\_\{i\}=R\(y^\{\(i\)\},z\)be the task reward, defined in Section[3\.2](https://arxiv.org/html/2609.38282#S3.SS2), and letR¯\\bar\{R\}andσR\\sigma\_\{R\}be its group mean and standard deviation\. The token\-level policy ratio and sequence advantage are

ρi,t​\(θ\)=πθ​\(yt\(i\)∣x,y<t\(i\)\)πθold​\(yt\(i\)∣x,y<t\(i\)\),Ai=Ri−R¯σR,Ai,t=Ai\.\\rho\_\{i,t\}\(\\theta\)=\\frac\{\\pi\_\{\\theta\}\(y\_\{t\}^\{\(i\)\}\\mid x,y\_\{<t\}^\{\(i\)\}\)\}\{\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(y\_\{t\}^\{\(i\)\}\\mid x,y\_\{<t\}^\{\(i\)\}\)\},\\qquad A\_\{i\}=\\frac\{R\_\{i\}\-\\bar\{R\}\}\{\\sigma\_\{R\}\},\\qquad A\_\{i,t\}=A\_\{i\}\.\(1\)WhenσR=0\\sigma\_\{R\}=0, we set all group advantages to zero\. GRPO\([Shao et al\., 2024](https://arxiv.org/html/2609.38282#bib.bib10)\)maximizes

𝒥GRPO\(θ\)=𝔼roll\[\\displaystyle\\mathcal\{J\}\_\{\\mathrm\{GRPO\}\}\(\\theta\)=\\mathbb\{E\}\_\{\\mathrm\{roll\}\}\\Biggl\[1G∑i=1G1Ti∑t=1Ti\\displaystyle\\frac\{1\}\{G\}\\sum\_\{i=1\}^\{G\}\\frac\{1\}\{T\_\{i\}\}\\sum\_\{t=1\}^\{T\_\{i\}\}\(2\)min\(ρi,t\(θ\)Ai,clip\(ρi,t\(θ\),1−ε,1\+ε\)Ai\)\],\\displaystyle\\min\\\!\\left\(\\rho\_\{i,t\}\(\\theta\)A\_\{i\},\\operatorname\{clip\}\\\!\\left\(\\rho\_\{i,t\}\(\\theta\),1\-\\varepsilon,1\+\\varepsilon\\right\)A\_\{i\}\\right\)\\Biggr\],whereε\\varepsilonis the clipping parameter\. Tokens are averaged within each response, and responses are averaged within each group\.

### 3\.1Mastery\-Gated Distillation

When to distill\.When the student can already generate a sequence that receives the maximum task reward for an input, that sequence requires no further correction under the current reward criterion\. Continuing to match the teacher distribution may then introduce unnecessary supervision noise and conflict with task\-reward optimization\([Han et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib15)\)\. We therefore disable group\-level distillation when any sampled response reaches the success threshold\. The gate is

g=𝕀\[maxi=1,…,GRi<τ\]\.g=\\mathbb\{I\}\\\!\\left\[\\max\_\{i=1,\\ldots,G\}R\_\{i\}<\\tau\\right\]\.\(3\)Here,𝕀⁡\[⋅\]\\mathbb\{I\}\[\\cdot\]is the indicator function, andτ\\tauis chosen to tolerate minor formatting differences such as whitespace and heading levels\. GRPO always uses all responses and reinforces those with above\-average rewards when group rewards differ\.

### 3\.2Reward\-Aware Attenuation

How strongly to distill\.A binary gate assigns the same weight to low\-performing and nearly solved groups retained for distillation\. Motivated by the lower directional supervisory SNR observed in higher\-reward groups within individual checkpoints \(Figure[4](https://arxiv.org/html/2609.38282#S5.F4)\), we use group\-average task reward to reduce guidance as performance improves, while the gate determines whether distillation remains active\.

For an outputyyand its GTzz, we combine normalized edit similarityReditR\_\{\\mathrm\{edit\}\}with exact perturbed\-word recallRrecallR\_\{\\mathrm\{recall\}\}:

R⁡\(y,z\)=\{Redit​\(y,z\),regular documents,η​Redit​\(y,z\)\+\(1−η\)​Rrecall​\(y,z\),text\-perturbed documents\.R\(y,z\)=\\begin\{cases\}R\_\{\\mathrm\{edit\}\}\(y,z\),&\\text\{regular documents\},\\\\ \\eta R\_\{\\mathrm\{edit\}\}\(y,z\)\+\(1\-\\eta\)R\_\{\\mathrm\{recall\}\}\(y,z\),&\\text\{text\-perturbed documents\}\.\\end\{cases\}Here,η∈\[0,1\]\\eta\\in\[0,1\]controls the balance between edit similarity and perturbed\-word recall\. Both components lie in\[0,1\]\[0,1\]; their definitions are given in Appendix[D](https://arxiv.org/html/2609.38282#A4)\.

The group\-mean reward isR¯=G−1​∑i=1GRi\\bar\{R\}=G^\{\-1\}\\sum\_\{i=1\}^\{G\}R\_\{i\}\. For groups withg=1g=1, we use

fκ​\(R¯\)=e−κ​R¯−e−κ1−e−κ\.f\_\{\\kappa\}\(\\bar\{R\}\)=\\frac\{e^\{\-\\kappa\\bar\{R\}\}\-e^\{\-\\kappa\}\}\{1\-e^\{\-\\kappa\}\}\.\(4\)The parameterκ\>0\\kappa\>0controls the decay rate\. Withfκ​\(0\)=1f\_\{\\kappa\}\(0\)=1andfκ​\(1\)=0f\_\{\\kappa\}\(1\)=0, this factor continuously weakens teacher guidance as group\-mean reward increases\. Gating and attenuation together give the group weightg​fκ​\(R¯\)gf\_\{\\kappa\}\(\\bar\{R\}\)\.

### 3\.3Student\-Adaptive Forward KL

For the local gradient comparison at a fixed generated prefix, write𝐩\\mathbf\{p\}and𝐪\\mathbf\{q\}for the student and teacher probability vectors\. Sampled\-token estimators of reverse KL,DKL\(𝐩∥𝐪\)D\_\{\\mathrm\{KL\}\}\(\\mathbf\{p\}\\\|\\mathbf\{q\}\), naturally use student rollouts to provide on\-policy token\-level credit\. Their direct signals act on sampled actions: they can penalize a student\-preferred error without explicitly specifying which unsampled alternative to promote\([Liu et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib14);[Han et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib15)\)\. FKL instead supplies teacher\-distribution targets at student\-generated prefixes\([Agarwal et al\., 2024](https://arxiv.org/html/2609.38282#bib.bib12)\), including teacher\-supported candidates absent from the sampled response\.

This direct correction can be strong under teacher–student disagreement\. Let𝐡\\mathbf\{h\}be the student logit vector, so𝐩=softmax⁡\(𝐡\)\\mathbf\{p\}=\\operatorname\{softmax\}\(\\mathbf\{h\}\)\. Letaadenote the sampled token andAAits fixed group\-relative advantage\. The negative unclipped token\-level GRPO surrogate, denotedℓGRPO\\ell\_\{\\mathrm\{GRPO\}\}, atθ=θold\\theta=\\theta\_\{\\mathrm\{old\}\}and the full\-vocabulary FKL lossℓFKL=DKL\(𝐪∥𝐩\)\\ell\_\{\\mathrm\{FKL\}\}=D\_\{\\mathrm\{KL\}\}\(\\mathbf\{q\}\\\|\\mathbf\{p\}\)have logit gradients

∇𝐡ℓGRPO\\displaystyle\\nabla\_\{\\mathbf\{h\}\}\\ell\_\{\\mathrm\{GRPO\}\}=A⁡\(𝐩−𝐞a\),\\displaystyle=A\(\\mathbf\{p\}\-\\mathbf\{e\}\_\{a\}\),\(5\)∇𝐡ℓFKL\\displaystyle\\nabla\_\{\\mathbf\{h\}\}\\ell\_\{\\mathrm\{FKL\}\}=𝐩−𝐪,\\displaystyle=\\mathbf\{p\}\-\\mathbf\{q\},where𝐞a\\mathbf\{e\}\_\{a\}is the one\-hot vector foraa\. A teacher\-supported candidate with very low student probability can receive a much stronger FKL signal than its policy\-gradient signal, particularly when the student is highly confident in its sampled token\. The FKL logit gradient is bounded, but its relative scale can be large\. Since GRPO clipping does not constrain an added FKL term, excessive auxiliary contributions may destabilize policy updates\.

We use Student\-Adaptive Forward KL \(SA\-FKL\) to moderate local updates when the student assigns low probability to the teacher’s preferred candidate\. Letbi,tb\_\{i,t\}denote the teacher’s Top\-1 token at the student\-generated prefix\. The probability weight and forward\-KL term are

bi,t\\displaystyle b\_\{i,t\}=arg⁡maxv∈𝒱​πθT​\(v∣xT,y<t\(i\)\),\\displaystyle=\\arg\\max\_\{v\\in\\mathcal\{V\}\}\\pi\_\{\\theta\_\{\\mathrm\{T\}\}\}\(v\\mid x\_\{\\mathrm\{T\}\},y\_\{<t\}^\{\(i\)\}\),\(6\)wi,t​\(θ\)\\displaystyle w\_\{i,t\}\(\\theta\)=sg⁡\[πθ​\(bi,t∣x,y<t\(i\)\)\],\\displaystyle=\\operatorname\{sg\}\\\!\\left\[\\pi\_\{\\theta\}\(b\_\{i,t\}\\mid x,y\_\{<t\}^\{\(i\)\}\)\\right\],di,t​\(θ\)\\displaystyle d\_\{i,t\}\(\\theta\)=DKL\(πθT\(⋅∣xT,y<t\(i\)\)∥πθ\(⋅∣x,y<t\(i\)\)\)\.\\displaystyle=D\_\{\\mathrm\{KL\}\}\\\!\\left\(\\pi\_\{\\theta\_\{\\mathrm\{T\}\}\}\(\\cdot\\mid x\_\{\\mathrm\{T\}\},y\_\{<t\}^\{\(i\)\}\)\\,\\middle\\\|\\,\\pi\_\{\\theta\}\(\\cdot\\mid x,y\_\{<t\}^\{\(i\)\}\)\\right\)\.Here,𝒱\\mathcal\{V\}is the vocabulary andsg\\operatorname\{sg\}stops gradient propagation through the weight\. The token contribution iswi,t​\(θ\)​di,t​\(θ\)w\_\{i,t\}\(\\theta\)d\_\{i,t\}\(\\theta\), so the student’s probability of the teacher’s Top\-1 token scales the full local KL gradient without changing its direction\. When the teacher favors a candidate with little student support, this weight attenuates the potentially strong FKL update\. The gradient is derived in Appendix[C\.2](https://arxiv.org/html/2609.38282#A3.SS2)\.

### 3\.4Joint optimization\.

We average the weighted KL terms over valid tokens in each group and apply the group controls inside the rollout expectation:

𝒥OPD​\(θ\)\\displaystyle\\mathcal\{J\}\_\{\\mathrm\{OPD\}\}\(\\theta\)=𝔼roll​\[g​fκ​\(R¯\)​∑i=1G∑t=1Tiwi,t​\(θ\)​di,t​\(θ\)∑i=1GTi\],\\displaystyle=\\mathbb\{E\}\_\{\\mathrm\{roll\}\}\\\!\\left\[gf\_\{\\kappa\}\(\\bar\{R\}\)\\frac\{\\sum\_\{i=1\}^\{G\}\\sum\_\{t=1\}^\{T\_\{i\}\}w\_\{i,t\}\(\\theta\)d\_\{i,t\}\(\\theta\)\}\{\\sum\_\{i=1\}^\{G\}T\_\{i\}\}\\right\],\(7\)𝒥⁡\(θ\)\\displaystyle\\mathcal\{J\}\(\\theta\)=𝒥GRPO​\(θ\)−λ​𝒥OPD​\(θ\)\.\\displaystyle=\\mathcal\{J\}\_\{\\mathrm\{GRPO\}\}\(\\theta\)\-\\lambda\\mathcal\{J\}\_\{\\mathrm\{OPD\}\}\(\\theta\)\.GAD\-RL maximizes𝒥\\mathcal\{J\}\. Since𝒥OPD\\mathcal\{J\}\_\{\\mathrm\{OPD\}\}is a distillation cost, it enters with a minus sign;λ\\lambdacontrols its strength\. The gate and attenuation are computed separately for each sampled group and held fixed during the update\. The vocabulary\-truncated approximation is specified in Appendix[D](https://arxiv.org/html/2609.38282#A4)\.

## 4Experiments

#### Experimental setup\.

For each backbone, the teacher and student start from the same base model\. The teacher is fine\-tuned on 200,000 transcription\-task examples\. Document\-parsing training uses a mixed dataset of 14,400 samples, with a 3:2 ratio of general to text\-perturbed documents\. General documents are sampled from Infinity\-Doc2\-5M\([Huang et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib5)\)and MonkeyDoc\([Li et al\., 2025a](https://arxiv.org/html/2609.38282#bib.bib6)\)\. We cross\-check their annotations against parsing outputs from PaddleOCR\-VL\-1\.6\([Zhang et al\., 2026b](https://arxiv.org/html/2609.38282#bib.bib7)\)and discard samples with text similarity below 0\.9\. All student training strategies start from the base checkpoint and run for 300 steps with a learning rate of10−610^\{\-6\}\. Unless otherwise specified, we report benchmark results from the checkpoint after 300 training steps\. Detailed training settings are provided in Appendix[D](https://arxiv.org/html/2609.38282#A4)\.

#### Baselines\.

We evaluate SFT, GRPO, and GRPO\+OPD \(fixed\-weight\) on both Qwen3\-VL\-2B and Qwen3\.5\-2B\.GRPO\+OPD \(fixed\-weight\)adds a fixed\-weight distillation loss directly to the GRPO objective on all responses\. The default distillation coefficient isλ=0\.005\\lambda=0\.005, used by both GAD\-RL and GRPO\+OPD \(fixed\-weight\) in the main experiments\.

### 4\.1Benchmarks and Metrics

*GlitchText*\([Yao et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib1)\)introduces controlled errors into familiar passages rendered on a plain background\. We report Identification Rate \(Ident\) and Correction Rate \(Cor\), averaged equally across Chinese and English\.*CHAOS\-Bench*\([Li et al\., 2026](https://arxiv.org/html/2609.38282#bib.bib27)\)introduces character corruptions into 500 academic\-paper page images with realistic layouts; we report Micro Recall\. For general document parsing, we use*OmniDocBench v1\.6*\([Ouyang et al\., 2025](https://arxiv.org/html/2609.38282#bib.bib20)\)and report its Overall score\. Metric definitions, matching rules, and aggregation details are provided in Appendix[E](https://arxiv.org/html/2609.38282#A5)\.

### 4\.2Main Results

Table[1](https://arxiv.org/html/2609.38282#S4.T1)compares the original model \(Baseline\), the post\-training baselines, and GAD\-RL for each student model\.

Table 1:Transcription faithfulness and general document parsing, grouped by student model\. Micro Recall is a ratio; other scores are percentages\. Bold marks the best result within each model block\.MethodCHAOS\-BenchGlitchTextOmniDocBench v1\.6Micro Recall↑\\uparrowIdent↑\\uparrowCor↓\\downarrowOverall↑\\uparrowStudent: Qwen3\.5\-2BBaseline0\.040276\.4814\.9480\.06SFT0\.376382\.2010\.7190\.29GRPO0\.514791\.304\.2190\.80GRPO\+OPD \(fixed\-weight\)0\.554991\.313\.1490\.90GAD\-RL0\.599293\.393\.4091\.18Student: Qwen3\-VL\-2BBaseline0\.023583\.918\.2846\.39SFT0\.380487\.757\.0087\.97GRPO0\.525588\.334\.3989\.86GRPO\+OPD \(fixed\-weight\)0\.416789\.264\.6987\.67GAD\-RL0\.563792\.153\.5990\.00On CHAOS\-Bench, GAD\-RL improves Micro Recall over GRPO by 8\.45 and 3\.82 percentage points on Qwen3\.5\-2B and Qwen3\-VL\-2B, respectively, and outperforms fixed\-weight GRPO\+OPD on both backbones\.

On GlitchText, GAD\-RL improves anomaly identification and reduces overcorrection relative to SFT and GRPO on both backbones; fixed\-weight OPD retains a lower correction rate on Qwen3\.5\-2B\. GAD\-RL also maintains comparable general document\-parsing performance on OmniDocBench\.

#### Training dynamics\.

GRPO leads early in training, but GAD\-RL overtakes it and finishes with higher mean reward and perturbed\-word Pass@8 \(Figure[2](https://arxiv.org/html/2609.38282#S4.F2)a,b\)\.

Figure 2:GRPO and GAD\-RL training dynamics\.\(a\) Mean reward\. \(b\) Fraction of groups with at least one of eight responses achieving perturbed\-word Recall of 1\. For GAD\-RL, \(c\) the fraction of groups with distillation disabled by the gate and \(d\) the mean attenuation factor among the remaining active groups\.The two controls play complementary roles over training\.The gate disables an increasing fraction of groups, while the mean attenuation factor among active groups falls from about 0\.12 to 0\.02 \(Figure[2](https://arxiv.org/html/2609.38282#S4.F2)c,d\)\. Teacher guidance therefore weakens even on groups that remain active\.

### 4\.3Ablation Studies

#### KL formulation in on\-policy distillation\.

Following prior work on divergence choices in on\-policy distillation\([Agarwal et al\., 2024](https://arxiv.org/html/2609.38282#bib.bib12)\), we compare K1 PG, reverse KL \(RKL\), Jensen–Shannon divergence \(JSD\), FKL, and SA\-FKL on Qwen3\.5\-2B within the GRPO\+OPD \(fixed\-weight\) setup, replacing only the distillation loss while keeping the coefficient atλ\\lambdaand all other settings fixed\.

Figure 3:FKL and SA\-FKL training dynamics\.With otherwise identical settings: \(a\) mean training reward \(EMA 0\.95\); \(b\) CHAOS\-Bench Micro Recall, with a step\-0 reference of 0\.0402\.SA\-FKL achieves the highest CHAOS\-Bench Recall after 100 steps \(Table[3](https://arxiv.org/html/2609.38282#S4.T3)\)\. FKL learns faster initially but shows a reward dip around steps 50–60; SA\-FKL is smoother over this interval and achieves higher Recall from step 40 onward \(Figure[3](https://arxiv.org/html/2609.38282#S4.F3)\)\.

#### Distillation coefficient\.

Among the four nonzero coefficients in Table[3](https://arxiv.org/html/2609.38282#S4.T3), onlyλ=0\.005\\lambda=0\.005outperforms GRPO\. Neither halving nor increasing this value helps, showing that fixed\-weight OPD is sensitive to coefficient selection\.

Table 2:Distillation objectives onCHAOS\-Benchat step 100\.
Table 3:Fixed\-weight distillation at different coefficients\.

#### Gating and attenuation\.

Table[4](https://arxiv.org/html/2609.38282#S4.T4)removes student\-probability weighting \(SW\), mastery\-gated distillation \(MGD\), or reward\-aware attenuation \(RAA\) atλ\\lambda\. Removing SW or RAA causes larger Recall losses than removing MGD\. The smaller incremental benefit of gating is consistent with RAA already suppressing high\-reward groups: atR¯=0\.9\\bar\{R\}=0\.9, it retains only1\.83%1\.83\\%of the base coefficient before token weighting\.

Table 4:Component ablations on CHAOS\-Bench atλ\\lambda\.

## 5Analysis

Building on Figure[1](https://arxiv.org/html/2609.38282#S1.F1), we examine whether the teacher prefers incorrect candidates at correct student outputs and how supervision quality varies with group reward within each checkpoint\. We then compare teacher\-intervention strategies through training experiments\.

### 5\.1Teacher Supervision as the Student Improves

Lowering a correct token’s probability does not necessarily imply an incorrect teacher Top\-1 prediction\. We therefore inspect teacher Top\-1 predictions at correct student outputs, reusing Figure[1](https://arxiv.org/html/2609.38282#S1.F1)’s 1,000 pages excluded from training, GRPO\-only student checkpoints, and frozen teacher\.

At the same student\-generated prefix, we count a teacher error when its Top\-1 token conflicts with the GT continuation\. Among correctly transcribed perturbed\-word\-associated tokens, this rate rises from 9\.52% at step 100 to 42\.88% at step 300 \(Table[5](https://arxiv.org/html/2609.38282#S5.T5)\), showing an increasing preference for incorrect candidates at positions the student already transcribes correctly\.

Table 5:Teacher Top\-1 errors at correct student outputs\. Counts refer to perturbed\-word\-associated tokens\. The final columns give teacher errors as percentages of all eligible tokens and of correct student tokens\.Token counts vary across checkpoints because the generated responses differ, resulting in different numbers of tokens aligned to perturbed\-word spans; detailed statistics and matching rules are provided in Appendix[B\.4](https://arxiv.org/html/2609.38282#A2.SS4)\.

#### Reward dependence within each checkpoint\.

Across\-checkpoint trends combine changes in training stage and task performance\. We therefore examine reward dependence within GRPO\+OPD checkpoints at steps 50 and 200, scoring current\-policy rollouts with the same frozen teacher\. Figure[4](https://arxiv.org/html/2609.38282#S5.F4)partitions 995 matched diagnostic page groups by mean task reward; scoring and bootstrap details are in Appendix[B\.5](https://arxiv.org/html/2609.38282#A2.SS5)\.

Figure 4:Teacher supervision becomes less favorable as group reward increases\.GRPO\+OPD groups at steps 50 and 200 are binned by mean rewardR¯\\bar\{R\}\. For perturbed\-word\-associated tokens: \(a\) beneficial\-to\-harmful count ratio \(log scale; dotted line: equal counts\); \(b\) beneficial and harmful fractions, with the remainder neutral\. The absolute probability\-change margin is 0\.1\.nngives group counts at step 50 / 200; error bars show 95% input\-group\-bootstrap confidence intervals\.Within both checkpoints, beneficial signals become less frequent as group reward increases, while harmful signals do not decline proportionally\. Consequently, useful teacher corrections become scarcer relative to conflicting supervision as the student performs better on an input\. This within\-checkpoint pattern supports using group reward to regulate teacher intervention: a training\-step schedule assigns the same coefficient to groups with different relative corrective opportunities\. Reward\-aware attenuation can instead reduce guidance on groups that already achieve high task reward\.

### 5\.2Distillation Control Strategies

How should teacher intervention be controlled?Section[5\.1](https://arxiv.org/html/2609.38282#S5.SS1)motivates comparing uniform changes in strength, response selection, and decay over training steps with state\-aware regulation\. Building on the coefficient ablation in Table[3](https://arxiv.org/html/2609.38282#S4.T3), Table[6](https://arxiv.org/html/2609.38282#S5.T6.fig1)compares filtering and decay strategies\. These Qwen3\.5\-2B comparisons retain the GRPO task loss and use a base coefficient ofλ=0\.005\\lambda=0\.005\. All distillation\-control comparisons use SA\-FKL; fixed\-weight OPD applies it to all responses\. Recall is evaluated on CHAOS\-Bench\. Low\-score OPD selects individual responses with total rewardRi<0\.5R\_\{i\}<0\.5; GAD\-RL gates whole groups and attenuates their weight using group\-mean reward\.

1The base coefficient isλ=0\.005\\lambda=0\.005; token\-level weights are omitted\. The gateggand attenuationfκ​\(R¯\)f\_\{\\kappa\}\(\\bar\{R\}\)are defined in Sections[3\.1](https://arxiv.org/html/2609.38282#S3.SS1)and[3\.2](https://arxiv.org/html/2609.38282#S3.SS2), respectively\. The factors⁡\(u\)s\(u\)decreases linearly with training stepuu\.

Table 6:Distillation decay and filtering strategies\. The best Recall is bold\.#### Response selection and temporal decay\.

Compared with GRPO\+OPD \(fixed\-weight\), low\-score OPD and linear\-decay OPD improve Recall by 2\.26 and 1\.44 percentage points, respectively\. These gains are consistent with the changing utility of teacher supervision: response selection concentrates guidance on outputs needing correction, while temporal decay relaxes teacher constraints as training progresses\.

#### GAD\-RL coefficient\.

Increasing GAD\-RL’s base coefficient tenfold lowers Recall by 1\.10 percentage points; both tested settings outperform GRPO\. The strong attenuation on active groups shown in Figure[2](https://arxiv.org/html/2609.38282#S4.F2)d may help explain this modest change: teacher guidance weakens as group rewards improve, even when the base coefficient is larger\.

## 6Conclusions

We presented GAD\-RL, which combines GRPO with group\-gated, reward\-attenuated, and student\-weighted on\-policy distillation\. Offline analysis shows that teacher–student disagreement increasingly affects correctly transcribed perturbed\-word\-associated tokens as the student improves, while favorable suppression of residual errors remains common\. GAD\-RL improves transcription faithfulness over SFT and GRPO on two student models while maintaining comparable OmniDocBench performance\. On Qwen3\.5\-2B, GAD\-RL outperforms GRPO at both tested distillation coefficients\. These findings support adapting teacher intervention to the student’s evolving task performance and local distributions\.

#### Limitations and future work\.

Our evaluation is limited to OCR transcription faithfulness in document parsing\. Future work will explore GAD\-RL on other tasks to assess whether policy\-state\-aware distillation control generalizes beyond OCR\.

### AI use statement

In this work, we used generative AI tools to assist with literature search and summarization, the design of mathematical notation, and formula verification\. AI tools also generated most of the code used for synthetic data generation, figure visualization, and method implementation\. The authors manually reviewed and validated the code and take responsibility for the final manuscript, mathematical claims, implementation, and reported results\.

### Reproducibility statement

Appendix[A](https://arxiv.org/html/2609.38282#A1)describes data synthesis, Appendix[B](https://arxiv.org/html/2609.38282#A2)details teacher\-signal analysis, and Appendix[C](https://arxiv.org/html/2609.38282#A3)provides the gradient derivations\. Training configurations and objectives are specified in Appendix[D](https://arxiv.org/html/2609.38282#A4), and evaluation metrics in Appendix[E](https://arxiv.org/html/2609.38282#A5)\. Upon acceptance, we will release the training, evaluation, and analysis code, the data\-synthesis pipeline, teacher\-training data, and the 1,000\-page analysis set\. Data will be distributed as files where their licenses permit, or through source manifests and reconstruction instructions where redistribution is restricted\.

### Ethics statement

This work studies faithful transcription using synthetic text perturbations and existing document benchmarks\. Synthetic documents are derived from arXiv LaTeX sources, whose reuse conditions depend on each source’s license\. Third\-party resources, including GlitchText, CHAOS\-Bench, and OmniDocBench, remain subject to their original licenses and terms\. Releases of derived data will preserve source attribution and respect redistribution restrictions\. Perturbed text is intended to test transcription fidelity and should not be treated as factual content\. Improved transcription faithfulness does not establish the truth of the underlying documents\.

## References

- Agarwalet al\.\(2024\)R\. Agarwal, N\. Vieillard, Y\. Zhou, P\. Stanczyk, S\. Ramos, M\. Geist, and O\. BachemOn\-policy distillation of language models: learning from self\-generated mistakes\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.38282#S1.p2.1),[§2](https://arxiv.org/html/2609.38282#S2.p3.1),[§3\.3](https://arxiv.org/html/2609.38282#S3.SS3.p1.1),[§4\.3](https://arxiv.org/html/2609.38282#S4.SS3.SSS0.Px1.p1.1)\.
- Akhondzadehet al\.\(2026\)M\. S\. Akhondzadeh, V\. Lingam, A\. Tejaswi, C\. Ekbote, S\. Sanghavi, and A\. BojchevskiReward\-gated on\-policy distillation\.arXiv preprint arXiv:2607\.04037\.External Links:[Link](https://arxiv.org/abs/2607.04037)Cited by:[§1](https://arxiv.org/html/2609.38282#S1.p3.1),[§2](https://arxiv.org/html/2609.38282#S2.p4.1)\.
- Baiet al\.\(2025\)S\. Bai, K\. Chen, X\. Liu, J\. Wang, W\. Ge, S\. Song, K\. Dang, P\. Wang, S\. Wang, J\. Tang,et al\.Qwen2\.5\-VL technical report\.arXiv preprint arXiv:2502\.13923\.External Links:[Link](https://arxiv.org/abs/2502.13923)Cited by:[§1](https://arxiv.org/html/2609.38282#S1.p1.1)\.
- Blecheret al\.\(2024\)L\. Blecher, G\. Cucurull, T\. Scialom, and R\. StojnicNougat: neural optical understanding for academic documents\.InInternational Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p1.1)\.
- Ding and Zhang \(2026\)Y\. Ding and R\. ZhangDoes on\-policy distillation really distill? From noisy teacher to self\-improvement\.arXiv preprint arXiv:2608\.31046\.External Links:[Link](https://arxiv.org/abs/2608.31046)Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p4.1)\.
- Hanet al\.\(2026\)Z\. Han, J\. Xiao, Z\. Lu, R\. Jin, Z\. Yao, Y\. Liu, H\. Hao, Y\. Sun, Y\. Yang, Q\. Gu, X\. Cai, and D\. XiongDistill where you fail: recovering learning signals of negative RL\-groups from adaptive teacher guidance\.arXiv preprint arXiv:2608\.00782\.External Links:[Link](https://arxiv.org/abs/2608.00782)Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p4.1),[§3\.1](https://arxiv.org/html/2609.38282#S3.SS1.p1.1),[§3\.3](https://arxiv.org/html/2609.38282#S3.SS3.p1.1)\.
- Huanget al\.\(2024\)Q\. Huang, X\. Dong, P\. Zhang, B\. Wang, C\. He, J\. Wang, D\. Lin, W\. Zhang, and N\. YuOPERA: alleviating hallucination in multi\-modal large language models via over\-trust penalty and retrospection\-allocation\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 13418–13427\.Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p2.1)\.
- Huanget al\.\(2026\)Z\. Huang, J\. Huang, K\. Ren, B\. Wang, W\. Li, J\. Feng, Y\. Wang, Y\. Yao, S\. Lin, Y\. Tang, C\. Peng, W\. Xu, W\. Chu, Y\. Xu, and Y\. QiInfinity\-Parser2 technical report\.arXiv preprint arXiv:2607\.07836\.External Links:[Link](https://arxiv.org/abs/2607.07836)Cited by:[§4](https://arxiv.org/html/2609.38282#S4.SS0.SSS0.Px1.p1.1)\.
- Janget al\.\(2026\)I\. Jang, J\. Yeom, J\. Yeo, H\. Lim, and T\. KimStable on\-policy distillation through adaptive target reformulation\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 42217–42227\.Cited by:[§1](https://arxiv.org/html/2609.38282#S1.p3.1),[§2](https://arxiv.org/html/2609.38282#S2.p3.1)\.
- Kimet al\.\(2022\)G\. Kim, T\. Hong, M\. Yim, J\. Nam, J\. Park, J\. Yim, W\. Hwang, S\. Yun, D\. Han, and S\. ParkOCR\-free document understanding transformer\.InEuropean Conference on Computer Vision,pp\. 498–517\.Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p1.1)\.
- Leeet al\.\(2026\)G\. G\. Lee, K\. E\. Ak, J\. Mohta, Y\. Xu, and D\. DimitriadisDo VLMs read or rewrite? On transcription faithfulness in vision\-language models\.arXiv preprint arXiv:2607\.21617\.External Links:[Link](https://arxiv.org/abs/2607.21617)Cited by:[§1](https://arxiv.org/html/2609.38282#S1.p1.1),[§2](https://arxiv.org/html/2609.38282#S2.p1.1)\.
- Leeet al\.\(2023\)K\. Lee, M\. Joshi, I\. R\. Turc, H\. Hu, F\. Liu, J\. M\. Eisenschlos, U\. Khandelwal, P\. Shaw, M\. Chang, and K\. ToutanovaPix2Struct: screenshot parsing as pretraining for visual language understanding\.InInternational Conference on Machine Learning,pp\. 18893–18912\.Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p1.1)\.
- Lenget al\.\(2024\)S\. Leng, H\. Zhang, G\. Chen, X\. Li, S\. Lu, C\. Miao, and L\. BingMitigating object hallucinations in large vision\-language models through visual contrastive decoding\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 13872–13882\.Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p2.1)\.
- Liet al\.\(2026\)G\. Li, X\. Wan, S\. Peng, W\. Wang, H\. Feng, Y\. Du, B\. Wu, Z\. Ruan, Z\. Lu, L\. Wu,et al\.HunyuanOCR\-1\.5: making lightweight OCR VLMs faster and better\.arXiv preprint arXiv:2607\.04884\.External Links:[Link](https://arxiv.org/abs/2607.04884)Cited by:[§1](https://arxiv.org/html/2609.38282#S1.p5.1),[§2](https://arxiv.org/html/2609.38282#S2.p1.1),[§4\.1](https://arxiv.org/html/2609.38282#S4.SS1.p1.1)\.
- Liet al\.\(2025a\)Z\. Li, Y\. Liu, Q\. Liu, Z\. Ma, Z\. Zhang, S\. Zhang, B\. Yang, Z\. Guo, J\. Zhang, X\. Wang, and X\. BaiMonkeyOCR: document parsing with a structure\-recognition\-relation triplet paradigm\.arXiv preprint arXiv:2506\.05218\.External Links:[Link](https://arxiv.org/abs/2506.05218)Cited by:[§4](https://arxiv.org/html/2609.38282#S4.SS0.SSS0.Px1.p1.1)\.
- Liet al\.\(2025b\)Z\. Li, H\. Shi, Y\. Gao, D\. Liu, Z\. Wang, Y\. Chen, T\. Liu, L\. Zhao, H\. Wang, and D\. N\. MetaxasThe hidden life of tokens: reducing hallucination of large vision\-language models via visual information steering\.InInternational Conference on Machine Learning,pp\. 35799–35819\.Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p2.1)\.
- Liuet al\.\(2026\)X\. Liu, K\. Jiao, C\. Xiao, R\. Zhao, J\. Ruan, B\. Li, J\. Liu, Q\. Wang, X\. Chen, J\. Wang, C\. Wang, T\. Xiao, and J\. ZhuTeacher\-guided policy optimization for on\-policy reasoning distillation under large policy divergence\.arXiv preprint arXiv:2605\.13230\.External Links:[Link](https://arxiv.org/abs/2605.13230)Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p3.1),[§3\.3](https://arxiv.org/html/2609.38282#S3.SS3.p1.1)\.
- Luet al\.\(2026\)S\. Lu, Y\. Li, Y\. Xia, Y\. Chen, A\. Ji, J\. Jiang, Q\. Chen, J\. Zhao, E\. Lin, H\. Li, C\. Qin, Z\. Xu, and W\. LuoOvisOCR2 technical report\.arXiv preprint arXiv:2607\.13639\.External Links:[Link](https://arxiv.org/abs/2607.13639)Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p1.1)\.
- Ouyanget al\.\(2025\)L\. Ouyang, Y\. Qu, H\. Zhou, J\. Zhu, R\. Zhang, Q\. Lin, B\. Wang, Z\. Zhao, M\. Jiang, X\. Zhao,et al\.OmniDocBench: benchmarking diverse PDF document parsing with comprehensive annotations\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 24838–24848\.Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p1.1),[§4\.1](https://arxiv.org/html/2609.38282#S4.SS1.p1.1)\.
- Poznanskiet al\.\(2025\)J\. Poznanski, L\. Soldaini, and K\. LoolmOCR 2: unit test rewards for document OCR\.arXiv preprint arXiv:2510\.19817\.External Links:[Link](https://arxiv.org/abs/2510.19817)Cited by:[§1](https://arxiv.org/html/2609.38282#S1.p1.1),[§1](https://arxiv.org/html/2609.38282#S1.p2.1),[§2](https://arxiv.org/html/2609.38282#S2.p1.1)\.
- Sarkaret al\.\(2025\)P\. Sarkar, S\. Ebrahimi, A\. Etemad, A\. Beirami, S\. Arik, and T\. PfisterMitigating object hallucination in MLLMs via data\-augmented phrase\-level alignment\.InInternational Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p2.1)\.
- Shaoet al\.\(2024\)Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. K\. Li, Y\. Wu,et al\.DeepSeekMath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.External Links:[Link](https://arxiv.org/abs/2402.03300)Cited by:[§1](https://arxiv.org/html/2609.38282#S1.p2.1),[§3](https://arxiv.org/html/2609.38282#S3.SS0.SSS0.Px1.p1.3)\.
- Sunet al\.\(2024\)Z\. Sun, S\. Shen, S\. Cao, H\. Liu, C\. Li, Y\. Shen, C\. Gan, L\. Gui, Y\. Wang, Y\. Yang,et al\.Aligning large multimodal models with factually augmented RLHF\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 13088–13110\.Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p2.1)\.
- Wanget al\.\(2026\)B\. Wang, B\. Wu, W\. Li, M\. Fang, Z\. Huang, J\. Huang, Y\. Liang, H\. Wang, L\. Chen, W\. Chu, and Y\. QiInfinity\-Parser: layout\-aware reinforcement learning with high\-quality document parsing dataset\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 1647–1667\.Cited by:[§1](https://arxiv.org/html/2609.38282#S1.p2.1),[§2](https://arxiv.org/html/2609.38282#S2.p1.1)\.
- Xuet al\.\(2025\)H\. Xu, Q\. Zhu, H\. Deng, J\. Li, L\. Hou, Y\. Wang, L\. Shang, R\. Xu, and F\. MiKDRL: post\-training reasoning LLMs via unified knowledge distillation and reinforcement learning\.arXiv preprint arXiv:2506\.02208\.External Links:[Link](https://arxiv.org/abs/2506.02208)Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p3.1)\.
- Xuet al\.\(2020\)Y\. Xu, M\. Li, L\. Cui, S\. Huang, F\. Wei, and M\. ZhouLayoutLM: pre\-training of text and layout for document image understanding\.New York, NY, USA,pp\. 1192–1200\.Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p1.1)\.
- Xuet al\.\(2026\)Y\. Xu, H\. Sang, Z\. Zhou, R\. He, and Z\. WangPACED: distillation and on\-policy self\-distillation at the frontier of student competence\.arXiv preprint arXiv:2603\.11178\.External Links:[Link](https://arxiv.org/abs/2603.11178)Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p4.1)\.
- Yaoet al\.\(2026\)Y\. Yao, M\. Liao, W\. Zhang, Z\. Li, and H\. ZhaoPAR: training\-free positional perturbation and attention recycling for faithful OCR\.InAnnual Meeting of the Association for Computational Linguistics,pp\. 23258–23273\.Cited by:[§1](https://arxiv.org/html/2609.38282#S1.p1.1),[§1](https://arxiv.org/html/2609.38282#S1.p5.1),[§2](https://arxiv.org/html/2609.38282#S2.p2.1),[§4\.1](https://arxiv.org/html/2609.38282#S4.SS1.p1.1)\.
- Yuet al\.\(2025\)Q\. Yu, Z\. Zhang, R\. Zhu, Y\. Yuan, X\. Zuo, Y\. Yue, W\. Dai, T\. Fan, G\. Liu, J\. Liu,et al\.DAPO: an open\-source LLM reinforcement learning system at scale\.InAdvances in Neural Information Processing Systems,Vol\.38\.Cited by:[§1](https://arxiv.org/html/2609.38282#S1.p2.1)\.
- Yuet al\.\(2024\)T\. Yu, Y\. Yao, H\. Zhang, T\. He, Y\. Han, G\. Cui, J\. Hu, Z\. Liu, H\. Zheng, M\. Sun,et al\.RLHF\-V: towards trustworthy MLLMs via behavior alignment from fine\-grained correctional human feedback\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 13807–13816\.Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p2.1)\.
- Zhanget al\.\(2026a\)Y\. Zhang, X\. Ma, Z\. Tan, and Z\. DongI\-SDPO: instance\-level adaptive self\-distillation policy optimization\.arXiv preprint arXiv:2608\.12957\.External Links:[Link](https://arxiv.org/abs/2608.12957)Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p4.1)\.
- Zhanget al\.\(2026b\)Z\. Zhang, H\. Liu, S\. Liang, Y\. Zhang, Y\. Xiang, J\. Liu, T\. Sun, M\. Lin, Y\. Zhang, C\. Zhou, T\. Gao, C\. Cui, Y\. Liu, D\. Yu, and Y\. MaPaddleOCR\-VL\-1\.6: expanding the frontier of document parsing with under\-optimized region refinement and progressive post\-training\.arXiv preprint arXiv:2606\.03264\.External Links:[Link](https://arxiv.org/abs/2606.03264)Cited by:[§4](https://arxiv.org/html/2609.38282#S4.SS0.SSS0.Px1.p1.1)\.
- Zhaoet al\.\(2026\)S\. Zhao, Z\. Xie, M\. Liu, J\. Huang, G\. Pang, F\. Chen, and A\. GroverSelf\-distilled reasoner: on\-policy self\-distillation for large language models\.arXiv preprint arXiv:2601\.18734\.External Links:[Link](https://arxiv.org/abs/2601.18734)Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p3.1)\.
- Zhaoet al\.\(2023\)Z\. Zhao, B\. Wang, L\. Ouyang, X\. Dong, J\. Wang, and C\. HeBeyond hallucinations: enhancing LVLMs through hallucination\-aware direct preference optimization\.arXiv preprint arXiv:2311\.16839\.External Links:[Link](https://arxiv.org/abs/2311.16839)Cited by:[§2](https://arxiv.org/html/2609.38282#S2.p2.1)\.

## Appendix AData Synthesis

### A\.1Text\-perturbed documents

We construct text\-perturbed documents from arXiv LaTeX projects\. Text, headings, formulas, and tables are reflowed into academic layouts to produce page images and matching Markdown targets\. The target preserves the anomalous text printed in the image\.

![Refer to caption](https://arxiv.org/html/2609.38282v1/figure_05.png)Figure 5:Example of a synthetically perturbed document page\. Boxes highlight character\-level substitutions in “noisy”, “model”, “robustness”, and “strongest”\.#### Page construction\.

Source extraction removes citations, cross\-references, bibliography entries, and figure content\. Tables and formulas retain their source\-derived content; tables use HTML in the target\. Pages preserve the association of tables with captions and headings with following content\. Unsupported structures and overflowing pages are rejected\.

#### Controlled perturbations\.

Candidate words contain at least four letters, occur uniquely on the page, and have unambiguous source positions\. Each edit replaces one lowercase letter without changing word length, using a shared weighted character\-pair table\. We request three edits per page with probability 0\.4 and four with probability 0\.6\. Headings, captions, tables, formulas, code, and numerical content are protected; replacements colliding with existing words are skipped\. Figure[5](https://arxiv.org/html/2609.38282#A1.F5)shows an example\.

#### Verification\.

Substitutions are applied consistently to the rendering source and target\. Extracted PDF text and positions verify consistency and detect overflow, without generating or repairing the target\. A separate analysis set contains 1,000 pages from 147 papers, including 140 table pages and 3,564 perturbed words\.

### A\.2Teacher rewrite data

The teacher is fine\-tuned on 200,000 transcription\-task examples\. We build copying and controlled\-format rewrite pairs from source\-derived transcriptions, without OCR or LLM content rewriting\. Samples contain 1,000–7,800 tokens with tables kept intact, and train/validation splits are made by source paper\. Approximately 10% of eligible word occurrences are perturbed, with at least three edits, using the same character\-pair table\.

Input and target preserve identical content, including anomalous words, numbers, formulas, and tables\. Only existing heading prefixes and an optional outer Markdown fence may change; heading\-level changes need not preserve hierarchy\. Validation requires exact target reconstruction using only these permitted edits\.

## Appendix BToken Classification and Teacher\-Signal Analysis

### B\.1Fixed\-trajectory scoring

For each page, we generate a response using the Qwen3\.5\-2B\-based student with thinking disabled and a maximum generation length of 8,192 tokens\. Student and teacher score the same emitted token ID at the same response prefix: the student is conditioned on the image, and the teacher on a heading\-rewrite prompt containing the reference text\. All probabilities are computed over the full vocabulary\.

### B\.2Token alignment, categories, and correctness

Decoded responses are aligned to the reference after Unicode NFKC normalization and removal of recognized markup and whitespace\. Annotated perturbed words must match unambiguous whole\-word reference spans\. Ambiguous or invalid alignments are excluded from correctness statistics\.

Each response token is assigned one category, in priority order:*perturbed\-word\-associated*if it overlaps the aligned full span of a perturbed word;*formatting*if it contains only whitespace or markup; and*body*otherwise\. Perturbed\-word association includes correct copies and overcorrections, not just the edited character\. Visible numbers, punctuation, formulas, table\-cell text, and code remain content; tokens mixing markup and content are body unless perturbed\-word overlap takes priority\.

Categories do not imply correctness\. A content token is correct only when all its aligned normalized content characters match the reference; a correct subtoken within an erroneous word can therefore remain correct\. Statistics count eligible response tokens equally, excluding prompts, trailing special tokens, and invalid alignments\.

### B\.3Teacher information reliability

We compare the original Qwen3\.5\-2B self\-teacher and the SFT teacher specified in Appendix[D](https://arxiv.org/html/2609.38282#A4)on the same 1,000 student responses\. Both receive identical reference text and the training\-time teacher prompt\. Correctness denominators exclude 5,179 records with inconsistent decoded\-token offsets\.

Figure 6:Teacher signals on erroneous tokens\. For the same emitted token,Δ​p<−0\.1\\Delta p<\-0\.1is beneficial suppression andΔ​p\>0\.1\\Delta p\>0\.1is harmful reinforcement\. Panel \(b\) uses an enlarged horizontal scale\.#### Error suppression\.

WithΔ​p=pT−pS\\Delta p=p^\{\\mathrm\{T\}\}\-p^\{\\mathrm\{S\}\}, the SFT teacher suppresses 69\.58% of the 37,631 erroneous content tokens by more than 0\.1, compared with 20\.54% for the self\-teacher\. Harmful reinforcement falls from 5\.88% to 1\.48% \(Figure[6](https://arxiv.org/html/2609.38282#A2.F6)\)\. On perturbed\-word errors, beneficial suppression reaches 99\.85%\. These directions concern the emitted token; suppression alone does not verify a faithful alternative\.

#### Teacher\-choice ablation\.

Table[7](https://arxiv.org/html/2609.38282#A2.T7)uses the same Qwen3\.5\-2B GRPO and default SFT\-teacher results as Table[1](https://arxiv.org/html/2609.38282#S4.T1)\. The SFT teacher improves Recall over the self\-teacher by 5\.51 percentage points, supporting the use of a teacher trained for transcription\.

Table 7:Teacher\-choice ablation for GAD\-RL on CHAOS\-Bench\.

### B\.4Frozen\-teacher supervision across GRPO checkpoints

Figure[1](https://arxiv.org/html/2609.38282#S1.F1)evaluates each GRPO\-only checkpoint on the same analysis set described in Section[5\.1](https://arxiv.org/html/2609.38282#S5.SS1), with teacher scoring performed offline\. Initialization, training data, and the teacher prompt match the main experiments \(Appendix[D](https://arxiv.org/html/2609.38282#A4)\)\. Base is the RL initial checkpoint \(step 0\); students receive no OPD updates\.

#### Signal definition\.

For an eligible perturbed\-word\-associated token, letc∈\{0,1\}c\\in\\\{0,1\\\}indicate correctness, defineΔ​p=pT−pS\\Delta p=p\_\{\\mathrm\{T\}\}\-p\_\{\\mathrm\{S\}\}, and seth=\(2​c−1\)​Δ​ph=\(2c\-1\)\\Delta p\. A signal is beneficial forh\>0\.1h\>0\.1, harmful forh<−0\.1h<\-0\.1, and neutral otherwise\. Invalid alignments and formatting\-only tokens are excluded\. With countsBB,HH, andUU, the plotted fractions use denominatorN=B\+H\+UN=B\+H\+U, and directional SNR isB/HB/H\.

Table 8:Fixed\-teacher signal counts and perturbed\-word Recall on the analysis set\. Signal statistics use eligible emitted tokens; Recall uses all 3,564 annotated words, including omissions\.
#### Token\-count variation\.

These counts follow the tokens each checkpoint generates\. Across the same 3,493 perturbed words represented at every checkpoint, the average number of tokens per word rises from 1\.22 at Base to 2\.94, 3\.65, and 3\.79 at steps 100, 200, and 300\. Finer output tokenization accounts for most of the increase, while the 1,000 input pages remain fixed\.

#### Preservation as the student improves\.

Recall uses normalized reference\-span matching\. From Base through steps 100, 200, and 300, the teacher assigns probabilities more than 0\.1 below the student’s to 28\.57/33\.88/69\.24/76\.86% of correctly transcribed perturbed\-word\-associated tokens\. Favorable suppression signals on remaining erroneous tokens occur at rates of 99\.85/97\.59/100/100%, with denominators 3,252/166/38/25\. At step 300 all potentially harmful signals occur on correct tokens\. The aggregate decline thus accompanies greater directional disagreement on currently correct output, while favorable signals on residual errors remain prevalent\.

#### Teacher Top\-1 at suppressed correct tokens\.

We split correct perturbed\-word\-associated tokens withΔ​p<−0\.1\\Delta p<\-0\.1by teacher Top\-1 compatibility \(Table[9](https://arxiv.org/html/2609.38282#A2.T9)\)\. Using Appendix[B](https://arxiv.org/html/2609.38282#A2)’s normalization, we require consecutive exact\-match GT anchors for the emitted token\. Nonempty Top\-1 content matching the GT continuation’s prefix is compatible, including alternative tokenizations; nonmatching content is incompatible\. Unavailable alignments and formatting\-only or undecodable candidates remain unresolved\. This tests local compatibility, not whole\-word correctness\.

Table 9:Teacher Top\-1 at correct perturbed\-word\-associated tokens withΔ​p<−0\.1\\Delta p<\-0\.1\. Entries are counts \(percent ofNN\), including unresolved cases in the denominator\.These diagnostics count tokens on the analysis set\. Checkpoints emit different token sets, and small remaining error subsets limit rate comparisons\.

### B\.5Reward\-stratified teacher supervision

Figure[4](https://arxiv.org/html/2609.38282#S5.F4)uses Qwen3\.5\-2B GRPO\+OPD checkpoints at steps 50 and 200 and the same frozen SFT teacher\. The two checkpoints share 995 matched diagnostic page groups\. Of the original 1,000 diagnostic pages, five are excluded because model inference repeatedly failed for these inputs despite retries;

We use the correctness\-aware signal definition in Appendix[B\.4](https://arxiv.org/html/2609.38282#A2.SS4)with an absolute probability\-change margin of 0\.1\. Within each bin, beneficial and harmful fractions divide the respective pooled token counts by all evaluable perturbed\-word\-associated tokens, including neutral tokens\. Directional SNR is the ratio of the pooled beneficial and harmful counts\. Error bars are 95% bootstrap confidence intervals from 2,500 resamples of input groups within each reward bin, keeping responses from the same input clustered\. Empty bins have no estimate\.

## Appendix CGradient Scale of Controlled Distillation

### C\.1Policy gradients and unweighted FKL

All gradients below are with respect to the student logits𝐡\\mathbf\{h\}at unit temperature\. Using Equation[5](https://arxiv.org/html/2609.38282#S3.E5), letp⁡\(a\)=1−ξp\(a\)=1\-\\xiandq⁡\(a\)≤1−δq\(a\)\\leq 1\-\\deltafor0<ξ<δ≤10<\\xi<\\delta\\leq 1:ξ\\xiis the student’s probability mass away from sampled tokenaa, andδ\\deltalower\-bounds the corresponding teacher mass\. Atθ=θold\\theta=\\theta\_\{\\mathrm\{old\}\}, before clipping and group controls,

‖∇ℓGRPO‖1=2​\|A\|​ξ,2​\(δ−ξ\)≤‖∇ℓFKL‖1≤2\.\\\|\\nabla\\ell\_\{\\mathrm\{GRPO\}\}\\\|\_\{1\}=2\|A\|\\xi,\\qquad 2\(\\delta\-\\xi\)\\leq\\\|\\nabla\\ell\_\{\\mathrm\{FKL\}\}\\\|\_\{1\}\\leq 2\.\(8\)For fixed disagreement and bounded nonzero advantages, FKL can be much stronger than PG at a high\-confidence error, although its logit gradient remains bounded\. This comparison holds at fixed prefixes and does not differentiate through the rollout distribution\.

### C\.2Teacher\-preferred\-token probability weighting

At a fixed prefix, letb=arg⁡maxv⁡q⁡\(v\)b=\\arg\\max\_\{v\}q\(v\)be the teacher’s Top\-1 token,w=sg⁡\[p⁡\(b\)\]w=\\operatorname\{sg\}\[p\(b\)\], andd=DKL\(𝐪∥𝐩\)d=D\_\{\\mathrm\{KL\}\}\(\\mathbf\{q\}\\\|\\mathbf\{p\}\)\. The weighted token loss isℓSA​\-​FKL=w​d\\ell\_\{\\mathrm\{SA\\text\{\-\}FKL\}\}=wd\. With the probability weight detached,

∇𝐡ℓSA​\-​FKL=p⁡\(b\)​\(𝐩−𝐪\)\.\\nabla\_\{\\mathbf\{h\}\}\\ell\_\{\\mathrm\{SA\\text\{\-\}FKL\}\}=p\(b\)\(\\mathbf\{p\}\-\\mathbf\{q\}\)\.\(9\)This scales the FKL gradient without changing its direction\. If the teacher prefers a candidate to which the student assigns little probability, the local update is attenuated even when the student is highly confident in its sampled tokenaa\. In the high\-confidence setting of Appendix[C\.1](https://arxiv.org/html/2609.38282#A3.SS1),b≠ab\\neq aimpliesp⁡\(b\)≤ξp\(b\)\\leq\\xi, hence

‖∇𝐡ℓSA​\-​FKL‖1≤2​p​\(b\)≤2​ξ\.\\\|\\nabla\_\{\\mathbf\{h\}\}\\ell\_\{\\mathrm\{SA\\text\{\-\}FKL\}\}\\\|\_\{1\}\\leq 2p\(b\)\\leq 2\\xi\.\(10\)This trades correction strength for update moderation when the student assigns low probability to the teacher’s preferred candidate\.

For the Top\-KKapproximation, let𝐪𝒦\\mathbf\{q\}^\{\\mathcal\{K\}\}retain the original teacher probabilities on𝒦\\mathcal\{K\}and be zero elsewhere, and sets=∑v∈𝒦q⁡\(v\)s=\\sum\_\{v\\in\\mathcal\{K\}\}q\(v\)\. Writingd𝒦d^\{\\mathcal\{K\}\}for the truncated KL term gives

∇𝐡\(w​d𝒦\)\\displaystyle\\nabla\_\{\\mathbf\{h\}\}\\\!\\left\(wd^\{\\mathcal\{K\}\}\\right\)=p⁡\(b\)​\(s​𝐩−𝐪𝒦\),\\displaystyle=p\(b\)\(s\\mathbf\{p\}\-\\mathbf\{q\}^\{\\mathcal\{K\}\}\),\(11\)‖∇𝐡\(w​d𝒦\)‖1\\displaystyle\\left\\\|\\nabla\_\{\\mathbf\{h\}\}\\\!\\left\(wd^\{\\mathcal\{K\}\}\\right\)\\right\\\|\_\{1\}≤2​s​p​\(b\)\.\\displaystyle\\leq 2sp\(b\)\.The scalar gategg, attenuationfκ​\(R¯\)f\_\{\\kappa\}\(\\bar\{R\}\), and coefficientλ\\lambdafurther scale these local contributions\. Actual parameter updates also depend on model Jacobians, batch aggregation, and the optimizer\.

## Appendix DTraining Configuration and Objectives

#### Objective averaging\.

Equation[2](https://arxiv.org/html/2609.38282#S3.E2)first averages valid tokens within each response, then averages responses within a group\. The OPD term in Equation[7](https://arxiv.org/html/2609.38282#S3.E7)instead divides the sum of weighted token KL terms by∑iTi\\sum\_\{i\}T\_\{i\}for each group, before taking the rollout expectation\. Groups withg=0g=0contribute zero to OPD and remain in the GRPO objective\. Rewards, advantages, and group controls are held fixed during policy updates\. The equivalent minimized training loss is−𝒥GRPO\+λ​𝒥OPD\-\\mathcal\{J\}\_\{\\mathrm\{GRPO\}\}\+\\lambda\\mathcal\{J\}\_\{\\mathrm\{OPD\}\}\.

Table 10:Shared GAD\-RL training configuration for Qwen3\.5\-2B and Qwen3\-VL\-2B\. Each student uses a teacher trained from the same backbone\.
#### Teacher input\.

The teacher receives reference text in a standalone heading\-rewrite prompt, with no image\. It changes existing heading prefixes, preserves all other characters, and outputs no outer Markdown fence\. The same teacher prompt is used for offline signal analysis\.

#### Training reward components\.

For text\-perturbed documents, we useη=0\.5\\eta=0\.5to give equal weight to edit similarity and perturbed\-word recall\. For outputyyand GTzz, the reward components in Section[3\.2](https://arxiv.org/html/2609.38282#S3.SS2)are

Redit​\(y,z\)\\displaystyle R\_\{\\mathrm\{edit\}\}\(y,z\)=1−Dlev​\(y,z\)max⁡\(\|y\|,\|z\|\),\\displaystyle=1\-\\frac\{D\_\{\\mathrm\{lev\}\}\(y,z\)\}\{\\max\(\|y\|,\|z\|\)\},\(12\)Rrecall​\(y,z\)\\displaystyle R\_\{\\mathrm\{recall\}\}\(y,z\)=∑wmin⁡\(cM⁡\(z\)​\(w\),cy​\(w\)\)\|M⁡\(z\)\|\.\\displaystyle=\\frac\{\\sum\_\{w\}\\min\\\!\\left\(c\_\{M\(z\)\}\(w\),c\_\{y\}\(w\)\\right\)\}\{\|M\(z\)\|\}\.Here,DlevD\_\{\\mathrm\{lev\}\}is the Levenshtein distance, and\|y\|\|y\|and\|z\|\|z\|are text lengths\.M⁡\(z\)M\(z\)is the multiset of annotated perturbed words as printed;cM⁡\(z\)​\(w\)c\_\{M\(z\)\}\(w\)andcy​\(w\)c\_\{y\}\(w\)count exact occurrences of wordwwin that multiset and the output\. The sum ranges over distinct annotated perturbed words, and\|M⁡\(z\)\|\|M\(z\)\|counts all annotated occurrences\. These training rewards are distinct from the benchmark metrics in Appendix[E](https://arxiv.org/html/2609.38282#A5)\.

#### Truncated FKL\.

For the vocabulary\-truncated approximation, let𝒦i,t\\mathcal\{K\}\_\{i,t\}contain the teacher’s Top\-KKcandidates, withK=32K=32\. We replacedi,td\_\{i,t\}in Equation[7](https://arxiv.org/html/2609.38282#S3.E7)with

di,t𝒦​\(θ\)\\displaystyle d^\{\\mathcal\{K\}\}\_\{i,t\}\(\\theta\)=∑v∈𝒦i,tπθT​\(v∣xT,y<t\(i\)\)\\displaystyle=\\sum\_\{v\\in\\mathcal\{K\}\_\{i,t\}\}\\pi\_\{\\theta\_\{\\mathrm\{T\}\}\}\(v\\mid x\_\{\\mathrm\{T\}\},y\_\{<t\}^\{\(i\)\}\)\(13\)×log⁡πθT​\(v∣xT,y<t\(i\)\)πθ​\(v∣x,y<t\(i\)\)\.\\displaystyle\\times\\log\\frac\{\\pi\_\{\\theta\_\{\\mathrm\{T\}\}\}\(v\\mid x\_\{\\mathrm\{T\}\},y\_\{<t\}^\{\(i\)\}\)\}\{\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}^\{\(i\)\}\)\}\.The teacher probabilities retain their original full\-vocabulary mass\. Without renormalization, this truncated surrogate can be negative\. The corresponding token loss is

ℓi,t𝒦​\(θ\)=wi,t​\(θ\)​di,t𝒦​\(θ\),\\ell^\{\\mathcal\{K\}\}\_\{i,t\}\(\\theta\)=w\_\{i,t\}\(\\theta\)d^\{\\mathcal\{K\}\}\_\{i,t\}\(\\theta\),\(14\)wherewi,t​\(θ\)=sg⁡\[πθ​\(bi,t∣x,y<t\(i\)\)\]w\_\{i,t\}\(\\theta\)=\\operatorname\{sg\}\[\\pi\_\{\\theta\}\(b\_\{i,t\}\\mid x,y\_\{<t\}^\{\(i\)\}\)\]andbi,tb\_\{i,t\}is the teacher’s Top\-1 token\. The weight uses the student’s full\-vocabulary probability of this teacher\-preferred token and is detached; gradients pass only through the KL term\. No additional rollout\-ratio weight is applied to OPD\. Teacher\-choice comparisons use the same objective and group controls; component comparisons remove the indicated control in Table[4](https://arxiv.org/html/2609.38282#S4.T4)\.

## Appendix EEvaluation Metrics

We use benchmark normalization and matching rules for the metrics in Section[4\.1](https://arxiv.org/html/2609.38282#S4.SS1)\.

#### GlitchText\.

ForKKannotated anomalies,KvisK\_\{\\mathrm\{vis\}\}counts predictions matching the visible perturbed form andKorigK\_\{\\mathrm\{orig\}\}counts restorations to the original form under official span alignment:

Ident=Kvis/K,Cor=Korig/K\.\\mathrm\{Ident\}=K\_\{\\mathrm\{vis\}\}/K,\\qquad\\mathrm\{Cor\}=K\_\{\\mathrm\{orig\}\}/K\.\(15\)Omissions and other errors count toward neither numerator\. Table[1](https://arxiv.org/html/2609.38282#S4.T1)reports equal\-weight means of the separate GlitchText\-ZH and GlitchText\-EN\-Padding200 scores, expressed as percentages\.

#### CHAOS\-Bench\.

For perturbed words𝒫i\\mathcal\{P\}\_\{i\}and predictiony^i\\hat\{y\}\_\{i\}on pageii, withhhindicating a case\-insensitive whole\-word match,

MicroRecallCHAOS=∑i∑w∈𝒫ih⁡\(w,y^i\)∑i\|𝒫i\|\.\\mathrm\{MicroRecall\}\_\{\\mathrm\{CHAOS\}\}=\\frac\{\\sum\_\{i\}\\sum\_\{w\\in\\mathcal\{P\}\_\{i\}\}h\(w,\\hat\{y\}\_\{i\}\)\}\{\\sum\_\{i\}\|\\mathcal\{P\}\_\{i\}\|\}\.\(16\)Each annotated target has equal weight across pages; unperturbed words are excluded\.

#### General parsing\.

We report the official OmniDocBench v1\.6 Overall score as a percentage\.

## Appendix FCase Study

Figure[7](https://arxiv.org/html/2609.38282#A6.F7)illustrates both useful and conflicting guidance\. In \(a\), the teacher suppresses an incorrect restoration to ordinary spelling; in \(b\), it reinforces a faithful perturbed\-word token\. In \(c\) and \(d\), it suppresses correctly transcribed perturbed\-word\-associated tokens\. Correctness follows the printed GT, including intentional spelling perturbations\. Thus privileged conditioning can produce both directionally beneficial and potentially conflicting signals, depending on token correctness and the direction of probability change\.

![Refer to caption](https://arxiv.org/html/2609.38282v1/figure_07.png)Figure 7:Beneficial and harmful teacher signals\. Panels show the page, enlarged perturbed\-word region, GT, student response, and token probabilities\. Yellow rows mark focal tokens;·denotes a space\. Student and teacher score the same emitted token at the same prefix, conditioned on image and privileged text, respectively\.

相似文章

基于噪声学生在线策略自蒸馏的自增强视觉语言模型

arXiv cs.LG

提出NOPD,一种自蒸馏方法,通过利用干净输入和损坏输入之间的预测差异,在没有外部监督的情况下改进视觉语言模型。在视觉推理任务上取得了显著提升,匹配或超过了强化学习和来自外部模型的蒸馏。

On-Policy 蒸馏中的 Off-Policy 教师问题

Hugging Face Daily Papers

该论文指出了 On-Policy 蒸馏中存在的一种 off-policy 不对称性:教师必须监督来自学生自身生成的、其未训练过的前缀。为此,论文提出了 SCOUT——一个通过可验证奖励的 RL 持续适应教师的协同训练框架,在不同配置、模型规模和推理领域上均一致地提升了蒸馏效果。