PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation

arXiv cs.CL Papers

Summary

This paper proposes PAMT, a process-aligned reinforcement learning framework for multi-domain machine translation that combines domain-aware long chain-of-thought supervision with step-level process rewards to improve domain-sensitive translation decisions.

arXiv:2608.03077v1 Announce Type: new Abstract: Multi-domain machine translation (MDMT) requires more than fluent generation: it demands domain-sensitive translation decisions such as domain disambiguation, terminology control, and stylistic adaptation. Large reasoning models (LRMs) make such decisions explicit through intermediate translation steps, but our analysis across 15 domains and four translation directions shows that this explicit reasoning is double-edged: it improves long-form and high-difficulty translation, yet often drifts in terminology-intensive and stylistically constrained settings. We trace this failure to a credit-assignment bottleneck: existing methods optimize final outputs or coarse trajectories, but cannot identify which translation steps actually help the final translation. To address this, we propose PAMT, a process-aligned training framework that combines cold-start domain-aware Long-CoT supervision with reinforcement learning. PAMT uses sequence-level format and outcome rewards for the final translation, together with a step-level process reward that measures how much each explicit translation step increases the likelihood of the reference translation. Across two backbones, PAMT improves over base models, outperforms MT-specialized baselines on average, and remains competitive with strong LLMs/LRMs across in-domain, OOD, and multilingual settings.
Original Article
View Cached Full Text

Cached at: 08/05/26, 07:43 AM

# Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation
Source: [https://arxiv.org/html/2608.03077](https://arxiv.org/html/2608.03077)
Yongshi Ye1,3, Biao Fu2,3,, Chongxuan Huang2,3, Yidong Chen2,3, Xiaodong Shi1,2,3,11footnotemark:1 1Institute of Artificial Intelligence, Xiamen University 2School of Informatics, Xiamen University 3Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan \(Xiamen University\), Ministry of Culture and Tourism \{yeyongshi,biaofu\}@stu\.xmu\.edu\.cn,mandel@xmu\.edu\.cn

###### Abstract

Multi\-domain machine translation \(MDMT\) requires more than fluent generation: it demands domain\-sensitive translation decisions such as domain disambiguation, terminology control, and stylistic adaptation\. Large reasoning models \(LRMs\) make such decisions explicit through intermediate translation steps, but our analysis across 15 domains and four translation directions shows that this explicit reasoning is double\-edged: it improves long\-form and high\-difficulty translation, yet often drifts in terminology\-intensive and stylistically constrained settings\. We trace this failure to a credit\-assignment bottleneck: existing methods optimize final outputs or coarse trajectories, but cannot identify which translation steps actually help the final translation\. To address this, we propose PAMT, a process\-aligned training framework that combines cold\-start domain\-aware Long\-CoT supervision with reinforcement learning\. PAMT uses sequence\-level format and outcome rewards for the final translation, together with a step\-level process reward that measures how much each explicit translation step increases the likelihood of the reference translation\. Across two backbones, PAMT improves over base models, outperforms MT\-specialized baselines on average, and remains competitive with strong LLMs/LRMs across in\-domain, OOD, and multilingual settings\.

PAMT: Process\-Aligned Reinforcement Learning for Multi\-Domain Machine Translation

Yongshi Ye1,3, Biao Fu2,3,††thanks:Corresponding authors\., Chongxuan Huang2,3, Yidong Chen2,3, Xiaodong Shi1,2,3,11footnotemark:11Institute of Artificial Intelligence, Xiamen University2School of Informatics, Xiamen University3Key Laboratory of Digital Protection and Intelligent Processing of Intangible CulturalHeritage of Fujian and Taiwan \(Xiamen University\), Ministry of Culture and Tourism\{yeyongshi,biaofu\}@stu\.xmu\.edu\.cn,mandel@xmu\.edu\.cn

## 1Introduction

![Refer to caption](https://arxiv.org/html/2608.03077v1/x1.png)Figure 1:Overview of vanilla reasoning\-augmented MT and PAMT\. PAMT aligns both the translation process and the final output, addressing the misaligned credit assignment of vanilla methods\.Multi\-domain machine translation \(MDMT\) requires more than semantic adequacy\. To translate faithfully across domains, a model must make domain\-sensitive decisions about ambiguity resolution, terminology, and styleSaunders \([2022](https://arxiv.org/html/2608.03077#bib.bib46)\); Jiang et al\. \([2020](https://arxiv.org/html/2608.03077#bib.bib21)\); Lai et al\. \([2022](https://arxiv.org/html/2608.03077#bib.bib27)\); Zheng et al\. \([2024a](https://arxiv.org/html/2608.03077#bib.bib66)\); Hu et al\. \([2024](https://arxiv.org/html/2608.03077#bib.bib20)\); Man et al\. \([2025](https://arxiv.org/html/2608.03077#bib.bib34)\)\. However, most translation methods based on large language models \(LLMs\) still treat translation as a direct sequence generation taskXu et al\. \([2024a](https://arxiv.org/html/2608.03077#bib.bib62)\); Guo et al\. \([2024](https://arxiv.org/html/2608.03077#bib.bib16)\); Pang et al\. \([2025](https://arxiv.org/html/2608.03077#bib.bib40)\), without explicitly modeling the intermediate reasoning process\. This makes reasoning\-oriented translation particularly appealing: explicit intermediate translation steps can, in principle, expose the decision process needed for domain\-faithful translationHe et al\. \([2024](https://arxiv.org/html/2608.03077#bib.bib19)\)\. Yet it remains unclear when these decisions help MDMT and when they fail\.

To investigate this question, we first conduct a systematic comparison of LLMs and LRMs across 15 domains and four translation directions \(Section[2](https://arxiv.org/html/2608.03077#S2)\)\. Our analysis reveals a clear split\. Explicit reasoning is most helpful on long\-context and high\-difficulty inputs \(Section[2\.1](https://arxiv.org/html/2608.03077#S2.SS1)\), where translation benefits from stepwise decomposition and refinement\. Yet it is less reliable in terminology\-intensive and stylistically constrained settings \(Section[2\.2](https://arxiv.org/html/2608.03077#S2.SS2)\), where seemingly plausible reasoning can drift away from domain\-specific conventions\. These results highlight the need to supervise and control the translation decisions within the translation process\.

Recent work has gradually moved from merely exposing the translation process to optimizing it\. Workflow\-based methods first make the process explicit, but only through a few coarse stagesChen et al\. \([2024a](https://arxiv.org/html/2608.03077#bib.bib5)\); Feng et al\. \([2025c](https://arxiv.org/html/2608.03077#bib.bib12)\); Wang et al\. \([2024c](https://arxiv.org/html/2608.03077#bib.bib61)\)\. Chain\-of\-Thought \(CoT\)\-based methods then provide more detailed process tracesHu et al\. \([2024](https://arxiv.org/html/2608.03077#bib.bib20)\); He et al\. \([2024](https://arxiv.org/html/2608.03077#bib.bib19)\); Wang et al\. \([2025a](https://arxiv.org/html/2608.03077#bib.bib53)\), yet still learn them mainly through offline imitation\. Reinforcement learning \(RL\)\-based reasoning\-augmented methods further optimize translation with reward signalsHe et al\. \([2025](https://arxiv.org/html/2608.03077#bib.bib18)\); Feng et al\. \([2025a](https://arxiv.org/html/2608.03077#bib.bib10)\); Wang et al\. \([2025c](https://arxiv.org/html/2608.03077#bib.bib55),[d](https://arxiv.org/html/2608.03077#bib.bib56)\); Yang et al\. \([2025](https://arxiv.org/html/2608.03077#bib.bib64)\); Li et al\. \([2025](https://arxiv.org/html/2608.03077#bib.bib31)\), but these signals still operate largely on final outputs or whole trajectories\. Thus, supervision remains misaligned with translation decisions \(Figure[1](https://arxiv.org/html/2608.03077#S1.F1)\)\. This raises a central question: How can intermediate translation decisions become creditable and optimizable for domain\-faithful translation?

We address this problem withPAMT, aProcess\-Aligned training framework for MDMT\. PAMT first performs cold\-start supervised fine\-tuning \(SFT\) on distilled domain\-aware Long\-CoT data to initialize an explicit translation process\. It then applies RL to align the translation process with the final translation outcome\. Specifically, final translation quality is optimized with sequence\-level format and outcome rewards, while the translation process is optimized with a step\-level reward defined by how much each explicit translation step increases the likelihood of the reference translation\. This process\-output alignment breaks the credit\-assignment bottleneck in reasoning\-augmented MT, allowing the model to optimize not only final translation quality which translation decisions support domain\-faithful translation\.

Table 1:Document\- and sentence\-level performance\.Our contributions are threefold\.\(1\)We systematically evaluate LLMs and LRMs for MDMT across 15 domains and four translation directions, showing that explicit translation reasoning helps in long\-form and high\-difficulty settings but often fails in terminology\-intensive and stylistically constrained ones\.\(2\)We identify a key bottleneck in reasoning\-augmented MT: supervision granularity is misaligned with decision granularity, making intermediate translation steps hard to credit and optimize, thereby leading to terminology and style drift\.\(3\)We propose PAMT, a process\-aligned training framework that mitigates this bottleneck by aligning translation processes with final translation outcomes, and show on two backbones that it is competitive with SOTA LLMs/LRMs while outperforming MT\-specialized baselines across in\-domain, out\-of\-domain, and multilingual settings\.

## 2Preliminary Evaluation

Setup\.We compare LLMs and LRMs across 15 domains and four translation directions \(En↔\\leftrightarrowZh, En↔\\leftrightarrowDe\), analyzing explicit translation processes along four dimensions: input length, difficulty, Multidimensional Quality Metrics \(MQM\)\-style quality, and terminology accuracy\. To compare each paradigm at its strongest, we use a best\-in\-group protocol instead of one\-to\-one model pairs; setup details are given in Appendix[B](https://arxiv.org/html/2608.03077#A2)\.

### 2\.1When Explicit Processes Help

Document\-level Evaluation\.On WMT22, we examine whether explicit reasoning becomes more useful with longer context \(Table[1](https://arxiv.org/html/2608.03077#S1.T1)\)\. At the sentence level, LRMs do not consistently outperform LLMs: the best LLM achieves the highest BLEU and COMET scores, while the best LRM leads only on COMETKiwi\. Still, sentence\-level evaluation can miss important errors in cross\-sentence consistency and discourse coherence\(Läubli et al\.,[2018](https://arxiv.org/html/2608.03077#bib.bib29); Voita et al\.,[2019](https://arxiv.org/html/2608.03077#bib.bib52)\)\. We therefore use BlonDe\(Jiang et al\.,[2022](https://arxiv.org/html/2608.03077#bib.bib22)\)to evaluate discourse phenomena such as entities, tense, pronouns, and discourse markers\. At the document level, the best LRM obtains the highest BlonDe score \(36\.17 vs\. 35\.79 for the best LLM\), suggesting that explicit reasoning may be more useful for discourse phenomena\.

Table 2:Performance by translation complexity for LRMs and traditional LLMs\.Impact of Translation Difficulty\.To examine how translation difficulty affects model behavior, we use DeepSeek\-V3 to assign source sentences from Multi\-Domain \(De→\\rightarrowEn\), WMT22 \(De↔\\leftrightarrowEn and Zh↔\\leftrightarrowEn\), and Guofeng WebNovel \(Zh→\\rightarrowEn\) into five difficulty levels \(Figure[6](https://arxiv.org/html/2608.03077#A6.F6); Table[2](https://arxiv.org/html/2608.03077#S2.T2)\)\. A clear pattern emerges: while traditional LLMs perform competitively on easy inputs, their performance drops sharply as difficulty increases, whereas LRMs surpass them from Level 2 onward and maintain consistent advantages on COMET and COMETKiwi\. This gap indicates that performance at higher difficulty is limited not by local generation, but by the translation process itself\. As inputs become more complex, success increasingly depends on process\-level decisions such as ambiguity resolution, compositional decomposition, and consistency\-preserving revision—capabilities that are better supported by explicit reasoning\.

### 2\.2When Explicit Processes Hurt

Table 3:MQM error analysis and terminology accuracy on De→\\rightarrowEn\. The top panel reports severity\-weighted GEMBA\-MQM scores, and the bottom panel reports MQM category ratios and terminology accuracy\.MQM Analysis\.We use GEMBA\-MQM\(Kocmi and Federmann,[2023](https://arxiv.org/html/2608.03077#bib.bib25)\)with DeepSeek\-V3 as the automatic MQM annotator to analyze both error severity and error types\. Table[3](https://arxiv.org/html/2608.03077#S2.T3)shows that the best LRM has the lowest severity\-weighted MQM score and error rate, indicating fewer severe errors overall\. We then follow MQM\-based expert evaluation practice\(Freitag et al\.,[2021](https://arxiv.org/html/2608.03077#bib.bib13)\)and inspect fine\-grained categories\. The strongest LRM still lags behind the strongest LLM on style and terminology errors \(29\.58/12\.22 vs\. 29\.43/10\.68\), suggesting that explicit reasoning can reduce overall error severity while remaining fragile on domain\-sensitive decisions such as register, stylistic convention, and terminology choice\.

Terminological Accuracy\.We further isolate term\-level behavior using the WMT23 Terminology Shared TaskSemenov et al\. \([2023](https://arxiv.org/html/2608.03077#bib.bib47)\), which provides bilingual sentence pairs with aligned source–target terms\. For each example, we provide only the source sentence and measure whether the expected target term is correctly produced in the translation \(Table[3](https://arxiv.org/html/2608.03077#S2.T3)\)\. The results reinforce the MQM findings: terminology accuracy is substantially less stable than the gains observed on general semantic metrics, and LRMs underperform in settings where translation depends more on faithful lexical realization\. This shows that stronger reasoning does not automatically yield stronger term control\. Instead, terminology translation is a process\-level decision that must remain aligned with domain constraints throughout generation\. Together with the MQM analysis, these results point to the same conclusion: the main bottleneck of LRMs in MDMT lies in how intermediate translation steps are supervised\.

## 3Related Work

Prior work makes the translation process explicit at different granularities to improve translation quality\. Workflow\-based methodsHe et al\. \([2024](https://arxiv.org/html/2608.03077#bib.bib19)\); Briakou et al\. \([2024](https://arxiv.org/html/2608.03077#bib.bib4)\); Chen et al\. \([2024b](https://arxiv.org/html/2608.03077#bib.bib6),[a](https://arxiv.org/html/2608.03077#bib.bib5)\); Wang et al\. \([2024c](https://arxiv.org/html/2608.03077#bib.bib61)\); Ki and Carpuat \([2024](https://arxiv.org/html/2608.03077#bib.bib23)\); Feng et al\. \([2025c](https://arxiv.org/html/2608.03077#bib.bib12)\)decompose translation into fixed stages such as drafting, evaluation, and refinement\. Later work introduces domain\-aware CoTHu et al\. \([2024](https://arxiv.org/html/2608.03077#bib.bib20)\)or multi\-agent trajectoriesWang et al\. \([2025b](https://arxiv.org/html/2608.03077#bib.bib54)\)for SFT, but still learns the process mainly by imitating traces constructed offline\. This further motivates RL\-based reasoning\-augmented MT, where the translation process and final output are optimized through outcome\-level feedback from translation metricsFeng et al\. \([2025a](https://arxiv.org/html/2608.03077#bib.bib10)\); He et al\. \([2025](https://arxiv.org/html/2608.03077#bib.bib18)\)or exemplar\-enhanced rewardsWang et al\. \([2025d](https://arxiv.org/html/2608.03077#bib.bib56)\)\. DeepTransWang et al\. \([2026](https://arxiv.org/html/2608.03077#bib.bib57)\)and TAT\-R1Li et al\. \([2025](https://arxiv.org/html/2608.03077#bib.bib31)\)further introduce process\- and output\-level rewards through external LLM scoring and terminology constraints\. Yet these rewards are still assigned over the whole process, making it difficult to identify which intermediate translation step causes an unfaithful translation output\.

## 4Method

To bridge the mismatch between explicit translation steps and optimization, we propose PAMT, a two\-stage process\-aligned framework: cold\-start SFT initializes the explicit process, and RL aligns the process with the final output\.

### 4\.1Cold Start

The goal of this stage is to activate the base model’s reasoning ability and ensure that it can produce an explicit translation process\. We therefore begin with a cold\-start SFT stage using a Long\-CoT translation dataset distilled from DeepSeek\-R1\. The dataset contains approximately 7k translation reasoning examples across 10 diverse domains, with around 700 examples per domain\. Each example follows the same structured format,<think\>\.\.\.</think\><answer\>\.\.\.</answer\>, where the model first generates the translation process and then outputs the final translation\.

### 4\.2RL Stage

Task Setup\.During RL, each training example provides a source sentencexix\_\{i\}for generation and a reference translationyi∗y\_\{i\}^\{\*\}only for reward computation\. For eachxix\_\{i\}, we sampleGGpolicy rollouts, with thegg\-th rollout structured as

oi,g=<think\>​zi,g​</think\><answer\>​yi,g​</answer\>,\\displaystyle o\_\{i,g\}=\\texttt\{<think\>\}z\_\{i,g\}\\texttt\{</think\><answer\>\}y\_\{i,g\}\\texttt\{</answer\>\},

\(1\)wherezi,gz\_\{i,g\}denotes the explicit translation process andyi,gy\_\{i,g\}is the final translation\.

For step\-level alignment, we split the reasoning span by\\n\\nintoKi,gK\_\{i,g\}reasoning steps:zi,g=\[zi,g\(1\),zi,g\(2\),…,zi,g\(Ki,g\)\]z\_\{i,g\}=\[z\_\{i,g\}^\{\(1\)\},z\_\{i,g\}^\{\(2\)\},\\dots,z\_\{i,g\}^\{\(K\_\{i,g\}\)\}\]\. LetTi,gT\_\{i,g\}be the response length andti,glastt\_\{i,g\}^\{\\mathrm\{last\}\}the last valid response token position\. Table[9](https://arxiv.org/html/2608.03077#A0.T9)summarizes the notation\.

Reward design\.We use sequence\-level rewards to supervise the final response and a step\-level reward to assess intermediate reasoning steps\.

For response format, we use a simple binary reward:ri,gfmt=1r\_\{i,g\}^\{\\mathrm\{fmt\}\}=1ifValid​\(oi,g\)=1\\mathrm\{Valid\}\(o\_\{i,g\}\)=1, andri,gfmt=−1r\_\{i,g\}^\{\\mathrm\{fmt\}\}=\-1otherwise, whereValid​\(⋅\)\\mathrm\{Valid\}\(\\cdot\)checks whether the output follows the structured format defined in Section[4\.1](https://arxiv.org/html/2608.03077#S4.SS1)\. For translation quality, we define the outcome reward as the equal\-weighted average of normalized BLEU, COMET, and COMETKiwi\. These metrics respectively capture lexical matching, reference\-based semantic quality, and reference\-free quality estimation, with COMETKiwi included to reduce over\-reliance on potentially noisy or domain\-specific references\. We divide sacreBLEU by 100 and use COMET and COMETKiwi in their native\[0,1\]\[0,1\]ranges:

ri,gout=1\|𝒬\|​∑q∈𝒬q~i,g,r\_\{i,g\}^\{\\mathrm\{out\}\}=\\frac\{1\}\{\|\\mathcal\{Q\}\|\}\\sum\_\{q\\in\\mathcal\{Q\}\}\\widetilde\{q\}\_\{i,g\},\(2\)where𝒬=\{BLEU,COMET,COMETKiwi\}\\mathcal\{Q\}=\\\{\\mathrm\{BLEU\},\\mathrm\{COMET\},\\mathrm\{COMETKiwi\}\\\}\.

These sequence\-level rewards evaluate the format and quality of the final response, but cannot attribute credit to individual reasoning steps\.

For process\-level alignment, we score each reasoning prefix with a frozen reference modelπref\\pi\_\{\\mathrm\{ref\}\}\. Specifically, for the prefix containing the firstkkreasoning steps in rolloutoi,go\_\{i,g\}, letzi,g,≤kz\_\{i,g,\\leq k\}denote their concatenation, and define the context as

ci,g,k=\[xi,<think\>,zi,g,≤k,</think\><answer\>\]\.\\displaystyle c\_\{i,g,k\}=\[x\_\{i\},\\texttt\{<think\>\},z\_\{i,g,\\leq k\},\\texttt\{</think\><answer\>\}\]\.

\(3\)We then define the process potential of this prefix as the teacher\-forced log\-likelihood of the reference translationyi∗y\_\{i\}^\{\*\}underπref\\pi\_\{\\mathrm\{ref\}\}:

ϕi,g,k=∑m=1\|yi∗\|log⁡πref​\(yi,m∗∣ci,g,k,yi,<m∗\)\.\\phi\_\{i,g,k\}=\\sum\_\{m=1\}^\{\|y\_\{i\}^\{\*\}\|\}\\log\\pi\_\{\\mathrm\{ref\}\}\\\!\\left\(y\_\{i,m\}^\{\*\}\\mid c\_\{i,g,k\},y\_\{i,<m\}^\{\*\}\\right\)\.\(4\)
Intuitively,ϕi,g,k\\phi\_\{i,g,k\}measures how well the reasoning prefix up to stepkksupports the reference translation\. If adding stepkkincreases this potential, the step makes the reference more predictable underπref\\pi\_\{\\mathrm\{ref\}\}; if it decreases the potential, the step makes the reference less predictable\. We capture this effect with the process gainri,g,kprocr\_\{i,g,k\}^\{\\mathrm\{proc\}\}:

ri,g,kproc=ϕi,g,k−ϕi,g,k−1,k=1,…,Ki,g,r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\\phi\_\{i,g,k\}\-\\phi\_\{i,g,k\-1\},\\quad k=1,\\dots,K\_\{i,g\},\(5\)whereϕi,g,0\\phi\_\{i,g,0\}denotes the potential of the empty reasoning prefix\.

Because optimization is token\-level, we uniformly distribute each step gain over that step’s tokens\. For any tokenttinzi,g\(k\)z\_\{i,g\}^\{\(k\)\}, we assign

ri,g,tproc=ri,g,kproc\|zi,g\(k\)\|\.r\_\{i,g,t\}^\{\\mathrm\{proc\}\}=\\frac\{r\_\{i,g,k\}^\{\\mathrm\{proc\}\}\}\{\|z\_\{i,g\}^\{\(k\)\}\|\}\.\(6\)This preserves the total gain of each step while avoiding a bias that would otherwise favor longer steps merely because they contain more tokens\.

Credit Assignment\.We map rewards to token positions based on their granularity\. Format and outcome rewards are sequence\-level signals placed on the last valid token, while the distributed process reward contributes at the corresponding positions:

ri,g,t=𝟏​\[t=ti,glast\]​\(ri,gfmt\+ri,gout\)\+λ​ri,g,tproc,r\_\{i,g,t\}=\\mathbf\{1\}\[t=t\_\{i,g\}^\{\\mathrm\{last\}\}\]\\bigl\(r\_\{i,g\}^\{\\mathrm\{fmt\}\}\+r\_\{i,g\}^\{\\mathrm\{out\}\}\\bigr\)\+\\lambda\\,r\_\{i,g,t\}^\{\\mathrm\{proc\}\},\(7\)whereλ\\lambdacontrols the weight of the process reward\.

For each valid generated tokent≤ti,glastt\\leq t\_\{i,g\}^\{\\mathrm\{last\}\}, its return\-to\-go is

Ri,g,t\\displaystyle R\_\{i,g,t\}=∑u=tTi,gri,g,u=ri,gfmt\+ri,gout\+λ​∑u=tTi,gri,g,uproc\.\\displaystyle=\\sum\_\{u=t\}^\{T\_\{i,g\}\}r\_\{i,g,u\}=r\_\{i,g\}^\{\\mathrm\{fmt\}\}\+r\_\{i,g\}^\{\\mathrm\{out\}\}\+\\lambda\\sum\_\{u=t\}^\{T\_\{i,g\}\}r\_\{i,g,u\}^\{\\mathrm\{proc\}\}\.\(8\)This decomposition shows that terminal format and outcome rewards are propagated to all generated tokens throughRi,g,tR\_\{i,g,t\}, while process rewards provide dense step\-level credit inside the reasoning trace\.

Optimization\.For theGGrollouts sampled for the same source sentencexix\_\{i\}, letRi,gtraj=Ri,g,1R\_\{i,g\}^\{\\mathrm\{traj\}\}=R\_\{i,g,1\}be the total return of rolloutgg\. We compute the group meanμi\\mu\_\{i\}and standard deviationσi\\sigma\_\{i\}over\{Ri,gtraj\}g=1G\\\{R\_\{i,g\}^\{\\mathrm\{traj\}\}\\\}\_\{g=1\}^\{G\}, and normalize each token\-level return as

Ai,g,t=Ri,g,t−μiσi\+ϵ\.A\_\{i,g,t\}=\\frac\{R\_\{i,g,t\}\-\\mu\_\{i\}\}\{\\sigma\_\{i\}\+\\epsilon\}\.\(9\)
We then optimize a GRPO\-style clipped surrogate objective with KL regularization:

ℒ​\(θ\)=\\displaystyle\\mathcal\{L\}\(\\theta\)=−1B​G∑i=1B∑g=1G1Ti,g∑t=1Ti,gmin\(ρi,g,t\(θ\)Ai,g,t,\\displaystyle\\,\-\\frac\{1\}\{BG\}\\sum\_\{i=1\}^\{B\}\\sum\_\{g=1\}^\{G\}\\frac\{1\}\{T\_\{i,g\}\}\\sum\_\{t=1\}^\{T\_\{i,g\}\}\\min\\\!\\Biggl\(\\rho\_\{i,g,t\}\(\\theta\)A\_\{i,g,t\},\(10\)clip\(ρi,g,t\(θ\),1−ϵ,1\+ϵ\)Ai,g,t\)\\displaystyle\\,\\operatorname\{clip\}\\\!\\bigl\(\\rho\_\{i,g,t\}\(\\theta\),1\-\\epsilon,1\+\\epsilon\\bigr\)A\_\{i,g,t\}\\Biggr\)\+β​KL​\(πθ∥πref\)\.\\displaystyle\\,\+\\beta\\,\\mathrm\{KL\}\\\!\\big\(\\pi\_\{\\theta\}\\,\\\|\\,\\pi\_\{\\mathrm\{ref\}\}\\big\)\.where

ρi,g,t​\(θ\)=πθ​\(oi,g,t∣xi,oi,g,<t\)πθold​\(oi,g,t∣xi,oi,g,<t\)\.\\rho\_\{i,g,t\}\(\\theta\)=\\frac\{\\pi\_\{\\theta\}\(o\_\{i,g,t\}\\mid x\_\{i\},o\_\{i,g,<t\}\)\}\{\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(o\_\{i,g,t\}\\mid x\_\{i\},o\_\{i,g,<t\}\)\}\.\(11\)
Notably, PAMT reusesπref\\pi\_\{\\mathrm\{ref\}\}for both process scoring and KL regularization, requiring no additional scoring model\. Algorithm[1](https://arxiv.org/html/2608.03077#algorithm1)summarizes the RL training procedure, and Appendix[H](https://arxiv.org/html/2608.03077#A8)derives the PAMT objective from standard GRPO\.

## 5Experiments

Table 4:In\-domain performance on 8 domains, measured by BLEU, COMET, and COMETKiwi \(KIWI\)\. Within MT\-Specialized Models and Our Models, the best score is shown in bold and the second\-best is underlined\.Table 5:Out\-of\-domain \(OOD\) performance on 5 domains, measured by BLEU, COMET, and KIWI\.Table 6:Multilingual performance on 5 language settings, measured by BLEU, COMET, and KIWI\.Table 7:Ablation on in\-domain and OOD test sets\.### 5\.1Experimental Settings

Datasets\.We use two training sets for the two stages of PAMT: a curated 7K domain\-aware Long\-CoT dataset for cold\-start SFT, and a separate 20K multi\-domain parallel dataset for RL\. Evaluation covers in\-domain, OOD, and multilingual test sets over seen and unseen language pairs\. Full details are provided in Appendix[B\.1](https://arxiv.org/html/2608.03077#A2.SS1)and Appendix[A\.1](https://arxiv.org/html/2608.03077#A1.SS1)\.

Metrics\.We report BLEU, COMETRei et al\. \([2020](https://arxiv.org/html/2608.03077#bib.bib44)\), and COMETKiwiRei et al\. \([2022b](https://arxiv.org/html/2608.03077#bib.bib45)\)for translation quality; metric and implementation details are in Appendices[B\.6](https://arxiv.org/html/2608.03077#A2.SS6)and[A\.2](https://arxiv.org/html/2608.03077#A1.SS2), respectively\.

Baselines\.We compare against three baseline groups\.LLMstreat MT mainly as direct sequence generation: DeepSeek\-V3DeepSeek\-AI et al\. \([2025](https://arxiv.org/html/2608.03077#bib.bib9)\), Gemini\-2\.0\-FlashDeepMind \([2024](https://arxiv.org/html/2608.03077#bib.bib7)\), GPT\-4oOpenAI et al\. \([2024](https://arxiv.org/html/2608.03077#bib.bib38)\), Gemma2\-9B\-ITTeam \([2024a](https://arxiv.org/html/2608.03077#bib.bib49)\), and Qwen2\.5\-7B\-InstructTeam \([2024b](https://arxiv.org/html/2608.03077#bib.bib50)\)\.LRMsexpose explicit reasoning but lack MT\-specific process alignment: DeepSeek\-R1Guo et al\. \([2025](https://arxiv.org/html/2608.03077#bib.bib15)\), Gemini\-2\.0\-Flash\-ThinkingDeepMind \([2025](https://arxiv.org/html/2608.03077#bib.bib8)\), and GPT\-5OpenAI \([2025](https://arxiv.org/html/2608.03077#bib.bib39)\)\.MT\-specialized modelsinclude non\-reasoning systems, namely TowerInstructAlves et al\. \([2024](https://arxiv.org/html/2608.03077#bib.bib2)\), ALMA\-RXu et al\. \([2024a](https://arxiv.org/html/2608.03077#bib.bib62),[b](https://arxiv.org/html/2608.03077#bib.bib63)\), SFT\-Parallel, Tower\-Plus\-9BRei et al\. \([2025](https://arxiv.org/html/2608.03077#bib.bib43)\), and MT\-RewardTreeFeng et al\. \([2025b](https://arxiv.org/html/2608.03077#bib.bib11)\), as well as reasoning\-augmented methods, namely CoT\-FTHu et al\. \([2024](https://arxiv.org/html/2608.03077#bib.bib20)\), MT\-R1\-ZeroFeng et al\. \([2025a](https://arxiv.org/html/2608.03077#bib.bib10)\), mExTransWang et al\. \([2025d](https://arxiv.org/html/2608.03077#bib.bib56)\), SSR\-X\-ZeroYang et al\. \([2025](https://arxiv.org/html/2608.03077#bib.bib64)\), and TAT\-R1Li et al\. \([2025](https://arxiv.org/html/2608.03077#bib.bib31)\)\.

### 5\.2Generalization Across Domains

A key question for process\-aligned MT is whether it learns reusable translation behavior or merely overfits to domain\-specific heuristics\. As shown in Tables[4](https://arxiv.org/html/2608.03077#S5.T4)and[5](https://arxiv.org/html/2608.03077#S5.T5), PAMT improves over its backbones and achieves the best average performance among MT\-specialized systems in both in\-domain and OOD settings\. In contrast, CoT\-FT relies on offline CoT imitation and therefore generalizes less reliably under domain shift, while TAT\-R1 improves robustness with terminology\-aware rewards but still operates mainly at the output level\. As a result, neither method can identify which intermediate translation step leads to correct domain\-sensitive choices\. By attributing changes in final translation likelihood to individual translation steps, PAMT enables step\-level credit assignment, reinforcing the decisions that transfer across domains and suppressing those that cause process drift\. Expert human evaluation is provided in Appendix[D](https://arxiv.org/html/2608.03077#A4)\.

### 5\.3Generalization Across Languages

Table[6](https://arxiv.org/html/2608.03077#S5.T6)reports the multilingual results, which test whether the model learns language\-pair\-specific patterns or a translation process that transfers across directions\. PAMT achieves the strongest average multilingual performance among MT\-specialized baselines and our models, with gains that are especially clear in the transfer\-heavy En\-X and X\-En settings\. This pattern suggests that PAMT learns more than language\-specific templates\. We attribute this advantage to step\-level process alignment: sequence\-level rewards collapse source interpretation, structural transfer, and target\-language realization into a single terminal signal, whereas PAMT credits each explicit translation step by its effect on the reference translation likelihood\. This allows the model to reinforce transferable intermediate decisions while suppressing brittle language\-specific shortcuts, leading to stronger generalization across language directions\.

### 5\.4Ablation Study

Table[7](https://arxiv.org/html/2608.03077#S5.T7)shows that all components of PAMT contribute to the final gains\. Removing credit assignment consistently degrades performance in both in\-domain and OOD settings, indicating that step\-level process gains must be propagated to the corresponding reasoning tokens to become effective optimization signals\. Removing the process reward also leads to clear drops, suggesting that sequence\-level outcome supervision alone is insufficient for improving intermediate translation reasoning\. The largest degradation comes from removing the quality reward, confirming that explicit supervision on the final translation remains indispensable\. Finally, removing RL yields the weakest overall results, showing that supervised initialization alone cannot fully translate explicit reasoning traces into translation gains\. Overall, these results validate that PAMT relies on the synergy of quality reward, process reward, and token\-level credit assignment\.

### 5\.5PAMT on Terminology and Style Drift\.

Table 8:MQM error rates and terminology accuracy; lower MQM and higher Term\. Acc\. are better\.Table[8](https://arxiv.org/html/2608.03077#S5.T8)directly evaluates whether process supervision mitigates terminology and style drift\. PAMT achieves the lowest error rates overall\. Compared with CoT\-FT, PAMT\-Gemma2\-9B\-IT reduces style and terminology errors by 3\.43 and 2\.63 points, showing that exposing the process alone is insufficient\. The contrast with TAT\-R1 is more striking: despite explicit terminology constraints, it exhibits substantially higher style and terminology errors \(19\.33/19\.19 vs\. 14\.15/14\.26\) and an extremely high non\-translation rate \(46\.15\), indicating that output\-level constraints cannot stabilize an unaligned process\. The ablation further confirms the role of process reward: removing it increases style and terminology errors \(from 14\.15/14\.26 to 16\.25/16\.64\)\. Notably, PAMT also achieves higher terminology accuracy than TAT\-R1 \(42\.97 vs\. 40\.95\)\. These results show that terminology and style drift arise from misaligned intermediate translation steps, and that PAMT mitigates this bottleneck through step\-level credit assignment\. We further provide qualitative case studies in Appendix[G](https://arxiv.org/html/2608.03077#A7)\.

![Refer to caption](https://arxiv.org/html/2608.03077v1/x2.png)Figure 2:Training dynamics of process rewards![Refer to caption](https://arxiv.org/html/2608.03077v1/x3.png)Figure 3:Training dynamics of reward signals and KL\.
### 5\.6Training Dynamics of Process Reward\.

Figure[2](https://arxiv.org/html/2608.03077#S5.F2)shows that PAMT induces highly consistent process\-level learning dynamics across two backbones\. As training proceeds, the fraction of positive\-gain reasoning steps steadily increases, while the fraction of negative\-gain steps decreases\. Meanwhile, the mean process gain also improves, indicating that intermediate reasoning steps become progressively less harmful and more supportive of the reference translation\. We further observe that the positive\-step fraction within each trajectory rises over time, suggesting that the improvement is distributed across the reasoning process rather than concentrated in only a few isolated steps\. Overall, these trends provide direct evidence that PAMT effectively optimizes intermediate translation reasoning, rather than only improving the final output\.

### 5\.7Process Reward Shape Translation Decisions\.

![Refer to caption](https://arxiv.org/html/2608.03077v1/x4.png)Figure 4:Decision\-type percentages in positive\- and negative\-gain reasoning steps across two PAMT backbones during training\.We further analyze which translation decisions are shaped by the changing process reward observed above\. Using keyword and pattern rules, we label each reasoning step as term selection, style calibration, or disambiguation\. Figure[4](https://arxiv.org/html/2608.03077#S5.F4)shows that, across both backbones, all three decision types become more frequent among positive\-gain steps and less frequent among negative\-gain steps during training\. This indicates that PAMT optimizes domain\-sensitive intermediate decisions, rather than only improving final\-output rewards\.

### 5\.8Frozen\-Model Overfitting Check\.

As an additional diagnostic, we check whether PAMT overfits to the frozen reference model\. Figure[3](https://arxiv.org/html/2608.03077#S5.F3)shows that outcome reward remains the dominant signal \(around 1\.7\), while process reward is much smaller in scale \(about \-0\.16\)\. The KL from the frozen reference model steadily increases during training\. These trends suggest that process reward shapes intermediate steps, while optimization remains driven by final translation quality\.

## 6Conclusion

In this work, we show that explicit translation reasoning is useful but fragile in MDMT: it helps long\-form and hard translation, yet can drift on terminology and style when intermediate decisions are not supervised\. PAMT addresses this credit\-assignment bottleneck by assigning step\-level process rewards based on each reasoning step’s contribution to the final translation\. Across two backbones, PAMT improves its base models, outperforms MT\-specialized baselines on average, and remains competitive with strong LLM/LRM systems across in\-domain, OOD, and multilingual settings\.

## Limitation

Our study has several limitations\. PAMT is designed for explicit translation processes, i\.e\., intermediate translation steps verbalized in the<think\>field, and therefore does not directly model latent internal reasoning that is not exposed in the output trace\. In addition, our process reward is reference\-based during training, which provides a stable signal for fine\-grained credit assignment but assumes access to parallel data\. The current implementation also uses simple delimiter\-based step segmentation and uniform reward distribution within each step, which is effective in practice but still a coarse approximation of step\-level contribution\. Finally, process\-aware optimization introduces additional training\-time cost because multiple step prefixes must be evaluated for each sampled trace, although this does not affect inference\. Extending PAMT to weaker\-supervision settings, more adaptive process segmentation, and more efficient training remains important future work\.

## References

- Aharoni and Goldberg \(2020\)Roee Aharoni and Yoav Goldberg\. 2020\.[Unsupervised domain clusters in pretrained language models](https://doi.org/10.18653/v1/2020.acl-main.692)\.In*Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics*, pages 7747–7763, Online\. Association for Computational Linguistics\.
- Alves et al\. \(2024\)Duarte M\. Alves, José Pombal, Nuno M\. Guerreiro, Pedro H\. Martins, João Alves, Amin Farajian, Ben Peters, Ricardo Rei, Patrick Fernandes, Sweta Agrawal, Pierre Colombo, José G\. C\. de Souza, and André F\. T\. Martins\. 2024\.[Tower: An open multilingual large language model for translation\-related tasks](https://arxiv.org/abs/2402.17733)\.*Preprint*, arXiv:2402\.17733\.
- Bawden et al\. \(2019\)Rachel Bawden, Kevin Bretonnel Cohen, Cristian Grozea, Antonio Jimeno Yepes, Madeleine Kittner, Martin Krallinger, Nancy Mah, Aurelie Neveol, Mariana Neves, Felipe Soares, Amy Siu, Karin Verspoor, and Maika Vicente Navarro\. 2019\.[Findings of the WMT 2019 biomedical translation shared task: Evaluation for MEDLINE abstracts and biomedical terminologies](https://doi.org/10.18653/v1/W19-5403)\.In*Proceedings of the Fourth Conference on Machine Translation \(Volume 3: Shared Task Papers, Day 2\)*, pages 29–53, Florence, Italy\. Association for Computational Linguistics\.
- Briakou et al\. \(2024\)Eleftheria Briakou, Jiaming Luo, Colin Cherry, and Markus Freitag\. 2024\.[Translating step\-by\-step: Decomposing the translation process for improved translation quality of long\-form texts](https://doi.org/10.18653/v1/2024.wmt-1.123)\.In*Proceedings of the Ninth Conference on Machine Translation*, pages 1301–1317, Miami, Florida, USA\. Association for Computational Linguistics\.
- Chen et al\. \(2024a\)Andong Chen, Lianzhang Lou, Kehai Chen, Xuefeng Bai, Yang Xiang, Muyun Yang, Tiejun Zhao, and Min Zhang\. 2024a\.[DUAL\-REFLECT: Enhancing large language models for reflective translation through dual learning feedback mechanisms](https://doi.org/10.18653/v1/2024.acl-short.64)\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\)*, pages 693–704, Bangkok, Thailand\. Association for Computational Linguistics\.
- Chen et al\. \(2024b\)Pinzhen Chen, Zhicheng Guo, Barry Haddow, and Kenneth Heafield\. 2024b\.[Iterative translation refinement with large language models](https://aclanthology.org/2024.eamt-1.17/)\.In*Proceedings of the 25th Annual Conference of the European Association for Machine Translation \(Volume 1\)*, pages 181–190, Sheffield, UK\. European Association for Machine Translation \(EAMT\)\.
- DeepMind \(2024\)Google DeepMind\. 2024\.Introducing gemini 2\.0: our new ai model for the agentic era\.[https://blog\.google/technology/google\-deepmind/google\-gemini\-ai\-update\-december\-2024/\#ceo\-message](https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/#ceo-message)\.Accessed: 2025\-04\-21\.
- DeepMind \(2025\)Google DeepMind\. 2025\.Gemini 2\.0 flash thinking\.[https://deepmind\.google/technologies/gemini/flash\-thinking](https://deepmind.google/technologies/gemini/flash-thinking)\.Accessed: 2025\-04\-21\.
- DeepSeek\-AI et al\. \(2025\)DeepSeek\-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, and 181 others\. 2025\.[Deepseek\-v3 technical report](https://arxiv.org/abs/2412.19437)\.*Preprint*, arXiv:2412\.19437\.
- Feng et al\. \(2025a\)Zhaopeng Feng, Shaosheng Cao, Jiahan Ren, Jiayuan Su, Ruizhe Chen, Yan Zhang, Jian Wu, and Zuozhu Liu\. 2025a\.[MT\-r1\-zero: Advancing LLM\-based machine translation via r1\-zero\-like reinforcement learning](https://doi.org/10.18653/v1/2025.findings-emnlp.1015)\.In*Findings of the Association for Computational Linguistics: EMNLP 2025*, pages 18685–18702, Suzhou, China\. Association for Computational Linguistics\.
- Feng et al\. \(2025b\)Zhaopeng Feng, Jiahan Ren, Jiayuan Su, Jiamei Zheng, Hongwei Wang, and Zuozhu Liu\. 2025b\.[Mt\-rewardtree: A comprehensive framework for advancing llm\-based machine translation via reward modeling](https://arxiv.org/abs/2503.12123)\.*Preprint*, arXiv:2503\.12123\.
- Feng et al\. \(2025c\)Zhaopeng Feng, Yan Zhang, Hao Li, Bei Wu, Jiayu Liao, Wenqiang Liu, Jun Lang, Yang Feng, Jian Wu, and Zuozhu Liu\. 2025c\.[TEaR: Improving LLM\-based machine translation with systematic self\-refinement](https://doi.org/10.18653/v1/2025.findings-naacl.218)\.In*Findings of the Association for Computational Linguistics: NAACL 2025*, pages 3922–3938, Albuquerque, New Mexico\. Association for Computational Linguistics\.
- Freitag et al\. \(2021\)Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey\. 2021\.[Experts, errors, and context: A large\-scale study of human evaluation for machine translation](https://doi.org/10.1162/tacl_a_00437)\.*Transactions of the Association for Computational Linguistics*, 9:1460–1474\.
- Freitag et al\. \(2024\)Markus Freitag, Nitika Mathur, Daniel Deutsch, Chi\-Kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Frederic Blain, Tom Kocmi, Jiayi Wang, David Ifeoluwa Adelani, Marianna Buchicchio, Chrysoula Zerva, and Alon Lavie\. 2024\.[Are LLMs breaking MT metrics? results of the WMT24 metrics shared task](https://doi.org/10.18653/v1/2024.wmt-1.2)\.In*Proceedings of the Ninth Conference on Machine Translation*, pages 47–81, Miami, Florida, USA\. Association for Computational Linguistics\.
- Guo et al\. \(2025\)Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z\. F\. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, and 180 others\. 2025\.[Deepseek\-r1: Incentivizing reasoning capability in llms via reinforcement learning](https://arxiv.org/abs/2501.12948)\.*Preprint*, arXiv:2501\.12948\.
- Guo et al\. \(2024\)Jiaxin Guo, Hao Yang, Zongyao Li, Daimeng Wei, Hengchao Shang, and Xiaoyu Chen\. 2024\.[A novel paradigm boosting translation capabilities of large language models](https://doi.org/10.18653/v1/2024.findings-naacl.42)\.In*Findings of the Association for Computational Linguistics: NAACL 2024*, pages 639–649, Mexico City, Mexico\. Association for Computational Linguistics\.
- He et al\. \(2020\)Jie He, Tao Wang, Deyi Xiong, and Qun Liu\. 2020\.[The box is in the pen: Evaluating commonsense reasoning in neural machine translation](https://doi.org/10.18653/v1/2020.findings-emnlp.327)\.In*Findings of the Association for Computational Linguistics: EMNLP 2020*, pages 3662–3672, Online\. Association for Computational Linguistics\.
- He et al\. \(2025\)Minggui He, Yilun Liu, Shimin Tao, Yuanchang Luo, Hongyong Zeng, Chang Su, Li Zhang, Hongxia Ma, Daimeng Wei, Weibin Meng, Hao Yang, Boxing Chen, and Osamu Yoshie\. 2025\.[R1\-t1: Fully incentivizing translation capability in llms via reasoning learning](https://arxiv.org/abs/2502.19735)\.*Preprint*, arXiv:2502\.19735\.
- He et al\. \(2024\)Zhiwei He, Tian Liang, Wenxiang Jiao, Zhuosheng Zhang, Yujiu Yang, Rui Wang, Zhaopeng Tu, Shuming Shi, and Xing Wang\. 2024\.[Exploring human\-like translation strategy with large language models](https://doi.org/10.1162/tacl_a_00642)\.*Transactions of the Association for Computational Linguistics*, 12:229–246\.
- Hu et al\. \(2024\)Tianxiang Hu, Pei Zhang, Baosong Yang, Jun Xie, Derek F\. Wong, and Rui Wang\. 2024\.[Large language model for multi\-domain translation: Benchmarking and domain CoT fine\-tuning](https://doi.org/10.18653/v1/2024.findings-emnlp.328)\.In*Findings of the Association for Computational Linguistics: EMNLP 2024*, pages 5726–5746, Miami, Florida, USA\. Association for Computational Linguistics\.
- Jiang et al\. \(2020\)Haoming Jiang, Chen Liang, Chong Wang, and Tuo Zhao\. 2020\.[Multi\-domain neural machine translation with word\-level adaptive layer\-wise domain mixing](https://doi.org/10.18653/v1/2020.acl-main.165)\.In*Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics*, pages 1823–1834, Online\. Association for Computational Linguistics\.
- Jiang et al\. \(2022\)Yuchen Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang, Jian Yang, Haoyang Huang, Rico Sennrich, Ryan Cotterell, Mrinmaya Sachan, and Ming Zhou\. 2022\.[BlonDe: An automatic evaluation metric for document\-level machine translation](https://doi.org/10.18653/v1/2022.naacl-main.111)\.In*Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies*, pages 1550–1565, Seattle, United States\. Association for Computational Linguistics\.
- Ki and Carpuat \(2024\)Dayeon Ki and Marine Carpuat\. 2024\.[Guiding large language models to post\-edit machine translation with error annotations](https://doi.org/10.18653/v1/2024.findings-naacl.265)\.In*Findings of the Association for Computational Linguistics: NAACL 2024*, pages 4253–4273, Mexico City, Mexico\. Association for Computational Linguistics\.
- Kocmi et al\. \(2022\)Tom Kocmi, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Novák, Martin Popel, and Maja Popović\. 2022\.[Findings of the 2022 conference on machine translation \(WMT22\)](https://aclanthology.org/2022.wmt-1.1/)\.In*Proceedings of the Seventh Conference on Machine Translation \(WMT\)*, pages 1–45, Abu Dhabi, United Arab Emirates \(Hybrid\)\. Association for Computational Linguistics\.
- Kocmi and Federmann \(2023\)Tom Kocmi and Christian Federmann\. 2023\.[GEMBA\-MQM: Detecting translation quality error spans with GPT\-4](https://doi.org/10.18653/v1/2023.wmt-1.64)\.In*Proceedings of the Eighth Conference on Machine Translation*, pages 768–775, Singapore\. Association for Computational Linguistics\.
- Kwon et al\. \(2023\)Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica\. 2023\.[Efficient memory management for large language model serving with pagedattention](https://doi.org/10.1145/3600006.3613165)\.In*Proceedings of the 29th Symposium on Operating Systems Principles*, SOSP ’23, page 611–626, New York, NY, USA\. Association for Computing Machinery\.
- Lai et al\. \(2022\)Wen Lai, Alexandra Chronopoulou, and Alexander Fraser\. 2022\.[m4adapter: Multilingual multi\-domain adaptation for machine translation with a meta\-adapter](https://doi.org/10.18653/v1/2022.findings-emnlp.315)\.In*Findings of the Association for Computational Linguistics: EMNLP 2022*, pages 4282–4296, Abu Dhabi, United Arab Emirates\. Association for Computational Linguistics\.
- Lai et al\. \(2024\)Xin Lai, Zhuotao Tian, Yukang Chen, Senqiao Yang, Xiangru Peng, and Jiaya Jia\. 2024\.[Step\-dpo: Step\-wise preference optimization for long\-chain reasoning of llms](https://arxiv.org/abs/2406.18629)\.*Preprint*, arXiv:2406\.18629\.
- Läubli et al\. \(2018\)Samuel Läubli, Rico Sennrich, and Martin Volk\. 2018\.[Has machine translation achieved human parity? a case for document\-level evaluation](https://doi.org/10.18653/v1/D18-1512)\.In*Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing*, pages 4791–4796, Brussels, Belgium\. Association for Computational Linguistics\.
- Li et al\. \(2023\)Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian\-Guang Lou, and Weizhu Chen\. 2023\.[Making language models better reasoners with step\-aware verifier](https://doi.org/10.18653/v1/2023.acl-long.291)\.In*Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 5315–5333, Toronto, Canada\. Association for Computational Linguistics\.
- Li et al\. \(2025\)Zheng Li, Mao Zheng, Mingyang Song, and Wenjie Yang\. 2025\.[Tat\-r1: Terminology\-aware translation with reinforcement learning and word alignment](https://arxiv.org/abs/2505.21172)\.*Preprint*, arXiv:2505\.21172\.
- Lightman et al\. \(2023\)Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe\. 2023\.[Let’s verify step by step](https://arxiv.org/abs/2305.20050)\.*Preprint*, arXiv:2305\.20050\.
- Luo et al\. \(2024\)Liangchen Luo, Yinxiao Liu, Rosanne Liu, Samrat Phatale, Meiqi Guo, Harsh Lara, Yunxuan Li, Lei Shu, Yun Zhu, Lei Meng, Jiao Sun, and Abhinav Rastogi\. 2024\.[Improve mathematical reasoning in language models by automated process supervision](https://arxiv.org/abs/2406.06592)\.*Preprint*, arXiv:2406\.06592\.
- Man et al\. \(2025\)Zhibo Man, Yuanmeng Chen, Yujie Zhang, Yufeng Chen, and Jinan Xu\. 2025\.Dmdteval: An evaluation and analysis of llms on disambiguation in multi\-domain translation\.*arXiv preprint arXiv:2504\.20371*\.
- Neves et al\. \(2018\)Mariana Neves, Antonio Jimeno Yepes, Aurélie Névéol, Cristian Grozea, Amy Siu, Madeleine Kittner, and Karin Verspoor\. 2018\.[Findings of the WMT 2018 biomedical translation shared task: Evaluation on Medline test sets](https://doi.org/10.18653/v1/W18-6403)\.In*Proceedings of the Third Conference on Machine Translation: Shared Task Papers*, pages 324–339, Belgium, Brussels\. Association for Computational Linguistics\.
- Ni et al\. \(2023\)Ansong Ni, Srini Iyer, Dragomir Radev, Veselin Stoyanov, Wen\-Tau Yih, Sida Wang, and Xi Victoria Lin\. 2023\.[LEVER: Learning to verify language\-to\-code generation with execution](https://proceedings.mlr.press/v202/ni23b.html)\.In*Proceedings of the 40th International Conference on Machine Learning*, volume 202 of*Proceedings of Machine Learning Research*, pages 26106–26128\. PMLR\.
- NLLB Team et al\. \(2024\)NLLB Team, Marta R\. Costa\-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, and 20 others\. 2024\.[Scaling neural machine translation to 200 languages](https://doi.org/10.1038/s41586-024-07335-x)\.*Nature*, 630\(8018\):841–846\.
- OpenAI et al\. \(2024\)OpenAI, :, Aaron Hurst, Adam Lerer, Adam P\. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander Mądry, Alex Baker\-Whitcomb, Alex Beutel, Alex Borzunov, Alex Carney, Alex Chow, Alex Kirillov, and 401 others\. 2024\.[Gpt\-4o system card](https://arxiv.org/abs/2410.21276)\.*Preprint*, arXiv:2410\.21276\.
- OpenAI \(2025\)OpenAI\. 2025\.Introducing GPT\-5\.[https://openai\.com/zh\-Hans\-CN/index/introducing\-gpt\-5/](https://openai.com/zh-Hans-CN/index/introducing-gpt-5/)\.
- Pang et al\. \(2025\)Jianhui Pang, Fanghua Ye, Derek Fai Wong, Dian Yu, Shuming Shi, Zhaopeng Tu, and Longyue Wang\. 2025\.[Salute the classic: Revisiting challenges of machine translation in the age of large language models](https://doi.org/10.1162/tacl_a_00730)\.*Transactions of the Association for Computational Linguistics*, 13:73–95\.
- Post \(2018\)Matt Post\. 2018\.[A call for clarity in reporting BLEU scores](https://doi.org/10.18653/v1/W18-6319)\.In*Proceedings of the Third Conference on Machine Translation: Research Papers*, pages 186–191, Brussels, Belgium\. Association for Computational Linguistics\.
- Rei et al\. \(2022a\)Ricardo Rei, José G\. C\. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André F\. T\. Martins\. 2022a\.[COMET\-22: Unbabel\-IST 2022 submission for the metrics shared task](https://aclanthology.org/2022.wmt-1.52/)\.In*Proceedings of the Seventh Conference on Machine Translation \(WMT\)*, pages 578–585, Abu Dhabi, United Arab Emirates \(Hybrid\)\. Association for Computational Linguistics\.
- Rei et al\. \(2025\)Ricardo Rei, Nuno M\. Guerreiro, José Pombal, João Alves, Pedro Teixeirinha, Amin Farajian, and André F\. T\. Martins\. 2025\.[Tower\+: Bridging generality and translation specialization in multilingual llms](https://arxiv.org/abs/2506.17080)\.*Preprint*, arXiv:2506\.17080\.
- Rei et al\. \(2020\)Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie\. 2020\.[COMET: A neural framework for MT evaluation](https://doi.org/10.18653/v1/2020.emnlp-main.213)\.In*Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*, pages 2685–2702, Online\. Association for Computational Linguistics\.
- Rei et al\. \(2022b\)Ricardo Rei, Marcos Treviso, Nuno M\. Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, José G\. C\. de Souza, Taisiya Glushkova, Duarte Alves, Luisa Coheur, Alon Lavie, and André F\. T\. Martins\. 2022b\.[CometKiwi: IST\-unbabel 2022 submission for the quality estimation shared task](https://aclanthology.org/2022.wmt-1.60/)\.In*Proceedings of the Seventh Conference on Machine Translation \(WMT\)*, pages 634–645, Abu Dhabi, United Arab Emirates \(Hybrid\)\. Association for Computational Linguistics\.
- Saunders \(2022\)Danielle Saunders\. 2022\.Domain adaptation and multi\-domain adaptation for neural machine translation: A survey\.*Journal of Artificial Intelligence Research*, 75:351–424\.
- Semenov et al\. \(2023\)Kirill Semenov, Vilém Zouhar, Tom Kocmi, Dongdong Zhang, Wangchunshu Zhou, and Yuchen Eleanor Jiang\. 2023\.[Findings of the WMT 2023 shared task on machine translation with terminologies](https://doi.org/10.18653/v1/2023.wmt-1.54)\.In*Proceedings of the Eighth Conference on Machine Translation*, pages 663–671, Singapore\. Association for Computational Linguistics\.
- Sheng et al\. \(2025\)Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu\. 2025\.[Hybridflow: A flexible and efficient rlhf framework](https://doi.org/10.1145/3689031.3696075)\.In*Proceedings of the Twentieth European Conference on Computer Systems*, EuroSys ’25, page 1279–1297, New York, NY, USA\. Association for Computing Machinery\.
- Team \(2024a\)Gemma Team\. 2024a\.[Gemma](https://doi.org/10.34740/KAGGLE/M/3301)\.
- Team \(2024b\)Qwen Team\. 2024b\.[Qwen2\.5: A party of foundation models](https://qwenlm.github.io/blog/qwen2.5/)\.
- Tian et al\. \(2014\)Liang Tian, Derek F\. Wong, Lidia S\. Chao, Paulo Quaresma, Francisco Oliveira, Yi Lu, Shuo Li, Yiming Wang, and Longyue Wang\. 2014\.[UM\-corpus: A large English\-Chinese parallel corpus for statistical machine translation](https://aclanthology.org/L14-1604/)\.In*Proceedings of the Ninth International Conference on Language Resources and Evaluation \(LREC’14\)*, pages 1837–1842, Reykjavik, Iceland\. European Language Resources Association \(ELRA\)\.
- Voita et al\. \(2019\)Elena Voita, Rico Sennrich, and Ivan Titov\. 2019\.[When a good translation is wrong in context: Context\-aware machine translation improves on deixis, ellipsis, and lexical cohesion](https://doi.org/10.18653/v1/P19-1116)\.In*Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics*, pages 1198–1212, Florence, Italy\. Association for Computational Linguistics\.
- Wang et al\. \(2025a\)Jiaan Wang, Fandong Meng, Yunlong Liang, and Jie Zhou\. 2025a\.[Drt: Deep reasoning translation via long chain\-of\-thought](https://arxiv.org/abs/2412.17498)\.*Preprint*, arXiv:2412\.17498\.
- Wang et al\. \(2025b\)Jiaan Wang, Fandong Meng, Yunlong Liang, and Jie Zhou\. 2025b\.[DRT: Deep reasoning translation via long chain\-of\-thought](https://doi.org/10.18653/v1/2025.findings-acl.351)\.In*Findings of the Association for Computational Linguistics: ACL 2025*, pages 6770–6782, Vienna, Austria\. Association for Computational Linguistics\.
- Wang et al\. \(2025c\)Jiaan Wang, Fandong Meng, and Jie Zhou\. 2025c\.[Deep reasoning translation via reinforcement learning](https://arxiv.org/abs/2504.10187)\.*Preprint*, arXiv:2504\.10187\.
- Wang et al\. \(2025d\)Jiaan Wang, Fandong Meng, and Jie Zhou\. 2025d\.[Extrans: Multilingual deep reasoning translation via exemplar\-enhanced reinforcement learning](https://arxiv.org/abs/2505.12996)\.*Preprint*, arXiv:2505\.12996\.
- Wang et al\. \(2026\)Jiaan Wang, Fandong Meng, and Jie Zhou\. 2026\.[DeepTrans: Deep reasoning translation via reinforcement learning](https://doi.org/10.1162/tacl.a.65)\.*Transactions of the Association for Computational Linguistics*, 14:47–63\.
- Wang et al\. \(2024a\)Longyue Wang, Siyou Liu, Chenyang Lyu, Wenxiang Jiao, Xing Wang, Jiahao Xu, Zhaopeng Tu, Yan Gu, Weiyu Chen, Minghao Wu, Liting Zhou, Philipp Koehn, Andy Way, and Yulin Yuan\. 2024a\.[Findings of the WMT 2024 shared task on discourse\-level literary translation](https://doi.org/10.18653/v1/2024.wmt-1.58)\.In*Proceedings of the Ninth Conference on Machine Translation*, pages 699–700, Miami, Florida, USA\. Association for Computational Linguistics\.
- Wang et al\. \(2023\)Longyue Wang, Zhaopeng Tu, Yan Gu, Siyou Liu, Dian Yu, Qingsong Ma, Chenyang Lyu, Liting Zhou, Chao\-Hong Liu, Yufeng Ma, Weiyu Chen, Yvette Graham, Bonnie Webber, Philipp Koehn, Andy Way, Yulin Yuan, and Shuming Shi\. 2023\.[Findings of the WMT 2023 shared task on discourse\-level literary translation: A fresh orb in the cosmos of LLMs](https://doi.org/10.18653/v1/2023.wmt-1.3)\.In*Proceedings of the Eighth Conference on Machine Translation*, pages 55–67, Singapore\. Association for Computational Linguistics\.
- Wang et al\. \(2024b\)Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui\. 2024b\.[Math\-shepherd: Verify and reinforce LLMs step\-by\-step without human annotations](https://doi.org/10.18653/v1/2024.acl-long.510)\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 9426–9439, Bangkok, Thailand\. Association for Computational Linguistics\.
- Wang et al\. \(2024c\)Yutong Wang, Jiali Zeng, Xuebo Liu, Fandong Meng, Jie Zhou, and Min Zhang\. 2024c\.[TasTe: Teaching large language models to translate through self\-reflection](https://doi.org/10.18653/v1/2024.acl-long.333)\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 6144–6158, Bangkok, Thailand\. Association for Computational Linguistics\.
- Xu et al\. \(2024a\)Haoran Xu, Young Jin Kim, Amr Sharaf, and Hany Hassan Awadalla\. 2024a\.[A paradigm shift in machine translation: Boosting translation performance of large language models](https://openreview.net/forum?id=farT6XXntP)\.In*The Twelfth International Conference on Learning Representations*\.
- Xu et al\. \(2024b\)Haoran Xu, Amr Sharaf, Yunmo Chen, Weiting Tan, Lingfeng Shen, Benjamin Van Durme, Kenton Murray, and Young Jin Kim\. 2024b\.[Contrastive preference optimization: Pushing the boundaries of LLM performance in machine translation](https://openreview.net/forum?id=51iwkioZpn)\.In*Forty\-first International Conference on Machine Learning*\.
- Yang et al\. \(2025\)Wenjie Yang, Mao Zheng, Mingyang Song, Zheng Li, and Sitong Wang\. 2025\.[Ssr\-zero: Simple self\-rewarding reinforcement learning for machine translation](https://arxiv.org/abs/2505.16637)\.*Preprint*, arXiv:2505\.16637\.
- Yao et al\. \(2024\)Binwei Yao, Ming Jiang, Tara Bobinac, Diyi Yang, and Junjie Hu\. 2024\.[Benchmarking machine translation with cultural awareness](https://doi.org/10.18653/v1/2024.findings-emnlp.765)\.In*Findings of the Association for Computational Linguistics: EMNLP 2024*, pages 13078–13096, Miami, Florida, USA\. Association for Computational Linguistics\.
- Zheng et al\. \(2024a\)Jiawei Zheng, Hanghai Hong, Feiyan Liu, Xiaoli Wang, Jingsong Su, Yonggui Liang, and Shikai Wu\. 2024a\.Fine\-tuning large language models for domain\-specific machine translation\.*arXiv preprint arXiv:2402\.15061*\.
- Zheng et al\. \(2024b\)Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, and Zheyan Luo\. 2024b\.[LlamaFactory: Unified efficient fine\-tuning of 100\+ language models](https://doi.org/10.18653/v1/2024.acl-demos.38)\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\)*, pages 400–410, Bangkok, Thailand\. Association for Computational Linguistics\.

Table 9:List of notation used in PAMT\.## Appendix ATraining Setup

### A\.1Training Data

#### Cold\-start SFT data\.

The cold\-start supervised fine\-tuning \(SFT\) stage uses a curated Long\-CoT translation dataset spanning ten domains and three translation directions: De→\\rightarrowEn, En→\\rightarrowZh, and Zh→\\rightarrowEn\. The dataset contains approximately 7K examples in total\. Each example follows the structured format<think\>\.\.\. </think\><answer\>\.\.\. </answer\>, where the<think\>segment describes an explicit translation process and the<answer\>segment provides the final translation\. These examples are distilled from strong teacher models to initialize the model’s ability to externalize translation reasoning\.

#### RL training data\.

For reinforcement learning \(RL\), we construct a multilingual multi\-domain machine translation training set covering three directions: De→\\rightarrowEn, En→\\rightarrowZh, and Zh→\\rightarrowEn\. The De→\\rightarrowEn portion is drawn from the multi\-domain dataset ofAharoni and Goldberg \([2020](https://arxiv.org/html/2608.03077#bib.bib1)\), covering IT, Law, Medical, Koran, and Subtitles\. The En→\\rightarrowZh portion is drawn from UM\-Corpus\(Tian et al\.,[2014](https://arxiv.org/html/2608.03077#bib.bib51)\), covering News, Laws, Subtitles, and Science\. The Zh→\\rightarrowEn Literary domain is taken from the GuoFeng\-Webnovel dataset used in the WMT23 and WMT24 literary translation tasks\(Wang et al\.,[2023](https://arxiv.org/html/2608.03077#bib.bib59),[2024a](https://arxiv.org/html/2608.03077#bib.bib58)\)\. From each domain, we randomly sample 2K sentence pairs and retain only examples with a minimum source length of 20 words, or 20 Chinese characters for Chinese inputs, to ensure sufficient room for explicit translation reasoning\. The final RL training set contains 20K examples\.

### A\.2Implementation Details

#### Cold\-start SFT\.

Our SFT implementation is based on LLaMA\-Factory111[https://github\.com/hiyouga/LLaMA\-Factory](https://github.com/hiyouga/LLaMA-Factory)Zheng et al\. \([2024b](https://arxiv.org/html/2608.03077#bib.bib67)\)\. We adopt Qwen\-2\.5\-7B\-Instruct and Gemma2\-9B\-IT as the base models and fine\-tune them on the 7K difficulty\-adaptive Long\-CoT examples with full\-parameter optimization\. Training is conducted for 2 epochs on 8 NVIDIA A100 80GB GPUs with a global batch size of 32\. We use the AdamW optimizer with a learning rate of1​e−51\\mathrm\{e\}\{\-5\}, a cosine learning rate scheduler, and a warm\-up ratio of 0\.1\. The maximum input sequence length is set to 4096 tokens\. We apply DeepSpeed ZeRO Stage 3 for memory\-efficient training\.

#### RL training\.

Our RL implementation is built onverl222[https://github\.com/volcengine/verl](https://github.com/volcengine/verl)Sheng et al\. \([2025](https://arxiv.org/html/2608.03077#bib.bib48)\)\. We train the model for 2 epochs on 8 NVIDIA A100 80GB GPUs with a global batch size of 128\. The rollout number is set to 8, and the rollout temperature is set to 1\.0\. For PPO updates, the optimization mini\-batch size is 16\. The learning rate is set to1​e−61\\mathrm\{e\}\{\-6\}, the KL loss coefficient isβ=1​e−3\\beta=1\\mathrm\{e\}\{\-3\}, and the process reward weight isλ=0\.1\\lambda=0\.1\. The maximum response length is set to 2048 tokens\. The entire RL training stage takes about 9 hours\.

#### Inference\.

During inference, we use thevLLM333[https://github\.com/vllm\-project/vllm](https://github.com/vllm-project/vllm)backend\(Kwon et al\.,[2023](https://arxiv.org/html/2608.03077#bib.bib26)\)for efficient decoding, with temperature set to 0\.0 and a repetition penalty of 1\.05\.

## Appendix BEvaluation Setup

We use the same evaluation setup for both the preliminary analysis in Section[2](https://arxiv.org/html/2608.03077#S2)and the main experiments in Section[5](https://arxiv.org/html/2608.03077#S5)\. This section describes the multi\-domain benchmark, the in\-domain and out\-of\-domain test sets, the multilingual evaluation protocol, and the post\-training data scale of MT\-specialized baselines\.

### B\.1Multi\-domain Evaluation Benchmark

#### Scope and design\.

We build a unified multi\-domain benchmark from publicly available corpora with explicit splits\. The benchmark serves two purposes\. First, it supports the diagnostic analyses in Section[2](https://arxiv.org/html/2608.03077#S2)by providing broad coverage across four translation directions: De⇒\\RightarrowEn, En⇒\\RightarrowDe, En⇒\\RightarrowZh, and Zh⇒\\RightarrowEn\. Second, it provides the in\-domain and out\-of\-domain evaluation sets used in the main experiments\. The benchmark spans both general\-purpose and domain\-specific text, including biomedical, legal, IT, science, subtitles, conversation, social media, cultural, commonsense, and literary content\. It is designed to cover high\-resource and low\-resource conditions, terminology\-intensive domains, context\-sensitive inputs, stylistically demanding text, and noisier informal genres, while minimizing potential data leakage and preserving comparability across systems\.

#### Data sources\.

For German⇔\\LeftrightarrowEnglish, the benchmark covers 11 domains\. Five domains—Medical, Law, IT, Koran, and Subtitles—are taken from the publicly available Multi\-Domain datasetAharoni and Goldberg \([2020](https://arxiv.org/html/2608.03077#bib.bib1)\)\. The Mixed domain is drawn from the WMT22 General Machine Translation TaskKocmi et al\. \([2022](https://arxiv.org/html/2608.03077#bib.bib24)\), whose subdomains include News, Social, E\-commerce, and Conversation\. Biomedical data is collected from the WMT18 and WMT19 Biomedical Machine Translation TasksNeves et al\. \([2018](https://arxiv.org/html/2608.03077#bib.bib35)\); Bawden et al\. \([2019](https://arxiv.org/html/2608.03077#bib.bib3)\)\.

For Chinese⇔\\LeftrightarrowEnglish, the benchmark covers 12 domains\. The Commonsense and Culture domains are taken from CommonMTHe et al\. \([2020](https://arxiv.org/html/2608.03077#bib.bib17)\)and CAMTYao et al\. \([2024](https://arxiv.org/html/2608.03077#bib.bib65)\), respectively, while the Literary domain is taken from the WMT23 Literary Machine Translation TaskWang et al\. \([2023](https://arxiv.org/html/2608.03077#bib.bib59)\)\. The general\-purpose Mixed domain, together with its News, Social, E\-commerce, and Conversation subdomains, is again sourced from the WMT22 General Machine Translation TaskKocmi et al\. \([2022](https://arxiv.org/html/2608.03077#bib.bib24)\)\. Biomedical data is collected from the WMT18 and WMT19 Biomedical Machine Translation TasksNeves et al\. \([2018](https://arxiv.org/html/2608.03077#bib.bib35)\); Bawden et al\. \([2019](https://arxiv.org/html/2608.03077#bib.bib3)\)\. In addition, the Laws, News, Science, and Subtitles domains are drawn from UM\-CorpusTian et al\. \([2014](https://arxiv.org/html/2608.03077#bib.bib51)\)\.

#### Detailed statistics\.

Tables[10](https://arxiv.org/html/2608.03077#A2.T10)–[13](https://arxiv.org/html/2608.03077#A2.T13)report the number of examples for each domain and direction\. No additional filtering or augmentation is applied to these evaluation sets\.

Table 10:Test sets and the number of samples for Zh⇒\\RightarrowEn translation tasks\. Bio denotes the Biomedical domain\.Table 11:Test sets and the number of samples for En⇒\\RightarrowZh translation tasks\.Table 12:Test sets and the number of samples for De⇒\\RightarrowEn translation tasks\.Table 13:Test sets and the number of samples for En⇒\\RightarrowDe translation tasks\.

### B\.2In\-Domain Test Sets

For in\-domain evaluation, we use the official test sets associated with the corpora used for RL training\. Specifically, these sets are drawn from the German\-English multi\-domain dataset\(Aharoni and Goldberg,[2020](https://arxiv.org/html/2608.03077#bib.bib1)\), UM\-Corpus\(Tian et al\.,[2014](https://arxiv.org/html/2608.03077#bib.bib51)\), and the GuoFeng\-Webnovel literary dataset\(Wang et al\.,[2023](https://arxiv.org/html/2608.03077#bib.bib59),[2024a](https://arxiv.org/html/2608.03077#bib.bib58)\)\. For the Literary domain, we merge thevalid\_1,valid\_2,test\_1, andtest\_2splits into a single evaluation set\. Table[14](https://arxiv.org/html/2608.03077#A2.T14)summarizes the in\-domain test sets used for the three training directions\.

Table 14:In\-domain test sets and the number of samples for En↔\\leftrightarrowZh and De→\\rightarrowEn translation tasks\.
### B\.3Out\-of\-Domain Test Sets

To evaluate cross\-domain generalization, we use public test sets from domains that are not included in the RL training data\. Specifically, the Conversation, E\-commerce, and Social domains are taken from the WMT22 shared tasksKocmi et al\. \([2022](https://arxiv.org/html/2608.03077#bib.bib24)\), the Culture domain is drawn from CAMTYao et al\. \([2024](https://arxiv.org/html/2608.03077#bib.bib65)\), and the Commonsense domain comes from CommonMTHe et al\. \([2020](https://arxiv.org/html/2608.03077#bib.bib17)\)\. These test sets complement the in\-domain evaluation by introducing more informal, culturally grounded, and knowledge\-sensitive inputs\. Table[15](https://arxiv.org/html/2608.03077#A2.T15)reports the corresponding statistics\.

Table 15:Out\-of\-domain test sets and sample counts for En↔\\leftrightarrowZh and De→\\rightarrowEn translation tasks\.
### B\.4Multilingual Evaluation

For unseen\-language evaluation, we use the FLORES\+ benchmark\(NLLB Team et al\.,[2024](https://arxiv.org/html/2608.03077#bib.bib37)\)and construct an*unseen*language set to minimize leakage from languages already covered by baseline post\-training data \(Table[6](https://arxiv.org/html/2608.03077#S5.T6)\)\. We first remove all languages appearing in baseline training coverage, including Chinese \(zh\), English \(en\), German \(de\), French \(fr\), Spanish \(es\), Portuguese \(pt\), Italian \(it\), Russian \(ru\), Korean \(ko\), Dutch \(nl\), Czech \(cs\), Icelandic \(is\), Ukrainian \(uk\), Hindi \(hi\), Japanese \(ja\), Polish \(pl\), Swedish \(sv\), Hungarian \(hu\), Romanian \(ro\), Danish \(da\), Norwegian \(no\), and Finnish \(fi\)\. We then further restrict the remaining FLORES\+ languages to those supported by both COMET and COMETKIWI, so that evaluation is consistent across all unseen directions\. After these two filtering steps, the final unseen\-language set contains 59 languages, listed in Table[16](https://arxiv.org/html/2608.03077#A2.T16)\.

The*seen*languages in our setting are German \(de\), English \(en\), and Chinese \(zh\), which appear in our training data through the German\-English multi\-domain dataset\(Aharoni and Goldberg,[2020](https://arxiv.org/html/2608.03077#bib.bib1)\), UM\-Corpus\(Tian et al\.,[2014](https://arxiv.org/html/2608.03077#bib.bib51)\), and the GuoFeng\-Webnovel literary dataset\.

Table 16:The 59 unseen languagesℒunseen\\mathcal\{L\}\_\{\\mathrm\{unseen\}\}used for En↔\\leftrightarrowX evaluation after filtering FLORES\+ by \(i\) post\-training language coverage of evaluated backbones and \(ii\) COMET/COMETKIWI language support\.
### B\.5Training Data Scale of MT Baselines

To facilitate fair comparison, we report the post\-training data scale used by each MT\-specialized baseline\. PAMT and most reasoning\-augmented baselines, including MT\-R1\-Zero\-7B, CoT\-FT\-7B, SFT\-Parallel, and mExTrans\-7B, are trained on roughly 27K examples\. In contrast, several other baselines use substantially larger corpora: TowerInstruct uses 637K examples, Tower\-Plus\-9B uses 286K, and ALMA\-R uses 21K\. SSR\-X\-Zero\-7B is trained on a smaller subset of 13K instances\.

### B\.6Evaluation Metrics

#### Automatic Metrics\.

We use three automatic evaluation metrics:

- •BLEU444[https://github\.com/mjpost/sacrebleu](https://github.com/mjpost/sacrebleu)Post \([2018](https://arxiv.org/html/2608.03077#bib.bib41)\), which evaluates surface\-level n\-gram overlap between system output and reference translations\.
- •COMET555Unbabel/wmt22\-comet\-daRei et al\. \([2022a](https://arxiv.org/html/2608.03077#bib.bib42)\), a reference\-based semantic metric trained on human quality judgments\.
- •CometKiwi666Unbabel/wmt22\-cometkiwi\-daRei et al\. \([2022b](https://arxiv.org/html/2608.03077#bib.bib45)\), a reference\-free version of COMET, useful when reference quality is poor or unavailable\.

#### MQM Evaluation Protocol\.

CategoryError TypeDescriptionAccuracyMistranslationInaccurate translation causing semantic distortion\.AdditionAdding extra information or emotions not in the source\.Under\-translationFailure to fully convey cultural or contextual nuances\.OmissionUnintentional exclusion of content from the source text\.UntranslatedRetaining source text without translation\.HallucinationGenerating content unrelated to the source text\.Off\-target TranslationTranslation misalignment caused by ambiguous input\.ContradictionContradicting itself or the source text\.FluencyGrammarErrors in sentence structure or syntax\.PunctuationIncorrect use of punctuation marks\.SpellingMisspelling of words\.Semantic RepetitionUnnecessary repetition of words or phrases\.Logical IncoherenceLack of logical flow or coherence in translation\.StyleAwkward ExpressionStilted or unnatural phrasing in the target language\.Unidiomatic UsageLiteral translation causes unnatural wording or grammar\.Style InconsistencyInconsistent stylistic choices within the translation\.Over\-localizationExcessive cultural adaptation leading to distortion\.TerminologyTerminology InconsistencyInconsistent translation of the same term\.Terminology MisuseUse of incorrect or inappropriate terms for the domain\.Cross\-domain ConfusionMisuse of terms due to domain shifts\.Incorrect Unit ConversionErrors in unit conversion \(e\.g\., metric to imperial\)\.FormattingErrors in domain\-specific formats \(e\.g\., legal, medical\)\.OthersAny other errors not covered in the above categories\.Source ErrorErrors present in the Source text itself\.Non\-translation ErrorTranslation is unassessable and unrelated to the Source\.Table 17:MQM Hierarchy\.To evaluate translation quality across domains and systems, we adopt an enhanced MQM hierarchy tailored to the characteristics of LLMs\. Based on the official MQM taxonomy, we introduce additional error types frequently observed in LLM outputs—such as hallucination, semantic repetition, and cross\-domain confusion\. Our final schema spans seven dimensions: Accuracy, Fluency, Style, Terminology, Others, Source Error, and Non\-translation Error, covering 25 fine\-grained error categories \(Table[17](https://arxiv.org/html/2608.03077#A2.T17)\)\.

For consistent and scalable annotation, we employ DeepSeek\-V3 as the automatic scoring model across all evaluations\. It is applied uniformly to both traditional and LRMs to ensure fair comparison\. The model follows a structured prompt designed to mimic human assessment while enforcing strict formatting and error attribution rules\. The full prompt is shown in Figure[5](https://arxiv.org/html/2608.03077#A6.F5)\.

### B\.7API and Implementation Details

The OpenAI, DeepSeek, and Gemini models used in this study are accessed via the following APIs: gpt\-4o\-2024\-11\-20, o1\-2024\-12\-17, o3\-mini\-2025\-01\-31, and gpt\-5\-2025\-08\-07 for OpenAI; deepseek\-chat\-2024\-12\-26 and deepseek\-reasoner for DeepSeek; and gemini\-2\.0\-flash and gemini\-2\.0\-flash\-thinking\-exp\-2025\-01\-21 for Gemini\.

## Appendix CHuman–V3 MQM Agreement

Automated MQM annotation may introduce bias, but full human MQM annotation is costly at the scale of our evaluation\. To validate the reliability of our automatic annotator, we randomly sample 3K examples from the human MQM annotations released by the WMT24 Metrics Shared Task\(Freitag et al\.,[2024](https://arxiv.org/html/2608.03077#bib.bib14)\)and compare them with DeepSeek\-V3 MQM labels under the same error schema\. The error\-presence agreement reaches 0\.8753, supporting the use of DeepSeek\-V3 for scalable MQM error analysis\. We therefore use automatic MQM as a proxy for aggregate and category\-level analysis, while not treating it as a replacement for expert MQM annotation\.

## Appendix DHuman Evaluation

To check whether PAMT’s gains are only artifacts of automatic metrics, we conduct a human preference evaluation on 60 examples\. Each example is annotated by three expert annotators and compares PAMT against one strong LLM, one strong LRM, and one strong MT baseline\. As shown in Table[18](https://arxiv.org/html/2608.03077#A4.T18), PAMT is competitive with GPT\-5 and DeepSeek\-V3, with more than half of the examples judged as ties, and is clearly preferred over CoT\-FT\. These results suggest that the improvements are not only due to metric overfitting\.

Table 18:Human preference evaluation on 60 examples\. Values are percentages\.
## Appendix EDiscussion on PRM\-Style Step Supervision

As discussed in Section[3](https://arxiv.org/html/2608.03077#S3), recent MT\-oriented process reward methods introduce process\-level feedback through external LLM scoring or terminology constraints, but they do not isolate the marginal contribution of each explicit translation step\. A broader line of PRM\-style supervision has been developed for reasoning tasks\(Lightman et al\.,[2023](https://arxiv.org/html/2608.03077#bib.bib32); Wang et al\.,[2024b](https://arxiv.org/html/2608.03077#bib.bib60); Lai et al\.,[2024](https://arxiv.org/html/2608.03077#bib.bib28); Luo et al\.,[2024](https://arxiv.org/html/2608.03077#bib.bib33); Li et al\.,[2023](https://arxiv.org/html/2608.03077#bib.bib30); Ni et al\.,[2023](https://arxiv.org/html/2608.03077#bib.bib36)\)\. These methods fit math and code reasoning, where intermediate states often admit relatively clear correctness signals\.

Such assumptions are weaker in MT\. Translation reasoning is not a chain of uniquely correct proof steps: multiple analyses may support valid translations, and step quality depends on terminology, style, discourse context, and lexical alternatives\. PAMT therefore avoids training a separate PRM or collecting human step labels\. Instead, it derives step\-level credit from the marginal change in reference likelihood under a frozen model, reusing the supervision already available in parallel MT data\.

## Appendix FPrompt Templates

You are a professional translation quality evaluator following the MQM \(Multidimensional Quality Metrics\) framework\.Task Instructions:1\.Compare the Prediction against both the Source and the Reference\.2\.Identify up to five of the most serious Errors for each translation sentence, using the MQM error types listed below\.3\.Assign exactly one severity level to each error\.4\.Special handling for two specific error types:•Source Error: Errors present in the Source text itself\.•Non\-translation Error: Translation is unassessable and unrelated to the Source\.5\.If no errors are found, return an empty JSON list \[\]\.6\.Only output a valid JSON object\. Do not include any additional text, comments, or explanations\.MQM Error Types \(See Table[17](https://arxiv.org/html/2608.03077#A2.T17)for the complete hierarchy\):•Mistranslation: Incorrect translation that alters or distorts meaning\.•Addition: Insertion of information or emotion not present in the source\.•Under\-translation: Partial omission of relevant or necessary information\.•Omission: Complete exclusion of content present in the source\.•Untranslated: Source text is copied without being translated\.Severity Levels:•Minor: Slight impact on readability or style; meaning remains clear\.•Major: Significant impact on usability, comprehension, or meaning\.Output Format:```
{
  "errors": [
    {
      "error_type": "Mistranslation",
      "severity": "Major",
      "explanation": "The correct translation should be ... "
    }
  ]
}
```

Source:\{source\_text\}Reference:\{reference\_text\}Prediction:\{prediction\_text\}Figure 5:Full prompt used to calculate MQM scores with DeepSeek\-V3\.Your task is to assess the difficulty of translating a given\{src\_lang\}sentence into\{tgt\_lang\}\. Please evaluate the difficulty based on the following criteria and output the result in JSON format, with the key "level":1\.Sentence complexity: Determine if the sentence is a simple sentence, a compound sentence, or includes subordinate clauses and other complex structures\.2\.Vocabulary difficulty: Assess whether the sentence contains commonly used words or specialized terms or slang\.3\.Grammar differences: Analyze if the sentence’s grammatical structure is similar to or differs significantly from\{tgt\_lang\}\.4\.Contextual understanding: Consider whether understanding specific cultural contexts or background knowledge is necessary for accurate translation\.Provide a difficulty level \(1\-5\), with 1 being the easiest and 5 being the most difficult\.And output the difficulty level in the following JSON format:```
{
    "level": "difficulty level"
}
```

Here is the\{src\_lang\}sentence:\{src\_text\}Figure 6:Full prompt used for evaluating translation difficulty with DeepSeek\-V3\.Translate the following\{src\_lang\}text into\{tgt\_lang\}while maintaining the domain style of the source text\.Source:\{src\_text\}Figure 7:Full prompt used for cold\-start data distillation\.Translate the following\{src\_lang\}text into\{tgt\_lang\}\.Source:\{src\_text\}Figure 8:Full prompt used for SFT and RL stage\.This section presents the prompts used in our experiments, including those for MQM\-style evaluation \(Figure[5](https://arxiv.org/html/2608.03077#A6.F5)\), difficulty assessment \(Figure[6](https://arxiv.org/html/2608.03077#A6.F6)\), cold\-start Long\-CoT data distillation \(Figure[7](https://arxiv.org/html/2608.03077#A6.F7)\), and model training in both SFT and RL stages \(Figure[8](https://arxiv.org/html/2608.03077#A6.F8)\)\.

## Appendix GCase Study: Resolving the Credit Assignment Bottleneck

Case Study: Trajectory of Process Potential in PAMT \(Law Domain\)Language Pair:German→\\rightarrowEnglish \(de\-en\)Scores:Format: 1\.0000 \| Answer: 2\.4624 \| Total: 3\.4624Source Input:\- "Übergangszeitraum": der Zeitraum zwischen dem 1\. Januar 1999 und dem 31, Dezember 2001\.Reference:\- "transitional period" shall mean the period between 1 January 1999 and 31 December 2001\.Prediction:\- "Transitional period": the period between 1 January 1999 and 31 December 2001\.Step\-by\-Step Reasoning Analysis:Initialϕ=−30\.7478\\phi=\-30\.7478•Step 1:Okay, let’s tackle this translation\. The original German sentence is about a term: "Übergangszeitraum": der Zeitraum zwischen dem 1\. Januar 1999 und dem 31\. Dezember 2001\.
Metrics:ϕ=−35\.9806\\phi=\-35\.9806ri,g,kproc=−5\.2327r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\-5\.2327•Step 2:First, translating the term "Übergangszeitraum" as "transitional period" makes sense in many contexts\. "Transitional" is straightforward here\. The phrase "der Zeitraum zwischen dem 1\. Januar 1999 und dem 31\. Dezember 2001" becomes "the period between 1 January 1999 and 31 December 2001\."
Metrics:ϕ=−32\.5788\\phi=\-32\.5788ri,g,kproc=\+3\.4018r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+3\.4018•Step 3:I should check the date format\. In European\-style writing, they use the ordinal, like "31\. December," but in English, we use numerals without the ordinal suffix, so "31 December\." Also, the years are written as 1999 and 2001\.
Metrics:ϕ=−32\.9279\\phi=\-32\.9279ri,g,kproc=−0\.3491r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\-0\.3491•Step 4:Putting it all together: "\- ’Transitional period’: the period between 1 January 1999 and 31 December 2001\."
Metrics:ϕ=−30\.3991\\phi=\-30\.3991ri,g,kproc=\+2\.5288r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+2\.5288•Step 5:I need to confirm if there’s any other terminology here\. The style is formal, so "period" is correct\. No need for extra words\. The original uses "der Zeitraum," which is straightforward, so "the period" works\.
Metrics:ϕ=−30\.0263\\phi=\-30\.0263ri,g,kproc=\+0\.3728r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+0\.3728•Step 6:I don’t see any technical jargon beyond this, so the translation should be straightforward\. No need for adjustments\.
Metrics:ϕ=−30\.7830\\phi=\-30\.7830ri,g,kproc=−0\.7567r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\-0\.7567•Step 7:Also, notice that in the original, there’s a hyphen at the start, which I should keep\. So the translation starts with "\-"\.
Metrics:ϕ=−30\.4701\\phi=\-30\.4701ri,g,kproc=\+0\.3129r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+0\.3129•Step 8:Double\-checking the dates: January 1 to December 31\. Yes, that’s the correct range\.
Metrics:ϕ=−30\.4014\\phi=\-30\.4014ri,g,kproc=\+0\.0687r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+0\.0687•Step 9:So the final translation should be:
Metrics:ϕ=−30\.7416\\phi=\-30\.7416ri,g,kproc=−0\.3401r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\-0\.3401•Step 10:\- "Transitional period": the period between 1 January 1999 and 31 December 2001\.
Metrics:ϕ=−29\.0100\\phi=\-29\.0100ri,g,kproc=\+1\.7316r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+1\.7316•Step 11:Yes, that’s concise and matches the formal style required\. The colon is kept as in the original\.
Metrics:ϕ=−28\.6008\\phi=\-28\.6008ri,g,kproc=\+0\.4091r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+0\.4091•Step 12:I think that’s all\. No other nuances in the German sentence\. The translation should be accurate and precise, which this seems to be\.
Metrics:ϕ=−28\.9584\\phi=\-28\.9584ri,g,kproc=−0\.3576r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\-0\.3576•Step 13:I don’t think there’s any idiomatic language here, so a direct translation works\. No need for adjustments to the structure\.
Metrics:ϕ=−28\.8792\\phi=\-28\.8792ri,g,kproc=\+0\.0792r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+0\.0792•Step 14:So the final translation is correct, and the style is formal and concise, as required\.
Metrics:ϕ=−27\.7316\\phi=\-27\.7316ri,g,kproc=\+1\.1476r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+1\.1476Figure 9:A detailed case study demonstrating the trajectory of process potentialϕ\\phiand step\-level gainri,g,kprocr\_\{i,g,k\}^\{\\mathrm\{proc\}\}during intermediate translation reasoning\. Positive values \(green\) indicate steps that increase the likelihood of the reference translation\.Case Study: Trajectory of Process Potential \(Terminological Rejection\)Language Pair:German→\\rightarrowEnglish \(de\-en\)Scores:Format: 1\.0000 \| Answer: 2\.0548 \| Total: 3\.0548Source Input:\- "Übergangszeitraum": der Zeitraum zwischen dem 1\. Januar 1999 und dem 31, Dezember 2001\.Ground Truth:\- "transitional period" shall mean the period between 1 January 1999 and 31 December 2001\.Final Answer:\- "Transitional period": the period from 1 January 1999 to 31 December 2001\.Step\-by\-Step Reasoning Analysis:Initialϕ=−30\.7478\\phi=\-30\.7478•Step 1:Okay, I need to translate the German sentence "\- ’Übergangszeitraum’: der Zeitraum zwischen dem 1\. Januar 1999 und dem 31, Dezember 2001\." into English\. First, let’s break down the sentence\.
Metrics:ϕ=−35\.0311\\phi=\-35\.0311ri,g,kproc=−4\.2833r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\-4\.2833•Step 2:"Übergangszeitraum" is "transition period" or "transitional period"\. "den Zeitraum zwischen dem 1\. Januar 1999 und dem 31\. Dezember 2001" translates to "the period from 1 January 1999 to 31 December 2001"\.
Metrics:ϕ=−34\.4548\\phi=\-34\.4548ri,g,kproc=\+0\.5764r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+0\.5764•Step 3:So the translation would be: "\- ’Transition period’: the period from 1 January 1999 to 31 December 2001\."
Metrics:ϕ=−35\.7834\\phi=\-35\.7834ri,g,kproc=−1\.3287r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\-1\.3287•Step 4:Now, I need to identify the domain\. The text is about a specific time frame, so it’s likely legal or governmental\. The style should be formal, so using "transition period" with a colon after the term is correct\.
Metrics:ϕ=−36\.7196\\phi=\-36\.7196ri,g,kproc=−0\.9362r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\-0\.9362•Step 5:Check for any nuances\. The German uses "bzw\." \(between\), so the translation is accurate\. Also, the dates are straightforward, no need for additional words\.
Metrics:ϕ=−36\.7160\\phi=\-36\.7160ri,g,kproc=\+0\.0036r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+0\.0036•Step 6:I should confirm that "Übergangszeitraum" is best as"transition period" rather than"transitory period" or"transitional period"\. Since the original uses "Übergangszeitraum", "transition" fits well\.
Metrics:ϕ=−37\.5387\\phi=\-37\.5387ri,g,kproc=−0\.8227r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\-0\.8227•Step 7:No other adjustments needed\. The sentence structure is simple, so the translation should mirror that simplicity\.
Metrics:ϕ=−36\.3380\\phi=\-36\.3380ri,g,kproc=\+1\.2007r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+1\.2007•Step 8:Final translation: "\- ’Transition period’: the period from 1 January 1999 to 31 December 2001\."
Metrics:ϕ=−36\.3168\\phi=\-36\.3168ri,g,kproc=\+0\.0212r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+0\.0212•Step 9:Yes, that’s concise and formal, which suits the domain style\. No extra fluff, just the necessary information\.
Metrics:ϕ=−35\.6590\\phi=\-35\.6590ri,g,kproc=\+0\.6578r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+0\.6578•Step 10:I think that’s all\. The domain is probably legal or financial, but since the text is a definition, the style is formal\. The translation seems to meet that requirement\.
Metrics:ϕ=−35\.1589\\phi=\-35\.1589ri,g,kproc=\+0\.5001r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+0\.5001•Step 11:I don’t see any other issues here\. The numbers and dates are correctly rendered, and the term is appropriately translated without any unnecessary complexity\.
Metrics:ϕ=−34\.6357\\phi=\-34\.6357ri,g,kproc=\+0\.5232r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+0\.5232•Step 12:So, the final answer is correct and the domain identification is accurate\. The translation reflects the formal style required by the context\.
Metrics:ϕ=−34\.4831\\phi=\-34\.4831ri,g,kproc=\+0\.1526r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+0\.1526Figure 10:Case study demonstrating terminological rejection\. In Step 6 \(ri,g,kproc=−0\.8227r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\-0\.8227\), the model explicitly considers the correct ground\-truth terminology \("transitional period"\) but erroneously talks itself out of it, dismissing it in favor of "transition period"\. The process potential accurately penalizes this active deviation from the reference terminology\.To elucidate how PAMT resolves the credit assignment bottleneck in reasoning\-augmented MT, we contrast two diverse reasoning trajectories sampled from the exact same source sentence during the RL rollout phase \(Figures[9](https://arxiv.org/html/2608.03077#A7.F9)and[10](https://arxiv.org/html/2608.03077#A7.F10)\)\. A persistent challenge in applying LRM to MDMT is terminological rejection—instances where a model successfully retrieves a correct domain constraint in its thought process but fails to execute it in the final translation\. Standard sequence\-level outcome rewards \(e\.g\., metric\-based reward\) struggle to optimize this, as they cannot isolate where the reasoning went wrong\. PAMT addresses this through dense, step\-level process gainsri,g,kprocr\_\{i,g,k\}^\{\\mathrm\{proc\}\}\. In the well\-aligned trajectory \(Figure[9](https://arxiv.org/html/2608.03077#A7.F9)\), the policy accurately hypothesizes and commits to the target\-domain term \("transitional period"\) early in the reasoning chain \(Step 2\)\. The process reward immediately isolates and validates this pivotal decision with a substantial gain \(ri,g,kproc=\+3\.4018r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\+3\.4018\), explicitly crediting the exact moment of domain alignment\. Conversely, the misaligned trajectory \(Figure[10](https://arxiv.org/html/2608.03077#A7.F10)\) exposes a critical reasoning flaw\. In Step 6, the model explicitly contemplates the ground\-truth term but actively talks itself out of it, dismissing it in favor of a sub\-optimal alternative \("transition period"\)\. While an outcome\-based reward would merely assign a slightly lower overall score to the final translation, PAMT’s process potential explicitly isolates this incorrect decision, applying a direct step\-level penalty \(ri,g,kproc=−0\.8227r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\-0\.8227\)\. Takeaway: This comparison confirms that PAMT does not merely optimize final string matching via trial and error\. By directly evaluating the marginal process gainri,g,kprocr\_\{i,g,k\}^\{\\mathrm\{proc\}\}of each intermediate step, PAMT explicitly steers the model’s translation decision, ensuring that domain\-faithful reasoning is reliably credited, sustained, and executed\.

1

2

Input :policy

πθ\\pi\_\{\\theta\}; frozen reference model

πref\\pi\_\{\\mathrm\{ref\}\}; training set

𝒟\\mathcal\{D\}; metric set

𝒬\\mathcal\{Q\}; group size

GG; process weight

λ\\lambda; clip ratio

ϵ\\epsilon; KL coefficient

β\\beta; number of GRPO optimization epochs

EoptE\_\{\\mathrm\{opt\}\}
Output :trained policy

πθ\\pi\_\{\\theta\}
3

4while*not converged*do

5

θold←θ\\theta\_\{\\mathrm\{old\}\}\\leftarrow\\theta
6Sample a prompt minibatch

ℬ=\{\(xi,yi∗\)\}i=1B\\mathcal\{B\}=\\\{\(x\_\{i\},y\_\{i\}^\{\*\}\)\\\}\_\{i=1\}^\{B\}from

𝒟\\mathcal\{D\}
7Initialize rollout buffer

ℛ←∅\\mathcal\{R\}\\leftarrow\\emptyset
8

//Phase 1: collect rollouts and compute rewards/advantages with fixedπθold\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}

9for*i←1i\\leftarrow 1toBB*do

10Sample

GGrollouts

\{oi,g\}g=1G∼πθold\(⋅∣xi\)\\\{o\_\{i,g\}\\\}\_\{g=1\}^\{G\}\\sim\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(\\cdot\\mid x\_\{i\}\)
11

12for*g←1g\\leftarrow 1toGG*do

13Parse

oi,go\_\{i,g\}into

<think\>​zi,g​</think\><answer\>​yi,g​</answer\>\\texttt\{<think\>\}z\_\{i,g\}\\texttt\{</think\><answer\>\}y\_\{i,g\}\\texttt\{</answer\>\}according to Eq\. \([1](https://arxiv.org/html/2608.03077#S4.E1)\)

14

15Split

zi,gz\_\{i,g\}by double newlines into steps

\[zi,g\(1\),…,zi,g\(Ki,g\)\]\[z\_\{i,g\}^\{\(1\)\},\\dots,z\_\{i,g\}^\{\(K\_\{i,g\}\)\}\]
16

17Cache old token log\-probabilities

ℓi,g,told←log⁡πθold​\(oi,g,t∣xi,oi,g,<t\)\\ell\_\{i,g,t\}^\{\\mathrm\{old\}\}\\leftarrow\\log\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(o\_\{i,g,t\}\\mid x\_\{i\},o\_\{i,g,<t\}\)for all response token positions

tt
18

19Compute format reward

ri,gfmtr\_\{i,g\}^\{\\mathrm\{fmt\}\}and outcome reward

ri,goutr\_\{i,g\}^\{\\mathrm\{out\}\}using Eq\. \([2](https://arxiv.org/html/2608.03077#S4.E2)\)

20

21Initialize

ri,g,tproc←0r\_\{i,g,t\}^\{\\mathrm\{proc\}\}\\leftarrow 0for all response token positions

tt
22

23Compute the initial process potential

ϕi,g,0\\phi\_\{i,g,0\}from the empty reasoning prefix

24

25for*k←1k\\leftarrow 1toKi,gK\_\{i,g\}*do

26Construct the prefix context

ci,g,kc\_\{i,g,k\}by Eq\. \([3](https://arxiv.org/html/2608.03077#S4.E3)\)

27

28Compute the process potential

ϕi,g,k\\phi\_\{i,g,k\}by Eq\. \([4](https://arxiv.org/html/2608.03077#S4.E4)\)

29

30Compute the step\-level process gain

ri,g,kproc←ϕi,g,k−ϕi,g,k−1r\_\{i,g,k\}^\{\\mathrm\{proc\}\}\\leftarrow\\phi\_\{i,g,k\}\-\\phi\_\{i,g,k\-1\}by Eq\. \([5](https://arxiv.org/html/2608.03077#S4.E5)\)

31

32Distribute

ri,g,kprocr\_\{i,g,k\}^\{\\mathrm\{proc\}\}uniformly to all tokens in step

zi,g\(k\)z\_\{i,g\}^\{\(k\)\}to obtain

ri,g,tprocr\_\{i,g,t\}^\{\\mathrm\{proc\}\}by Eq\. \([6](https://arxiv.org/html/2608.03077#S4.E6)\)

33

34

35Compute the token\-level reward sequence

\{ri,g,t\}\\\{r\_\{i,g,t\}\\\}by Eq\. \([7](https://arxiv.org/html/2608.03077#S4.E7)\)

36

37Compute return\-to\-go

\{Ri,g,t\}\\\{R\_\{i,g,t\}\\\}by Eq\. \([8](https://arxiv.org/html/2608.03077#S4.E8)\) and set

Ri,gtraj←Ri,g,1R\_\{i,g\}^\{\\mathrm\{traj\}\}\\leftarrow R\_\{i,g,1\}
38

39

40Compute group statistics

μi,σi\\mu\_\{i\},\\sigma\_\{i\}from

\{Ri,gtraj\}g=1G\\\{R\_\{i,g\}^\{\\mathrm\{traj\}\}\\\}\_\{g=1\}^\{G\}
41

42for*g←1g\\leftarrow 1toGG*do

43foreach*response token positionttin rolloutoi,go\_\{i,g\}*do

44Compute the token\-level advantage

Ai,g,t←Ri,g,t−μiσi\+ϵA\_\{i,g,t\}\\leftarrow\\frac\{R\_\{i,g,t\}\-\\mu\_\{i\}\}\{\\sigma\_\{i\}\+\\epsilon\}according to Eq\. \([9](https://arxiv.org/html/2608.03077#S4.E9)\)

45

46

47Add

\(xi,oi,g,\{ℓi,g,told\}t,\{Ai,g,t\}t\)\(x\_\{i\},o\_\{i,g\},\\\{\\ell\_\{i,g,t\}^\{\\mathrm\{old\}\}\\\}\_\{t\},\\\{A\_\{i,g,t\}\\\}\_\{t\}\)to rollout buffer

ℛ\\mathcal\{R\}
48

49

50

//Phase 2: optimizeπθ\\pi\_\{\\theta\}on the fixed rollout buffer

51for*e←1e\\leftarrow 1toEoptE\_\{\\mathrm\{opt\}\}*do

52foreach*optimization minibatchℬ~⊂ℛ\\widetilde\{\\mathcal\{B\}\}\\subset\\mathcal\{R\}*do

53Compute the importance ratio in Eq\. \([11](https://arxiv.org/html/2608.03077#S4.E11)\) using

πθ\\pi\_\{\\theta\}and cached

\{ℓi,g,told\}\\\{\\ell\_\{i,g,t\}^\{\\mathrm\{old\}\}\\\}
54

55Compute the GRPO objective on

ℬ~\\widetilde\{\\mathcal\{B\}\}by Eq\. \([10](https://arxiv.org/html/2608.03077#S4.E10)\)

56

57Take one optimizer step on

θ\\theta
58

59

60

Algorithm 1RL training ofPAMT
## Appendix HFrom GRPO to Process\-Aware Credit Assignment

For clarity, we first rewrite vanilla GRPO as terminal\-only token credit assignment, and then show how process reward changes the return\-to\-go at each token position\. We use the same rollout notation as in the main text:iiindexes the source sentence,ggindexes the rollout sampled for that source sentence, andttindexes response tokens\.

#### Vanilla GRPO\.

In vanilla GRPO, the sequence\-level reward is placed only on the last valid token:

ri,g,t=𝟏​\[t=ti,glast\]​\(ri,gfmt\+ri,gout\),\(λ=0\)\.r\_\{i,g,t\}=\\mathbf\{1\}\[t=t\_\{i,g\}^\{\\mathrm\{last\}\}\]\\bigl\(r\_\{i,g\}^\{\\mathrm\{fmt\}\}\+r\_\{i,g\}^\{\\mathrm\{out\}\}\\bigr\),\\qquad\(\\lambda=0\)\.\(12\)The return\-to\-go is therefore

Ri,g,t=∑u=tTi,gri,g,u=ri,gfmt\+ri,gout,t≤ti,glast\.R\_\{i,g,t\}=\\sum\_\{u=t\}^\{T\_\{i,g\}\}r\_\{i,g,u\}=r\_\{i,g\}^\{\\mathrm\{fmt\}\}\+r\_\{i,g\}^\{\\mathrm\{out\}\},\\qquad t\\leq t\_\{i,g\}^\{\\mathrm\{last\}\}\.\(13\)Thus, reward\-to\-go propagates the same terminal reward to all previous tokens, and the trajectory return isRi,gtraj=Ri,g,1R\_\{i,g\}^\{\\mathrm\{traj\}\}=R\_\{i,g,1\}\. For theGGrollouts sampled from the same source sentencexix\_\{i\}, we compute

μi=1G​∑g=1GRi,gtraj,\\mu\_\{i\}=\\frac\{1\}\{G\}\\sum\_\{g=1\}^\{G\}R\_\{i,g\}^\{\\mathrm\{traj\}\},\(14\)σi=1G​∑g=1G\(Ri,gtraj−μi\)2\.\\sigma\_\{i\}=\\sqrt\{\\frac\{1\}\{G\}\\sum\_\{g=1\}^\{G\}\\left\(R\_\{i,g\}^\{\\mathrm\{traj\}\}\-\\mu\_\{i\}\\right\)^\{2\}\}\.\(15\)Hence every valid token in the same rollout shares the same advantage:

Ai,g,t=Ri,g,t−μiσi\+ϵ\.A\_\{i,g,t\}=\\frac\{R\_\{i,g,t\}\-\\mu\_\{i\}\}\{\\sigma\_\{i\}\+\\epsilon\}\.\(16\)

#### Step\-level process reward\.

To obtain process\-level supervision, we first define a zero\-step prefix:

ci,g,0=\[xi,<think\></think\>,<answer\>\]\.c\_\{i,g,0\}=\[x\_\{i\},\\texttt\{<think\>\}\\texttt\{</think\>\},\\texttt\{<answer\>\}\]\.\(17\)Fork=1,…,Ki,gk=1,\\dots,K\_\{i,g\}, let

ci,g,k=\[xi,<think\>,zi,g,≤k,</think\>,<answer\>\]\.c\_\{i,g,k\}=\[x\_\{i\},\\texttt\{<think\>\},z\_\{i,g,\\leq k\},\\texttt\{</think\>\},\\texttt\{<answer\>\}\]\.\(18\)We define the process potential of thekk\-step prefix as

ϕi,g,k=∑m=1\|yi∗\|log⁡πref​\(yi,m∗∣ci,g,k,yi,<m∗\)\.\\phi\_\{i,g,k\}=\\sum\_\{m=1\}^\{\|y\_\{i\}^\{\*\}\|\}\\log\\pi\_\{\\mathrm\{ref\}\}\\bigl\(y\_\{i,m\}^\{\*\}\\mid c\_\{i,g,k\},y\_\{i,<m\}^\{\*\}\\bigr\)\.\(19\)The step\-level process gain is then

ri,g,kproc=ϕi,g,k−ϕi,g,k−1,k=1,…,Ki,g\.r\_\{i,g,k\}^\{\\mathrm\{proc\}\}=\\phi\_\{i,g,k\}\-\\phi\_\{i,g,k\-1\},\\qquad k=1,\\dots,K\_\{i,g\}\.\(20\)A positiveri,g,kprocr\_\{i,g,k\}^\{\\mathrm\{proc\}\}means that adding stepzi,g\(k\)z\_\{i,g\}^\{\(k\)\}makes the reference translation more predictable underπref\\pi\_\{\\mathrm\{ref\}\}\.

To optimize at the token level, we distribute each step gain uniformly over the tokens in that step:

ri,g,tproc=\{ri,g,kproc\|zi,g\(k\)\|,if token​t​belongs to step​zi,g\(k\),0,otherwise\.r\_\{i,g,t\}^\{\\mathrm\{proc\}\}=\\begin\{cases\}\\dfrac\{r\_\{i,g,k\}^\{\\mathrm\{proc\}\}\}\{\|z\_\{i,g\}^\{\(k\)\}\|\},&\\text\{if token \}t\\text\{ belongs to step \}z\_\{i,g\}^\{\(k\)\},\\\\\[8\.0pt\] 0,&\\text\{otherwise\.\}\\end\{cases\}\(21\)This preserves the total gain of each reasoning step while avoiding a length bias toward longer steps\.

#### Unified token reward\.

After adding process reward, the token\-level reward becomes

ri,g,t=𝟏​\[t=ti,glast\]​\(ri,gfmt\+ri,gout\)\+λ​ri,g,tproc\.r\_\{i,g,t\}=\{\\color\[rgb\]\{0\.109375,0\.46875,0\.2734375\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.109375,0\.46875,0\.2734375\}\\mathbf\{1\}\[t=t\_\{i,g\}^\{\\mathrm\{last\}\}\]\\bigl\(r\_\{i,g\}^\{\\mathrm\{fmt\}\}\+r\_\{i,g\}^\{\\mathrm\{out\}\}\\bigr\)\}\+\{\\color\[rgb\]\{0\.70703125,0\.25390625,0\.15625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.70703125,0\.25390625,0\.15625\}\\lambda\\,r\_\{i,g,t\}^\{\\mathrm\{proc\}\}\}\.\(22\)Compared with vanilla GRPO, thegreen termis unchanged, while thered termis the newly introduced dense process supervision\.

The corresponding return\-to\-go is

Ri,g,t\\displaystyle R\_\{i,g,t\}=∑u=tTi,gri,g,u\\displaystyle=\\sum\_\{u=t\}^\{T\_\{i,g\}\}r\_\{i,g,u\}\(23\)=ri,gfmt\+ri,gout\+λ​∑u=tTi,gri,g,uproc\.\\displaystyle=\{\\color\[rgb\]\{0\.109375,0\.46875,0\.2734375\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.109375,0\.46875,0\.2734375\}r\_\{i,g\}^\{\\mathrm\{fmt\}\}\+r\_\{i,g\}^\{\\mathrm\{out\}\}\}\+\{\\color\[rgb\]\{0\.70703125,0\.25390625,0\.15625\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0\.70703125,0\.25390625,0\.15625\}\\lambda\\sum\_\{u=t\}^\{T\_\{i,g\}\}r\_\{i,g,u\}^\{\\mathrm\{proc\}\}\}\.Equation \([23](https://arxiv.org/html/2608.03077#A8.E23)\) reduces to Eq\. \([13](https://arxiv.org/html/2608.03077#A8.E13)\) whenλ=0\\lambda=0, i\.e\., standard GRPO\.

More explicitly, suppose tokenttis theℓ\\ell\-th token in reasoning stepzi,g\(k\)z\_\{i,g\}^\{\(k\)\}, whereℓ∈\{1,…,\|zi,g\(k\)\|\}\\ell\\in\\\{1,\\dots,\|z\_\{i,g\}^\{\(k\)\}\|\\\}\. Then

Ri,g,t=ri,gfmt\+ri,gout\+λ\(\\displaystyle R\_\{i,g,t\}=r\_\{i,g\}^\{\\mathrm\{fmt\}\}\+r\_\{i,g\}^\{\\mathrm\{out\}\}\+\\lambda\\Biggl\(\|zi,g\(k\)\|−ℓ\+1\|zi,g\(k\)\|​ri,g,kproc\\displaystyle\\frac\{\|z\_\{i,g\}^\{\(k\)\}\|\-\\ell\+1\}\{\|z\_\{i,g\}^\{\(k\)\}\|\}r\_\{i,g,k\}^\{\\mathrm\{proc\}\}\(24\)\+∑j=k\+1Ki,gri,g,jproc\)\.\\displaystyle\+\\sum\_\{j=k\+1\}^\{K\_\{i,g\}\}r\_\{i,g,j\}^\{\\mathrm\{proc\}\}\\Biggr\)\.Hence, different reasoning tokens aggregate different suffixes of future process reward\. Tokens whose remaining reasoning suffix is more helpful receive larger returns, while tokens followed by harmful steps receive smaller returns\. In contrast, for tokens in the<answer\>span, all process rewards lie in the past, so

Ri,g,t=ri,gfmt\+ri,gout,t∈<answer\>span\.R\_\{i,g,t\}=r\_\{i,g\}^\{\\mathrm\{fmt\}\}\+r\_\{i,g\}^\{\\mathrm\{out\}\},\\qquad t\\in\\texttt\{<answer\>\}\\text\{ span\}\.\(25\)
The trajectory return now becomes

Ri,gtraj=Ri,g,1=ri,gfmt\+ri,gout\+λ​∑k=1Ki,gri,g,kproc\.R\_\{i,g\}^\{\\mathrm\{traj\}\}=R\_\{i,g,1\}=r\_\{i,g\}^\{\\mathrm\{fmt\}\}\+r\_\{i,g\}^\{\\mathrm\{out\}\}\+\\lambda\\sum\_\{k=1\}^\{K\_\{i,g\}\}r\_\{i,g,k\}^\{\\mathrm\{proc\}\}\.\(26\)

#### Advantage normalization\.

As in GRPO, we normalize the token return using group statistics computed from the updated trajectory returns:

μi=1G​∑g=1GRi,gtraj,\\mu\_\{i\}=\\frac\{1\}\{G\}\\sum\_\{g=1\}^\{G\}R\_\{i,g\}^\{\\mathrm\{traj\}\},\(27\)σi=1G​∑g=1G\(Ri,gtraj−μi\)2\.\\sigma\_\{i\}=\\sqrt\{\\frac\{1\}\{G\}\\sum\_\{g=1\}^\{G\}\\left\(R\_\{i,g\}^\{\\mathrm\{traj\}\}\-\\mu\_\{i\}\\right\)^\{2\}\}\.\(28\)The token\-level advantage is then

Ai,g,t=Ri,g,t−μiσi\+ϵ\.A\_\{i,g,t\}=\\frac\{R\_\{i,g,t\}\-\\mu\_\{i\}\}\{\\sigma\_\{i\}\+\\epsilon\}\.\(29\)Unlike vanilla GRPO,Ai,g,tA\_\{i,g,t\}now depends on the token position through the remaining future process reward\.

#### Optimization\.

The importance ratio is

ρi,g,t​\(θ\)=πθ​\(oi,g,t∣xi,oi,g,<t\)πθold​\(oi,g,t∣xi,oi,g,<t\)\.\\rho\_\{i,g,t\}\(\\theta\)=\\frac\{\\pi\_\{\\theta\}\(o\_\{i,g,t\}\\mid x\_\{i\},o\_\{i,g,<t\}\)\}\{\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(o\_\{i,g,t\}\\mid x\_\{i\},o\_\{i,g,<t\}\)\}\.\(30\)We further define the clipped ratio

ρ¯i,g,t​\(θ\)=clip⁡\(ρi,g,t​\(θ\),1−ϵ,1\+ϵ\)\.\\bar\{\\rho\}\_\{i,g,t\}\(\\theta\)=\\operatorname\{clip\}\\bigl\(\\rho\_\{i,g,t\}\(\\theta\),\\,1\-\\epsilon,\\,1\+\\epsilon\\bigr\)\.\(31\)and the clipped token objective

ℓi,g,t​\(θ\)=min⁡\(ρi,g,t​\(θ\)​Ai,g,t,ρ¯i,g,t​\(θ\)​Ai,g,t\)\.\\ell\_\{i,g,t\}\(\\theta\)=\\min\\Bigl\(\\rho\_\{i,g,t\}\(\\theta\)A\_\{i,g,t\},\\bar\{\\rho\}\_\{i,g,t\}\(\\theta\)A\_\{i,g,t\}\\Bigr\)\.\(32\)The GRPO\-style clipped loss is

ℒclip​\(θ\)=−1B​G​∑i=1B∑g=1G1Ti,g​∑t=1Ti,gℓi,g,t​\(θ\),\\mathcal\{L\}\_\{\\mathrm\{clip\}\}\(\\theta\)=\-\\frac\{1\}\{BG\}\\sum\_\{i=1\}^\{B\}\\sum\_\{g=1\}^\{G\}\\frac\{1\}\{T\_\{i,g\}\}\\sum\_\{t=1\}^\{T\_\{i,g\}\}\\ell\_\{i,g,t\}\(\\theta\),\(33\)and the full objective is

ℒ​\(θ\)=ℒclip​\(θ\)\+β​KL​\(πθ∥πref\)\.\\mathcal\{L\}\(\\theta\)=\\mathcal\{L\}\_\{\\mathrm\{clip\}\}\(\\theta\)\+\\beta\\,\\mathrm\{KL\}\\bigl\(\\pi\_\{\\theta\}\\,\\\|\\,\\pi\_\{\\mathrm\{ref\}\}\\bigr\)\.\(34\)
In this view, vanilla GRPO is simply the special caseλ=0\\lambda=0, where reward\-to\-go propagates the same terminal reward to all previous tokens\. PAMT keeps the same GRPO optimization backbone, but changes the token return by adding dense process rewards, making token credit assignment position\-dependent\.

#### Toy example\.

Consider one rollout with three reasoning steps\. Suppose their step\-level gains are

\(ri,g,kproc\)k=13=\(0\.4,−0\.3,0\.2\),\\bigl\(r\_\{i,g,k\}^\{\\mathrm\{proc\}\}\\bigr\)\_\{k=1\}^\{3\}=\(0\.4,\\,\-0\.3,\\,0\.2\),and their lengths are

\(\|zi,g\(1\)\|,\|zi,g\(2\)\|,\|zi,g\(3\)\|\)=\(2,1,1\)\.\(\|z\_\{i,g\}^\{\(1\)\}\|,\\,\|z\_\{i,g\}^\{\(2\)\}\|,\\,\|z\_\{i,g\}^\{\(3\)\}\|\)=\(2,\\,1,\\,1\)\.Intuitively, the first step is helpful, the second step introduces process drift, and the third step partially corrects it\.

By distributing each step gain uniformly over its tokens, the token\-level process rewards become

\(ri,g,tproc\)t=14=\(0\.2,0\.2,−0\.3,0\.2\)\.\\bigl\(r\_\{i,g,t\}^\{\\mathrm\{proc\}\}\\bigr\)\_\{t=1\}^\{4\}=\(0\.2,\\,0\.2,\\,\-0\.3,\\,0\.2\)\.Supposeri,gfmt\+ri,gout=0\.8r\_\{i,g\}^\{\\mathrm\{fmt\}\}\+r\_\{i,g\}^\{\\mathrm\{out\}\}=0\.8andλ=1\\lambda=1\. Then the return\-to\-go on the four reasoning tokens is

\(Ri,g,1,Ri,g,2,Ri,g,3,Ri,g,4\)=\(1\.1,0\.9,0\.7,1\.0\),\(R\_\{i,g,1\},\\,R\_\{i,g,2\},\\,R\_\{i,g,3\},\\,R\_\{i,g,4\}\)=\(1\.1,\\,0\.9,\\,0\.7,\\,1\.0\),because each token receives the sequence\-level reward plus the suffix sum of the remaining process rewards\. For example,

Ri,g,2=0\.8\+\(0\.2−0\.3\+0\.2\)=0\.9,R\_\{i,g,2\}=0\.8\+\(0\.2\-0\.3\+0\.2\)=0\.9,while

Ri,g,3=0\.8\+\(−0\.3\+0\.2\)=0\.7\.R\_\{i,g,3\}=0\.8\+\(\-0\.3\+0\.2\)=0\.7\.For tokens in the<answer\>span, all process rewards already lie in the past, so

Ri,g,t=0\.8\.R\_\{i,g,t\}=0\.8\.
Under vanilla GRPO, all of these positions would instead receive the same return0\.80\.8\. This example shows that process\-aware credit assignment is not simply “earlier is better”; rather, a token receives a higher return when the remaining reasoning suffix is more helpful for producing the reference translation\.

Similar Articles

Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation

arXiv cs.CL

This paper introduces Translation with Thought (TwT), a resource-rational framework for multi-domain machine translation that adaptively modulates reasoning effort based on input difficulty, trained via supervised fine-tuning on difficulty-aware reasoning traces and reinforcement learning. TwT-7B and TwT-14B outperform larger SOTA reasoning models while reducing token usage by 32–60%.

Translate-R1: Cost-Aware Translation Tool Use via Reinforcement Learning

arXiv cs.CL

Translate-R1 introduces a reinforcement learning approach for cost-aware translation tool use in LLMs, where the model learns to decide when to translate inputs based on its own comprehension and a cost-sensitivity parameter, achieving Pareto-optimal trade-offs across multiple languages.

Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR

Hugging Face Daily Papers

This paper proposes Transfer-Aware Curriculum (TAC), a bandit-style online curriculum for multi-domain RLVR that prioritizes domains whose updates benefit other domains using gradient-geometry alignment. TAC improves macro-averaged accuracy on Qwen3-1.7B and Llama3.2-3B over fixed and learnability-only curricula.

Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation

Hugging Face Daily Papers

The paper proposes SMART, a self-evolving multi-agent system for long-form subtitle translation that uses test-time training with a dynamic router, Mixture-of-Agents layer, and judge-refiner loop to build series-level memory and adapt agent prompts without retraining the underlying LLMs. It also introduces the SubtitleArena benchmark (14 genres, 15 target locales) and the SubMQM evaluation framework, achieving a 6.9% average MQM penalty reduction over competing agent systems and top results on the MuSC benchmark.