StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation

arXiv cs.CL Papers

Summary

StalePO introduces a token-level preference optimization method for machine translation that uses legacy post-edits to improve model performance, showing significant gains in quality metrics.

arXiv:2609.16340v1 Announce Type: new Abstract: Machine translation systems are periodically upgraded to stronger models, but the available preference signal is human post-edits of an older system's outputs, which the newer model may already surpass. Moreover, collecting fresh post-edits for every new model is prohibitively expensive. We call this the Stale Preference problem. Standard DPO can fail in this setting: it may increase the likelihood of inferior post-edits, erode the model's existing quality, and fail to provide the per-token control needed to correct localized errors. We introduce StalePO, an objective derived from three requirements this regime imposes. Likelihood movement must be downward on both responses, the policy must be anchored to its own base response, and the KL constraint must apply at the token level. These requirements are jointly necessary. In ablations, each mechanism in isolation leaves the model's performance indistinguishable from the base model, and only their combination converts stale feedback into gains. On English-to-Hindi and English-to-Turkish localization data, StalePO improves the fraction of segments passing all LLM-as-judge MQM quality checks by 14.9 and 4.6 percentage points, respectively, with gains concentrated on style and fluency. A human evaluation under the same framework confirms these gains on English-to-Hindi, raising the fraction of segments passing all seven human checks by 13.8 percentage points.
Original Article
View Cached Full Text

Cached at: 09/16/26, 08:45 AM

# StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation
Source: [https://arxiv.org/abs/2609.16340](https://arxiv.org/abs/2609.16340)
[View PDF](https://arxiv.org/pdf/2609.16340)

> Abstract:Machine translation systems are periodically upgraded to stronger models, but the available preference signal is human post\-edits of an older system's outputs, which the newer model may already surpass\. Moreover, collecting fresh post\-edits for every new model is prohibitively expensive\. We call this the Stale Preference problem\. Standard DPO can fail in this setting: it may increase the likelihood of inferior post\-edits, erode the model's existing quality, and fail to provide the per\-token control needed to correct localized errors\. We introduce StalePO, an objective derived from three requirements this regime imposes\. Likelihood movement must be downward on both responses, the policy must be anchored to its own base response, and the KL constraint must apply at the token level\. These requirements are jointly necessary\. In ablations, each mechanism in isolation leaves the model's performance indistinguishable from the base model, and only their combination converts stale feedback into gains\. On English\-to\-Hindi and English\-to\-Turkish localization data, StalePO improves the fraction of segments passing all LLM\-as\-judge MQM quality checks by 14\.9 and 4\.6 percentage points, respectively, with gains concentrated on style and fluency\. A human evaluation under the same framework confirms these gains on English\-to\-Hindi, raising the fraction of segments passing all seven human checks by 13\.8 percentage points\.

## Submission history

From: Rohit Dhaipule \[[view email](https://arxiv.org/show-email/33f30615/2609.16340)\] **\[v1\]**Mon, 14 Sep 2026 20:56:35 UTC \(3,442 KB\)

Similar Articles

Token-weighted Direct Preference Optimization with Attention

arXiv cs.CL

Proposes AttentionPO, a token-weighted direct preference optimization method that uses attention from the LLM itself to estimate token weights, improving alignment performance on AlpacaEval, MT-Bench, and ArenaHard without requiring a separate reward model.

StoicLLM: Preference Optimization for Philosophical Alignment in Small Language Models

arXiv cs.CL

This research paper investigates using preference optimization (ORPO, AlphaPO) on small language models like Llama-3.2-3B and Qwen-3-4B to align them with Stoic philosophy using micro-datasets. The study finds that while 300 examples can effectively encode Stoic virtues, small models still struggle with outward-facing cosmopolitan duties.