APEX-VW: A Document-Level English-Spanish Post-Editing Dataset in the Healthcare Domain
Summary
This paper introduces APEX-VW, a new document-level English-Spanish post-editing dataset built from NHS virtual-ward documents and professional post-editing in Trados Studio. The corpus is designed to support research on terminology normalisation, correction propagation, and human-in-the-loop translation support.
View Cached Full Text
Cached at: 08/11/26, 08:06 AM
# APEX-VW: A Document-Level English-Spanish Post-Editing Dataset in the Healthcare Domain
Source: [https://arxiv.org/html/2608.08059](https://arxiv.org/html/2608.08059)
\(2026\)
###### Abstract\.
Post\-Editing \(PE\) of Machine Translation \(MT\) output often involves repeating the same lexical and terminological corrections across many segments, especially in specialised and highly repetitive documents\. Despite substantial work on Automatic Post\-Editing \(APE\), most available corpora operate at the sentence level, others are synthetic, and overall not designed to study how corrections propagate in realistic Computer\-Assisted Translation \(CAT\) workflows\. This paper presents the APEX\-VW \(Automatic Post\-Editing eXperiments on Virtual Wards\) Corpus, a new open English\-Spanish \(EN\-ES\) dataset built from recent NHS virtual\-ward documents and professional PE in Trados Studio, with controlled MT, terminology, and quality assurance settings\. The corpus contains seven document\-coherent source texts totalling 42k words, translated with four MT systems representing different paradigms and then post\-edited by professional translators\. Unlike prior resources such as WMT APE corpora, eSCAPE, MLQE\-PE, or LangMark, the dataset preserves document order and CAT\-tool context, making it suitable for research on terminology normalisation, correction propagation, and human\-in\-the\-loop translation support\. The paper describes the corpus design, data preparation, PE setup, and initial corpus statistics, and positions the resource as a benchmark for document\-level APE and propagation\-aware assistive tools\.
post\-editing, automatic post\-editing, post\-editing dataset, machine translation, document\-level machine translation, terminology, healthcare
††conference:Conference on Information and Knowledge Management; 2026; Rome, Italy††journalyear:2026††ccs:Information systems Machine translation††ccs:Computing methodologies Natural language processing## 1\.Introduction
Post\-editing \(PE\) is now a routine part of professional translation workflows, but it often requires translators to correct repeated Machine Translation \(MT\) errors segment by segment\(Domingoet al\.,[2017](https://arxiv.org/html/2608.08059#bib.bib24)\)\. This is particularly problematic in specialised domains, where a suboptimal translation of a key term may recur many times within and across documents, creating unnecessary effort and increasing the risk of inconsistency if some instances are overlooked\(Čuloet al\.,[2017](https://arxiv.org/html/2608.08059#bib.bib25)\)\. Existing Computer\-Assisted Translation \(CAT\) tools – such as Trados Studio or memoQ – partly address this through Translation Memories \(TMs\) and Term Bases \(TBs\), yet they generally do not propagate newly made PE corrections beyond exact or near\-exact repetition patterns\.
Previous work on adaptive, interactive, and online PE has repeatedly explored the idea that systems should learn from user feedback during the translation process rather than treat each segment as an isolated unit\(Knight and Chander,[1994](https://arxiv.org/html/2608.08059#bib.bib2); Ortiz\-Martínezet al\.,[2010](https://arxiv.org/html/2608.08059#bib.bib4); Simard and Foster,[2013](https://arxiv.org/html/2608.08059#bib.bib5); Lagardaet al\.,[2015](https://arxiv.org/html/2608.08059#bib.bib6); Chatterjeeet al\.,[2017](https://arxiv.org/html/2608.08059#bib.bib7); Negriet al\.,[2018a](https://arxiv.org/html/2608.08059#bib.bib8)\)\. From the professional perspective, user\-oriented studies likewise suggest that translators value tools that can remember and reuse prior decisions, particularly in repetitive workflows where consistency is critical\(Lagoudaki,[2008](https://arxiv.org/html/2608.08059#bib.bib10); Escribe and Candel\-Mora,[2024](https://arxiv.org/html/2608.08059#bib.bib11)\)\. At the same time, currently available PE and Automatic Post\-Editing \(APE\) datasets do not align well with this use case\. Existing resources have been crucial for the development of APE and related evaluation tasks\(Bojaret al\.,[2015](https://arxiv.org/html/2608.08059#bib.bib12); do Carmoet al\.,[2020](https://arxiv.org/html/2608.08059#bib.bib1); Fomichevaet al\.,[2022](https://arxiv.org/html/2608.08059#bib.bib15); Velazquezet al\.,[2025](https://arxiv.org/html/2608.08059#bib.bib16)\), but they are predominantly sentence\-level and mostly designed for benchmarking rather than for studying document\-level correction behaviour in realistic CAT environments\. This leaves a gap for openly licensed, professionally post\-edited, document\-coherent datasets that make it possible to investigate correction propagation, terminology normalisation, and consistency support across a translation project\.
Therefore, the present study introduces such a resource: an open English–Spanish \(EN–ES\) PE dataset of healthcare service documents, which are high in repetitive terminology\. The corpus was designed to support research on term normalisation, repeated correction patterns, and propagation\-aware assistance in realistic CAT settings\. Its main contributions are: \(1\) a document\-coherent, openly reusable EN–ES PE corpus in a high\-risk healthcare domain; \(2\) a controlled Trados\-based workflow with four MT paradigms and professional PE; and \(3\) detailed corpus and setup documentation to support future work on document\-level APE, interactive PE, and consistency support for MT\.
## 2\.Related Work
APE is commonly defined as the task of improving raw MT output by learning from triplets composed of source text, raw MT output, and a corresponding post\-edited version\(do Carmoet al\.,[2020](https://arxiv.org/html/2608.08059#bib.bib1)\)\. As MT, APE has evolved from rule\-based approaches to statistical and neural methods\(do Carmoet al\.,[2020](https://arxiv.org/html/2608.08059#bib.bib1)\)\. The broader idea that MT should work in partnership with human revisers is anchored in early reflections on human–machine cooperation in translation\(Bar\-Hillel,[1960](https://arxiv.org/html/2608.08059#bib.bib9)\)\. This idea was the foundations of early implementations of APE, including\(Knight and Chander,[1994](https://arxiv.org/html/2608.08059#bib.bib2)\)and\(Allen and Hogan,[2000](https://arxiv.org/html/2608.08059#bib.bib3)\)\. These early approaches are particularly relevant to the present work because they already foregrounded the basic intuition behind propagation: repeated human corrections should become reusable knowledge for subsequent segments\. This intuition was further developed in interactive MT research, with models designed to update incrementally from user intervention, allowing the system to adapt during the translation process rather than only offline\(Ortiz\-Martínezet al\.,[2010](https://arxiv.org/html/2608.08059#bib.bib4)\)\.
More direct work on propagation emerged in PE itself\. PEPr explicitly proposed propagation as a way of learning from corrections on the fly and applying them to similar future segments\(Simard and Foster,[2013](https://arxiv.org/html/2608.08059#bib.bib5)\)\. This work paved the way for subsequent online APE systems\(Lagardaet al\.,[2015](https://arxiv.org/html/2608.08059#bib.bib6); Chatterjeeet al\.,[2017](https://arxiv.org/html/2608.08059#bib.bib7); Negriet al\.,[2018a](https://arxiv.org/html/2608.08059#bib.bib8)\)\. Together, these studies show that adaptive PE is most meaningful when repeated correction patterns can be captured and reused over time\.
The practical importance of such assistance is also supported from the post\-editor’s perspective\. Survey\-based evidence suggests that translators value tools that can remember and reuse previous decisions\(Lagoudaki,[2008](https://arxiv.org/html/2608.08059#bib.bib10)\)\. More recently, a survey reported that correction propagation is among the features practitioners most want in CAT environments for PE projects\(Escribe and Candel\-Mora,[2024](https://arxiv.org/html/2608.08059#bib.bib11)\)\. This indicates that propagation is not merely an algorithmic convenience, but a practical requirement linked to productivity, consistency, and reduced cognitive burden in repetitive projects\.
Despite this motivation, currently available corpora only partially support research in this direction\. WMT APE shared\-task datasets were central to the development of the field, but they were primarily designed for benchmarking and differ substantially across editions in domain, MT quality, and PE conditions\(Bojaret al\.,[2015](https://arxiv.org/html/2608.08059#bib.bib12); do Carmoet al\.,[2020](https://arxiv.org/html/2608.08059#bib.bib1)\)\. Large synthetic resources such as eSCAPE greatly expand the amount of training data available for APE, but because they use human references as proxy post\-edits, thus not capturing authentic human PE behaviour\(Negriet al\.,[2018b](https://arxiv.org/html/2608.08059#bib.bib13)\)\. Newer resources have further enriched the landscape: SubEdits adds professionally post\-edited subtitle data\([Chollampattet al\.,](https://arxiv.org/html/2608.08059#bib.bib14)\), MLQE\-PE combines post\-edits with multilingual Quality Estimation \(QE\) annotations\(Fomichevaet al\.,[2022](https://arxiv.org/html/2608.08059#bib.bib15)\), and LangMark expands multilingual APE benchmarking with modern MT outputs and curated triplets\(Velazquezet al\.,[2025](https://arxiv.org/html/2608.08059#bib.bib16)\)\. However, these resources remain predominantly sentence\-level and do not preserve contiguous document structure or CAT\-workflow context, which limits their usefulness for studying terminology normalisation and correction propagation across a document\.
The dataset presented in this paper is intended to complement, rather than replace, those resources\. Instead of maximising scale or multilingual breadth, it prioritises document coherence, open licensing, and workflow realism through professional EN–ES PE in Trados Studio over a set of recent healthcare documents from a public entity\. This design makes it particularly suitable for studying terminology normalisation, repeated corrections, and propagation\-aware assistance in conditions that are difficult to capture in shuffled sentence collections\.
## 3\.Dataset Construction
This section presents how the dataset was created, from the selection of source texts and the PE setup, to its availability, and also discusses ethical considerations\.
### 3\.1\.Source Selection and Preprocessing
The dataset focuses on healthcare, specifically virtual wards and related service\-delivery texts produced by the National Health Service \(NHS\)\. This domain was chosen because it combines high practical relevance with dense and repeated terminology around care pathways, hospital\-at\-home models, monitoring, referral, diagnostics, and organisational procedures\. It also offers recent public\-sector documents with clear licensing and stable provenance, making open redistribution feasible\. After considering product documentation, instructions for use and technical documentation, the corpus design prioritised openly licensed NHS and related UK public\-sector materials\. This decision trades broader genre coverage for legal clarity, accessibility, and domain coherence\. The final source set contains seven documents published between 2022 and 2025 and totals roughly 42k words after cleaning\. Because the documents come from the same institutional ecosystem, they share conceptual framing and terminology, which is important for studying propagation not only within documents but also across the corpus as a whole\. All source texts were converted from PDF or HTML into DOCX and cleaned conservatively for use in Trados Studio\. The cleaning process preserved linguistically meaningful content and document structure while removing layout artefacts such as page furniture, navigation elements, duplicated headers, reference lists, and low\-value numeric tables\. This process produced source files that resemble realistic translation assignments and support stable segmentation and alignment for downstream MT and PE analysis\.
### 3\.2\.MT and Post\-Editing Setup
The corpus was processed in Trados Studio 2024 under controlled conditions designed to approximate professional PE workflows while keeping the setup reproducible\. Standard sentence\-based segmentation was used, segment split/merge was allowed for local corrections, and source editing was restricted to fixing obvious artefacts inherited from source conversion\. Quality assurance checks were kept close to Trados Studio defaults with minor adjustments \(including terminology consistency, numeric agreement, and simple punctuation mismatches\)\. No pre\-existing TM was used at project creation, although an empty project TM was populated during work\. An indicative EN\-ES TB was provided to support consistency without constraining translators too rigidly\. Four MT systems were selected to represent major paradigms currently used in localisation workflows: DeepL, ModernMT, Language Weaver, and an OpenAI\-based GPT\-5 system \(all available as plug\-ins through the RWS Appstore\)\. All MT outputs were generated on 18 April 2026, using the then\-current production models exposed through each provider’s integrations in Trados Studio\. Each document was associated with exactly one MT system\.
Three professional linguists were involved in this phase: two carried out full PE following expectations of publishable quality, and a third focused on proofreading the resulting translations and ensuring terminological consistency\. The linguists were between 25 and 40 years old, had academic training in translation and/or linguistics, and had professional experience in EN¿ES translation\. All had native\-level competence in the target language\. They were compensated according to their agreed professional rates and informed that the work formed part of a research project, but they were instructed to post\-edit as they normally would in response to the translation brief\. Access to the original NHS materials was provided for context and coherence checking\. The assignment was organised by document batch rather than by isolated segments\. Linguist\_1 post\-edited the DeepL and ModernMT batches, corresponding to Source Text \(ST\) ST1 and ST2, while Linguist\_2 post\-edited the Language Weaver and OpenAI batches, corresponding to ST3–ST5 and ST6–ST7, respectively\. This document\-level allocation preserved MT consistency patterns and document flow, which is important for studying propagation\-related phenomena over time\.
### 3\.3\.Availability and Ethical Considerations
The design and release of APEX\-VW explicitly follow the FAIR guiding principles for scientific data management and stewardship\(Wilkinsonet al\.,[2016](https://arxiv.org/html/2608.08059#bib.bib20)\)\.
- •Findable:The dataset and accompanying terminology resources are deposited in Zenodo with a persistent DOI, rich metadata \(including domain, language pair, MT systems, licences, and version\), and standardised keywords\.
- •Accessible:All files are downloadable without registration under a clear open licence\.
- •Interoperable:The dataset is released in a widely used, open format \(CSV\) and encoded in UTF\-8, with explicit field descriptions, enabling use with standard MT, APE, and CAT\-tool pipelines\.
- •Reusable:Provenance is documented in detail \(source selection, cleaning, MT configuration, PE setup\), licences are stated explicitly, and a datasheet describes intended use cases, limitations, and known risks, supporting reliable reuse in future work\.
Following the Datasheets for Datasets recommendations\(Gebruet al\.,[2021](https://arxiv.org/html/2608.08059#bib.bib21)\), the repository includes a datasheet that documents the motivation, composition, collection process, preprocessing, uses, and ethical considerations of APEX\-VW\. The datasheet details, among others, the institutional sources and licences of the NHS materials, the profiles and compensation of the post\-editors, the MT and CAT configurations, and known limitations and appropriate use cases\.
From an ethical and legal perspective, the dataset is based exclusively on publicly available institutional documents and professional translation work\. No patient records, user\-generated content, or other personal or sensitive data were collected or released\.
Table 1\.Overview of the dataset by source text, showing corpus size, lexical profile, MT system, post\-editor allocation, and MT\-to\-PE quality metrics\.
## 4\.Initial Statistics and Relevance
Table[1](https://arxiv.org/html/2608.08059#S3.T1)summarises the main properties of the dataset at the document level, including corpus size, segment counts, lexical profile, MT system assignment, post\-editor allocation, and PE outcome metrics\. In total, the corpus contains 42,108 words and 2,712 analysed segments across seven document\-coherent source texts, with an average segment length of 15\.53 words, an overall Repetition Rate \(RR\) of 0\.19, a vocabulary size of 3,686, a lexical density of 0\.61, and a type–token ratio \(TTR\) ranging from 0\.13 to 0\.29111The corpus\-level TTR value \(0\.09\) is lower than any per\-document TTR because TTR decreases as text length increases, therefore the corpus\-level value is not directly comparable to the individual document values\.\. These figures confirm that the dataset combines moderate lexical variety with substantial repetition, making it suitable for studying terminology normalisation and propagation across segments and documents\.
Inspired by the use of RR proposed by\(Bertoldiet al\.,[2013](https://arxiv.org/html/2608.08059#bib.bib23)\)and later used as a complexity indicator in the WMT APE shared tasks\(Bhattacharyyaet al\.,[2023](https://arxiv.org/html/2608.08059#bib.bib22)\), we define RR for APEX\-VW as a measure of the repetitiveness of a text based on the rate of non\-singletonnn\-gram types forn=1…4n=1\\dots 4and combining them using the geometric mean\. Formally, for eachn∈\{1,2,3,4\}n\\in\\\{1,2,3,4\\\}, letVn\>1V\_\{n\}^\{\>1\}be the set ofnn\-gram types that occur more than once in the text andVnV\_\{n\}the set of allnn\-gram types\. The RR is then defined as
RR=\(∏n=14\|Vn\>1\|\|Vn\|\)1/4\.\\mathrm\{RR\}=\\Bigg\(\\prod\_\{n=1\}^\{4\}\\frac\{\\lvert V\_\{n\}^\{\>1\}\\rvert\}\{\\lvert V\_\{n\}\\rvert\}\\Bigg\)^\{\\\!1/4\}\.
A clear variation appears across documents\. RR values indeed range from 0\.07 to 0\.21, with the highest value observed in ST1 and the lowest in ST7, while average segment length ranges from 13\.94 to 17\.69 words\. This variation is useful because it creates different propagation conditions, from shorter, more repetitive texts to longer and lexically denser segments that may require more context\-sensitive corrections\.
The PE metrics suggest substantial differences in the amount of intervention required for different MT outputs, although these values should not be interpreted as a controlled ranking because MT systems were assigned by document rather than evaluated on the same source text\. At the corpus level, PE yielded TER 29\.28\(Snoveret al\.,[2006](https://arxiv.org/html/2608.08059#bib.bib18)\), BLEU 62\.38\(Papineniet al\.,[2002](https://arxiv.org/html/2608.08059#bib.bib17)\), and COMET 0\.88\(Reiet al\.,[2020](https://arxiv.org/html/2608.08059#bib.bib19)\)\. DeepL and ModernMT show lower TER and higher BLEU/COMET on their respective texts, whereas Language Weaver and OpenAI required more extensive editing on the documents assigned to them\. For the intended use of the dataset, the important point is not system comparison as such, but the fact that the corpus captures a range of PE effort profiles and therefore a range of propagation\-relevant correction patterns\.
Another contribution of the resource is the expansion of the terminology material during PE\. The initial indicative TB was deliberately lightweight \(21 terms, 29 acronyms\), but the PE process produced an updated TB with 215 term entries and a separate acronym resource with 79 entries, which were consolidated and validated during the final consistency\-checking step by the proofreading linguist\. These expanded resources are valuable in their own right because they record terminology decisions that emerged during the workflow and can support future research on terminology management, consistency modelling, and terminology\-aware MT and APE\. In addition, a further analysis of edit statistics shows that PE involved substantial rewriting rather than only minimal surface correction\. Across the corpus, 6,790 insertions, 1,285 deletions, and 9,523 replacements were implemented, with a mean edit distance of 6\.49 per segment\. The number of PE tokens \(57,659\) is also higher than the number of MT tokens \(52,154\), which suggests that many edits involved explicitation, restructuring, or expansion rather than simple substitution\. This is relevant for propagation research because it shows that repeated interventions are not limited to single\-word replacements but may involve richer correction patterns\. Finally, the corpus supports at least two propagation scenarios identified during dataset preparation\. First, in some cases, MT is internally consistent but repeatedly wrong, so the same incorrect translation must be corrected multiple times across a document \(Scenario A\)\. Second, in other cases, MT itself is inconsistent, and the post\-editor must normalise competing translations of the same concept across the document set \(Scenario B\)\. To illustrate Scenario A, we inspected how the English term*virtual ward*was translated within ST2\. In this document,*virtual ward*\(and its plural form\) occurs 35 times in the source across 31 segments, and has one main rendering in Spanish, namely*sala virtual*\(or*salas virtuales*\)\. In other words, MT is lexically consistent but systematically chooses a suboptimal term \(as the linguists chose to use*unidad de hospitalización virtual*, so corrections had to be implemented repeatedly wherever the term appears\. Scenario B is exemplified by ST1, where*virtual ward\(s\)*occurs frequently \(254 source instances\) but the MT system distributes these occurrences across several competing Spanish translations\. The most common options include*sala virtual*\(219\),*unidad virtual*\(15\) and*servicio de hospitalización virtual*\(6\), alongside a longer tail of rarer variants such as*distrito virtual*,*sistema virtual*and*unidad de atención a distancia*\. In this setting, the post\-editor is not only correcting individual instances, but also normalising terminology by converging these alternatives onto a preferred form throughout the document\.
## 5\.Limitations
It should be acknowledged that the dataset is restricted to one language direction, one domain cluster, a modest number of documents, and a small number of post\-editors\. It also reflects a snapshot of MT systems and PE decisions at a specific moment in time rather than a longitudinal record of evolving edits\. In addition, because different source documents were assigned to different MT systems, the current corpus is better suited to resource release and methodological study than to strict cross\-system benchmarking\. Moreover, the CAT setup targets a specific tool \(Trados Studio\) and configuration, which may limit direct transfer of observations to other environments\. These limits are balanced by the dataset’s main strength: it provides an openly reusable, professionally post\-edited, document\-level EN\-ES resource specifically designed for correction propagation research in realistic CAT conditions\. By combining document coherence, open licensing, MT diversity, and detailed workflow documentation, it fills a gap left by sentence\-level datasets and creates a basis for work on propagation\-aware PE assistance, document\-level MT and APE, and consistency support in specialised translation\.
## 6\.Future Work
Several avenues for future work are opened by this dataset\. First, the corpus can support the design and evaluation of propagation\-aware assistance in CAT tools, for example by learning when and how to propose document\-level term propagation or pattern\-based suggestions during PE\. Second, the resource enables systematic experiments on document\-level and online APE, including models that exploit repeated correction patterns and terminology edits to adapt over the course of a project\. Third, CAT environments could incorporate internal adaptive models that learn not only from term\-level corrections but also from recurring stylistic and register\-related edits\. Fourth, extending the dataset to additional domains, language pairs, and CAT environments would make it possible to test how robust propagation strategies are across different workflows\. Finally, combining the existing corpus with richer annotation—such as error typologies, fine\-grained effort measures, or explicit propagation events—would support more detailed studies of how human post\-editors manage consistency over time\.
## 7\.Conclusion
This study has presented APEX\-VW, an openly available EN–ES PE dataset in the healthcare domain, constructed from recent NHS virtual\-ward and related service documents\. By preserving document structure, CAT\-workflow context, and detailed MT and PE statistics, the corpus fills a gap left by predominantly sentence\-level APE resources\. The dataset is specifically designed to support research on correction propagation, terminology normalisation, and consistency support in realistic professional settings, and is intended to serve as a basis for new propagation\-aware models and tools in document\-level translation workflows\.
## Acknowledgements
This project is funded by the European Association for Machine Translation \(EAMT\) through its Sponsorship of Activities programme\. We would also like to express our sincere gratitude to Paloma Vega Centeno for her precious help during the post\-editing phase\.
## GenAI Usage Disclosure
The present paper and the accompanying dataset documentation were prepared with the assistance of Generative AI tools\. These tools were used to help with language polishing, restructuring of existing text, and formatting suggestions, based on content and instructions provided by the authors\. All writing suggestions were reviewed and approved by the authors\. All dataset design decisions, experimental configurations and analyses are original work by the authors\.
## References
- J\. Allen and C\. Hogan \(2000\)Toward the development of a post\-editing module for raw machine translation output: a controlled language perspective\.InProceedings of the Third International Controlled Language Applications Workshop \(CLAW\-00\),Cited by:[§2](https://arxiv.org/html/2608.08059#S2.p1.1)\.
- Y\. Bar\-Hillel \(1960\)The present status of automatic translation of languages\.Advances in Computers1,pp\. 91–163\.Cited by:[§2](https://arxiv.org/html/2608.08059#S2.p1.1)\.
- N\. Bertoldi, M\. Cettolo, and M\. Federico \(2013\)Cache\-based online adaptation for machine translation enhanced computer assisted translation\.InProceedings of Machine Translation Summit XIV: Papers,A\. Way, K\. Sima’an, and M\. L\. Forcada \(Eds\.\),Nice, France\.External Links:[Link](https://aclanthology.org/2013.mtsummit-papers.5/)Cited by:[§4](https://arxiv.org/html/2608.08059#S4.p2.7)\.
- P\. Bhattacharyya, R\. Chatterjee, M\. Freitag, D\. Kanojia, M\. Negri, and M\. Turchi \(2023\)Findings of the WMT 2023 shared task on automatic post\-editing\.InProceedings of the Eighth Conference on Machine Translation \(WMT 2023\),Singapore,pp\. 672–681\.External Links:[Link](https://aclanthology.org/2023.wmt-1.55)Cited by:[§4](https://arxiv.org/html/2608.08059#S4.p2.7)\.
- O\. Bojar, R\. Chatterjee, C\. Federmann, B\. Haddow, M\. Huck, C\. Hokamp, P\. Koehn, V\. Logacheva, C\. Monz, M\. Negri, M\. Post, C\. Scarton, L\. Specia, and M\. Turchi \(2015\)Findings of the 2015 workshop on statistical machine translation\.InProceedings of the Tenth Workshop on Statistical Machine Translation,O\. Bojar, R\. Chatterjee, C\. Federmann, B\. Haddow, C\. Hokamp, M\. Huck, V\. Logacheva, and P\. Pecina \(Eds\.\),Lisbon, Portugal,pp\. 1–46\.External Links:[Link](https://aclanthology.org/W15-3001/),[Document](https://dx.doi.org/10.18653/v1/W15-3001)Cited by:[§1](https://arxiv.org/html/2608.08059#S1.p2.1),[§2](https://arxiv.org/html/2608.08059#S2.p4.1)\.
- R\. Chatterjee, G\. Gebremelak, M\. Negri, and M\. Turchi \(2017\)Online automatic post\-editing for multi\-domain neural machine translation\.InProceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers,M\. Lapata, P\. Blunsom, and A\. Koller \(Eds\.\),Valencia, Spain,pp\. 525–535\.External Links:[Link](https://aclanthology.org/E17-1050/)Cited by:[§1](https://arxiv.org/html/2608.08059#S1.p2.1),[§2](https://arxiv.org/html/2608.08059#S2.p2.1)\.
- \[7\]S\. Chollampatt, R\. H\. Susanto, L\. Tan, and E\. SzymanskaCan automatic post\-editing improve nmt? the subedits corpus and experiments\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 2736–2746\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.217/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.217)Cited by:[§2](https://arxiv.org/html/2608.08059#S2.p4.1)\.
- O\. Čulo, S\. Hansen\-Schirra, J\. Nitzke, G\. De Sutter, M\. Lefer, and I\. Delaere \(2017\)Contrasting terminological variation in post\-editing and human translation of texts from the technical and medical domain\.InEmpirical Translation Studies: New Methodological and Theoretical Traditions,G\. De Sutter, M\. Lefer, and I\. Delaere \(Eds\.\),Benjamins Translation Library, Vol\.300,pp\. 183–206\.External Links:[Link](https://books.google.es/books?id=u2vNDgAAQBAJ)Cited by:[§1](https://arxiv.org/html/2608.08059#S1.p1.1)\.
- F\. do Carmo, D\. Shterionov, J\. Moorkens, J\. Wagner, M\. Hossari, E\. Paquin, D\. Schmidtke, D\. Groves, and A\. Way \(2020\)A review of the state\-of\-the\-art in automatic post\-editing\.Machine Translation35\(1–2\),pp\. 1–39\.External Links:[Document](https://dx.doi.org/10.1007/s10590-020-09252-y),[Link](https://api.semanticscholar.org/CorpusID:234437292)Cited by:[§1](https://arxiv.org/html/2608.08059#S1.p2.1),[§2](https://arxiv.org/html/2608.08059#S2.p1.1),[§2](https://arxiv.org/html/2608.08059#S2.p4.1)\.
- M\. Domingo, Á\. Peris, and F\. Casacuberta \(2017\)Segment\-based interactive\-predictive machine translation\.Machine Translation31\(4\),pp\. 163–185\.External Links:[Document](https://dx.doi.org/10.1007/s10590-017-9213-3),[Link](https://link.springer.com/article/10.1007/s10590-017-9213-3)Cited by:[§1](https://arxiv.org/html/2608.08059#S1.p1.1)\.
- M\. Escribe and M\. Á\. Candel\-Mora \(2024\)Optimising translation tools for post\-editing: results of a user survey\.InProceedings of New Trends in Translation and Technology \(NeTTT 2024\),Varna, Bulgaria,pp\. 63–75\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.26615/issn.2815-4711.2024%5F006)Cited by:[§1](https://arxiv.org/html/2608.08059#S1.p2.1),[§2](https://arxiv.org/html/2608.08059#S2.p3.1)\.
- M\. Fomicheva, S\. Sun, E\. Fonseca, C\. Zerva, F\. Blain, V\. Chaudhary, F\. Guzmán, N\. Lopatina, L\. Specia, and A\. F\. T\. Martins \(2022\)MLQE\-pe: a multilingual quality estimation and post\-editing dataset\.InProceedings of the Thirteenth Language Resources and Evaluation Conference,N\. Calzolari, F\. Béchet, P\. Blache, K\. Choukri, C\. Cieri, T\. Declerck, S\. Goggi, H\. Isahara, B\. Maegaard, J\. Mariani, H\. Mazo, J\. Odijk, and S\. Piperidis \(Eds\.\),Marseille, France,pp\. 4963–4974\.External Links:[Link](https://aclanthology.org/2022.lrec-1.530/)Cited by:[§1](https://arxiv.org/html/2608.08059#S1.p2.1),[§2](https://arxiv.org/html/2608.08059#S2.p4.1)\.
- T\. Gebru, J\. Morgenstern, B\. Vecchione, J\. W\. Vaughan, H\. Wallach, H\. D\. III, and K\. Crawford \(2021\)Datasheets for datasets\.Commun\. ACM64\(12\),pp\. 86–92\.External Links:ISSN 0001\-0782,[Link](https://doi.org/10.1145/3458723),[Document](https://dx.doi.org/10.1145/3458723)Cited by:[§3\.3](https://arxiv.org/html/2608.08059#S3.SS3.p4.1)\.
- K\. Knight and I\. Chander \(1994\)Automated postediting of documents\.Proceedings of AAAI,pp\. 779–784\.Cited by:[§1](https://arxiv.org/html/2608.08059#S1.p2.1),[§2](https://arxiv.org/html/2608.08059#S2.p1.1)\.
- A\. L\. Lagarda, D\. Ortiz\-Martínez, V\. Alabau, and F\. Casacuberta \(2015\)Translating without in\-domain corpus: machine translation post\-editing with online learning techniques\.Comput\. Speech Lang\.3229\(13–4\),pp\. 109––134\.External Links:ISSN 0885\-2308,[Link](https://doi.org/10.1016/j.csl.2014.10.004),[Document](https://dx.doi.org/10.1016/j.csl.2014.10.004)Cited by:[§1](https://arxiv.org/html/2608.08059#S1.p2.1),[§2](https://arxiv.org/html/2608.08059#S2.p2.1)\.
- E\. Lagoudaki \(2008\)The value of machine translation for the professional translator\.InProceedings of the 8th Conference of the Association for Machine Translation in the Americas \(AMTA\),Waikiki, USA,pp\. 262–269\.External Links:[Link](https://aclanthology.org/2008.amta-srw.4/)Cited by:[§1](https://arxiv.org/html/2608.08059#S1.p2.1),[§2](https://arxiv.org/html/2608.08059#S2.p3.1)\.
- M\. Negri, M\. Turchi, N\. Bertoldi, and M\. Federico \(2018a\)Online neural automatic post\-editing for neural machine translation\.InProceedings of the Fifth Italian Conference on Computational Linguistics \(CLiC\-it 2018\),E\. Cabrio, A\. Mazzei, and F\. Tamburini \(Eds\.\),Turin, Italy,pp\. 289–294\.External Links:[Link](https://aclanthology.org/2018.clicit-1.51/)Cited by:[§1](https://arxiv.org/html/2608.08059#S1.p2.1),[§2](https://arxiv.org/html/2608.08059#S2.p2.1)\.
- M\. Negri, M\. Turchi, R\. Chatterjee, and N\. Bertoldi \(2018b\)ESCAPE: a large\-scale synthetic corpus for automatic post\-editing\.InProceedings of the Eleventh International Conference on Language Resources and Evaluation \(LREC 2018\),N\. Calzolari, K\. Choukri, C\. Cieri, T\. Declerck, S\. Goggi, K\. Hasida, H\. Isahara, B\. Maegaard, J\. Mariani, H\. Mazo, A\. Moreno, J\. Odijk, S\. Piperidis, and T\. Tokunaga \(Eds\.\),Miyazaki, Japan\.External Links:[Link](https://aclanthology.org/L18-1004/)Cited by:[§2](https://arxiv.org/html/2608.08059#S2.p4.1)\.
- D\. Ortiz\-Martínez, I\. García\-Varea, and F\. Casacuberta \(2010\)Online learning for interactive statistical machine translation\.InHuman Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics,R\. Kaplan, J\. Burstein, M\. Harper, and G\. Penn \(Eds\.\),Los Angeles, California,pp\. 546–554\.External Links:[Link](https://aclanthology.org/N10-1079/)Cited by:[§1](https://arxiv.org/html/2608.08059#S1.p2.1),[§2](https://arxiv.org/html/2608.08059#S2.p1.1)\.
- K\. Papineni, S\. Roukos, T\. Ward, and W\. Zhu \(2002\)BLEU: a method for automatic evaluation of machine translation\.InProceedings of the 40th Annual Meeting of the Association for Computational Linguistics,P\. Isabelle, E\. Charniak, and D\. Lin \(Eds\.\),Philadelphia, Pennsylvania, USA,pp\. 311–318\.External Links:[Link](https://aclanthology.org/P02-1040/),[Document](https://dx.doi.org/10.3115/1073083.1073135)Cited by:[§4](https://arxiv.org/html/2608.08059#S4.p4.1)\.
- R\. Rei, C\. Stewart, A\. C\. Farinha, and A\. Lavie \(2020\)COMET: a neural framework for mt evaluation\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 2685–2702\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.213/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.213)Cited by:[§4](https://arxiv.org/html/2608.08059#S4.p4.1)\.
- M\. Simard and G\. Foster \(2013\)PEPr: post\-edit propagation using online learning\.InProceedings of the XIV Machine Translation Summit,pp\. 191–198\.External Links:[Link](https://aclanthology.org/2013.mtsummit-papers.24.pdf)Cited by:[§1](https://arxiv.org/html/2608.08059#S1.p2.1),[§2](https://arxiv.org/html/2608.08059#S2.p2.1)\.
- M\. Snover, B\. Dorr, R\. Schwartz, L\. Micciulla, and J\. Makhoul \(2006\)A study of translation edit rate with targeted human annotation\.InProceedings of the 7th Conference of the Association for Machine Translation in the Americas: Technical Papers,Cambridge, Massachusetts, USA,pp\. 223–231\.External Links:[Link](https://aclanthology.org/2006.amta-papers.25/)Cited by:[§4](https://arxiv.org/html/2608.08059#S4.p4.1)\.
- D\. Velazquez, M\. Grace, K\. Karageorgos, L\. Carin, A\. Schliem, D\. Zaikis, and R\. Wechsler \(2025\)LangMark: a multilingual benchmark for automatic post\-editing in the era of large language models\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 32653–32667\.External Links:[Link](https://aclanthology.org/2025.acl-long.1569/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1569)Cited by:[§1](https://arxiv.org/html/2608.08059#S1.p2.1),[§2](https://arxiv.org/html/2608.08059#S2.p4.1)\.
- M\. D\. Wilkinson, M\. Dumontier, I\. J\. Aalbersberg, G\. Appleton, M\. Axton, A\. Baak, N\. Blomberg, J\. W\. Boiten, L\. B\. da Silva Santos, P\. E\. Bourne, J\. Bouwman, A\. J\. Brookes, T\. Clark, M\. Crosas, I\. Dillo, O\. Dumon, S\. Edmunds, C\. T\. Evelo, R\. Finkers, A\. Gonzalez\-Beltran, A\. J\. G\. Gray, P\. Groth, C\. Goble, J\. S\. Grethe, J\. Heringa, P\. A\. C\. ’t Hoen, R\. Hooft, T\. Kuhn, R\. Kok, J\. Kok, S\. J\. Lusher, M\. E\. Martone, A\. Mons, A\. L\. Packer, B\. Persson, P\. Rocca\-Serra, M\. Roos, A\. J\. van Schaik, S\. Sansone, E\. Schultes, T\. Sengstag, T\. Slater, G\. Strawn, M\. A\. Swertz, M\. Thompson, J\. van der Lei, E\. van Mulligen, J\. Velterop, A\. Waagmeester, P\. Wittenburg, K\. Wolstencroft, J\. Zhao, and B\. Mons \(2016\)The FAIR guiding principles for scientific data management and stewardship\.Scientific Data3,pp\. 160018\.External Links:[Document](https://dx.doi.org/10.1038/sdata.2016.18)Cited by:[§3\.3](https://arxiv.org/html/2608.08059#S3.SS3.p2.1)\.Similar Articles
AI_LectureNote: A Retrospective Pilot Study of a Post-ASR Workflow for English-Script Rendering and Semantic Drift in Korean-English Medical Lectures
A retrospective pilot study evaluating AI_LectureNote, a post-ASR workflow for Korean-English medical lectures, showing that while it improves English-script rendering, it introduces semantic drift and polarity failures in the transcripts.
Enhancing Scientific Discourse: Machine Translation for the Scientific Domain
This paper presents the development of parallel and monolingual corpora for scientific machine translation across Spanish-English, French-English, and Portuguese-English, targeting four domains: Cancer Research, Energy Research, Neuroscience, and Transportation. The corpora are used to fine-tune neural machine translation systems, addressing challenges of specialized vocabulary and syntax in scientific text.
VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?
Introduces VisEditBench, a benchmark of 1,395 human-annotated visualization code-editing tasks for vision-language models, covering feedback-guided repair and reference-guided restyling. Evaluates 20 VLMs and proposes VisEditAgent, a render-grounded editing framework that improves pass rates from 55.75% to 67.99%.
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
VEFX-Bench introduces a large-scale human-annotated video editing dataset (5,049 examples) with multi-dimensional quality labels and a specialized reward model for standardized evaluation of video editing systems. The paper addresses the lack of comprehensive benchmarks in AI-assisted video creation by providing VEFX-Dataset, VEFX-Reward, and a 300-video-prompt benchmark that reveals gaps in current editing models.
AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
AVE-Compass is a new benchmark for holistic evaluation of audio-video editing abilities, with 145 videos, 196 editing instructions, and 2,688 checklist items. It also proposes AVE-Agent, a modular agent framework that improves cross-modal editing via self-reflection and evaluator feedback.