Contrastive ESA: Human Evaluation of Multiple Translations at Once

arXiv cs.CL Papers

Summary

Introduces Contrastive Error Span Annotation (cESA), a protocol for human evaluation of multiple translations simultaneously, reducing annotation time and noise compared to standard methods.

arXiv:2607.26640v1 Announce Type: new Abstract: Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost. We introduce Contrastive Error Span Annotation (cESA), a protocol that presents multiple translations of the source input (text, video, audio, image). In cESA, the annotator sees multiple translations of the same document, marks major and minor error spans, and then assigns a score from 0% to 100% on absolute scale. By allowing annotators to access the shared context across multiple outputs, cESA facilitates more consistent and efficient judgments. We validate cESA using a large-scale human evaluation of English->Japanese translations of 12 models, demonstrating reductions in annotation time and noise compared to standard pointwise evaluation. Unlike existing contrastive ranking methods, cESA yields absolute quality judgments that enable simple, interpretable non-parametric model rankings without the need for post-hoc corrections.
Original Article
View Cached Full Text

Cached at: 07/30/26, 09:58 AM

# Contrastive ESA: Human Evaluation of Multiple Translations at Once
Source: [https://arxiv.org/abs/2607.26640](https://arxiv.org/abs/2607.26640)
[View PDF](https://arxiv.org/pdf/2607.26640)

> Abstract:Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost\. We introduce Contrastive Error Span Annotation \(cESA\), a protocol that presents multiple translations of the source input \(text, video, audio, image\)\. In cESA, the annotator sees multiple translations of the same document, marks major and minor error spans, and then assigns a score from 0% to 100% on absolute scale\. By allowing annotators to access the shared context across multiple outputs, cESA facilitates more consistent and efficient judgments\. We validate cESA using a large\-scale human evaluation of English\-\>Japanese translations of 12 models, demonstrating reductions in annotation time and noise compared to standard pointwise evaluation\. Unlike existing contrastive ranking methods, cESA yields absolute quality judgments that enable simple, interpretable non\-parametric model rankings without the need for post\-hoc corrections\.

## Submission history

From: Vilém Zouhar \[[view email](https://arxiv.org/show-email/5bb1ecb7/2607.26640)\] **\[v1\]**Wed, 29 Jul 2026 09:01:59 UTC \(5,306 KB\)

Similar Articles