Contrastive ESA: Human Evaluation of Multiple Translations at Once
Summary
Introduces Contrastive Error Span Annotation (cESA), a protocol for human evaluation of multiple translations simultaneously, reducing annotation time and noise compared to standard methods.
View Cached Full Text
Cached at: 07/30/26, 09:58 AM
# Contrastive ESA: Human Evaluation of Multiple Translations at Once Source: [https://arxiv.org/abs/2607.26640](https://arxiv.org/abs/2607.26640) [View PDF](https://arxiv.org/pdf/2607.26640) > Abstract:Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost\. We introduce Contrastive Error Span Annotation \(cESA\), a protocol that presents multiple translations of the source input \(text, video, audio, image\)\. In cESA, the annotator sees multiple translations of the same document, marks major and minor error spans, and then assigns a score from 0% to 100% on absolute scale\. By allowing annotators to access the shared context across multiple outputs, cESA facilitates more consistent and efficient judgments\. We validate cESA using a large\-scale human evaluation of English\-\>Japanese translations of 12 models, demonstrating reductions in annotation time and noise compared to standard pointwise evaluation\. Unlike existing contrastive ranking methods, cESA yields absolute quality judgments that enable simple, interpretable non\-parametric model rankings without the need for post\-hoc corrections\. ## Submission history From: Vilém Zouhar \[[view email](https://arxiv.org/show-email/5bb1ecb7/2607.26640)\] **\[v1\]**Wed, 29 Jul 2026 09:01:59 UTC \(5,306 KB\)
Similar Articles
Evaluating Multilingual Sentence Embeddings for Translation Error Detection:An English--Greek Contrastive Study
This study evaluates multilingual sentence embeddings for distinguishing correct English–Greek translations from erroneous ones, finding that embeddings provide useful semantic signals but are better integrated into broader translation evaluation frameworks.
When to Use Extra Context: Evidence-Grounded Terminology Adaptation for Simultaneous Speech Translation
This paper proposes EGTA, an Evidence-Grounded Terminology Adaptation framework for simultaneous speech translation that selectively uses document-specific terminology to improve translation of rare terms, achieving significant gains in named-entity and acronym recall without full-model fine-tuning.
A Practical Evaluation Method for Long-Form Simultaneous Speech-to-Speech Translation
This paper proposes a practical evaluation method for long-form simultaneous speech-to-speech translation that uses ASR, forced alignment, and sentence embedding alignment to compute latency and quality metrics on continuous speech, overcoming limitations of prior approaches.
SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition
SESSE is a training-free framework that decomposes holistic LLM-as-judge evaluations into structured sub-questions, enabling better interpretability and diagnosis of label ambiguity while achieving competitive performance with fine-tuned models.
DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
DocOCR-Eval proposes an annotation-free framework that uses a correction and ranking strategy to evaluate and select OCR tools without ground truth labels, showing that aggregating multiple multimodal large language models improves alignment with human rankings.