machine-translation

Tag

Cards List
#machine-translation

Last Translation Benchmark

Hugging Face Daily Papers · 5d ago Cached

The Last Translation Benchmark introduces a live dataset of peer-reviewed, multimodal examples designed to evaluate and break leading machine translation models, with handcrafted verification rules for reliable assessment. It addresses the saturation of current benchmarks and the unreliability of automatic metrics.

0 favorites 0 likes
#machine-translation

The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space

arXiv cs.CL · 5d ago Cached

The paper proposes the interlingua hypothesis, suggesting that large language models perform translation by encoding source text into a latent task-agnostic feature space and decoding from it, supported by empirical evidence on variance, causal influence, and monolingual fine-tuning.

0 favorites 0 likes
#machine-translation

Evaluating Multilingual Sentence Embeddings for Translation Error Detection:An English--Greek Contrastive Study

arXiv cs.CL · 6d ago Cached

This study evaluates multilingual sentence embeddings for distinguishing correct English–Greek translations from erroneous ones, finding that embeddings provide useful semantic signals but are better integrated into broader translation evaluation frameworks.

0 favorites 0 likes
#machine-translation

Length-Adaptive Decoding for Masked Diffusion Machine Translation

Hugging Face Daily Papers · 2026-08-23 Cached

This paper introduces Entropy-Valley, a training-free length selector for masked diffusion machine translation that uses predictive entropy to improve adequacy, showing that length choice matters more than unmasking order.

0 favorites 0 likes
#machine-translation

TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

arXiv cs.CL · 2026-08-20 Cached

The paper introduces TranslatePsy-AfriSLM, an open-source collection of machine translation resources for 19 Sub-Saharan African languages, demonstrating that fine-tuned small language models with filtered synthetic data outperform much larger models like TranslateGemma-27B and Qwen3.5-122B-A10B.

0 favorites 0 likes
#machine-translation

SuTRA : Structurally-Unified Tokenization with Root Awareness

arXiv cs.CL · 2026-08-20 Cached

SuTRA is a morphology-aware tokenization algorithm that preserves akshara indivisibility for Indic languages, reducing morphological shattering and achieving improvements in machine translation metrics over standard BPE methods.

0 favorites 0 likes
#machine-translation

A Pilot Study of Autocompleting Tokenizers

arXiv cs.CL · 2026-08-18 Cached

This paper proposes a compression scheme for byte-level tokenization using an autocomplete model to remove predictable bytes from input sequences, reducing sequence length while maintaining machine translation performance across diverse languages.

0 favorites 0 likes
#machine-translation

Poly-Dialectal Neural Machine Translation System for Bangla Regional Dialects

arXiv cs.CL · 2026-08-13 Cached

This paper presents a unified poly-dialectal neural machine translation system for 12 Bangla regional dialects, introducing the largest multi-dialect parallel corpus to date and achieving state-of-the-art BLEU scores with a fine-tuned BanglaT5 model using DoRA.

0 favorites 0 likes
#machine-translation

When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation

arXiv cs.CL · 2026-08-13 Cached

This paper measures the self-referential evaluation loop when a knowledge base is used as the gold standard for entity-level machine translation in low-resource historical domains, showing that gains from KB injection are confined to overlapping segments and do not reflect true translation quality.

0 favorites 0 likes
#machine-translation

OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents

arXiv cs.CL · 2026-08-11 Cached

This paper introduces OmnilingualGAIA2, a multilingual expansion of the GAIA2 agentic benchmark across ten languages, revealing a universal cross-lingual performance gap of 8.8–18.4 pass@3 points that is model-driven and persists with scale. The authors argue that multilingual agentic evaluation should become standard for globally deployed agents.

0 favorites 0 likes
#machine-translation

Mitigating Gender Bias in English to Romanian Machine Translation

arXiv cs.CL · 2026-08-11 Cached

This paper proposes a hybrid pipeline combining fine-tuned LLaMA-based gender classification with tag-aware neural machine translation to mitigate gender bias in English-to-Romanian MT, introducing new datasets and improving gender accuracy by over 40 points on benchmarks.

0 favorites 0 likes
#machine-translation

Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?

arXiv cs.CL · 2026-08-11 Cached

This paper investigates whether automatic evaluation metrics for machine translation are reliable for Classical Chinese to English translation, using a diagnostic framework based on minimal pairs. It finds all metrics have blind spots, with MetricX24 performing best overall.

0 favorites 0 likes
#machine-translation

APEX-VW: A Document-Level English-Spanish Post-Editing Dataset in the Healthcare Domain

arXiv cs.CL · 2026-08-11 Cached

This paper introduces APEX-VW, a new document-level English-Spanish post-editing dataset built from NHS virtual-ward documents and professional post-editing in Trados Studio. The corpus is designed to support research on terminology normalisation, correction propagation, and human-in-the-loop translation support.

0 favorites 0 likes
#machine-translation

Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation

arXiv cs.CL · 2026-08-11 Cached

This paper presents an embedding initialization method for adding unseen low-resource languages to multilingual NMT models, evaluated on Limbum-English translation. The averaged multi-language embedding matches the best single proxy and drastically outperforms zero-shot and from-scratch baselines.

0 favorites 0 likes
#machine-translation

Analysis of Numerical Localisation in LLM Translations

arXiv cs.CL · 2026-08-07 Cached

This paper analyses the capability of five large language models to localise times, numbers, and dates when translating between English and German, and tests strategies to improve accuracy—finding that embedding localisation principles into the prompt context yields statistically significant improvements.

0 favorites 0 likes
#machine-translation

Pun Intended: Multi-Agent Translation of Wordplay with Contrastive Learning and Phonetic-Semantic Embeddings

arXiv cs.CL · 2026-08-06 Cached

This paper explores three LLM-based approaches for translating puns from English to French, combining contrastive learning and phonetic-semantic embeddings. Their multi-agent and guided chain-of-thought systems ranked first and second in the CLEF JOKER 2025 Task 2 competition under expert human evaluation.

0 favorites 0 likes
#machine-translation

Towards End-to-End Multilingual Metaphor Processing: Integrating Detection, Translation, and Evaluation

arXiv cs.CL · 2026-08-06 Cached

A PhD proposal outlining a unified end-to-end framework for multilingual metaphor processing, integrating metaphor detection, translation evaluation, and joint modeling using linguistic theory and large language models.

0 favorites 0 likes
#machine-translation

Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment $\unicode{x2013}$ Is English Enough?

arXiv cs.CL · 2026-08-05 Cached

The paper compares 27 cross-lingual alignment (CLA) score variants for predicting LLM performance on multilingual classification and translation tasks, and proposes a PMI-based translation metric. It finds that CLA with English predicts translation quality comparably to or better than source-target CLA, supporting the view that LLMs use English as an internal pivot language.

0 favorites 0 likes
#machine-translation

PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation

arXiv cs.CL · 2026-08-05 Cached

This paper proposes PAMT, a process-aligned reinforcement learning framework for multi-domain machine translation that combines domain-aware long chain-of-thought supervision with step-level process rewards to improve domain-sensitive translation decisions.

0 favorites 0 likes
#machine-translation

TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation

arXiv cs.CL · 2026-08-05 Cached

Introduces TQLite, a distillation framework that uses a multi-LRM jury to train small language models for real-time MQM-based translation quality evaluation, achieving performance far exceeding off-the-shelf SLMs while remaining cost-effective.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback