machine-translation

Tag

Cards List
#machine-translation

Conversational Domain Adaptation of IndicTrans2 across 21 Indic Languages via Experience Replay and Model Soups

arXiv cs.CL · 2026-06-30 Cached

This paper adapts IndicTrans2-1B to conversational register across 21 Indic languages using experience replay and model souping, achieving conversational gains without sacrificing general-domain performance, though human evaluation shows the metric-based gains may not reflect perceived quality improvements.

0 favorites 0 likes
#machine-translation

NagaTranslate: Building a translation and voice pipeline for low-resource Nagaland creoles (Whisper, VITS, LLMs) [P]

Reddit r/MachineLearning · 2026-06-28

Introduces NagaTranslate, a pipeline for translation and voice synthesis for low-resource Nagaland creoles using Whisper, VITS, and LLMs.

0 favorites 0 likes
#machine-translation

Neural Machine Translation for Low-Resource Tangkhul--English

arXiv cs.CL · 2026-06-25 Cached

Presents a neural machine translation system for the severely under-resourced Tangkhul–English language pair, achieving strong BLEU, chrF++, BERTScore, and COMET scores using fine-tuned ByT5-large and mT5-small models.

0 favorites 0 likes
#machine-translation

AI translation of literary texts is "fine", but readers still prefer human translations

Hugging Face Daily Papers · 2026-06-24 Cached

A study comparing human and AI translations of literary works shows that while machine translations are deemed 'fine', readers still prefer human translations for their immersiveness and clarity. Automatic metrics fail to capture reader preferences.

0 favorites 0 likes
#machine-translation

Translating the Untranslatable: An Operationalizable Ontology for Untranslatability

arXiv cs.CL · 2026-06-17 Cached

This paper introduces a structured ontology for untranslatability in machine translation, along with a taxonomy of compensation strategies and a multilingual dataset. Human preference studies show translator quality depends on the strategy used, with a preference for explanatory translations.

0 favorites 0 likes
#machine-translation

Speaking in Self-Assessing Tongues: On the Verbalized Confidence of LLMs in Machine Translation

arXiv cs.CL · 2026-06-17 Cached

This paper investigates verbalized methods for extracting LLM confidence in machine translation outputs, comparing them with internal token probabilities. The study finds that while both approaches perform similarly in error detection and calibration, there is little correlation between internal and verbalized confidence measures.

0 favorites 0 likes
#machine-translation

Emotion Profiling in LLM-Based Literary Translation: Systematic Shifts Across MT and Post-Editing

arXiv cs.CL · 2026-06-10 Cached

This paper investigates whether LLM translations exhibit identifiable emotional profiles and how post-editing reshapes them toward human-like norms, using a comparative study of Margaret Atwood's 'Oryx and Crake' translated to Italian.

0 favorites 0 likes
#machine-translation

mmPISA-bench: Do LLMs Reason Equally Well Across 43 Languages?

arXiv cs.CL · 2026-06-08 Cached

Introduces mmPISA-bench, a compact multilingual reasoning benchmark derived from PISA, and evaluates proprietary LLMs across 43 languages, finding that they reason effectively with some performance variations, and that machine-translated questions do not degrade accuracy.

0 favorites 0 likes
#machine-translation

Multilingual Coreference Resolution via Cycle-Consistent Machine Translation

arXiv cs.CL · 2026-06-05 Cached

This paper proposes a novel pipeline for multilingual coreference resolution that uses cycle-consistent machine translation from English to low-resource languages to generate training data, validated by back-translation and BERT similarity. Experiments on four low-resource languages show significant performance gains, enabling accurate coreference resolution where no prior corpora existed.

0 favorites 0 likes
#machine-translation

ComplexityMT: Benchmarking the Interaction Between Text Complexity and Machine Translation

arXiv cs.CL · 2026-06-05 Cached

Introduces ComplexityMT, a benchmark for evaluating the interaction between text complexity and machine translation across six languages using CEFR levels, showing that higher complexity makes translation harder and that MT shifts complexity levels.

0 favorites 0 likes
#machine-translation

Beyond "To whom it may concern": Tailoring Machine Translation to Audience and Intent

arXiv cs.CL · 2026-06-03 Cached

This paper systematically evaluates purpose-driven machine translation using LLMs across 50 languages, finding that explicit instructions significantly improve adaptation quality, particularly for informal domains and larger models, while traditional metrics fail to capture adaptation quality.

0 favorites 0 likes
#machine-translation

G^2C-MT: Graph-Guided Context Selection for Document-Level Machine Translation

arXiv cs.CL · 2026-06-03 Cached

Proposes G²C-MT, a graph-guided context selection framework for document-level machine translation that models structured discourse dependencies via a lightweight discourse graph and depth-biased random walk, outperforming baselines on multiple LLMs.

0 favorites 0 likes
#machine-translation

Translating Classical Poetry into Modern Prose

arXiv cs.CL · 2026-06-03 Cached

Introduces Padyam2Gadyam, a dataset for translating 13th-17th century Telugu classical poetry into modern Telugu and English prose, and evaluates five LLMs on the task.

0 favorites 0 likes
#machine-translation

KletterMix: Climbing Toward High-Quality German Pretraining Data

Hugging Face Daily Papers · 2026-06-02

KletterMix is a high-quality German pretraining corpus built by translating a state-of-the-art English pretraining dataset into German while preserving structure and diversity. Controlled experiments show models trained on KletterMix achieve measurable improvements on German-language benchmarks.

0 favorites 0 likes
#machine-translation

Model-Based Quality Assessment for Massively Multilingual Parallel Data

Hugging Face Daily Papers · 2026-05-29 Cached

This paper proposes a model-based approach to assess massively multilingual parallel data by decomposing it into parallelism assessment and reference-free quality estimation, finding that no single universal metric works across all language directions.

0 favorites 0 likes
#machine-translation

ROC Analysis for Evaluating Translation Quality Estimation Systems

arXiv cs.CL · 2026-05-26 Cached

This paper proposes using Receiver Operating Characteristic (ROC) analysis to evaluate translation quality estimation (QE) systems, demonstrating that it offers actionable performance insights for business decision-making and is consistent with current methods.

0 favorites 0 likes
#machine-translation

Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection

arXiv cs.CL · 2026-05-25 Cached

This paper introduces a principled approach to multilingual language steering using sparse autoencoders (SAEs) trained on multilingual data and a novel layer selection rule based on the intersection of multilingual alignment and language separability, evaluated on LLaMA-3.1-8B and Gemma-2-9B for machine translation and cross-lingual summarization.

0 favorites 0 likes
#machine-translation

Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild

arXiv cs.CL · 2026-05-22 Cached

Hy-MT2 is a family of fast, efficient multilingual translation models from Tencent, available in 1.8B, 7B, and 30B-A3B sizes, supporting 33 languages and outperforming previous open-source and commercial models.

0 favorites 0 likes
#machine-translation

@FeitengLi: Hy-MT2 - a new open-source multilingual translation model that matches top-tier large models in capability, supports translation between 33 languages, and offers flexible instruction capabilities. It achieves 2-bit quantization under 500MB, making it well-suited for on-device deployment. https://modelsc…

X AI KOLs Timeline · 2026-05-21 Cached

Hy-MT2 is a new open-source multilingual translation model from Tencent Hy that supports 33 languages, offers flexible instruction capabilities, and achieves 2-bit quantization under 500MB for on-device deployment.

1 favorites 1 likes
#machine-translation

@AdinaYakup: Hy-MT2 New translation model family from @TencentHunyuan 1.8B / 7B / 30B-A3B MoE Supports 33 languages 1.8B > 440MB wit…

X AI KOLs Following · 2026-05-21 Cached

Tencent Hunyuan released Hy-MT2, a family of translation models up to 30B parameters with MoE, supporting 33 languages and quantized for on-device use.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback