Tag
This paper adapts IndicTrans2-1B to conversational register across 21 Indic languages using experience replay and model souping, achieving conversational gains without sacrificing general-domain performance, though human evaluation shows the metric-based gains may not reflect perceived quality improvements.
Introduces NagaTranslate, a pipeline for translation and voice synthesis for low-resource Nagaland creoles using Whisper, VITS, and LLMs.
Presents a neural machine translation system for the severely under-resourced Tangkhul–English language pair, achieving strong BLEU, chrF++, BERTScore, and COMET scores using fine-tuned ByT5-large and mT5-small models.
A study comparing human and AI translations of literary works shows that while machine translations are deemed 'fine', readers still prefer human translations for their immersiveness and clarity. Automatic metrics fail to capture reader preferences.
This paper introduces a structured ontology for untranslatability in machine translation, along with a taxonomy of compensation strategies and a multilingual dataset. Human preference studies show translator quality depends on the strategy used, with a preference for explanatory translations.
This paper investigates verbalized methods for extracting LLM confidence in machine translation outputs, comparing them with internal token probabilities. The study finds that while both approaches perform similarly in error detection and calibration, there is little correlation between internal and verbalized confidence measures.
This paper investigates whether LLM translations exhibit identifiable emotional profiles and how post-editing reshapes them toward human-like norms, using a comparative study of Margaret Atwood's 'Oryx and Crake' translated to Italian.
Introduces mmPISA-bench, a compact multilingual reasoning benchmark derived from PISA, and evaluates proprietary LLMs across 43 languages, finding that they reason effectively with some performance variations, and that machine-translated questions do not degrade accuracy.
This paper proposes a novel pipeline for multilingual coreference resolution that uses cycle-consistent machine translation from English to low-resource languages to generate training data, validated by back-translation and BERT similarity. Experiments on four low-resource languages show significant performance gains, enabling accurate coreference resolution where no prior corpora existed.
Introduces ComplexityMT, a benchmark for evaluating the interaction between text complexity and machine translation across six languages using CEFR levels, showing that higher complexity makes translation harder and that MT shifts complexity levels.
This paper systematically evaluates purpose-driven machine translation using LLMs across 50 languages, finding that explicit instructions significantly improve adaptation quality, particularly for informal domains and larger models, while traditional metrics fail to capture adaptation quality.
Proposes G²C-MT, a graph-guided context selection framework for document-level machine translation that models structured discourse dependencies via a lightweight discourse graph and depth-biased random walk, outperforming baselines on multiple LLMs.
Introduces Padyam2Gadyam, a dataset for translating 13th-17th century Telugu classical poetry into modern Telugu and English prose, and evaluates five LLMs on the task.
KletterMix is a high-quality German pretraining corpus built by translating a state-of-the-art English pretraining dataset into German while preserving structure and diversity. Controlled experiments show models trained on KletterMix achieve measurable improvements on German-language benchmarks.
This paper proposes a model-based approach to assess massively multilingual parallel data by decomposing it into parallelism assessment and reference-free quality estimation, finding that no single universal metric works across all language directions.
This paper proposes using Receiver Operating Characteristic (ROC) analysis to evaluate translation quality estimation (QE) systems, demonstrating that it offers actionable performance insights for business decision-making and is consistent with current methods.
This paper introduces a principled approach to multilingual language steering using sparse autoencoders (SAEs) trained on multilingual data and a novel layer selection rule based on the intersection of multilingual alignment and language separability, evaluated on LLaMA-3.1-8B and Gemma-2-9B for machine translation and cross-lingual summarization.
Hy-MT2 is a family of fast, efficient multilingual translation models from Tencent, available in 1.8B, 7B, and 30B-A3B sizes, supporting 33 languages and outperforming previous open-source and commercial models.
Hy-MT2 is a new open-source multilingual translation model from Tencent Hy that supports 33 languages, offers flexible instruction capabilities, and achieves 2-bit quantization under 500MB for on-device deployment.
Tencent Hunyuan released Hy-MT2, a family of translation models up to 30B parameters with MoE, supporting 33 languages and quantized for on-device use.