Tag
This paper audits a multilingual affective generation benchmark and reveals that system rankings are driven by measurement artifacts rather than genuine performance differences, with annotator variability playing a key role.
This paper presents COILD, an Indic-centric parallel corpus with over 1.16 million sentence pairs for 20 Indian language pairs and a domain-centric benchmark, demonstrating improvements in machine translation when fine-tuning multilingual models.
This paper presents a bilingual evaluation suite for cross-lingual legal question answering in Vietnamese labour law, assessing retrieval, translation, and verifier-guided correction to enhance answer faithfulness and address unsupported claims.
AfriSyCo studies answer switching in AI models around African-language factual content, showing how assertive framing and verification prompts significantly affect accuracy across languages and model checkpoints.
The paper proposes a novel scope-conditioned generation framework that integrates structured stereotype characteristics into Large Language Model prompts for effective multilingual counterspeech, validated on a human-curated dataset with significant improvements in factuality and effectiveness.
This paper investigates whether transformer models internalize the same linguistic features as traditional models for multilingual readability assessment, using SHAP and TCAV across five languages.
The paper proposes SALT, a lightweight post-training method that injects span-level supervision into cross-lingual sentence encoders to improve token representations, achieving top results on multilingual token-level benchmarks and enhancing sentence-level performance.
This paper analyzes how LLMs' faithfulness to provided context depends on perceived plausibility, using factual, counterfactual, and fictional RDF triples in multiple languages. It finds a weak context–memory conflict and emphasizes that the choice of LLM judge can overestimate its strength.
The paper introduces VakyArth, the first pragmatic benchmark for Indic languages, evaluating LLMs on cultural and contextual reasoning across Hindi, Punjabi, Tamil, and Malayalam. It reveals consistent failures in models on pragmatic meanings, with systematic differences across languages and tasks.
This paper introduces Centroid Intervention Fusion (CIF), a projection fusion framework that unifies multilingual intervention operators to enhance cross-lingual representation learning in large language models, achieving performance gains especially for low-resource languages.
The paper investigates how prompt and response languages affect LLM content generation, finding that prompt language significantly influences output length while maintaining semantic fidelity through conceptual paraphrasing.
This paper audits extractive prompt compressors across ten languages, revealing that English-trained models exhibit significant performance gaps on non-English text at high compression rates and proposes a translate-then-compress pipeline as a more effective alternative.
The paper investigates whether large language models can phonetically decode encoded languages, such as German written in Cyrillic characters, to assess their abstraction capabilities and performance beyond standard Latin-based training data.
Cross-lingual Ranking Preference Optimization (CRPO) is a novel framework that enhances multilingual LLM alignment by transferring English preference knowledge to target languages through hierarchical ranking optimization, demonstrating improved performance in instruction-following and knowledge utilization across multiple languages.
The paper describes a submission to the WMT 2026 MIST shared task, using the Tiny Aya Global model with task-specialized QLoRA adapters for multilingual summarization and question answering.
This paper presents a trilingual topic modeling framework for analyzing Sri Lankan parliamentary debates, using LLM-based text extraction and multilingual embeddings to handle code-mixed text and achieve superior cluster purity over traditional methods.
This paper investigates the tokens learned when tokenization is optimized jointly with language modeling, comparing tokenizer-free methods across multiple languages and finding that they produce distinct, efficient vocabularies for NLP.
This paper proves that there is no theoretical curse of multilinguality for embedding space structure, showing that the minimum dimensionality required grows only logarithmically with the number of languages, suggesting empirical issues stem from data and training conditions.
This paper defines and quantifies sentiment drift in RLHF-trained summarization models, proposes a Policy Attribution framework to identify causes, and introduces a Sentiment-Aware KL Regularization method to reduce drift.
A PhD proposal outlining a unified end-to-end framework for multilingual metaphor processing, integrating metaphor detection, translation evaluation, and joint modeling using linguistic theory and large language models.