cross-lingual

Tag

Cards List
#cross-lingual

Cross-Lingual Exploration for Parametric Knowledge

arXiv cs.CL · 2026-06-24 Cached

This paper explores cross-lingual prompting strategies to improve access to parametric knowledge in large language models, demonstrating significant gains in knowledge transfer and factual recall across 17 languages on multilingual benchmarks.

0 favorites 0 likes
#cross-lingual

MMed-Bench-IR: A Heterogeneous Benchmark for Multilingual Medical Information Retrieval

arXiv cs.CL · 2026-06-24 Cached

MMed-Bench-IR is a heterogeneous benchmark for multilingual medical information retrieval across six languages, evaluating cross-lingual alignment, concept discrimination, and evidence retrieval. It reveals severe performance drops for non-English queries, highlighting gaps in existing English-only evaluations.

0 favorites 0 likes
#cross-lingual

Data Scale, Not Latency, Shapes Cross-Lingual Encoder Transfer in Streaming ASR

arXiv cs.AI · 2026-06-24 Cached

This paper investigates the impact of data scale versus latency on cross-lingual transfer for streaming ASR, finding that multilingual initialization benefits are data-limited, not latency-limited, and diminish as target-language data increases.

0 favorites 0 likes
#cross-lingual

G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment

arXiv cs.CL · 2026-06-18 Cached

G-IdiomAlign is a gloss-pivoted benchmark for evaluating cross-lingual idiom alignment in LLMs, featuring controlled multiple-choice and gloss-contrastive protocols to diagnose literal translation bias and the effect of semantic pivots.

0 favorites 0 likes
#cross-lingual

LLM Parameters for Math Across Languages: Shared or Separate?

arXiv cs.CL · 2026-06-18 Cached

This paper presents a cross-lingual mechanistic analysis of mathematical reasoning in LLMs, finding partial overlap of math-associated parameters across languages, concentrated in intermediate layers. English has the largest set of math-relevant parameters, while lower-resource languages have smaller sets.

0 favorites 0 likes
#cross-lingual

When English Isn't the Best Teacher: Source Language Effects in Cross-Lingual In-Context Learning

arXiv cs.CL · 2026-06-17 Cached

This paper empirically studies cross-lingual transfer in in-context learning across seven tasks, six models, and typologically diverse languages, showing that fine-tuning based expectations do not consistently apply and offering new heuristics for source language selection.

0 favorites 0 likes
#cross-lingual

Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation

arXiv cs.CL · 2026-06-17 Cached

This study evaluates bilingual fine-tuning with language identification tokens for improving ASR in low-resource languages across nine diverse language pairs, finding that high LID accuracy is beneficial and that providing the LID token at inference can boost performance when LID accuracy is low.

0 favorites 0 likes
#cross-lingual

NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama

arXiv cs.CL · 2026-06-17 Cached

This paper introduces NarrativeWorldBench, a benchmark for evaluating long-horizon narrative consistency in audio dramas, and N-VSSM, a latent state-space model that outperforms frontier LLMs across multiple horizons and languages.

0 favorites 0 likes
#cross-lingual

Beyond English: Uncovering the Multilingual Gap in Vision-Language-Action Models

arXiv cs.CL · 2026-06-16 Cached

This paper presents the first systematic study of multilingual instruction following in Vision-Language-Action (VLA) models, revealing significant performance degradation when models trained on English are evaluated on other languages. The authors propose Multilingual Principal Component Alignment (MPCA) to reduce the multilingual performance gap.

0 favorites 0 likes
#cross-lingual

Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus

arXiv cs.CL · 2026-06-16 Cached

Introduces XBCP (Cross-lingual BrowseComp-Plus), a benchmark for evaluating deep research agents and retrievers in cross-lingual and multilingual settings. Results show significant performance degradation when evidence is in a different language from the query, highlighting both retrieval failures and agent-side difficulty in integrating language-mismatched evidence.

0 favorites 0 likes
#cross-lingual

Evaluating and Preserving Lexical Stress in English-to-Chinese Speech-to-Speech Translation

arXiv cs.CL · 2026-06-16 Cached

A research paper proposing a new metric and stress-aware system for evaluating and preserving lexical stress in English-to-Chinese speech-to-speech translation, demonstrating significant improvements over existing approaches while maintaining translation quality.

0 favorites 0 likes
#cross-lingual

Does the Judge Prefer English? Evaluating Language-Switching Invariance in LLM-as-a-Judge

arXiv cs.CL · 2026-06-15 Cached

This paper proposes Judge-LS, a protocol to evaluate whether LLM-as-a-judge models are invariant to language switching between English and Chinese. It finds that switching languages causes 10.7-14.4% preference flips and that judges achieve their highest accuracy in English.

0 favorites 0 likes
#cross-lingual

When Similar Means Different: Evaluating LLMs on Arabic--Hebrew Cognates

arXiv cs.CL · 2026-06-12 Cached

This paper introduces SemCog Bench, a curated benchmark of 1,858 Arabic-Hebrew word pairs with sentence-level annotations, to evaluate LLMs' ability to distinguish true cognates from false friends and loanwords. Results show high accuracy on true cognates but sharp drops on false friends, highlighting a key limitation in cross-lingual semantic reasoning.

0 favorites 0 likes
#cross-lingual

One Jailbreak, Many Tongues: Learning Language-Insensitive Intention Representations for Multilingual Jailbreak Detection

arXiv cs.CL · 2026-06-11 Cached

This paper proposes MLJailDe, a multilingual jailbreak detection framework that uses back-translation data augmentation and relative-distance constraints to improve cross-lingual generalization and robustness, achieving 98.5% F1 score across 11 languages.

0 favorites 0 likes
#cross-lingual

Improving Cross-Lingual Factual Recall via Consistency-Driven Reinforcement Learning

arXiv cs.CL · 2026-06-08 Cached

This paper introduces PolyFact, a large-scale multilingual factual QA dataset, and demonstrates that reinforcement learning via GRPO significantly improves cross-lingual factual consistency in LLMs compared to supervised fine-tuning, by reorganizing multilingual representations.

0 favorites 0 likes
#cross-lingual

Domain-Aware Mispronunciation Detection and Diagnosis Using Language-Specific Statistical Graphs

arXiv cs.CL · 2026-06-05 Cached

提出一种利用语言特定统计图构建的领域感知发音错误检测与诊断方法,在L2-ARCTIC基准上达到59.52%的F1分数,优于多个基线模型。

0 favorites 0 likes
#cross-lingual

DuDi: Dual-Signal Distillation with Cross-Lingual Verbalizer

arXiv cs.CL · 2026-06-04 Cached

DuDi is a dual-signal multilingual distillation framework combining sequence-level and token-level signals with a cross-lingual verbalizer to improve small language models' performance on Southeast Asian languages. Experiments on SEA-HELM show DuDi consistently outperforms competitive distillation baselines across multiple model families and scales.

0 favorites 0 likes
#cross-lingual

How Far Do Auto-Interpretation Labels Generalize: A Controlled Study Across Languages, Scripts, and Rewordings

arXiv cs.CL · 2026-06-02 Cached

This paper investigates whether auto-generated labels for sparse autoencoder features generalize across languages and scripts, using Serbian digraphia as a controlled testbed. It finds that while feature sets show substantial overlap across languages, the labels often fail to track the same concept in non-English inputs, particularly in less represented scripts.

0 favorites 0 likes
#cross-lingual

Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit

arXiv cs.CL · 2026-06-01 Cached

This paper audits six large language models for gender stereotyping across English, Korean, Chinese, and Japanese, anchoring against human baselines. It finds that LLM stereotyping often exceeds human cross-country variation and can compound across languages, introducing a four-pattern framework to characterize such behaviors.

0 favorites 0 likes
#cross-lingual

XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks

arXiv cs.CL · 2026-06-01 Cached

XLGoBench introduces a synthetic benchmark of algorithmic tasks to detect cross-lingual skill gaps in LLMs, demonstrating persistent gaps across multiple state-of-the-art models.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback