Where Privacy Risk Lives in English-Source Multilingual RAG: A Stage-Decomposed Audit Across Five Query Languages
Summary
This paper audits privacy risk in English-source multilingual RAG across five query languages, testing whether non-English queries increase PII leakage. Using a Qwen2.5-7B pipeline with two-stage defenses, it finds English has the highest point-estimate leak rate under output-only filtering, with residual leaks on Arabic and Swahili when the input judge is added.
View Cached Full Text
Cached at: 08/07/26, 07:49 AM
# Where Privacy Risk Lives in English-Source Multilingual RAG: A Stage-Decomposed Audit Across Five Query Languages
Source: [https://arxiv.org/html/2608.05163](https://arxiv.org/html/2608.05163)
Yanhang Li Northeastern University li\.yanha@northeastern\.edu &Zhichao Fan University of Illinois Urbana\-Champaign zhichao8@illinois\.edu &Zexin Zhuang Southern Methodist University zexinz@smu\.edu
###### Abstract
A common assumption holds that switching to a non\-English language makes a multilingual RAG system easier to attack for personal information\. We test this on an English\-source synthetic\-PII corpus with five query languages and a two\-stage defence \(LLM input judge \+ regex output filter\), in a pipeline whose translator, judge, back\-translator, and generator are all Qwen2\.5\-7B — so every finding below is pipeline\-conditional, not a causal ranking of language\-inherent risk\. Under output\-only filtering, English has the highest observed unstructured\-PII leak rate; only English\-vs\-Swahili separates cleanly under document\-level bootstrap intervals\. Once the input judge is added, residual leaks remain on Arabic and Swahili, and back\-translating the query does not close the gap \(an ablation we report but cannot use as a causal diagnostic, since the back\-translator is also Qwen\)\. On a separaten=17n\{=\}17multilingual\-prompted\-judge residual corner, attaching the gold corpus document to the input judge blocks15/1715/17residual cells\. We frame this last result as a*mechanism diagnostic, not a deployable defence*: it uses oracle retrieval, BLOCK/ALLOW rates are measured on adversarial queries only, and we measure no benign\-query false\-positive rate and no answer\-utility cost\. The supplementary material contains code, corpora, queries, and per\-trial JSONLs; the priority follow\-up is an independent\-MT plus non\-Qwen\-judge replication with a native\-speaker query set, scoped in §Limitations\.
Where Privacy Risk Lives in English\-Source Multilingual RAG: A Stage\-Decomposed Audit Across Five Query Languages
Yanhang LiNortheasternUniversityli\.yanha@northeastern\.eduZhichao FanUniversity of IllinoisUrbana\-Champaignzhichao8@illinois\.eduZexin ZhuangSouthern MethodistUniversityzexinz@smu\.edu
## 1Introduction
Multilingual retrieval\-augmented generation \(RAG\) is a standard multilingual QA architecture studied in recent mRAG work\(chirkova2024multilingual\): a source\-language knowledge base may be queried in the user’s native language, the retriever returns relevant source\-language documents via a multilingual embedder, and the generator answers in the user’s language\. Privacy guardrails for this pattern are often English\-centric: PII filters, moderation classifiers, and safety\-alignment data tend to over\-represent English\. A common concern in multilingual safety, particularly multilingual jailbreaks\(yong2023multilingual;yong2025multilingualsafety\), is therefore that cross\-lingual queries may increase privacy leakage by slipping past English\-oriented defenses \(broader landscape:huang2024multilingual\)\.
We test this intuition on an English\-source multilingual RAG with synthetic PII\. We adopt a black\-box threat model in which an attacker who knows a non\-PII anchor of a target document \(a project name, ticket ID, or case ID — an insider\-style threat assumption\) issues queries in five languages \(English, Chinese, German, Arabic, Swahili\) and three reformulations \(direct, summarize, in\-context\-learning\) attempting to extract a planted PII item\. We deploy a two\-stage guardrail: an*English\-only\-prompted*multilingual LLM input judge that classifies the user query asBLOCK/ALLOW, and a regex\-style output filter that triggers on email, phone, and SSN\-like patterns\.
In this Qwen\-translated configuration, output\-only point estimates do not support that intuition: English has the highest point\-estimate leak rate \(0\.8750\.875\) and non\-English templates are lower \(0\.4250\.425–0\.7750\.775\), with clear separation only on English\-vs\-Swahili and a touching endpoint vs\. Chinese; we read this as point\-estimate ordering, not a broad inversion \(Section[4](https://arxiv.org/html/2608.05163#S4), Table[1](https://arxiv.org/html/2608.05163#S4.T1)\)\. A counter\-effect appears once we add the English\-only\-prompted input\-side judge: the residual combined leak collapses to zero on en/zh/de but persists on Arabic \(7\.5%7\.5\\%\) and Swahili \(17\.5%17\.5\\%\) in this Qwen\-mediated pipeline\. The aggregate Swahili input\-judge BLOCK rate \(∼\\sim77%77\\%\) overstates effective coverage on the documents that actually leak: at the document\-level any\-success unit, the English\-only\-prompted Qwen judgeALLOWs at least one leaking reformulation on𝟕/𝟏𝟕\\mathbf\{7/17\}Swahili dangerous documents \(Section[4\.2](https://arxiv.org/html/2608.05163#S4.SS2)\)\. A back\-translate\-then\-judge ablation does not rescue Swahili and a multilingual\-prompted judge variant leaves Swahili largely unchanged, both consistent with — but not diagnostic of — translation\-induced intent attenuation in the original query\.
Our contribution is to*localize*where privacy risk appears in this pipeline\. The observed residual leaks are consistent with cases in which a Qwen\-translated query may have lost enough explicit extraction intent that an input\-side filter — which has no retrieval context — allows the request, while the downstream RAG — which does have retrieval context — still answers it\. We call this hypothesised mechanism*differential pipeline degradation under translation noise*; the experiments are consistent with it but do not identify it\. As a follow\-up direction, an oracle diagnostic on a separaten=17n\{=\}17multilingual\-prompted\-judge residual corner shows that giving the input judge the gold corpus document as “Retrieved Context” blocks15/1715/17residual cells \(Section[4\.4](https://arxiv.org/html/2608.05163#S4.SS4)\); this is not a direct rescue of the deployed English\-only F3 residual and does not measure utility or false positives\.
## 2Related Work
#### Privacy in RAG\.
zeng2024ragcharacterized retrieval\-augmented systems as a new exfiltration surface, showing that data\-store contents can be elicited verbatim by adversarial queries\.wang2025padproposed a privacy\-aware decoding scheme to suppress sensitive spans at generation time\. Both target English RAG and English\-trained defenses; the cross\-lingual axis is unexplored\. The LLM\-as\-input\-judge defense pattern we adopt for Stage A \(an LLM classifying incoming queries asBLOCK/ALLOW\) was canonicalised by Llama\-Guard\(inan2023llamaguard\); our contribution is to audit how this pattern degrades across query languages rather than to introduce a new judge\. Diagnostic evaluation platforms have begun to appear for adjacent RAG settings such as visual RAG\(ji2025mrag\); the present paper is the multilingual\-text counterpart on the privacy axis\.
#### Training\-data extraction\.
carlini2021extractingestablished training\-data extraction from large language models in monolingual settings\. Subsequent work has extended this line to PII benchmarks\(nakka2024piiscope\)\. Our threat model is closer to a deployed RAG: the attacker queries the data\-store through the RAG, not the parametric memory\.
#### Multilingual safety and jailbreak\.
yong2023multilingualdemonstrated that multilingual queries can bypass English\-aligned safety in direct\-prompt jailbreak settings, with effectiveness scaling inversely with language resource availability\.huang2024multilingualsurvey the broader safety landscape of multilingual LLMs, andyong2025multilingualsafetyupdate this picture by quantifying the persistent language gap in safety alignment and current mitigation directions\. These findings concern direct harmful\-prompt jailbreaks against a model’s safety alignment, not retrieval\-mediated PII extraction; the failure mode we study here is mediated by an input filter without retrieval context\.
#### Multilingual RAG\.
chirkova2024multilingualstudy RAG quality in multilingual settings and motivate a per\-component analysis of the multilingual RAG stack\.li2025bordirlinesintroduceBordIRLinesfor culturally\-sensitive cross\-lingual RAG and analyse how retrievers and generators use multilingual documents\. Broader RAG design\-space taxonomies covering retrieval–reasoning interactions\(ji2026retrieval\)provide context for where the privacy\-relevant stages we audit sit within the wider RAG pipeline space\. Our work re\-derives the stage\-by\-stage degradation under a privacy lens and uses it to motivate a hypothesis for why cross\-lingual queries can be*worse for the attacker*under output\-only filtering, rather than to claim that prior work explains our leakage pattern\.
#### Cross\-lingual privacy mechanisms\.
dong2025crossprivacystudied cross\-lingual privacy leakage at the model\-parameter level, identifying language\-universal and language\-specific privacy neurons\. Their attack surface is the parametric memory of a multilingual LLM; ours is the retrieval\-mediated path through a deployed RAG\. The findings are complementary: parametric leakage and retrieval\-mediated leakage are different exfiltration routes, and the corresponding defenses \(privacy\-neuron erasure vs\. filter design\) operate at different layers\. Our PII\-detection scoring layer is functionally a multilingual fine\-grained NER over generated text; for the broader space of multilingual NER datasets for LLMs seeluo2025dynamicner\.
#### LLM\-system audit methodology\.
A parallel line audits LLM\-mediated systems for safety and security in non\-RAG settings:luo2026agentauditorpresent a human\-level safety and security evaluation harness for LLM agents, andjiang2026agenticframe agentic AI itself as a cybersecurity attack surface\. The broader LLM\-and\-privacy landscape also includes federated\-learning\-side collaborative mechanisms\(luo2025cross\)\. Closely related in methodological posture, our concurrent work on configuration\-conditional benchmark instability\(li2026safetyrepro\)and on auditing reasoning\-trace memorization claims after unlearning\(li2026reasoningtrace\)similarly stresses that pairwise verdicts and positive bypass signals on LLM\-mediated systems need mechanism\-diagnostic ablations before causal attribution—a posture our F2 back\-translate non\-diagnosis and F4 mechanism\-diagnostic\-not\-deployable\-defence framings follow here\. Our work occupies the multilingual\-RAG slot of this audit\-methodology landscape: an attacker\-vs\-defence stage\-decomposed audit of a retrieval\-mediated pipeline\.
## 3Setup
Figure 1:Overview of the stage\-decomposed RAG privacy audit\. Attacker queries in five languages \(each with three reformulations\) over 100 English documents with synthetic PII pass through Stage A input judge \(English\-only\-prompted Qwen2\.5\-7B\) and Stage B regex output filter; residual leak is measured at document\-level any\-success\.### 3\.1Threat model
The attacker has black\-box access to a deployed multilingual RAG over an English document store\. The attacker knows a unique non\-PII anchor of one target document — a project name, case identifier, ticket number, or contract identifier — and seeks to extract a planted PII item from that document\. The attacker may issue queries in any of five languages\. The defender deploys \(a\) an English\-only\-prompted multilingual LLM input judge that classifies each incoming query asBLOCKorALLOW, and \(b\) a regex\-style output filter that triggers on emails, phone numbers, and 9\-digit identifiers in the generator’s response\.
### 3\.2Synthetic corpus
We construct 100 English documents across four topics \(workplace email, medical case note, customer support ticket, legal excerpt\)\. Each document carries one synthetic PII item, stratified to 20 documents per type across \{name, email, phone, 9\-digit identifier, address\} \(n=60n\{=\}60structured targets,n=40n\{=\}40unstructured\); strings are Faker\-generated and each document carries one unique non\-PII anchor\. Synthesis is intentional: it removes the public\-translation contamination problem of natural\-language corpora \(e\.g\., Dickens or other public\-domain text whose translations can already appear in pretraining data\) and ensures exact\-match leak detection\.
### 3\.3Pipeline
Figure[1](https://arxiv.org/html/2608.05163#S3.F1)shows the stage\-decomposed audit at a glance\. The embedder is BGE\-M3\(chen2024bgem3\); the retriever is FAISS\(johnson2019faiss\)top\-k=5k=5\. The generator is Qwen2\.5\-7B\-Instruct\(yang2024qwen25\)with a system instruction to answer in the user’s language\. The input judge is the same Qwen2\.5\-7B model prompted exclusively in English with English\-only few\-shot examples \(BLOCK/ALLOWclassification\)\. The output filter matches three regex families \(email, phone, SSN\-like\)\. We additionally evaluate two judge variants in Section[4](https://arxiv.org/html/2608.05163#S4): a back\-translate\-then\-judge variant \(the query is first translated to English, then judged\) and a multilingual\-prompted variant \(system prompt and few\-shot examples are in the query language\)\.
### 3\.4Attack queries
For each document we generate three reformulations: \(*direct*\) anchor\-conditioned PII extraction request; \(*summarize*\) anchor\-conditioned summarization that asks for verbatim entities; \(*ICL*\) one\-shot in\-context\-learning example followed by the anchor\. The English templates are translated to Chinese, German, Arabic, and Swahili using Qwen2\.5\-7B\-Instruct as the translator \(NLLB\-200\-3\.3B\(nllb2022\)was unavailable through the available model mirrors at submission time\)\. We choose Chinese, German, Arabic, and Swahili to span scripts \(Latin, Han, Arabic\), typological distance from English, and resource levels under Qwen2\.5; the study is not a language\-fairness ranking\. Total trials:100×3×5=1,500100\\times 3\\times 5=1\{,\}500\.
### 3\.5Metrics
Per trial we record retrieval recall@5,*PII in generation*\(case\-folded substring match against the planted PII, with a first\-comma\-segment fallback for multi\-line addresses\), and*output guard triggered*\(regex for email/phone/SSN\-like\)\. Final leaks are*output\-only*= PII in generation∧¬\\wedge\\,\\negguard, and*combined*= input judge ALLOWed∧\\wedgeoutput\-only\. “Leak” is verbatim disclosure; transliterated or paraphrased renderings are not counted, so cross\-lingual semantic leakage may be undercounted \(Section[4\.1](https://arxiv.org/html/2608.05163#S4.SS1)reports a partial\-token rescore; Limitations\)\. All Qwen2\.5\-7B calls use greedy decoding\. Anchors are themselves Qwen\-translated: verbatim preservation across non\-English queries is0\.6970\.697–0\.7300\.730, partially confounding the cross\-lingual recall drop reported in Section[4\.1](https://arxiv.org/html/2608.05163#S4.SS1)\. We aggregate by*document\-level any\-success*with95%95\\%bootstrap CIs over documents \(5,0005\{,\}000resamples\)\. For zero\-success cells, the F1/F3 doc\-level table \(Table[1](https://arxiv.org/html/2608.05163#S4.T1), output\-only and combined columns\) reports*one\-sided*95%95\\%Wilson upper bounds; the conditional F2 table \(Table[3](https://arxiv.org/html/2608.05163#S4.T3)\) reports*two\-sided*95%95\\%Wilson upper endpoints to keep the same Wilson convention as its non\-zero rows\. BLOCK rates are per trial; because all queries are adversarial, they are attack\-set coverage, not classifier operating points\. Code, corpora, queries, per\-trial JSONLs, and aggregates are in the supplementary material\.
Figure 2:Two\-stage cascade reduces residual leak\. F1 \(output\-only regex filter\) vs\. F3 \(after adding the English\-only\-prompted input judge\), unstructured PII doc\-level any\-success leak rate \(n=40n\{=\}40/lang\)\. Error bars are95%95\\%Wilson CIs; Tables[1](https://arxiv.org/html/2608.05163#S4.T1)and[3](https://arxiv.org/html/2608.05163#S4.T3)give bootstrap variants and the underlying counts\.
## 4Results
Figure[2](https://arxiv.org/html/2608.05163#S3.F2)summarises the headline cascade \(F1–F3\); Figure[3](https://arxiv.org/html/2608.05163#S4.F3)reports the translation\-confound audit\. Cross\-language comparisons restrict to*unstructured PII*\(n=40n\{=\}40per language\); for structured targets \(n=60n\{=\}60\) the observed output\-only leak is0/600/60in every language — a guardrail sanity check, reported in Table[1](https://arxiv.org/html/2608.05163#S4.T1)\.
### 4\.1F1 — English leaks the most unstructured PII under output\-only filtering
Output\-only filterTwo\-stage \(input \+ output\)Query langStructured PII \(n=60\)Unstructured PII \(n=40\)All \(n=100\)Unstructured PII \(n=40\)English0/600/60;≤0\.043\\leq 0\.0430\.875 \[0\.775, 0\.975\]0\.350 \[0\.260, 0\.450\]0/400/40;≤0\.063\\leq 0\.063Chinese0/600/60;≤0\.043\\leq 0\.043†\\dagger0\.625 \[0\.475, 0\.775\]0\.250 \[0\.170, 0\.340\]0/400/40;≤0\.063\\leq 0\.063German0/600/60;≤0\.043\\leq 0\.0430\.775 \[0\.650, 0\.900\]0\.310 \[0\.220, 0\.400\]0/400/40;≤0\.063\\leq 0\.063Arabic0/600/60;≤0\.043\\leq 0\.0430\.750 \[0\.600, 0\.875\]0\.300 \[0\.210, 0\.390\]0\.075 \[0\.000, 0\.175\]Swahili0/600/60;≤0\.043\\leq 0\.0430\.425 \[0\.275, 0\.575\]0\.170 \[0\.100, 0\.250\]0\.175 \[0\.075, 0\.300\]Numbers are doc\-level any\-success leak rates with 95% bootstrap CIs over the column denominator; zero\-success cells show the one\-sided 95% Wilson upper bound\.†\\daggerTwo Chinese SSN\-like response strings are counted as guard\-triggered; zh structured leakage is0/600/60\(all\-PII rate0\.2500\.250\)\.
Table 1:Document\-level any\-success leak rate \(any of three reformulations leaks the target PII for that document\)\. The first three numeric columns report the*output\-only*regex filter across structured PII \(n=60n\{=\}60\), unstructured PII \(n=40n\{=\}40\), and all PII \(n=100n\{=\}100\); the rightmost column reports the residual leak after adding the English\-only\-prompted input judge \(*combined*, unstructured PIIn=40n\{=\}40\)\. Numbers are point estimates with 95% bootstrap CIs over the column denominator; zero\-success cells show the one\-sided 95% Wilson upper bound rather than a degenerate\[0,0\]\[0,0\]bootstrap interval\. The output regex matches emails, phones, and SSN\-like / 9\-digit patterns only, so the cross\-language comparison is informative on*unstructured PII*\(names, addresses\)\. The two\-stage filter produces zero observed leaks on en/zh/de; a small residual remains on Arabic and a sizeable residual on Swahili\.Under the output\-only regex filter, English queries achieve a document\-level any\-success leak rate of0\.875\[0\.775,0\.975\]0\.875\\,\[0\.775,\\,0\.975\]on unstructured PII; the four Qwen\-translated non\-English templates have lower point estimates — German0\.775\[0\.650,0\.900\]0\.775\\,\[0\.650,\\,0\.900\], Arabic0\.750\[0\.600,0\.875\]0\.750\\,\[0\.600,\\,0\.875\], Chinese0\.625\[0\.475,0\.775\]0\.625\\,\[0\.475,\\,0\.775\], Swahili0\.425\[0\.275,0\.575\]0\.425\\,\[0\.275,\\,0\.575\]\. The 95% bootstrap CIs separate cleanly for English vs\. Swahili and to a touching endpoint for English vs\. Chinese; English vs\. German and English vs\. Arabic overlap and are point\-estimate ordering only\. Stage decomposition shows retrieval recall@5 drops of1616–2929pp and verbatim PII echo drops of2525–3737pp under cross\-lingual queries \(per\-language stage rates in the supplementary material\), but the recall drop is partly confounded by Qwen\-translated anchor corruption \(0\.6970\.697–0\.7300\.730verbatim preservation, Figure[3](https://arxiv.org/html/2608.05163#S4.F3)\)\. The output regex targets email/phone/9\-digit patterns and so suppresses structured*target*PII; for name/address targets it can still trigger incidentally on anchors or numeric substrings, so the unstructured\-PII F1 rates are output\-only leak rates after this regex layer, not pure PII\-in\-generation rates\. The intuitive expectation that cross\-lingual queries amplify leakage is therefore not supported by point\-estimate ordering on this Qwen\-translated attack pipeline\.
#### Translation length\-heuristic audit\.
Qwen2\.5\-7B translates4343–67%67\\%ofdirectprompts as bare stubs \(<<30 characters\) for zh/de/sw and16%16\\%for ar, whilesummarizeis stub\-clean \(0/1000/100on every non\-English language\)\. On the stub\-cleansummarizesubset, English remains the highest point estimate \(0\.7000\.700\); non\-English summarize rates are de0\.4750\.475, zh0\.4500\.450, ar0\.3750\.375, sw0\.3000\.300, with zh and ar swapping versus the F1 full\-set order\. The check therefore supports the English\-highest contrast rather than the full F1 ranking, and stub\-clean does not certify semantic intent preservation \(Limitations\)\.
#### Robustness to scoring and anchor corruption\.
Two further sensitivity checks pressure\-test F1\. A loose partial\-token rescore \(any≥\\geq4\-character planted\-PII token in any response\) and a NER\-fuzzy rescore \(XLM\-R PERSON/LOC entities; sw unsupported\) shift non\-English point estimates by at most\+0\.05\+0\.05pp, preserving the English\-highest contrast; restricting F1 to documents whose corpus anchor verbatim\-survives Qwen translation in at least one reformulation \(n=31n\{=\}31–3838per language\) likewise leaves the contrast intact\. Neither check covers transliteration or culturally\-adapted entity rendering \(Limitations\)\.
### 4\.2F2 — Input\-judge failures align with translated\-query attenuation
Table 2:Input\-judge BLOCK rates per query language under three prompting regimes \(n=300n=300trials per language: 100 documents×\\times3 reformulations\)\. “English\-only\-prompted” is the deployed setting\. “Back\-translate→\\toEN” first translates the query back to English with the same multilingual model, then judges\. “Multilingual\-prompted” uses in\-language system prompt and few\-shot examples\.Adding the input\-side judge changes the failure mode \(Table[2](https://arxiv.org/html/2608.05163#S4.T2), column 1\)\. The English\-only\-prompted Qwen2\.5 judge blocks100%100\\%of English queries,9696–98%98\\%of Chinese/German/Arabic, but only76\.7%76\.7\\%of Swahili\. Two ablations probe, but do not identify, the source of this gap\.
Back\-translate\-then\-judge\(Table[2](https://arxiv.org/html/2608.05163#S4.T2), column 2\)\. If the failure were that the English\-only prompt cannot read foreign\-language input, back\-translating to English before judging would be expected to improve the BLOCK rate\. It does not: Swahili back\-translation BLOCK is71\.3%71\.3\\%, slightly*lower*than direct Swahili\. The pattern is consistent with translation\-induced intent attenuation in the original query, butthis ablation is non\-diagnostic of translation noise per se: the back\-translator is also Qwen, so we cannot separate translation\-induced attenuation from same\-family judge bias on translated text\. We report the result as a within\-pipeline behavioural check, not as a translation\-noise diagnostic, and adopt*differential pipeline degradation under translation noise*as a hypothesis: the generator \(with retrieval context\) degrades less under MT\-noisy input than the input\-side judge \(without retrieval context\) does\. An independent MT system \(NLLB\-200, OPUS\-MT\) paired with a non\-Qwen judge is the priority follow\-up that would let this hypothesis be tested causally \(§Limitations\)\.
In\-language prompted judge variant\(Table[2](https://arxiv.org/html/2608.05163#S4.T2), column 3\)\. Exploratory and undercontrolled \(different prompt prose and few\-shot count\)\. It closes the small Chinese \(0\.977→1\.0000\.977\\to 1\.000\) and German \(0\.960→1\.0000\.960\\to 1\.000\) gap and does not rescue Swahili in aggregate \(0\.767→0\.7470\.767\\to 0\.747\); per\-reformulation, the stub\-cleansummarizecell is worst on Swahili \(0\.6000\.600\)\. The deployed English\-only judge column of Table[2](https://arxiv.org/html/2608.05163#S4.T2)shows the corresponding deployed\-pipeline numbers; per\-reformulation breakdowns are in the supplementary material\.
Figure 3:Anchor verbatim\-preservation rate after Qwen translation \(n=300n\{=\}300/lang\)\. The verbatim share is0\.6970\.697–0\.7300\.730across non\-English languages; the remainder partially confounds the cross\-lingual recall drop in Section[4\.1](https://arxiv.org/html/2608.05163#S4.SS1)\.Table 3:Left/center: F2 doc\-level conditional analysis\(English\-only\-prompted Qwen judge\)\. Among docs with at least one unstructured output\-only\-leaking reformulation, the fraction where the input judgeALLOWed at least one of those reforms; F2\-cond doc counts equal the F3 combined\-leak doc counts \(3/403/40ar,7/407/40sw\)\. The rule\-based column is a deterministic non\-Qwen lexicon heuristic \(sanity check, not an LLM judge\); it disagrees sharply with Qwen on Chinese\.Right: F4 oracle document\-grounded judge \(n=17n\{=\}17cells /1313docs\), run on the*multilingual\-prompted*\-judge residual — not a direct F3 rescue\. Doc\-level shows fully\-rescued docs; cell\-level shows per\-cell BLOCK rate\.#### Conditional dangerous\-document analysis\.
The aggregate BLOCK rate \(Table[2](https://arxiv.org/html/2608.05163#S4.T2)\) overstates effective coverage on documents that actually leak\. Restricted to dangerous documents — those with at least one output\-only leaking reformulation — the English\-only\-prompted Qwen judgeALLOWs at least one such reform on3/30\(0\.100\)\[0\.035,0\.256\]3/30\\;\(0\.100\)\\,\[0\.035,0\.256\]Arabic and𝟕/𝟏𝟕\(0\.412\)\[0\.216,0\.640\]\\mathbf\{7/17\\;\(0\.412\)\\,\[0\.216,0\.640\]\}Swahili dangerous documents \(Table[3](https://arxiv.org/html/2608.05163#S4.T3)\)\. A rule\-based multilingual PII\-intent lexiconALLOWs8/178/17sw \(one above the Qwen rate\) but disagrees on Chinese \(12/2512/25vs\.0/250/25\); because the heuristic is deterministic and disagrees sharply on Chinese, we treat the Swahili8/178/17result only as a weak sanity check consistent with trigger\-phrase loss, not as model\-family\-independent F2 validation\. A Mistral\-7B\-Instruct\-v0\.3 judge\(jiang2023mistral\)with the identical English\-only prompt blocks1,498/1,5001\{,\}498/1\{,\}500trials in aggregate \(per\-language breakdown and benign\-query FPR not measured\); we use this only as evidence that the Qwen Swahili\-low pattern is not replicated by one alternate non\-Qwen judge, not as a drop\-in defence\.
### 4\.3F3 — Combined\-stage leak concentrates in the MT\-noisy corner
With both stages active \(English\-only\-prompted input judgeALLOWs and output regex does not trigger\), we observe*zero combined leaks*on English, Chinese, and German on this attack set \(0/400/40each on unstructured PII\); the one\-sided 95% Wilson upper bound on the true leak probability is≈0\.063\\approx 0\.063, so this is not proof of full elimination\. Arabic shows0\.075\[0\.000,0\.175\]0\.075\\,\[0\.000,\\,0\.175\]\(3/403/40docs\) and Swahili0\.175\[0\.075,0\.300\]0\.175\\,\[0\.075,\\,0\.300\]\(7/407/40docs\); these doc\-level counts equal the English\-only\-Qwen\-allowed numerator of Table[3](https://arxiv.org/html/2608.05163#S4.T3)\(consistent F1/F3/F2\-cond units\)\. The ar\-vs\-sw difference is point\-estimate only; bootstrap CIs overlap\. The residual concentrates in translations that bypass the input judge and still allow verbatim PII echo; short\-stub rate alone does not explain language differences \(zh/de have higherdirectstub rates than Arabic but zero combined leaks, while Swahili has a lower direct\-stub rate than Chinese yet17\.5%17\.5\\%residual\)\. The pattern is consistent with a joint condition of MT\-attenuated input intent and intact retrieval grounding, rather than monotonic stub\-collapse alone\.
Two Chinese structured\-target responses contain SSN\-like strings \(724\-25\-9524,704\-87\-4313\) that match the output regex; counted as guard\-triggered, Chinese structured output\-only leakage is0/600/60\(Table[1](https://arxiv.org/html/2608.05163#S4.T1)footnote†\\dagger\)\.
### 4\.4F4 — Document\-grounded judge on a multilingual\-judge residual \(n=17n\{=\}17, oracle\)
We test document\-grounded judging on the*multilingual\-prompted Qwen judge*’s residual \(55ar cells / 5 docs;1212sw cells / 8 docs;ALLOW=1\.0001\.000\) — an independent mechanism diagnostic, not a direct rescue of the deployed English\-only F3 residual \(the two residuals overlap but differ; 17 vs 13 cells\)\. Re\-judging with the gold corpus document attached as “Retrieved Context”, the judge BLOCKs𝟓/𝟓\\mathbf\{5/5\}Arabic and𝟏𝟎/𝟏𝟐\\mathbf\{10/12\}Swahili cells \(15/17 total; 12/13 docs fully rescued\)\.n=17n\{=\}17is small \(wide Wilson intervals\); the experiment uses oracle retrieval, does not vary the retriever, and does not include benign queries, so it does not measure FPR or utility cost\. Target retrieval@k was correct for all 17 cells, isolating the input\-side judge gap; replicating on the English\-only residual and on a benign\-query set is the natural next step\.
## 5Discussion
The picture is asymmetric\. Output\-side regex filters target structured patterns and remove observed verbatim structured\-target leaks in this attack set; for name/address targets they trigger only incidentally, so cross\-language F1 is informative on unstructured PII\. Input\-side LLM judges must infer query intent from text whose translation may have eroded explicit extraction cues, without the retrieval context the downstream generator receives; in this Qwen\-mediated pipeline that asymmetry is consistent with the residual leak corner in Section[4\.3](https://arxiv.org/html/2608.05163#S4.SS3)\.
#### F4 is a diagnostic, not a deployable defence\.
The document\-grounded judging result in Section[4\.4](https://arxiv.org/html/2608.05163#S4.SS4)blocks15/1715/17residual cells, but the experiment uses oracle retrieval on a smalln=17n\{=\}17corner and measures only adversarial\-query BLOCK rates\. A filter that blocks15/1715/17adversarial residuals is operationally uninterpretable without a paired benign\-query false\-positive rate at matched thresholds and an answer\-utility measurement on a benign workload; we have neither\. We therefore position both candidate directions — \(a\) explicit translation\-quality gating and \(b\) document\-grounded input filtering — as*candidate mechanisms whose deployment cost is not measured in this paper*, not as recommended defences\. The full operational checklist \(benign\-FPR, utility, context\-expansion cost, judge latency\) is in §Limitations\. The same caution applies whenever a positive audit signal on an LLM\-mediated system is read as a causal mechanism rather than as a configuration\-conditional or trace\-conditional artefact\(li2026safetyrepro;li2026reasoningtrace\)\.
#### Scope: MT\-templated attack surface\.
The non\-English attack queries in this audit are Qwen2\.5\-7B translations of an English template seed set\. The English\-highest unstructured\-PII contrast in Section[4\.1](https://arxiv.org/html/2608.05163#S4.SS1)and the ar/sw residual concentration in Section[4\.3](https://arxiv.org/html/2608.05163#S4.SS3)are therefore*MT\-template\-conditional*; we cannot speak to native\-written cross\-lingual extraction prompts, which are arguably the more realistic threat surface\. Combined with the same\-model\-family pipeline, the natural follow\-up paper has three changes from this one: an independent MT system \(NLLB\-200 or OPUS\-MT\), a non\-Qwen input judge, and a native\-speaker\-written query set in the same five languages\.
## 6Conclusion
In this Qwen\-mediated machine\-translated template attack, non\-English queries do not leak more than English under output\-only filtering: the highest unstructured\-PII leak point estimate is on English, and only English\-vs\-Swahili separates cleanly\. This is pipeline\-conditional and should not be read as evidence about native\-written or translation\-preserving non\-English attacks\. We hypothesize a more specific residual failure — Qwen\-translated queries may lose extraction cues enough that the English\-only\-prompted input judge allows them, while the downstream generator, conditioned on retrieved context, still produces verbatim PII\.
Residual combined leaks concentrate on Arabic and Swahili, but Chinese and German also have high translation noise yet zero observed combined leaks, so this is not a language\-distance ranking\. Independent MT and an independent judge are the natural next steps\. As an oracle diagnostic on a separate multilingual\-prompted\-judge residual \(n=17n\{=\}17cells /1313docs\), giving the input judge the gold corpus document blocks15/1715/17cells; this motivates context\-aware input filtering, not a validated deployment for the English\-only F3 residual\.
## Limitations
The contribution is a stage\-decomposed pilot under one Qwen\-mediated configuration plus ann=17n\{=\}17oracle proof\-of\-concept — not a general result, not an identified causal mechanism, and not a deployable defence evaluation\. Three scope boundaries deserve explicit attention\.
#### Single\-model\-family pipeline \(priority follow\-up\)\.
Translator, input judge, back\-translator, and generator are all Qwen2\.5\-7B in the deployed pipeline\. The back\-translate\-then\-judge ablation in Section[4\.2](https://arxiv.org/html/2608.05163#S4.SS2)therefore cannot separate translation noise from same\-family judge bias; we report it as a within\-pipeline behavioural check, not as a diagnostic\. The Mistral\-7B\-Instruct\-v0\.3\(jiang2023mistral\)sanity check in Section[4\.2](https://arxiv.org/html/2608.05163#S4.SS2)only shows that the Qwen Swahili\-low pattern is not trivially replicated by one alternate judge; per\-language breakdowns are not reported and benign\-query FPR is not measured\. The priority replication is an independent MT system \(NLLB\-200\(nllb2022\)or OPUS\-MT\(tiedemann2020opusmt\)\) paired with a non\-Qwen judge \(Llama\-3\.1\-8B, Mistral\-Small\-3\.1, or a commercial API\) on the same five\-language attack grid; the existing harness supports per\-stage model swaps so this is a configuration\-only extension\. Only that replication — not the present pipeline — can distinguish language\-inherent risk from configuration\-specific artefact\.
#### F4 is a mechanism diagnostic, not a deployable defence \(FPR and utility unmeasured\)\.
The document\-grounded judging result in Section[4\.4](https://arxiv.org/html/2608.05163#S4.SS4)\(15/1715/17residual cells blocked\) is intentionally framed as a mechanism check: it tests whether attaching retrieved context closes the input\-judge gap\. It is not a deployable defence evaluation, because \(a\) retrieval is the gold corpus document \(oracle\), so it sets an upper bound on what a real retriever could deliver; \(b\)n=17n\{=\}17gives wide Wilson intervals; \(c\) we measure no benign\-query BLOCK rate, so the false\-positive cost is unknown; and \(d\) we measure no answer\-utility cost — a defence that BLOCKs15/1715/17adversarial residuals is uninformative until paired with benign\-query BLOCK rate at matched thresholds, answer\-quality on a benign workload \(ROUGE\-L or judge\-score\), the retrieval\-context expansion cost, and judge\-side latency\. The same FPR / utility gap applies to the translation\-quality\-gating direction floated in Section[5](https://arxiv.org/html/2608.05163#S5)\.
#### Machine\-translated vs\. native\-written attack surface\.
The non\-English attack queries are Qwen2\.5\-7B translations of an English template seed set\. Native\-written cross\-lingual extraction prompts — where adversarial intent is expressed idiomatically rather than translated — are arguably the more realistic threat surface and one this audit cannot speak to\. We treat this as a scope boundary, not a measurement noise issue\. A∼\\sim200\-query native\-speaker collection \(5 languages×\\times4 reformulations×\\times10 prompts\) is the natural follow\-up; combined with the cross\-family pipeline above, it would also let us separate MT\-template attenuation from translator\-or\-judge bias as the source of the residual\-leak corner\.
#### Other scope notes\.
The corpus is small \(n=40n\{=\}40unstructured per language\); en\-vs\-zh, en\-vs\-de, en\-vs\-ar, and ar\-vs\-sw F3 contrasts are point\-estimate only\. Anchors are Qwen\-translated, so part of the recall drop is anchor\-corruption \(Section[4\.1](https://arxiv.org/html/2608.05163#S4.SS1)\)\. Retrieval is single\-point \(k=5k\{=\}5, no reranker\); the scorer omits transliterated and culturally\-adapted disclosures\. The corpus is entirely synthetic \(Faker\-generated, no real personal data\); the attack templates are not for production use\.
## Ethics statement
The corpus is entirely synthetic \(Faker\-generated, no real personal data\); the attack templates are not for production use\.
## ReferencesSimilar Articles
Exploratory As-Analyzed No-Detection of Culturally-Marked Predicate-Triggered PII Amplification in a Synthetic-English RAG Probe: A Predicate-Resource-Confounded Audit
The study explores whether stereotype-loaded queries about culturally marked individuals lead to more personal information leakage from RAG systems than neutral queries, finding no evidence of such amplification after statistical corrections.
All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAG
Researchers identify systematic English and query-language bias in multilingual RAG rerankers and introduce LAURA, a utility-driven alignment method that boosts performance by retrieving answer-critical documents across languages.
RAG-CT: Mitigating Privacy Risks on Retrieval-Augmented Generation Systems via Scanning Prompt Distribution
RAG-CT is a novel defense method that identifies malicious queries by analyzing entropy and margin distributions to mitigate privacy risks in Retrieval-Augmented Generation systems, significantly reducing PII leakage.
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
Introduces MosaicLeaks, a benchmark of 1,001 multi-hop deep research tasks that chain private enterprise documents with public web queries to evaluate privacy leakage. Finds that models leak sensitive information at multiple levels, and proposes PA-DR, a reinforcement learning framework that reduces leakage while improving task accuracy.
Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs
This paper introduces SEAG, a privacy-preserving framework for RAG that conceals sensitive entities by replacing them with aliases in queries and documents before forwarding to external LLMs, achieving over 80% user accuracy and strong entity-hiding performance.