Lost in Translation: Measuring the Effect of Non-Native English on End User Performance of Large Language Models

arXiv cs.CL Papers

Summary

This paper introduces Fable, a controlled dataset of 190,911 English prompt variants derived from 174K real user prompts, and evaluates 34 open-weight LLMs to show that models mirror higher-level rhetorical and lexical qualities of non-native speakers' prompts, producing lower-quality and less-fluent responses for less fluent users.

arXiv:2609.36214v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used by people whose first language is not English, yet these users have been shown to receive systematically lower-quality responses than fluent speakers. Which specific features of non-native English drive this gap remains unclear, because fluency is itself a composite of mechanical accuracy, vocabulary use, organization, and discourse coherence. Here, we introduce FABLE, a controlled dataset of 190,911 English prompt variants derived from 174K real user prompts for writing-related tasks. Evaluating responses from 34 open-weight LLMs, we find a clear asymmetry; while models do not propagate surface errors such as misspellings into their outputs, models do mirror higher-level rhetorical and lexical qualities present in the user's prompt. Further, the overall quality of responses differs substantially between the least- and most-fluent prompts. These results highlight a key LLM performance disparity for non-native English LLM users, resulting in both lower-quality and less-fluent answers.
Original Article
View Cached Full Text

Cached at: 09/30/26, 09:51 AM

# Lost in Translation: Measuring the Effect of Non-Native English on End User Performance of Large Language Models
Source: [https://arxiv.org/html/2609.36214](https://arxiv.org/html/2609.36214)
Eleanor Lin11footnotemark:1David JurgensAffiliation:Electrical Engineering and Computer Science DepartmentAffiliation:University of MichiganAffiliation:Ann Arbor, MI 48105, USAEmail:[\{yszhou,elealin,jurgens\}@umich\.edu](mailto:)

###### Abstract

Large language models \(LLMs\) are increasingly used by people whose first language is not English, yet these users have been shown to receive systematically lower\-quality responses than fluent speakers\. Which specific features of non\-native English drive this gap remains unclear, because fluency is itself a composite of mechanical accuracy, vocabulary use, organization, and discourse coherence\. Here, we introduceFable, a controlled dataset of 190,911 English prompt variants derived from 174K real user prompts for writing\-related tasks\. Evaluating responses from 34 open\-weight LLMs, we find a clear asymmetry; while models do not propagate surface errors such as misspellings into their outputs, models do mirror higher\-level rhetorical and lexical qualities present in the user’s prompt\. Further, the overall quality of responses differs substantially between the least\- and most\-fluent prompts\. These results highlight a key LLM performance disparity for non\-native English LLM users, resulting in both lower\-quality and less\-fluent answers\.

## 1Introduction

Non\-native English speakers represent a critical and rapidly growing user base for large language models \(LLMs\)\([Liu et al\., 2025a](https://arxiv.org/html/2609.36214#bib.bib6)\)\. However, LLMs have demonstrated performance degradation for non\-native English speakers\([Reusens et al\., 2025](https://arxiv.org/html/2609.36214#bib.bib41)\)\. Intuitively, this performance degradation may stem from differences in the language produced by non\-native speakers of different fluency levels, compared to the native English that LLMs are trained on\. Fluency, however, is not a single property: it is a composite of features ranging from surface mechanics such as spelling and grammar to higher\-level qualities such as organization and discourse coherence\([Chambers, 1997](https://arxiv.org/html/2609.36214#bib.bib25)\)\. Without knowing which of these features matter for LLM behavior, it is difficult to diagnose why non\-native speakers are disadvantaged and to design effective interventions to improve LLM performance for these users\. Here, we systematically test which features of non\-native English drive performance degradation for non\-native speakers\.

Existing work on prompt sensitivity has examined how surface\-level variations such as formatting\([He et al\., 2024](https://arxiv.org/html/2609.36214#bib.bib30)\), example ordering\([Guo et al\., 2024](https://arxiv.org/html/2609.36214#bib.bib46);[Lu et al\., 2022](https://arxiv.org/html/2609.36214#bib.bib31)\), length\([Liu et al\., 2025b](https://arxiv.org/html/2609.36214#bib.bib32)\), tone\([Dobariya and Kumar, 2025](https://arxiv.org/html/2609.36214#bib.bib33)\), and politeness\([Yin et al\., 2024](https://arxiv.org/html/2609.36214#bib.bib34)\)affect model behavior\. A smaller line of work has begun to consider specific linguistic properties such as tense, mood, and voice\([Leidinger et al\., 2023](https://arxiv.org/html/2609.36214#bib.bib35)\)and the broad effect of speaker nativeness\([Reusens et al\., 2025](https://arxiv.org/html/2609.36214#bib.bib41)\)\. However, these studies typically treat fluency as a single binary or scalar property and rely on benchmark\-style prompts that do not reflect how everyday users actually interact with LLMs\. As a result, it remains unclear which specific aspects of non\-native English are most consequential for downstream task performance, or whether their effects are consistent across realistic, open\-ended user tasks\.

We address this gap by introducingFable\(Fluency\-Adjusted Benchmark for LLM Generation Evaluation\), a controlled dataset of 190,911 English prompt variants derived from real\-world user interactions for writing\-related tasks\. Starting from 174K filtered prompts in English and Chinese collected from WildChat\([Zhao et al\., 2024](https://arxiv.org/html/2609.36214#bib.bib5)\)and ShareLM\([Don\-Yehiya et al\., 2025](https://arxiv.org/html/2609.36214#bib.bib20)\)—a language pair representing one of the largest populations of bilingual LLM users\([Yang, 2006](https://arxiv.org/html/2609.36214#bib.bib17)\)—we generate multiple English translations of each Chinese prompt that conform to specified profiles along the five dimensions of the ESL Composition Profile\([Jacobs et al\., 1981](https://arxiv.org/html/2609.36214#bib.bib27)\): Content, Organization, Vocabulary, Language Use, and Mechanics\. This design disentangles the contribution of each fluency dimension while holding the prompt’s semantic intent approximately constant\. We evaluate 34 open\-weight LLMs on a shared sample of 5,000 prompt variants, and evaluate their responses using both rule\-based linguistic metrics and LLM\-as\-judge ratings of fluency and overall response quality\.

Our analysis reveals a striking asymmetry between Surface Form Features and higher\-order qualities\. LLMs do not substantially propagate surface errors \(e\.g\., misspellings, grammatical mistakes\) from prompts into their responses, but they do mirror higher\-level rhetorical and lexical qualities: prompts with weaker Content, Organization, or Vocabulary yield responses with correspondingly weaker proficiency along those same dimensions\. More importantly, prompt fluency has a substantial effect on response quality: a one\-point increase in each retained prompt\-fluency dimension on the four\-point scale leads to a 0\.04–0\.14 point increase in overall response quality on a five\-point scale\. This effect is consistent across writing\-focused task categories, although the relative importance of individual fluency dimensions varies by task\. Together, these findings suggest that users who turn to LLMs to compensate for limited English fluency may still receive systematically lower\-quality output, with implications for equitable access to LLM technology\.

Our paper makes the following three contributions\. \(1\) We introduceFable, a controlled dataset of 190,911 prompt translations spanning the ESL Composition Profile that enables fine\-grained study of how prompt fluency affects LLM behavior\. \(2\) We provide an evaluation methodology that jointly quantifies surface\-level and higher\-order linguistic properties of both prompts and responses\. \(3\) We demonstrate empirical evidence that LLM response quality is shaped more by the rhetorical and lexical qualities of a prompt and that, across 34 models, LLM responses to less\-fluent prompts are themselves less fluent and lower quality\. Our results point to an important source of inequality in access to fluent, high\-quality outputs by all users\. All data and code will be made available upon publication \(CC\-by\-SA\-4\.0\)\.

## 2Related Work

##### Language Fluency and Language Technology

Language fluency has been shown to play a critical role across multiple NLP fields\. In Machine Translation, prior work demonstrates that noisy input significantly degrades translation quality for both neural MT systems and large language models\([Popović et al\., 2024](https://arxiv.org/html/2609.36214#bib.bib36);[Pan et al\., 2024](https://arxiv.org/html/2609.36214#bib.bib37)\)\. This phenomenon can be viewed as a manifestation of reduced language fluency of input\. Meanwhile, “fluency” is also recognized as a fundamental dimension in the evaluation of translation quality, with researchers proposing structured frameworks, such as MQM[Park and Padó \(2024\)](https://arxiv.org/html/2609.36214#bib.bib38), to systematically assess fluency alongside other aspects of translation quality\. Similarly, in Information Retrieval, the effectiveness of retrieval systems is closely tied to the quality of user queries: well\-formed queries tend to achieve better performance[Chikkamath et al\. \(2024\)](https://arxiv.org/html/2609.36214#bib.bib39), while typos or other linguistic errors can substantially reduce system accuracy[Zhuang and Zuccon \(2022\)](https://arxiv.org/html/2609.36214#bib.bib40)\. In this work, we seek to understand the impact of language fluency on LLM responses to non\-native English speakers\.

##### Native v\. Non\-native Use of LLMs

Prior studies about contrastive rhetoric have shown that first language and cultural background influence how people write in a second language[Kaplan and others \(1966\)](https://arxiv.org/html/2609.36214#bib.bib43)\. In particular, native and non\-native English speakers show systematically different usage of linguistic structure beyond semantic content[Rabinovich et al\. \(2016\)](https://arxiv.org/html/2609.36214#bib.bib42)\. Such variation naturally carries over to interactions with LLMs; e\.g\.,[Reusens et al\. \(2025\)](https://arxiv.org/html/2609.36214#bib.bib41)demonstrate that the nativeness of prompts significantly influences model performance\. Our work identifies the specific aspects of non\-native prompt fluency that cause differences in model behavior\.

##### Effects of Prompt Design

Prompts that are semantically equivalent can exhibit a substantial model performance gap solely due to variations in their surface form\([Cao et al\., 2024](https://arxiv.org/html/2609.36214#bib.bib28)\)\. Even when the semantic content remains identical, different formatting can lead to different performance\([He et al\., 2024](https://arxiv.org/html/2609.36214#bib.bib30);[Sclar et al\., 2023](https://arxiv.org/html/2609.36214#bib.bib29)\)\. Example ordering also plays an important role in designing prompts for in\-context learning\([Lu et al\., 2022](https://arxiv.org/html/2609.36214#bib.bib31);[Guo et al\., 2024](https://arxiv.org/html/2609.36214#bib.bib46)\)\. Prompt length also serves as a decisive factor in domain\-specific tasks\([Liu et al\., 2025b](https://arxiv.org/html/2609.36214#bib.bib32)\)\. Additionally, in real\-world interactions, research indicates that the tone\([Dobariya and Kumar, 2025](https://arxiv.org/html/2609.36214#bib.bib33)\)and politeness\([Yin et al\., 2024](https://arxiv.org/html/2609.36214#bib.bib34)\)of a prompt significantly influence how models respond\. This extensive evidence of LLM sensitivity to variation in prompts’ surface form provides the motivation for our own comparison of semantically equivalent prompts at different fluency levels\.

While these studies provide valuable insights into different dimensions of prompts, a systematic evaluation of how the language fluency level of prompts influences model behaviors is still missing\. Moreover, prior studies often evaluate prompt effectiveness using benchmark tasks that may not reflect real\-world user interactions\. In contrast, our work systematically investigates how variations in prompt fluency influence LLM responses on real\-world interaction data\.

## 3Curating a Dataset of Realistic, Translated User Prompts

To investigate the impact of fluency in prompts on model responses and downstream task performance, we curate a set of prompts from real\-world interactions in the WildChat\([Zhao et al\., 2024](https://arxiv.org/html/2609.36214#bib.bib5)\)and ShareLM\([Don\-Yehiya et al\., 2025](https://arxiv.org/html/2609.36214#bib.bib20)\)datasets\. Both datasets consist of conversations of diverse domains collected from real human\-AI interactions\. These prompts allow us to better match the experience of everyday users in realistic workflows \(e\.g\., summarizing an email\) when compared with prompts for multiple choice questions traditionally seen in NLP\.

##### How Are People Using LLMs?

Adapting the methodology and taxonomy proposed in[Chatterji et al\. \(2025\)](https://arxiv.org/html/2609.36214#bib.bib26), we categorize each prompt into one of 24 predefined task categories\. The distribution of prompts across task categories is shown in Fig[1](https://arxiv.org/html/2609.36214#S3.F1)\. We observed that among the ten most common user task categories, writing\-related queries account for the majority, with “Computer Programming” being the only category not directly related to writing\. Moreover, we believe that LLM responses in these writing\-related tasks such as “argument or summary generation,” “creative ideation,” and “edit or critique provided text” may be influenced by users’ fluency\.

##### Data Filtering

To ensure data quality and experimental consistency, we apply a sequence of data filtering steps\. For multi\-turn conversations, we only keep the first user prompt\. While we recognize the importance of multi\-turn conversations, our focus here is on how differences in linguistic proficiency affect the initial model response, rather than whether users caneventuallyachieve comparable outcomes to native speakers after multiple rounds of repair\. The first\-turn model response is consequential because it determines what the user must subsequently correct, clarify, or refine\. A systematic quality deficit at the first turn means that non\-native users begin the interaction at a disadvantage and must expend additional turns, time, and effort to reach parity; this additional interaction burden is itself an equity cost\. Additionally, effectively simulating realistic interactions across multiple turns remains a challenging problem\([Ivey et al\., 2024](https://arxiv.org/html/2609.36214#bib.bib7)\), and there are currently no curated datasets that are easily amenable to testing multi\-turn effects of fluency\. Therefore, we leave exploration of multi\-turn scenarios to future work\.

Additionally, since it is difficult to evaluate the fluency of very short prompts, and extremely long prompts are less representative of typical user queries, we perform length\-based filtering\. We remove both the shortest and longest 5% of prompts and randomly retain 30% of the prompts shorter than 30 tokens to control the prompt length distribution\.

To filter out prompts that contain multiple languages, we perform sentence\-level language identification using the fastText language identification model[Joulin et al\. \(2017\)](https://arxiv.org/html/2609.36214#bib.bib23)\. We segment each prompt into sentences and keep prompts which have all sentences classified as the same language with confidence scores above 0\.6\.

In the following experiments, we focus on prompts in English and Chinese\. English is the dominant language used in LLM training and inference\([Qin et al\., 2024](https://arxiv.org/html/2609.36214#bib.bib18)\), while Chinese represents one of the most\-spoken non\-English languages used with LLMs\([Liu et al\., 2024](https://arxiv.org/html/2609.36214#bib.bib19)\)\. Further, Chinese\-English bilinguals worldwide make up a large, natural population of speakers with varying degrees of English fluency who are therefore affected by models’ sensitivity to English fluency\. Studying these two languages also reflects common real\-world scenarios where bilingual users may tend to interact with LLMs in English even when it is not their first language\([Dedema et al\., 2026](https://arxiv.org/html/2609.36214#bib.bib45)\)\. We do recognize that non\-native English speakers have many other first languages, each producing characteristic interference patterns\([Kobayashi and Rinnert, 1992](https://arxiv.org/html/2609.36214#bib.bib9)\)\. The specific fluency effects we observe may not generalize to other source languages, particularly those with substantially different typological properties \(e\.g\., morphologically rich languages, or languages with writing systems beyond Han characters\)\. Nonetheless, the fluency taxonomy we use is language\-agnostic\. We therefore expect aspects of our findings to generalize beyond Chinese\-speaking users\.

We restrict the prompts to those focusing on writing, as these represent the most common forms of real\-world LLM usage according to our analysis of WildChat and ShareLM and prior studies on real\-world usage data, e\.g\.[Chatterji et al\. \(2025\)](https://arxiv.org/html/2609.36214#bib.bib26)\. Specifically, we only include prompts from the following task types: “argument or summary generation,” “edit or critique provided text,” “personal writing or communication,” “translation,” and “write fiction\.” These prompts are well\-suited for studying the relationship between prompt fluency and model behavior\. After the above processing steps, our final dataset consists of 173,726 prompts\. WildChat contributes 21,734 English and 27,532 Chinese prompts, while ShareLM contributes 90,121 English and 34,339 Chinese prompts\. To provide a reference for downstream response quality evaluation, we use DeepSeek\-V4\.1\-Flash\([DeepSeek\-AI, 2026](https://arxiv.org/html/2609.36214#bib.bib4)\)to translate the original Chinese prompts into English, and treat the resulting outputs as ground\-truth translations\.

Figure 1:The top\-10 task categories for Chinese prompts in WildChat and ShareLM, categorized using the taxonomy proposed by[Chatterji et al\. \(2025\)](https://arxiv.org/html/2609.36214#bib.bib26), shows that many real\-world user prompts involve writing and information\-seeking tasks rather than other questions\. Bold denotes categories of prompts we use\.

## 4Introducing Translation Artifacts

Language fluency is a multidimensional spectrum, with many features of language contributing to the perception of fluency, such as grammatical structure, word choice, discourse coherence, or even capitalization\. To more precisely control for these different factors, we introduceFable, which systematically constructs prompts using heterogeneous sources of disfluencies and provides multiple variants of the same prompt intent across the fluency spectrum\.Fableallows testing how specific types of linguistic behavior influence a model’s downstream behavior\. Here, we first describe how we measure fluency and then how we generate the data forFable\.

### 4\.1Quantifying Fluency

We quantify language fluency from two complementary perspectives:Surface Form Featuresthat measure specific properties like mechanical competence \(spelling, capitalization\), andRhetorical and Lexical Qualitiesfeatures that measure higher\-level constructs like discourse coherence and vocabulary use\. Surface Form Features mark the writer’s ability to control the form of the language\([Witte and Faigley, 1981](https://arxiv.org/html/2609.36214#bib.bib13)\), while rhetorical and lexical qualities measure compositional competence\([Siekmann et al\., 2022](https://arxiv.org/html/2609.36214#bib.bib21)\)\.

ForSurface Form Features, we build a rule\-based evaluation pipeline based on LanguageTool\([Miłkowski, 2010](https://arxiv.org/html/2609.36214#bib.bib1);[LanguageTool contributors, 2026](https://arxiv.org/html/2609.36214#bib.bib2)\)to detect spelling, grammatical, lexical, and stylistic errors\. We also compute Flesch Reading Ease to evaluate readability\([Flesch, 1948](https://arxiv.org/html/2609.36214#bib.bib22)\)\.

ForRhetorical and Lexical Qualities, we adopt a widely used analytic ESL writing rubric which evaluates a given text along five dimensions: Content, Organization, Vocabulary, Language Use, and Mechanics\([Jacobs et al\., 1981](https://arxiv.org/html/2609.36214#bib.bib27)\)\. A brief description of all dimensions is presented in Table[1](https://arxiv.org/html/2609.36214#S4.T1)\. We collect these proficiency dimensions using an LLM\-as\-a\-judge framework, where the judging model assigns an integer score from 1 to 4 for each dimension independently according to the rubric definitions\. We usegoogle/gemma\-4\-26B\-A4Bas judge\. To validate our pipeline, we use the ICNALE dataset\([Ishikawa, 2018](https://arxiv.org/html/2609.36214#bib.bib24)\), consisting of 656 learner\-written texts annotated with four\-level CEFR\-aligned writer proficiency labels\. We apply our judge template and take the average score from all dimensions; this score has a correlation of 0\.47, suggesting the judge is moderately well calibrated on what is a hard task for humans\.

Table 1:The five dimensions in the adapted ESL Composition Profile\([Jacobs et al\., 1981](https://arxiv.org/html/2609.36214#bib.bib27)\)\.
### 4\.2Controlled Generation of Translation

To systematically study the effect of prompt fluency on models’ responses, we constructFableto contain a range of translations with controlled fluency levels\. Specifically, we create prompts that instruct according to the five dimensions ofRhetorical and Lexical Qualities\. Each dimension is assigned a discrete level from 1 to 4, corresponding to increasing levels of proficiency in this specific dimension: Level 1 denotesvery poorproficiency, level 2 representsfair to poorproficiency, level 3 meansgood to averageproficiency, and level 4 stands forexcellent to very good proficiency\. Each combination of these levels defines atranslation profile\. For each reference Chinese\-language prompt, we include \(i\) 10 translations produced by random samples of valid profiles and \(ii\) translations from four uniform proficiency profiles where all dimensions have the same level\. We then use Qwen3\-32B to generate translations that conform to the specified ESL composition profile\. Appendix Table[2](https://arxiv.org/html/2609.36214#A1.T2)presents an example of translated prompts under different instructed levels\. This process yields multiple translation variants per prompt, enabling a systematic exploration of translation\-induced prompt variation across different dimensions of fluency\.

Because LLM\-generated translations may contain extraneous text or generation artifacts, we apply post\-generation quality control before constructing the final dataset\. This procedure combines rule\-based checks for generation\-related text, repeated trigram detection, language identification, and paragraph\-count checks\. After post\-generation quality control,Fablecontains 190,911 English prompt variants spanning the five selected writing\-related task categories\.

Fableis not intended to reproduce every idiosyncratic pattern found in non\-native writing\. Rather, it operationalizes fluency using a theory\-grounded taxonomy that decomposes this inherently difficult\-to\-define construct into five established and complementary dimensions: Content, Organization, Vocabulary, Language Use, and Mechanics\. Taken together, these dimensions provide a robust and comprehensive representation of fluency that extends beyond surface\-level grammatical correctness to include higher\-level rhetorical, lexical, and organizational properties\.

### 4\.3Evaluating Translated Prompts

To analyze the effectiveness of our approach for generating prompts with controlled fluency levels, we evaluate the translated prompts along two dimensions: language fluency and semantic preservation\. For language fluency, we apply the evaluation pipeline described above in Section[4\.1](https://arxiv.org/html/2609.36214#S4.SS1)to assess theSurface Form FeaturesandRhetorical and Lexical Qualitiesof each translated prompt—i\.e\., does instructing a model to introduce specific aspects of fluency lead to those appearing in the translation? In addition, we consider semantic preservation during translation, as semantic distortion may affect model behavior and downstream task performance\.

We found that higher instructed proficiency levels generally yield corresponding improvements in measured fluency, although cross\-dimensional effects remain\. Semantic similarity is only weakly associated with instructed fluency, and human evaluation of 100 randomly sampled translations yields a mean semantic\-preservation score of 4\.13 out of 5, suggesting that the translations generally preserve the original intent\. Additionally, instructed levels across all five fluency dimensions show little association with semantic preservation\. Detailed methods and results are provided in Appendix[A\.3](https://arxiv.org/html/2609.36214#A1.SS3)\.

## 5Quantifying the Impact of Prompt Fluency

Given two prompts with the same intent, but which differ in English fluency, how different will LLMs’ responses be? We evaluateresponsesfor their fluency \(surface form features and rhetorical and lexical qualities\), as well as how well they address the user query \(response quality\)\.

### 5\.1Experimental Setup

To generate responses to our fluency\-controlled prompts, we evaluate 34 LLMs spanning multiple model families and sizes, including Qwen, Gemma, Llama, Nemotron, OLMo, and others\. The complete list of models is provided in Appendix[A\.4](https://arxiv.org/html/2609.36214#A1.SS4)\. We randomly sample 5,000 English prompt variants fromFableand use the same sampled set for all models\.

Drawing on prior work on LLM\-as\-judge evaluation\([Zheng et al\., 2023](https://arxiv.org/html/2609.36214#bib.bib8);[Liu et al\., 2023](https://arxiv.org/html/2609.36214#bib.bib12)\), we develop a response\-quality evaluation protocol using Qwen3\.8\-27B\([Qwen Team, 2026](https://arxiv.org/html/2609.36214#bib.bib3)\)\. To mitigate the potential inconsistencies caused by varying translation quality in the user prompts, we provide the golden translations obtained earlier in Section[3](https://arxiv.org/html/2609.36214#S3)to the judging LLM\. This design allows the evaluation to focus on differences in model responses rather than artifacts introduced by translation noise\. \(See Appendix[A\.6](https://arxiv.org/html/2609.36214#A1.SS6)for more details\.\)

To examine the evaluation protocol, we measure the consistency between the scores of the LLM judge and two human annotators on a random sample of 100 model responses\. For overall response quality, human interannotator agreement is Krippendorff’sα=0\.71\\alpha=0\.71, and the judge’s scores achieve a Spearman correlation ofρ=0\.55\\rho=0\.55with the mean human scores\. Given the challenging and subjective nature of evaluating open\-ended responses, this moderate positive correlation provides encouraging evidence that the judge captures meaningful aspects of human assessments of response quality\.

### 5\.2Results

The effects of prompt fluency are assessed in terms of the generated language and task accuracy\.

#### 5\.2\.1Response Fluency

LLMs are known to accommodate to users’ style\([Durandard et al\., 2025](https://arxiv.org/html/2609.36214#bib.bib47)\)\. For less fluent users, such accommodation is potentially disadvantageous in writing tasks as LLMs may mirror undesired errors\. For example, a user asking to help draft an email with a less\-fluent prompt may receive a version that has grammatical mistakes\.

![Refer to caption](https://arxiv.org/html/2609.36214v1/prompt_fluency_vs_response_language_correctness.png)Figure 2:Prompt rhetorical and lexical qualities do not exhibit a systematic impact on model responses’ surface form features\. This figure presents the regression results between different linguistic properties of responses and different fluency metrics of prompts\.##### Response Surface Form Features

To assess the impact of the prompt’s fluency, we first fit separate linear regressions to predict the Surface Form Features for each response from the input prompt’s rhetorical and lexical quality levels\.

Prompt fluency shows relatively weak associations with the surface form features of LLM responses\. As shown in Figure[2](https://arxiv.org/html/2609.36214#S5.F2), LLMs did not substantially accommodate to reproduce errors from the prompt\. The largest changes were in the reading ease, but changes to other qualities did not produce a systematic trend of introducing more/fewer mechanistic errors\. These results suggest that the LLMs we test exhibit relatively stable performance at the lower linguistic level \(e\.g\., few grammar errors\) and are largely robust to variations in prompt fluency\. Details of the experiment are provided in Appendix[A\.7\.1](https://arxiv.org/html/2609.36214#A1.SS7.SSS1)\.

![Refer to caption](https://arxiv.org/html/2609.36214v1/prompt_fluency_vs_response_fluency_heatmap.png)Figure 3:The prompt’s qualities significantly influence the LLM’s response qualities, shown by regression coefficients of the prompt’s rhetorical/lexical quality levels \(cols\) on the response’s \(row\)\.
##### Response Rhetorical and Lexical Qualities

Do the rhetorical and lexical qualities of the prompt influence these qualities in the LLM’s response? To answer this, we score responses using the same LLM\-as\-judge and then, similar to Surface Form Features, we fit a linear regression to predict each of the response’s quality levels from the prompt’s fluency profile\.

LLM responses tend to mirror the rhetorical and lexical qualities of users’ prompts, as shown in Figure[3](https://arxiv.org/html/2609.36214#S5.F3)\. In particular, higher levels of Content, Organization, and Vocabulary show positive associations with response proficiency, whereas Language Use and Mechanics exhibit negative associations\. Thus, users who write with lower levels of Organization or less diverse Vocabulary receive LLM responses matching those qualities\. Perceived language proficiency is known to be associated with peer esteem\([Dev and Qiqieh, 2016](https://arxiv.org/html/2609.36214#bib.bib10)\)and positive social outcomes\([Pandey and Pandey, 2014](https://arxiv.org/html/2609.36214#bib.bib11), e\.g\., hiring; \)\. Our results suggest that users who are less fluent and attempting to use these models to make up for their gap in fluency will still suffer a penalty due to these LLMs’ propagation of errors\. Details of the experiment are provided in Appendix[A\.7\.2](https://arxiv.org/html/2609.36214#A1.SS7.SSS2)\.

#### 5\.2\.2Response Correctness

##### Response Quality

Figure 4:Higher prompt fluency leads to better response quality, with Content and Mechanics showing larger positive correlation than Vocabulary and Language Use\. Points represent regression coefficients, and error bars indicate 95% confidence intervals\.Is the correctness of the models’ output sensitive to the rhetorical and lexical qualities of the prompt, i\.e\., do models give less\-accurate outputs when the prompt is less fluent? To answer this question, we evaluate using the protocol described in Section[5\.1](https://arxiv.org/html/2609.36214#S5.SS1)and obtain an overall quality score for each response and analyze its relationship with prompt fluency using regression\.

Prompt fluency has a notable impact on model responses, more fluent prompts generally leading to higher quality outputs, as seen in Figure[4](https://arxiv.org/html/2609.36214#S5.F4)\. More specifically, a one\-point increase in a prompt fluency dimension on the four\-point scale is associated with an estimated 0\.04–0\.14 point increase in overall response quality on a five\-point scale\. For reference, a tenfold increase in model size \(going from a 1B to 10B model\) is associated with an approximately 0\.21 increase additional response\-quality points, which suggests less fluent LLM users experience substantially worse responses, akin to those given by much less capable models\. Among all fluency metrics, Content and Mechanics show a more notable impact on response quality compared to Vocabulary or Language Use\. One possible explanation is that Content and Mechanics capture how clearly a request is communicated, whereas richer vocabulary and more proficient grammar may provide smaller additional benefits when the intended request is already understandable\. Details of the experiment are provided in Appendix[A\.7\.3](https://arxiv.org/html/2609.36214#A1.SS7.SSS3)\.

##### Semantic Matched Pair Analysis

Figure 5:Different prompt fluency dimensions exhibit different impact on response quality in semantically matched comparisons\. Points represent regression coefficients relating within\-pair differences in each fluency dimension to differences in overall response quality\.To further examine the relationship between prompt fluency and response quality while reducing variation in semantic content, we conduct a complementary analysis using semantically matched pairs of Chinese prompts collected from real\-world user prompts\. For each pair, we translate two prompts with different fluency profiles and then regress differences in overall response quality on differences in the five prompt fluency dimensions\.

Even with similar semantic intent, more fluent prompts generally receive higher\-quality responses, as shown in Figure[5](https://arxiv.org/html/2609.36214#S5.F5)\. Content, Vocabulary, and Mechanics retain positive impact on response quality\. Organization has a coefficient close to zero, whereas Language Use shows a negative association\. These results further support our previous findings by being broadly consistent with the preceding analysis\. For prompts with similar semantic intent, simply improving the fluency level can significantly improve the response quality\. This suggests that refining how a request is expressed can help users obtain better responses in everyday interactions with LLMs\. Details of the experiment are provided in Appendix[A\.7\.4](https://arxiv.org/html/2609.36214#A1.SS7.SSS4)\.

##### Does prompt fluency impact the output of some tasks more?

![Refer to caption](https://arxiv.org/html/2609.36214v1/prompt_fluency_vs_response_quality_heatmap_by_task.png)Figure 6:Prompt fluency shows broadly consistent positive associations with response quality across task types, with variation in individual dimensions\. Colors indicate coefficients from separate regressions within each task\.Given that models produce less fluent and lower quality responses for less fluent prompts, here we assess whether this behavior is more pronounced in certain tasks\. We apply a similar regression analysis to measure the relationship between response quality and different prompt fluency dimensions by task type, shown in Figure[6](https://arxiv.org/html/2609.36214#S5.F6)\.

More fluent prompts generally receive higher\-quality responses across writing tasks, although the strength of this relationship varies by task types\. Specifically, “translation” exhibits relatively weak associations with Language Use and Mechanics, while “editing or critiquing provided text” and “personal writing or communication” show less sensitivity to Content and Vocabulary\. “Fiction writing” presents a notable exception to the generally positive pattern, with Vocabulary negatively associated with response quality\. Details of the experiment are provided in Appendix[A\.7\.5](https://arxiv.org/html/2609.36214#A1.SS7.SSS5)\.

Overall, our results suggest that prompt fluency has a substantial influence on model response quality\. More fluent prompts generally lead to responses with higher quality across different task types, although the magnitude of this effect varies slightly depending on the task\. Among different fluency dimensions, Content and Mechanics generally exhibit stronger influences on response quality\. These findings suggest that the effectiveness of prompts depends not only on their semantic intent, but also on how fluently and coherently users express their intentions\. More broadly, our results provide practical implications for real\-world interactions with LLMs, indicating that improving prompt fluency may serve as a simple yet effective strategy for obtaining high quality model responses\.

## 6Conclusion

We investigate how the fluency of non\-native English prompts shapes the responses that large language models produce\. To enable controlled experimentation, we curateFable, a collection of 190,911 writing task–related prompts that vary systematically in fluency while preserving semantic intent\. In experiments with responses from 34 popular open\-source models, we find that LLMs mirror the rhetorical and lexical qualities of the prompt, and their overall response quality is consistently lower for less fluent prompts\. More specifically, a one\-point increase in a retained fluency dimension on the four\-point scale corresponds to an estimated 0\.04–0\.14 point increase in response quality on the five\-point scale\. These effects are broadly consistent across writing tasks\. Taken together, our findings indicate that non\-native English\-speaking users of LLMs are likely disadvantaged through their degree of fluency in writing prompts\. This disparity in quality is driven less by surface mistakes than by the higher\-level discourse and lexical qualities of their prompts, and we hopeFableenables future work on diagnosing and mitigating this disparity\.

### AI use statement

In this work, we used generative AI tools for generating synthetic datasets, assisting with translation, cleaning and reformatting data, designing and providing feedback on research methodology and experiments, implementing methods, supporting qualitative and thematic data analysis, creation of artifacts, discovering research topics and identifying gaps, brainstorming, sourcing/searching for information, identifying relevant literature, summarizing and analyzing existing literature, and proposing a title and keywords for the research paper\. We have not used generative AI tools for help developing theoretical models or conceptual frameworks, formulating mathematical claims, providing critical ingredients for proving mathematical claims, assisting in the writing of proofs, proposing or refining hypotheses, interpreting results, formulating questions for surveys or interviews, creating or modifying scientific figures or images, suggesting experimental parameters, formatting references, suggesting a structure for the research paper, or transcribing recordings of research material\. Additionally, we used generative AI tools for creating and editing software code, drafting parts of this research paper, and editing to improve readability\. We have reviewed all AI\-assisted work, including manually reviewing model outputs and proposed edits to writing\. We take responsibility for the final content of this work, including text, claims or artifacts produced with the aid of generative AI\.

### Reproducibility statement

To support reproducibility of our results, we will release downloadable source code and theFabledata upon publication\. We also provide a complete description of data processing steps in the main text of the paper and Appendix, including prompts used for data synthesis and response rating\.

## References

- Caoet al\.\(2024\)B\. Cao, D\. Cai, Z\. Zhang, Y\. Zou, and W\. LamOn the worst prompt performance of large language models\.Advances in Neural Information Processing Systems37,pp\. 69022–69042\.Cited by:[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px3.p1.1)\.
- Chambers \(1997\)F\. ChambersWhat do we mean by fluency?\.System25\(4\),pp\. 535–544\.External Links:ISSN 0346\-251X,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/S0346-251X%2897%2900046-8),[Link](https://www.sciencedirect.com/science/article/pii/S0346251X97000468)Cited by:[§1](https://arxiv.org/html/2609.36214#S1.p1.1)\.
- Chatterjiet al\.\(2025\)A\. Chatterji, T\. Cunningham, D\. J\. Deming, Z\. Hitzig, C\. Ong, C\. Y\. Shan, and K\. WadmanHow people use chatgpt\.NBER Working PaperTechnical Report34255,National Bureau of Economic Research\.External Links:[Document](https://dx.doi.org/10.3386/w34255),[Link](https://www.nber.org/papers/w34255)Cited by:[Figure 1](https://arxiv.org/html/2609.36214#S3.F1),[§3](https://arxiv.org/html/2609.36214#S3.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2609.36214#S3.SS0.SSS0.Px2.p5.1)\.
- Chenet al\.\(2024\)J\. Chen, S\. Xiao, P\. Zhang, K\. Luo, D\. Lian, and Z\. LiuBge m3\-embedding: multi\-lingual, multi\-functionality, multi\-granularity text embeddings through self\-knowledge distillation\.arXiv preprint arXiv:2402\.032164\(5\)\.Cited by:[§A\.3](https://arxiv.org/html/2609.36214#A1.SS3.SSS0.Px3.p1.1)\.
- Chikkamathet al\.\(2024\)R\. Chikkamath, D\. Rastogi, M\. Maan, and M\. EndresIs your search query well\-formed? a natural query understanding for patent prior art search\.World Patent Information76,pp\. 102254\.Cited by:[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px1.p1.1)\.
- Dedemaet al\.\(2026\)M\. Dedema, C\. X\. Goh, and P\. ZhangWhen generative ai mixes languages: multilingual users’ code\-switching behavior in human\-llm interaction\.InProceedings of the Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems,CHI EA ’26,New York, NY, USA\.External Links:ISBN 9798400722813,[Link](https://doi.org/10.1145/3772363.3798492),[Document](https://dx.doi.org/10.1145/3772363.3798492)Cited by:[§3](https://arxiv.org/html/2609.36214#S3.SS0.SSS0.Px2.p4.1)\.
- DeepSeek\-AI \(2026\)DeepSeek\-AIDeepSeek\-V4\.1\-Flash: pushing the limits of KV cache compression\.External Links:2609\.19969,[Link](https://arxiv.org/abs/2609.19969)Cited by:[§3](https://arxiv.org/html/2609.36214#S3.SS0.SSS0.Px2.p5.1)\.
- Dev and Qiqieh \(2016\)S\. Dev and S\. QiqiehThe relationship between english language proficiency, academic achievement and self\-esteem of non\-native\-english\-speaking students\.\.International Education Studies9\(5\),pp\. 147–155\.Cited by:[§5\.2\.1](https://arxiv.org/html/2609.36214#S5.SS2.SSS1.Px2.p2.1)\.
- Dobariya and Kumar \(2025\)O\. Dobariya and A\. KumarMind your tone: investigating how prompt politeness affects llm accuracy \(short paper\)\.arXiv preprint arXiv:2510\.04950\.Cited by:[§1](https://arxiv.org/html/2609.36214#S1.p2.1),[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px3.p1.1)\.
- Don\-Yehiyaet al\.\(2025\)S\. Don\-Yehiya, L\. Choshen, and O\. AbendThe sharelm collection and plugin: contributing human\-model chats for the benefit of the community\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\),Vienna, Austria,pp\. 167–177\.Cited by:[§1](https://arxiv.org/html/2609.36214#S1.p3.1),[§3](https://arxiv.org/html/2609.36214#S3.p1.1)\.
- Durandardet al\.\(2025\)N\. Durandard, S\. Dhawan, and T\. PoibeauLanguage style matching in large language models\.InProceedings of the 26th Annual Meeting of the Special Interest Group on Discourse and Dialogue,pp\. 620–636\.Cited by:[§5\.2\.1](https://arxiv.org/html/2609.36214#S5.SS2.SSS1.p1.1)\.
- Flesch \(1948\)R\. FleschA new readability yardstick\.\.Journal of applied psychology32\(3\),pp\. 221\.Cited by:[§4\.1](https://arxiv.org/html/2609.36214#S4.SS1.p2.1)\.
- Guoet al\.\(2024\)Q\. Guo, L\. Wang, Y\. Wang, W\. Ye, and S\. ZhangWhat makes a good order of examples in in\-context learning\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 14892–14904\.External Links:[Link](https://aclanthology.org/2024.findings-acl.884/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.884)Cited by:[§1](https://arxiv.org/html/2609.36214#S1.p2.1),[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px3.p1.1)\.
- Heet al\.\(2024\)J\. He, M\. Rungta, D\. Koleczek, A\. Sekhon, F\. X\. Wang, and S\. HasanDoes prompt formatting have any impact on llm performance?\.arXiv preprint arXiv:2411\.10541\.Cited by:[§1](https://arxiv.org/html/2609.36214#S1.p2.1),[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px3.p1.1)\.
- Ishikawa \(2018\)S\. IshikawaThe icnale edited essays; a dataset for analysis of l2 english learner essays based on a new integrative viewpoint\.English Corpus Studies25,pp\. 117–130\.Cited by:[§4\.1](https://arxiv.org/html/2609.36214#S4.SS1.p3.1)\.
- Iveyet al\.\(2024\)J\. Ivey, S\. Kumar, J\. Liu, H\. Shen, S\. Rakshit, R\. Raju, H\. Zhang, A\. Ananthasubramaniam, J\. Kim, B\. Yi, D\. Wright, A\. Israeli, A\. G\. Møller, L\. Zhang, and D\. JurgensReal or robotic? assessing whether llms accurately simulate qualities of human responses in dialogue\.arXiv preprint arXiv:2409\.08330\.Cited by:[§3](https://arxiv.org/html/2609.36214#S3.SS0.SSS0.Px2.p1.1)\.
- Jacobset al\.\(1981\)H\. L\. Jacobs, S\. A\. Zinkgraf, D\. R\. Wormuth, V\. F\. Hartfiel, and J\. B\. HugheyTesting esl composition: a practical approach\.Newbury House\.Cited by:[§1](https://arxiv.org/html/2609.36214#S1.p3.1),[§4\.1](https://arxiv.org/html/2609.36214#S4.SS1.p3.1),[Table 1](https://arxiv.org/html/2609.36214#S4.T1)\.
- Joulinet al\.\(2017\)A\. Joulin, E\. Grave, P\. Bojanowski, and T\. MikolovBag of tricks for efficient text classification\.InProceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers,pp\. 427–431\.Cited by:[§3](https://arxiv.org/html/2609.36214#S3.SS0.SSS0.Px2.p3.1)\.
- Kaplanet al\.\(1966\)R\. B\. Kaplanet al\.Cultural thought patterns in inter\-cultural education\.Language learning16\(1\)\.Cited by:[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px2.p1.1)\.
- Kobayashi and Rinnert \(1992\)H\. Kobayashi and C\. RinnertEffects of first language on second language writing: translation versus direct composition\.Language learning42\(2\),pp\. 183–209\.Cited by:[§3](https://arxiv.org/html/2609.36214#S3.SS0.SSS0.Px2.p4.1)\.
- Kwonet al\.\(2023\)W\. Kwon, Z\. Li, S\. Zhuang, Y\. Sheng, L\. Zheng, C\. H\. Yu, J\. E\. Gonzalez, H\. Zhang, and I\. StoicaEfficient memory management for large language model serving with pagedattention\.InProceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles,Cited by:[§A\.8](https://arxiv.org/html/2609.36214#A1.SS8.p1.1)\.
- LanguageTool contributors \(2026\)LanguageTool contributorsLanguageTool: style and grammar checker\.Note:[https://github\.com/languagetool\-org/languagetool](https://github.com/languagetool-org/languagetool)Version 6\.8, accessed 2026\-09\-26Cited by:[§4\.1](https://arxiv.org/html/2609.36214#S4.SS1.p2.1)\.
- Leidingeret al\.\(2023\)A\. Leidinger, R\. Van Rooij, and E\. ShutovaThe language of prompting: what linguistic properties make a prompt successful?\.InFindings of the association for computational linguistics: EMNLP 2023,pp\. 9210–9232\.Cited by:[§1](https://arxiv.org/html/2609.36214#S1.p2.1)\.
- Liuet al\.\(2024\)C\. Liu, L\. Yu, J\. Li, R\. Jin, Y\. Huang, L\. Shi, J\. Zhang, X\. Ji, T\. Cui, T\. Liu, J\. Song, H\. Zan, S\. Li, and D\. XiongOpenEval: benchmarking Chinese LLMs across capability, alignment and safety\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\),Y\. Cao, Y\. Feng, and D\. Xiong \(Eds\.\),Bangkok, Thailand,pp\. 190–210\.External Links:[Link](https://aclanthology.org/2024.acl-demos.19/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-demos.19)Cited by:[§3](https://arxiv.org/html/2609.36214#S3.SS0.SSS0.Px2.p4.1)\.
- Liuet al\.\(2025a\)J\. Liu, Y\. He, Z\. Zheng, Y\. Bu, and C\. NiAI\-assisted writing is growing fastest among non\-english\-speaking and less established scientists\.arXiv preprint arXiv:2511\.15872\.Cited by:[§1](https://arxiv.org/html/2609.36214#S1.p1.1)\.
- Liuet al\.\(2025b\)Q\. Liu, W\. Wang, and J\. WillardEffects of prompt length on domain\-specific tasks for large language models\.External Links:2502\.14255,[Link](https://arxiv.org/abs/2502.14255)Cited by:[§1](https://arxiv.org/html/2609.36214#S1.p2.1),[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px3.p1.1)\.
- Liuet al\.\(2023\)Y\. Liu, D\. Iter, Y\. Xu, S\. Wang, R\. Xu, and C\. ZhuG\-eval: nlg evaluation using gpt\-4 with better human alignment\.InProceedings of the 2023 conference on empirical methods in natural language processing,pp\. 2511–2522\.Cited by:[§5\.1](https://arxiv.org/html/2609.36214#S5.SS1.p2.1)\.
- Luet al\.\(2022\)Y\. Lu, M\. Bartolo, A\. Moore, S\. Riedel, and P\. StenetorpFantastically ordered prompts and where to find them: overcoming few\-shot prompt order sensitivity\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 8086–8098\.Cited by:[§1](https://arxiv.org/html/2609.36214#S1.p2.1),[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px3.p1.1)\.
- Miłkowski \(2010\)M\. MiłkowskiDeveloping an open\-source, rule\-based proofreading tool\.Software: Practice and Experience40\(7\),pp\. 543–566\.External Links:[Document](https://dx.doi.org/10.1002/spe.971)Cited by:[§4\.1](https://arxiv.org/html/2609.36214#S4.SS1.p2.1)\.
- Panet al\.\(2024\)L\. Pan, Y\. Leng, and D\. XiongCan large language models learn translation robustness from noisy\-source in\-context demonstrations?\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),pp\. 2798–2808\.Cited by:[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px1.p1.1)\.
- Pandey and Pandey \(2014\)M\. Pandey and P\. PandeyBetter english for better employment opportunities\.International journal of multidisciplinary approach and studies1\(4\),pp\. 93–100\.Cited by:[§5\.2\.1](https://arxiv.org/html/2609.36214#S5.SS2.SSS1.Px2.p2.1)\.
- Park and Padó \(2024\)D\. Park and S\. PadóMulti\-dimensional machine translation evaluation: model evaluation and resource for korean\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),pp\. 11723–11744\.Cited by:[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px1.p1.1)\.
- Paszkeet al\.\(2019\)A\. Paszke, S\. Gross, F\. Massa, A\. Lerer, J\. Bradbury, G\. Chanan, T\. Killeen, Z\. Lin, N\. Gimelshein, L\. Antiga, A\. Desmaison, A\. Köpf, E\. Yang, Z\. DeVito, M\. Raison, A\. Tejani, S\. Chilamkurthy, B\. Steiner, L\. Fang, J\. Bai, and S\. ChintalaPyTorch: an imperative style, high\-performance deep learning library\.InAdvances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8\-14, 2019, Vancouver, BC, Canada,H\. M\. Wallach, H\. Larochelle, A\. Beygelzimer, F\. d’Alché\-Buc, E\. B\. Fox, and R\. Garnett \(Eds\.\),pp\. 8024–8035\.External Links:[Link](https://proceedings.neurips.cc/paper/2019/hash/bdbca288fee7f92f2bfa9f7012727740-Abstract.html)Cited by:[§A\.8](https://arxiv.org/html/2609.36214#A1.SS8.p1.1)\.
- Popovićet al\.\(2024\)M\. Popović, E\. Lapshinova\-Koltunski, and M\. KoponenEffects of different types of noise in user\-generated reviews on human and machine translations including chatgpt\.InProceedings of the Ninth Workshop on Noisy and User\-Generated Text \(W\-NUT 2024\),pp\. 17–30\.Cited by:[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px1.p1.1)\.
- Qinet al\.\(2024\)L\. Qin, Q\. Chen, Y\. Zhou, Z\. Chen, Y\. Li, L\. Liao, M\. Li, W\. Che, and P\. S\. YuMultilingual large language model: a survey of resources, taxonomy and frontiers\.External Links:2404\.04925,[Link](https://arxiv.org/abs/2404.04925)Cited by:[§3](https://arxiv.org/html/2609.36214#S3.SS0.SSS0.Px2.p4.1)\.
- Qwen Team \(2026\)Qwen TeamQwen3\.8\-Max: a new bar for coding and cowork\.External Links:[Link](https://qwen.ai/blog?id=qwen3.8)Cited by:[§5\.1](https://arxiv.org/html/2609.36214#S5.SS1.p2.1)\.
- Rabinovichet al\.\(2016\)E\. Rabinovich, S\. Nisioi, N\. Ordan, and S\. WintnerOn the similarities between native, non\-native and translated texts\.InProceedings of the 54th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 1870–1881\.Cited by:[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px2.p1.1)\.
- Reusenset al\.\(2025\)M\. Reusens, P\. Borchert, J\. De Weerdt, and B\. BaesensNative design bias: studying the impact of english nativeness on language model performance\.InProceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics,pp\. 1195–1215\.Cited by:[§1](https://arxiv.org/html/2609.36214#S1.p1.1),[§1](https://arxiv.org/html/2609.36214#S1.p2.1),[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px2.p1.1)\.
- Sclaret al\.\(2023\)M\. Sclar, Y\. Choi, Y\. Tsvetkov, and A\. SuhrQuantifying language models’ sensitivity to spurious features in prompt design or: how i learned to start worrying about prompt formatting\.arXiv preprint arXiv:2310\.11324\.Cited by:[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px3.p1.1)\.
- Siekmannet al\.\(2022\)L\. Siekmann, J\. M\. Parr, and V\. BusseStructure and coherence as challenges in composition: a study of assessing less proficient efl writers’ text quality\.Assessing Writing54,pp\. 100672\.Cited by:[§4\.1](https://arxiv.org/html/2609.36214#S4.SS1.p1.1)\.
- Witte and Faigley \(1981\)S\. P\. Witte and L\. FaigleyCoherence, cohesion, and writing quality\.College Composition & Communication32\(2\),pp\. 189–204\.Cited by:[§4\.1](https://arxiv.org/html/2609.36214#S4.SS1.p1.1)\.
- Wolfet al\.\(2020\)T\. Wolf, L\. Debut, V\. Sanh, J\. Chaumond, C\. Delangue, A\. Moi, P\. Cistac, T\. Rault, R\. Louf, M\. Funtowicz, J\. Davison, S\. Shleifer, P\. von Platen, C\. Ma, Y\. Jernite, J\. Plu, C\. Xu, T\. Le Scao, S\. Gugger, M\. Drame, Q\. Lhoest, and A\. RushTransformers: state\-of\-the\-art natural language processing\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,Q\. Liu and D\. Schlangen \(Eds\.\),Online,pp\. 38–45\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-demos.6),[Link](https://aclanthology.org/2020.emnlp-demos.6)Cited by:[§A\.8](https://arxiv.org/html/2609.36214#A1.SS8.p1.1)\.
- Yang \(2006\)J\. YangLearners and users of english in china\.English today22\(2\),pp\. 3–10\.Cited by:[§1](https://arxiv.org/html/2609.36214#S1.p3.1)\.
- Yinet al\.\(2024\)Z\. Yin, H\. Wang, K\. Horio, D\. Kawahara, and S\. SekineShould we respect llms? a cross\-lingual study on the influence of prompt politeness on llm performance\.InProceedings of the Second Workshop on Social Influence in Conversations \(SICon 2024\),pp\. 9–35\.Cited by:[§1](https://arxiv.org/html/2609.36214#S1.p2.1),[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px3.p1.1)\.
- Zhaoet al\.\(2024\)W\. Zhao, X\. Ren, J\. Hessel, C\. Cardie, Y\. Choi, and Y\. DengWildChat: 1m chatgpt interaction logs in the wild\.External Links:2405\.01470,[Link](https://arxiv.org/abs/2405.01470)Cited by:[§1](https://arxiv.org/html/2609.36214#S1.p3.1),[§3](https://arxiv.org/html/2609.36214#S3.p1.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing,et al\.Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.Advances in neural information processing systems36,pp\. 46595–46623\.Cited by:[§5\.1](https://arxiv.org/html/2609.36214#S5.SS1.p2.1)\.
- Zhuang and Zuccon \(2022\)S\. Zhuang and G\. ZucconCharacterbert and self\-teaching for improving the robustness of dense retrievers on queries with typos\.InProceedings of the 45th international ACM SIGIR conference on research and development in information retrieval,pp\. 1444–1454\.Cited by:[§2](https://arxiv.org/html/2609.36214#S2.SS0.SSS0.Px1.p1.1)\.

## Appendix AAppendix

### A\.1Translated Prompt Synthesizing Details

We use the below prompt to synthesize prompt translations according to specific fluency levels in the ESL Composition Profile\.

CONTENT\_SPECS=\{

4:”Contentisrichandfullydevelopedwithspecific,relevantdetailsthatclearlyaddressthetopic\.”,

3:”Contentshowsadequateknowledgeofthetopicandgenerallyrelevantideas,thoughsomepointslackdepth\.”,

2:”Contentshowslimitedknowledgeofthetopic,withfewsupportingdetailsandunderdevelopedideas\.”,

1:”Contentshowsverylittlegraspofthetopic;ideasarevague,off\-topic,orinsufficienttoevaluate\.”\}

ORGANIZATION\_SPECS=\{

4:”Ideasareclearlyorganizedwithlogicalsequencing,clearparagraphing,andstrongcohesion\.”,

3:”Organizationisgenerallyclear,thoughsomepartsmaybelooselyconnectedoruneveninsequencing\.”,

2:”Organizationisweak;ideasaresometimesconfusedordisconnectedandlackclearprogression\.”,

1:”Organizationisminimal;ideasarehardtofollowandoverallstructureisunclearormissing\.”\}

VOCAB\_SPECS=\{

4:”Usesawiderangeofvocabularywithaccurateandappropriatewordchoiceandidiomaticexpressions\.”,

3:”Usesanadequaterangeofvocabularywithoccasionalword\-choiceerrors,butmeaningisusuallyclear\.”,

2:”Usesalimitedrangeofvocabularywithfrequentword\-choiceproblems;meaningissometimesobscured\.”,

1:”Vocabularyisverylimited,withfrequentinappropriateorincorrectwordchoices;meaningisoftenunclear\.”\}

LANGUSE\_SPECS=\{

4:”Useseffectivesimpleandcomplexsentenceconstructionswithfewerrorsintense,agreement,andwordorder\.”,

3:”Usesmostlycorrectsimplesentencesandsomecomplexones;grammarerrorsoccurbutrarelyblockunderstanding\.”,

2:”Showsmajorproblemswithgrammarandsentencestructure;errorsarefrequentandsometimesobscuremeaning\.”,

1:”Hasverylittlecontrolofsentenceconstruction;grammarerrorsdominateandcommunicationoftenbreaksdown\.”\}

MECH\_SPECS=\{

4:”Demonstratesmasteryofspelling,punctuation,capitalization,andparagraphingwithalmostnoerrors\.”,

3:”Showsoccasionalerrorsinspelling,punctuation,capitalization,orparagraphing,butmeaningisnotobscured\.”,

2:”Showsfrequenterrorsinspelling,punctuation,capitalization,orparagraphing;meaningmaybeaffected\.”,

1:”Showspervasiveerrorsinspelling,punctuation,capitalization,andparagraphing;textishardtoread\.”\}

user\_message=f”’

Whenproducingtext,youmustsimulatethefollowingESLcompositionprofile:

Content\(level\):\{content\_desc\}

Organization\(level\):\{org\_desc\}

Vocabulary\(level\):\{vocab\_desc\}

Languageuse\(level\):\{lang\_desc\}

Mechanics\(level\):\{mech\_desc\}

TranslatethefollowingquestionintoEnglishthatmatchesthisprofile:\{original\_prompt\}

”’

### A\.2Translation Examples

Table[2](https://arxiv.org/html/2609.36214#A1.T2)shows examples of the same prompt translated with different uniform rating levels for each quality in instructions \(e\.g\., all rated 4\)\. Prompts retain much of their semantic equivalence, while prompts from instructions with lower ratings have visibly more errors\.

Table 2:Example translations generated from the same Chinese prompt under different instructed levels\. While the semantic intent remains largely preserved, the fluency shows a notable gap across different levels\.
### A\.3Evaluation of Translated Prompts

##### Surface Form Features

Higher instructed proficiency levels generally yield fewer linguistic errors and improved readability as shown in Figure[7](https://arxiv.org/html/2609.36214#A1.F7)\. In particular, higher Mechanics levels are associated with fewer spelling and grammar–syntax errors, while higher Language Use levels are associated with improved readability\. These relationships are not uniform: higher Vocabulary levels are associated with more detected spelling errors\.

![Refer to caption](https://arxiv.org/html/2609.36214v1/predefined_level_vs_linguistic_error.png)Figure 7:Standardized regression coefficients relating instructed proficiency levels to surface form features of translated prompts\. Negative coefficients indicate fewer errors for error metrics; positive coefficients indicate greater reading ease for readability\. Stars denote statistical significance: \*p<0\.1p<0\.1, \*\*p<0\.05p<0\.05, \*\*\*p<0\.01p<0\.01\.
##### Rhetorical and Lexical Qualities

Higher instructed levels generally correspond to higher judged proficiency scores as shown in Figure[8](https://arxiv.org/html/2609.36214#A1.F8)\. The instructed level of a dimension generally has the strongest association with its corresponding measured score, particularly for Vocabulary, Language Use, and Mechanics\. Cross\-dimensional associations are also present, indicating that the instructions produce targeted fluency variation without fully independent control of the five dimensions\.

![Refer to caption](https://arxiv.org/html/2609.36214v1/predefined_level_vs_actual_fluency_heatmap.png)Figure 8:Standardized regression coefficients relating the five instructed proficiency levels to judged proficiency scores of translated prompts\. Results show both targeted and cross\-dimensional associations\. Stars denote statistical significance: \*p<0\.1p<0\.1, \*\*p<0\.05p<0\.05, \*\*\*p<0\.01p<0\.01\.
##### Semantic Preservation

Cross\-lingual cosine similarity, measured usingBAAI/bge\-m3\([Chen et al\., 2024](https://arxiv.org/html/2609.36214#bib.bib44)\), varies little with instructed fluency\. The Spearman correlation between the average instructed level and similarity to the original Chinese prompt is 0\.06\.

Human evaluation of 100 randomly sampled translations yields a mean semantic\-preservation score of 4\.13 on a five\-point scale, where 1 denotes severe information loss and 5 denotes no loss\. This suggests that the translations generally preserve the original semantics with limited information loss on average\. Additionally, Pearson correlations between individual instructed levels and the same\-intent score are presented in Table[3](https://arxiv.org/html/2609.36214#A1.T3)\. Together, these results support generally preserved intent and little association between instructed fluency and semantic preservation, while not ruling out semantic changes in individual translations\.

Table 3:Pearson correlations between instructed proficiency levels and the same\-intent score\.

### A\.4Evaluated Models

Table[4](https://arxiv.org/html/2609.36214#A1.T4)lists the 34 models evaluated in our experiments\. All models receive the same set of 5,000 English prompt variants randomly sampled fromFable, and their responses are assessed using the same evaluation pipeline\.

Table 4:The 34 evaluated models, grouped by model family\. Version and release identifiers are retained where applicable\. IT denotes instruction\-tuned models\.
### A\.5Fluency Evaluation

We use the below prompt to evaluate the fluency of LLM responses using the ESL Composition Profile\.

user\_message=f”’

Evaluatewritingstrictlyaccordingtotheprovidedrubric:\{rubric\}\.

Evaluatethefollowingtexton5dimensions:Content,Organization,Vocabulary,LanguageUse,Mechanics\.

Eachscoremustbeanintegerfrom1to4\.

”’

### A\.6Response Quality Evaluation

We use the below prompt to evaluate the quality of LLM responses to instructions of varying fluency\.

YouareanimpartialevaluatorofAIassistantresponses\.

Evaluatetheassistantresponserelativetotheuser’squestion\.

Scoreeachevaluationdimensionindependently\.

Beobjective,consistent,andappropriatelycritical\.

Donotletthescoreforonedimensionautomaticallydetermineanotherdimension\.

OutputonlyvalidJSON\.

Foreachdimension,provideabriefjustificationbeforeassigningthescore\.

Eachjustificationmustbeconciseandnomorethan30words\.

User message

Evaluatetheassistantresponseonthefollowingfivedimensions:

1\.Helpfulness:Howusefultheresponseisinhelpingtheuserachievetheirgoal\.

2\.Relevance:Howdirectlytheresponseaddressestheuser’squestionwithoutunnecessaryoroff\-topiccontent\.

3\.Accuracy:Howfactuallycorrectandlogicallysoundtheresponseis\.

4\.Creativity:Howeffectivelytheresponsedemonstratesoriginality,insight,orcreativeproblem\-solvingwhenappropriate\.

5\.LevelOfDetail:Whethertheresponseprovidesanappropriateamountofspecificandusefuldetail\.

AlsoprovideanOverallscorereflectingtheoverallqualityoftheresponse\.

TheOverallscoreshouldbeaholisticjudgmentandshouldnotsimplybethearithmeticaverageofthefivedimensionscores\.

ForeverydimensionandtheOverallscore,usethesameratingscale:

1=VeryPoor

2=Poor

3=Average

4=Good

5=VeryGood

Evaluateeachdimensionindependentlybasedonlyonthatdimension\.

OutputexactlyinthefollowingJSONformat:

\{

"Helpfulness":\{

"reason":"briefjustification",

"score":1

\},

"Relevance":\{

"reason":"briefjustification",

"score":1

\},

"Accuracy":\{

"reason":"briefjustification",

"score":1

\},

"Creativity":\{

"reason":"briefjustification",

"score":1

\},

"LevelOfDetail":\{

"reason":"briefjustification",

"score":1

\},

"Overall":\{

"reason":"briefjustification",

"score":1

\}

\}

\[Question\]

\{reference\_translation\}

\[TheStartofAssistant’sAnswer\]

\{model\_response\}

\[TheEndofAssistant’sAnswer\]

### A\.7Additional Experimental Details

#### A\.7\.1Response Surface Form Features

For each response surface form measure, we estimate

Y=α\+βC​C\+βV​V\+βL​L\+βM​M\+𝜸⊤​𝐗\+ϵ,Y=\\alpha\+\\beta\_\{C\}C\+\\beta\_\{V\}V\+\\beta\_\{L\}L\+\\beta\_\{M\}M\+\\bm\{\\gamma\}^\{\\top\}\\mathbf\{X\}\+\\epsilon,whereCC,VV,LL, andMMdenote prompt Content, Vocabulary, Language Use, and Mechanics scores\.𝐗\\mathbf\{X\}includes log\-transformed model size, prompt length, response length, and model\-family fixed effects\. The outcome and all continuous predictors are standardized\. We exclude Organization because its VIF exceeded 5, to reduce collinearity\. The responses used in this experiment were generated by 34 models mentioned in Appendix[A\.4](https://arxiv.org/html/2609.36214#A1.SS4)\.

#### A\.7\.2Rhetorical and Lexical Quality of Response

We fit a separate regression for each metric of response rhetorical and lexical quality:

Yk=αk\+βk,C​C\+βk,O​O\+βk,V​V\+βk,L​L\+βk,M​M\+ϵk,Y\_\{k\}=\\alpha\_\{k\}\+\\beta\_\{k,C\}C\+\\beta\_\{k,O\}O\+\\beta\_\{k,V\}V\+\\beta\_\{k,L\}L\+\\beta\_\{k,M\}M\+\\epsilon\_\{k\},whereYkY\_\{k\}is the response score for metrickk, andCC,OO,VV,LL, andMMare the prompt’s Content, Organization, Vocabulary, Language Use, and Mechanics scores, respectively\. The responses used in this experiment were generated by Qwen3\-14B, Llama\-3\.1\-8B, and GPT\-OSS\-20B\.

#### A\.7\.3Response Quality

We fit an OLS regression:

Q=α\+βC​C\+βV​V\+βL​L\+βM​M\+βS​log10⁡\(S\)\+𝜸⊤​𝐗\+ϵ,Q=\\alpha\+\\beta\_\{C\}C\+\\beta\_\{V\}V\+\\beta\_\{L\}L\+\\beta\_\{M\}M\+\\beta\_\{S\}\\log\_\{10\}\(S\)\+\\bm\{\\gamma\}^\{\\top\}\\mathbf\{X\}\+\\epsilon,whereQQis overall response quality on a five\-point scale;CC,VV,LL, andMMare the prompt’s Content, Vocabulary, Language Use, and Mechanics scores on a four\-point scale; andSSis model size in billions of parameters\.𝐗\\mathbf\{X\}includes prompt length, response length, and model\-family fixed effects\. We exclude Organization to reduce collinearity because we identify its VIF\>5\>5\. The responses used in this experiment were generated by 34 models mentioned in Appendix[A\.4](https://arxiv.org/html/2609.36214#A1.SS4)\.

#### A\.7\.4Semantically Matched Pair Analysis

For semantically matched prompt pairs, we fit:

Δ​Q=α\+βC​Δ​C\+βO​Δ​O\+βV​Δ​V\+βL​Δ​L\+βM​Δ​M\+ϵ,\\Delta Q=\\alpha\+\\beta\_\{C\}\\Delta C\+\\beta\_\{O\}\\Delta O\+\\beta\_\{V\}\\Delta V\+\\beta\_\{L\}\\Delta L\+\\beta\_\{M\}\\Delta M\+\\epsilon,whereΔ\\Deltadenotes the difference within each pair,QQis overall response quality, andCC,OO,VV,LL, andMMare the five prompt fluency scores\. We estimate the regression using OLS on unstandardized differences and report 95% confidence intervals\. The responses used in this experiment were generated by Qwen3\-14B, Llama\-3\.1\-8B, and GPT\-OSS\-20B\.

#### A\.7\.5Analysis by Task Type

For each task categorytt, we fit a separate OLS regression:

Q=αt\+βt,C​C\+βt,V​V\+βt,L​L\+βt,M​M\+𝜸t⊤​𝐗\+ϵ,Q=\\alpha\_\{t\}\+\\beta\_\{t,C\}C\+\\beta\_\{t,V\}V\+\\beta\_\{t,L\}L\+\\beta\_\{t,M\}M\+\\bm\{\\gamma\}\_\{t\}^\{\\top\}\\mathbf\{X\}\+\\epsilon,whereQQis overall response quality;CC,VV,LL, andMMare the four retained prompt fluency scores; and𝐗\\mathbf\{X\}contains prompt and response lengths\. All variables retain their original scales and we exclude Organization to reduce collinearity because we identify its VIF\>5\>5\. The responses used in this experiment were generated by 34 models mentioned in Appendix[A\.4](https://arxiv.org/html/2609.36214#A1.SS4)\.

### A\.8Computational Environment

We run all our experiments on 4 NVIDIA\-L40S\-48GB GPUs\. All LLM inferences are powered by vLLM 0\.5\.4[Kwon et al\. \(2023\)](https://arxiv.org/html/2609.36214#bib.bib14), Huggingface Transformers 4\.43\.3[Wolf et al\. \(2020\)](https://arxiv.org/html/2609.36214#bib.bib15)and PyTorch 2\.4\.0[Paszke et al\. \(2019\)](https://arxiv.org/html/2609.36214#bib.bib16)on a CUDA 12\.4 environment\. Temperatures are set to 0\.0 to minimize the effect of randomness\.

### A\.9Licenses

All data and code will be publicly released under the CC BY\-SA 4\.0 license\.

Similar Articles

When English Rewrites Local Knowledge: Global Narrative Dominance in Large Language Models

arXiv cs.CL

This paper introduces CulturalNB, a dataset of Bengali cultural question-answer pairs, and evaluates nine LLMs for cross-lingual cultural bias. Findings show that English prompting increases global narrative substitution and reduces local perspectives, revealing that cultural failures in LLMs are grounding and prioritization issues, not just missing knowledge.