How Robust Is Multimodal Claim Verification to LLM Rewriting?

arXiv cs.CL Papers

Summary

This paper studies how multimodal claim verification models respond to stylistic text changes induced by LLM rewriting. Evaluating 11 open-weight VLMs (2B–38B), the authors find accuracy is largely robust to natural rewriting and controlled LLM-word injection, though hedging-oriented modifications cause consistent probability shifts across nearly all models.

arXiv:2610.02841v1 Announce Type: new Abstract: LLMs are known to introduce stylistic changes into generated text, yet how these stylistic shifts affect model decisions on scientific tasks remains underexplored. In this paper, we focus on multimodal claim verification, where the goal is to determine whether a textual claim is grounded in a given piece of evidence. We apply two rewriting strategies: natural rewriting, which simulates how researchers routinely use LLMs to polish academic text, and controlled injection, which inserts a single LLM-associated word to isolate the effect of vocabulary choice. We evaluate 11 open-weight models spanning five VLM families and ranging from 2B to 38B parameters. We find that models are robust to these modifications: most show no significant drop in accuracy, and compared to prior work on review-score manipulation, verification appears far more stable. However, consistent probability shifts do occur. Hedging-oriented conditions produce significant shifts across nearly all models, while boosting conditions show a weaker effect and general polishing conditions (e.g., grammar correction, fluency improvement) have little effect.
Original Article
View Cached Full Text

Cached at: 10/05/26, 10:03 AM

# How Robust Is Multimodal Claim Verification to LLM Rewriting?
Source: [https://arxiv.org/html/2610.02841](https://arxiv.org/html/2610.02841)
Xanh HoAffiliation:NII LLMC, JapanEmail:[xanh@nii\.ac\.jp](mailto:[email protected])Andre Greiner\-PetterAffiliation:NII LLMC, JapanEmail:[aizawa@nii\.ac\.jp](mailto:[email protected])Sunisth KumarAffiliation:National Institute of Informatics, JapanAffiliation:University of Göttingen, GermanyEmail:[greinerpetter@gipplab\.org](mailto:[email protected])Tian Cheng XiaEmail:[tiancheng\.xia@studio\.unibo\.it](mailto:[email protected])Florian BoudinAffiliation:University of Bologna, ItalyEmail:[sunisth@g\.ecc\.u\-tokyo\.ac\.jp](mailto:[email protected])Email:[florian\.boudin@univ\-nantes\.fr](mailto:[email protected])Akiko AizawaAffiliation:NII LLMC, JapanAffiliation:National Institute of Informatics, JapanAffiliation:The University of Tokyo, JapanAffiliation:Inria, LS2N, Nantes Université, France

###### Abstract

LLMs are known to introduce stylistic changes into generated text, yet how these stylistic shifts affect model decisions on scientific tasks remains underexplored\. In this paper, we focus on multimodal claim verification, where the goal is to determine whether a textual claim is grounded in a given piece of evidence\. We apply two rewriting strategies: natural rewriting, which simulates how researchers routinely use LLMs to polish academic text, and controlled injection, which inserts a single LLM\-associated word to isolate the effect of vocabulary choice\. We evaluate 11 open\-weight models spanning five VLM families and ranging from 2B to 38B parameters\. We find that models are robust to these modifications: most show no significant drop in accuracy, and compared to prior work on review\-score manipulation, verification appears far more stable\. However, consistent probability shifts do occur\. Hedging\-oriented conditions produce significant shifts across nearly all models, while boosting conditions show a weaker effect and general polishing conditions \(e\.g\., grammar correction, fluency improvement\) have little effect\.111[https://github\.com/yunangwu/llm\-rewrite\-claim\-verification](https://github.com/yunangwu/llm-rewrite-claim-verification)

## 1Introduction

Figure 1:A claim verification task requires deciding whether a claim is supported or refuted by evidence\. This figure shows how a rational decision\-maker should adjustp⁡\(support\)p\(\\text\{support\}\)when the claim’s intensity modifier changes\. A stronger modifier \(e\.g\., “considerably” to “dramatically”\) makes a supported claim harder to verify and a refuted claim easier to refute, both loweringp⁡\(support\)p\(\\text\{support\}\)\. However, this reasoning only applies when the modification is semantically meaningful\. Our central question: when the modification is stylistic and does not alter the claim’s semantic content, does the model still exhibit the same behavioral shift?LLMs are quietly shaping how scientific papers are written and reviewed\. Several studies\([Kobak et al\., 2025](https://arxiv.org/html/2610.02841#bib.bib9);[Matsui, 2025](https://arxiv.org/html/2610.02841#bib.bib14);[Cunningham et al\., 2025](https://arxiv.org/html/2610.02841#bib.bib4)\)report that the quality and style of scientific writing have changed since the appearance of ChatGPT\. At the same time, conferences such as ICLR have started exploring LLM\-assisted review pipelines\([Thakkar et al\., 2025](https://arxiv.org/html/2610.02841#bib.bib19)\), suggesting that LLMs may increasingly be involved in evaluating text that LLMs helped produce\.

These stylistic shifts may have consequences beyond presentation, affecting how downstream systems evaluate scientific text\. Indeed, recent work shows that adversarial paraphrasing and textual adversarial attacks can inflate paper\-level review scores\([Kaneko, 2026](https://arxiv.org/html/2610.02841#bib.bib8);[Lin et al\., 2025](https://arxiv.org/html/2610.02841#bib.bib13)\)\. But paper\-level reviewing conflates many factors \(e\.g\., argument structure, novelty, presentation\), making it hard to isolate what drives the effect\. We instead study single\-claim verification against visual evidence, where the only variable is how the claim is worded\.

We study this through multimodal claim verification\([Ho et al\., 2026](https://arxiv.org/html/2610.02841#bib.bib6);[Wang et al\., 2025](https://arxiv.org/html/2610.02841#bib.bib22);[Lal et al\., 2025](https://arxiv.org/html/2610.02841#bib.bib10)\), where a textual claim is checked against a table or figure from a scientific paper\. Using SciClaimEval\([Ho et al\., 2026](https://arxiv.org/html/2610.02841#bib.bib6)\), we rewrite claims under two strategies:natural rewriting, which paraphrases claims using LLM prompts, andcontrolled injection, which inserts individual words into otherwise unchanged claims\. Natural rewriting simulates five scenarios that reflect how researchers routinely use LLMs \(e\.g\., grammar correction, fluency improvement\)\. Controlled injection isolates the effect of single words, showing that even one vocabulary insertion can shift a model’s prediction\.

We measure how model confidence changes by modelingp⁡\(support\)p\(\\text\{support\}\)as the average of 10 runs using self\-consistency\([Wang et al\., 2023](https://arxiv.org/html/2610.02841#bib.bib23)\)and test on 11 models across 5 model families\. We find that models are robust to stylistic change in terms of accuracy, regardless of family or size\. On the other hand, we observe meaningful probability shifts in a consistent direction under different rewriting strategies\.

ConditionClaimOriginalAdditionally, 46\.6% of participants were primary school graduates, 9\.4% reported alcohol use, and 13% were smokers, with an average duration of smoking of 22\.83 ± 6\.94 years\.Natural RewritingLanguage Polishing \(LP\)Additionally,46\.6 %of participants wereprimary\-schoolgraduates,9\.4 %reported alcohol use, and13 %were smokers, with an averagesmokingduration of 22\.83 ± 6\.94 years\.Fluency Improvement \(FI\)Moreover, 46\.6 %oftheparticipantshad completedprimaryschool; 9\.4 %reportedconsuming alcohol;and13 %were smokers, with an average smokinghistoryof 22\.83 ± 6\.94 years\.Academic Style Alignment \(AS\)Furthermore, 46\.6 %of participantshad completedprimaryeducation, 9\.4 %reported alcoholconsumption,and13 %were smokers, witha mean smokingduration of 22\.83 ± 6\.94 years\.Outcome\-Aware \(OA\)Notably, 46\.6 %ofthe cohortwereprimary\-schoolgraduates,9\.4 %reported alcoholconsumption,and13 %were smokers, witha meansmokinghistoryof 22\.83 ± 6\.94 years\.Cautious Tone \(CT\)Additionally,approximately46\.6% of participantshad completedprimaryschool, about9\.4% reported alcoholconsumption,androughly13%identified assmokers, witha reportedaveragesmokingduration of 22\.83 ± 6\.94 years\.Controlled InjectionConservative \(C1\)Additionally,approximately46\.6% of participants were primary school graduates, 9\.4% reported alcohol use, and 13% were smokers, with an average duration of smoking of 22\.83 ± 6\.94 years\.Decorative \(D1\)Additionally,noteworthy46\.6% of participants were primary school graduates, 9\.4% reported alcohol use, and 13% were smokers, with an average duration of smoking of 22\.83 ± 6\.94 years\.Table 1:Rewrite examples for claimval\_tab\_0484from SciClaimEval\([Ho et al\., 2026](https://arxiv.org/html/2610.02841#bib.bib6)\)\. Natural rewriting paraphrases claims using LLM prompts that simulate common academic editing scenarios, while controlled injection inserts a single LLM\-associated word into the original claim\.The core contributions of this paper are listed as follows:

- •A systematic empirical study of stylistic sensitivity in multimodal claim verification\. We evaluate 11 open\-weight VLMs across five families under two rewriting strategies \(natural LLM paraphrasing and controlled single\-word injection\), controlling for stylistic variation as the primary experimental variable\.
- •We show that several rewriting conditions, especially cautious\-tone and conservative\-word injection, induce consistent probability shifts even when accuracy remains stable, revealing a more nuanced vulnerability than what has been previously documented for other scientific tasks such as paper reviewing\([Kaneko, 2026](https://arxiv.org/html/2610.02841#bib.bib8);[Lin et al\., 2025](https://arxiv.org/html/2610.02841#bib.bib13)\)\.

## 2Related Work

### 2\.1LLM\-Assisted Scientific Writing

Several studies have examined the use of LLMs as writing tools and the traits of LLM\-assisted text\.[Liang et al\. \(2024a\)](https://arxiv.org/html/2610.02841#bib.bib11)proposed a method for estimating the fraction of LLM\-modified content in a corpus and applied it to peer reviews from AI conferences held after the release of ChatGPT, estimating, for example, that 16\.9% of review sentences at EMNLP 2023 were substantially modified by LLMs\.[Liang et al\. \(2024b\)](https://arxiv.org/html/2610.02841#bib.bib12)applied this framework to 950k papers from arXiv, bioRxiv, and Nature portfolio journals published between 2020 and 2024, observing a steady increase in LLM usage, with computer science showing the largest growth\.

[Kobak et al\. \(2025\)](https://arxiv.org/html/2610.02841#bib.bib9)examined 15 million PubMed abstracts and concluded that at least 13\.5 percent of those published in 2024 were generated or written with the assistance of an LLM, with certain style words such as “crucial,” “delve,” and “potential” being used excessively\.[Matsui \(2025\)](https://arxiv.org/html/2610.02841#bib.bib14)compiled a list of LLM style words from previous research and compared them against a control group, finding that 103 out of 135 words showed meaningful uplift in the post\-LLM era\.[Cunningham et al\. \(2025\)](https://arxiv.org/html/2610.02841#bib.bib4)analyzed 17 million papers from Semantic Scholar and reported that in fields such as computer science and engineering, writing styles have trended toward “hyperbole” since 2023\.

### 2\.2LLM\-Assisted Adversarial Rewriting

There is also work showing that rewriting or injecting words, even without changing its factual content, can influence the model’s review score\.[Lin et al\. \(2025\)](https://arxiv.org/html/2610.02841#bib.bib13)proposed that textual adversarial attacks can artificially inflate LLM review scores while keeping the semantic content of the paper unchanged\. Different attack methods \(character\-level, word\-level, and sentence\-level\) all demonstrate some degree of effectiveness\.[Kaneko \(2026\)](https://arxiv.org/html/2610.02841#bib.bib8)instead employed an LLM\-based paraphrasing approach as the main attack method and reported that this method can optimize the review score without changing the claims of the paper\.

### 2\.3Multimodal Claim Verification

Natural RewritingControlled RewritingModelBaseLPFIASOACTD1C1gemma\-4\-26B\-A4B\-itP\-Acc0\.5480\.5460\.5370\.5290\.5320\.5240\.5600\.538F10\.7890\.7850\.7820\.7730\.7700\.7900\.7920\.790gemma\-4\-31B\-itP\-Acc0\.6100\.5850\.579\*0\.5810\.565\*\*\*0\.557\*\*\*0\.6060\.586F10\.8240\.8140\.809\*0\.8100\.798\*\*\*0\.806\*0\.8220\.816GLM\-4\.6V\-FlashP\-Acc0\.5710\.5660\.5630\.5650\.5630\.5390\.5650\.568F10\.7910\.7910\.7880\.7900\.7850\.7900\.7920\.797InternVL3\-2BP\-Acc0\.0410\.0570\.0520\.0520\.0600\.0370\.0510\.045F10\.6580\.6640\.6640\.6650\.6650\.6670\.6660\.667InternVL3\-8BP\-Acc0\.2080\.1900\.1900\.1870\.1930\.135\*\*\*0\.2030\.173F10\.6990\.6910\.6960\.6920\.6920\.6870\.7000\.695InternVL3\-38BP\-Acc0\.3380\.2960\.3070\.2890\.2940\.2730\.3390\.329F10\.7180\.7040\.7010\.6990\.6920\.7090\.7120\.718Kimi\-VL\-A3BP\-Acc0\.3530\.3620\.401\*\*\*0\.3780\.3530\.3400\.3580\.340F10\.7300\.7330\.7400\.7390\.7250\.7270\.7250\.726Qwen3\-VL\-2B\-InstructP\-Acc0\.2590\.2860\.2730\.2600\.2740\.2540\.2640\.257F10\.6580\.6670\.6630\.6580\.6690\.6630\.6610\.664Qwen3\-VL\-4B\-InstructP\-Acc0\.4520\.4530\.4620\.4470\.4730\.4870\.4680\.464F10\.7100\.7170\.7190\.7070\.7160\.745\*\*\*0\.7250\.725Qwen3\-VL\-8B\-InstructP\-Acc0\.5160\.5290\.5180\.5130\.5130\.5430\.5190\.501F10\.7390\.7480\.7400\.7400\.7250\.7650\.7410\.738Qwen3\-VL\-32B\-InstructP\-Acc0\.6170\.6090\.6180\.5890\.5770\.6040\.6100\.619F10\.8050\.8000\.8030\.7880\.771\*\*\*0\.8080\.7980\.810

Table 2:Claim verification performance under different rewriting conditions\.Base: original unmodified claims\. Significance levels: \*q<0\.05q<0\.05, \*\*q<0\.01q<0\.01, \*\*\*q<0\.001q<0\.001\(BH\-corrected\)\.333Due to computational constraints, dev set results for InternVL3\-38B were not available\.Claim verification is the task of determining whether a claim is supported by given evidence \(classification\)\. Datasets such as FEVER and SciFact\([Thorne et al\., 2018](https://arxiv.org/html/2610.02841#bib.bib20);[Wadden et al\., 2020](https://arxiv.org/html/2610.02841#bib.bib21)\)are well\-known benchmarks for this task\. In this paper, we focus on a specific subset: multimodal scientific claim verification, where the evidence is a figure or table from a paper and the model must ground a claim using the image\. SciVer\([Wang et al\., 2025](https://arxiv.org/html/2610.02841#bib.bib22)\)collects figures from arXiv computer science papers, and experts were recruited to create both supported and refuted claims based on the evidence\. MuSciClaim\([Lal et al\., 2025](https://arxiv.org/html/2610.02841#bib.bib10)\)gathers evidence and supported claims from several open\-access sources across multiple domains, including physics and biology, and generates refuted claims through manual perturbation\. SciClaimEval\([Ho et al\., 2026](https://arxiv.org/html/2610.02841#bib.bib6)\)is another multimodal claim verification dataset that collects diverse evidence from various sources; however, it creates refuted pairs by modifying the evidence rather than the claim\.

### 2\.4Sensitivity to Surface\-Level Perturbations

LLMs are well known to be affected by surface\-level linguistic traits, which can influence their performance on downstream tasks\.[Xu et al\. \(2024\)](https://arxiv.org/html/2610.02841#bib.bib24)report that LLMs can rewrite a sentence in a meaning\-preserving way that changes the predicted label in a classification task, such as sentiment analysis\.[Sclar et al\. \(2024\)](https://arxiv.org/html/2610.02841#bib.bib16)also claim that LLMs are sensitive to prompt format, with different models reacting differently to variations in prompts across tasks\. Furthermore,[Chen et al\. \(2024\)](https://arxiv.org/html/2610.02841#bib.bib3)show that LLM outputs can be affected by semantic\-agnostic biases, such as the inclusion of fake references or superficial stylistic changes\. Our work differs from the existing literature in that we focus on how non\-adversarial, naturalistic rewriting affects LLMs’ claim verification ability\.

## 3Methodology

### 3\.1Multimodal Claim Verification

Claim verification is the task of determining whether a claim is supported by a given piece of evidence\. In the multimodal setting, the evidence is a table or figure image from a scientific paper\. Given a claimccand evidenceee, the goal is to predict the verdictV⁡\(c,e\)∈\{support,refute\}V\(c,e\)\\in\\\{\\text\{support\},\\text\{refute\}\\\}\. In this work, we perturb the claim by rewriting it using an LLM with a specific rewriting procedure, resulting in a rewritten claimc′c^\{\\prime\}\. We then compare the verdictsV⁡\(c,e\)V\(c,e\)andV⁡\(c′,e\)V\(c^\{\\prime\},e\)to measure how rewriting affects verification outcomes\.

### 3\.2Rewriting Strategies

To investigate how LLM\-assisted rewriting affects downstream claim verification, we design two complementary rewriting strategies\. The first isnatural rewriting, which rewrites a claim using a collection of prompts designed to simulate realistic rewriting use cases\. The second iscontrolled rewriting, which injects a single word into the claim to study the effect of a small, controlled rewriting operation\. An example of how a claim is modified using these rewriting strategies is shown in Table[1](https://arxiv.org/html/2610.02841#S1.T1)\.

#### 3\.2\.1Natural Rewriting

Natural RewritingControlled RewritingModelBaseLPFIASOACTD1C1gemma\-4\-26B\-A4B\-itpsupp\_\{\\text\{sup\}\}0\.851Δ​psup\\Delta p\_\{\\text\{sup\}\}\-0\.011\-0\.024\*\*\*\-0\.024\*\*\*\-0\.045\*\*\*\+0\.037\*\*\*\-0\.011\*\+0\.016\*prefp\_\{\\text\{ref\}\}0\.321Δ​pref\\Delta p\_\{\\text\{ref\}\}\+0\.007\-0\.004\+0\.002\-0\.019\*\+0\.070\*\*\*\-0\.013\*\+0\.021\*\*\*gemma\-4\-31B\-itpsupp\_\{\\text\{sup\}\}0\.919Δ​psup\\Delta p\_\{\\text\{sup\}\}\-0\.002\-0\.016\*\*\*\-0\.020\*\*\-0\.040\*\*\*\+0\.019\*\*\-0\.007\+0\.005prefp\_\{\\text\{ref\}\}0\.315Δ​pref\\Delta p\_\{\\text\{ref\}\}\+0\.024\*\*\*\+0\.022\*\*\*\+0\.015\*\+0\.006\+0\.078\*\*\*\+0\.010\+0\.036\*\*\*GLM\-4\.6V\-Flashpsupp\_\{\\text\{sup\}\}0\.823Δ​psup\\Delta p\_\{\\text\{sup\}\}\+0\.009\-0\.004\-0\.005\-0\.017\+0\.056\*\*\*\+0\.007\+0\.022\*\*\*prefp\_\{\\text\{ref\}\}0\.283Δ​pref\\Delta p\_\{\\text\{ref\}\}\+0\.013\*\+0\.013\+0\.015\+0\.006\+0\.080\*\*\*\+0\.011\+0\.026\*\*\*InternVL3\-2Bpsupp\_\{\\text\{sup\}\}0\.840Δ​psup\\Delta p\_\{\\text\{sup\}\}\-0\.001\+0\.002\+0\.008\-0\.004\+0\.051\*\*\*\-0\.002\+0\.018\*\*\*prefp\_\{\\text\{ref\}\}0\.810Δ​pref\\Delta p\_\{\\text\{ref\}\}\-0\.013\*\-0\.009\-0\.010\-0\.017\*\+0\.040\*\*\*\-0\.003\+0\.014\*InternVL3\-8Bpsupp\_\{\\text\{sup\}\}0\.877Δ​psup\\Delta p\_\{\\text\{sup\}\}\+0\.003\+0\.007\+0\.005\-0\.004\+0\.037\*\*\*\+0\.005\+0\.025\*\*\*prefp\_\{\\text\{ref\}\}0\.693Δ​pref\\Delta p\_\{\\text\{ref\}\}\+0\.006\+0\.000\+0\.014\+0\.010\+0\.071\*\*\*\+0\.003\+0\.037\*\*\*InternVL3\-38Bpsupp\_\{\\text\{sup\}\}0\.816Δ​psup\\Delta p\_\{\\text\{sup\}\}\+0\.002\-0\.004\-0\.002\-0\.026\*\*\+0\.046\*\*\*\-0\.017\*\*\+0\.009prefp\_\{\\text\{ref\}\}0\.520Δ​pref\\Delta p\_\{\\text\{ref\}\}\+0\.026\*\+0\.007\+0\.025\*\+0\.005\+0\.090\*\*\*\-0\.003\+0\.027\*\*\*Kimi\-VL\-A3Bpsupp\_\{\\text\{sup\}\}0\.834Δ​psup\\Delta p\_\{\\text\{sup\}\}\+0\.001\-0\.006\-0\.000\-0\.005\+0\.018\*\*\*\-0\.002\+0\.004prefp\_\{\\text\{ref\}\}0\.531Δ​pref\\Delta p\_\{\\text\{ref\}\}\-0\.001\-0\.014\-0\.005\-0\.004\+0\.024\*\*\*\+0\.005\+0\.022\*\*\*Qwen3\-VL\-2B\-Instructpsupp\_\{\\text\{sup\}\}0\.701Δ​psup\\Delta p\_\{\\text\{sup\}\}\-0\.005\-0\.009\-0\.006\-0\.005\+0\.029\*\*\*\-0\.005\+0\.008prefp\_\{\\text\{ref\}\}0\.484Δ​pref\\Delta p\_\{\\text\{ref\}\}\+0\.010\-0\.006\+0\.003\-0\.000\+0\.046\*\*\*\+0\.009\+0\.025\*\*\*Qwen3\-VL\-4B\-Instructpsupp\_\{\\text\{sup\}\}0\.684Δ​psup\\Delta p\_\{\\text\{sup\}\}\-0\.001\+0\.000\-0\.013\-0\.026\*\+0\.056\*\*\*\+0\.002\+0\.017\*prefp\_\{\\text\{ref\}\}0\.282Δ​pref\\Delta p\_\{\\text\{ref\}\}\+0\.006\-0\.006\-0\.006\-0\.024\*\*\*\+0\.040\*\*\*\-0\.008\+0\.019\*\*\*Qwen3\-VL\-8B\-Instructpsupp\_\{\\text\{sup\}\}0\.679Δ​psup\\Delta p\_\{\\text\{sup\}\}\+0\.014\*\-0\.003\-0\.006\-0\.036\*\*\*\+0\.062\*\*\*\-0\.007\+0\.017\*\*prefp\_\{\\text\{ref\}\}0\.212Δ​pref\\Delta p\_\{\\text\{ref\}\}\+0\.014\*\*\*\+0\.005\+0\.002\-0\.013\+0\.045\*\*\*\+0\.002\+0\.024\*\*\*Qwen3\-VL\-32B\-Instructpsupp\_\{\\text\{sup\}\}0\.782Δ​psup\\Delta p\_\{\\text\{sup\}\}\+0\.001\-0\.018\*\*\-0\.027\*\*\*\-0\.065\*\*\*\+0\.049\*\*\*\-0\.016\*\*\*\+0\.011prefp\_\{\\text\{ref\}\}0\.198Δ​pref\\Delta p\_\{\\text\{ref\}\}\+0\.015\*\*\*\-0\.001\+0\.002\-0\.012\+0\.061\*\*\*\-0\.014\*\*\*\+0\.019\*\*\*

Table 3:Changes inp⁡\(support\)p\(\\text\{support\}\)on the supported and refuted subsets after claim rewriting\. Values under each rewriting condition representΔ​p=p′−p\\Delta p=p^\{\\prime\}\-p, whereppis the base prediction andp′p^\{\\prime\}is the prediction after rewriting\. The values ofp⁡\(support\)p\(\\text\{support\}\)are computed separately for the supported subset \(denoted aspsupp\_\{\\text\{sup\}\}\) and the refuted subset \(denoted asprefp\_\{\\text\{ref\}\}\)\. We observe statistically consistent effects for CT and C1 across models\. Significance levels: \*q<0\.05q<0\.05, \*\*q<0\.01q<0\.01, \*\*\*q<0\.001q<0\.001\(BH\-corrected\)\.The goal of natural rewriting is to demonstrate that even in unintentional, realistic settings, stylistic changes introduced by LLM rewriting still occur and affect downstream tasks\. We define five rewriting categories, each reflecting a common motivation for using LLMs to polish academic text:

- •Language Polishing \(LP\)Rewrite or fix the text to correct grammatical errors \(e\.g\.,“Can you fix the grammar of this text?”\)\.
- •Fluency Improvement \(FI\)Rewrite the text to improve fluency and readability \(e\.g\.,“Make this paragraph flow more naturally\.”\)\.
- •Academic Style Alignment \(AS\)Rewrite the text to match the conventions of a target venue \(e\.g\.,“Rewrite this to match the style of an ACL paper\.”\)\. We select five major ML conferences: ACL, AAAI, NeurIPS, ICLR, and ICML\.
- •Outcome\-Aware \(OA\)Rewrite the text to maximize perceived quality by reviewers, in terms of review scores and acceptance likelihood \(e\.g\.,“Polish this so it sounds more convincing to reviewers\.”\)\.
- •Cautious Tone \(CT\)Rewrite the text to make claims more measured and conservative, avoiding overstatement or unsupported certainty \(e\.g\.,“Rewrite this in a more cautious and conservative tone\.”\)\.

For each rewriting category, we design five prompt variants for each rewriting category, shown in Table[9](https://arxiv.org/html/2610.02841#A9.T9)in the Appendix, to reduce dependence on a single prompt formulation\. Instead of applying all variants to every data point, which would increase computational cost by a factor of five, we randomly assign one variant to each data point\. This assignment is fixed across all experimental settings to ensure comparability\. We also append an additional suffix prompt \(see Figure[4](https://arxiv.org/html/2610.02841#A3.F4)\) to encourage the model to preserve the underlying meaning of the claim during rewriting\.

##### Human Validation

To validate the quality of the generated claims, we conducted an annotation study in which four NLP researchers judged whether rewrites were semantically consistent with the original, covering 282 claims across conditions OA, CT, C1, and D1 \(Appendix[H](https://arxiv.org/html/2610.02841#A8)\)\. Annotators judged 81\.8% of rewrites as consistent \(from 74\.5% for OA to 90\.0% for D1\), though inter\-annotator agreement was low \(κ=0\.25\\kappa=0\.25\)\. This lowκ\\kappais largely a consequence of label skew, and Gwet’s AC1, which is robust to skew, indicates substantial agreement \(AC1=0\.71\\text\{AC1\}=0\.71\)\. Together, these suggest that while annotators agree semantic changes are uncommon, they disagree on*which*claims cross the boundary\.

#### 3\.2\.2Controlled Rewriting

While natural rewriting captures realistic use cases, a key observation is that the primary stylistic trait of LLM\-rewritten text is the frequent use of characteristic vocabulary\. The goal of controlled rewriting is to isolate this effect by selectively injecting LLM\-associated words into original claims\.

##### Word sources\.

We draw from two established LLM\-vocabulary studies:[Kobak et al\. \(2025\)](https://arxiv.org/html/2610.02841#bib.bib9)and[Matsui \(2025\)](https://arxiv.org/html/2610.02841#bib.bib14)\. We categorize words into three groups \(decorative,conservative, andneutral\) based on LLM classification using human\-written definitions, reviewed and validated by the author\.

We define two controlled rewriting conditions based on the injected word pool:Dfor decorative words andCfor conservative words\. The decorative set serves as a surrogate for rewriting styles such as OA, whereas the conservative set is expected to induce an effect closer to CT\. Neutral words are excluded because they are not expected to shift the style signal\. This contrast allows us to test whether decorative/conservative vocabulary is a causal factor in downstream performance changes, rather than merely a side effect of the rewriting process\. Since neither source contains sufficient hedge words after categorization, we supplement the conservative pool with 31 additional hedging terms identified through LLM\-assisted generation\. This results in two word pools with 39 conservative words and 114 decorative words\.

##### Word Injection\.

For each claim, the model injects exactly one word from the corresponding pool into the original claim\. For each injection, we provide the model with the original claim, the injection instructions, and a randomly sampled subset of candidate words from the selected pool \(15 for C, 30 for D\)\. This subsampling encourages diversity across data points while allowing the model to choose a grammatically natural insertion\. The prompt requires the model to preserve all original words, avoid deletion, replacement, or reordering, and keep the factual content unchanged \(see Figure[5](https://arxiv.org/html/2610.02841#A4.F5)\)\.

Figure 2:Shift decomposition compass\. Each point represents one rewriting condition, averaged over all models\. Thexx\-axis shows the mean change inp⁡\(support\)p\(\\text\{support\}\)on supporting evidence pairs \(Δ​psup\\Delta p\_\{\\text\{sup\}\}\) and theyy\-axis on refuting evidence pairs \(Δ​pref\\Delta p\_\{\\text\{ref\}\}\)\. Circles denote different rewriting conditions\. Points near the origin indicate robustness to rewriting\.

## 4Experimental Setup

### 4\.1Datasets

Dataset\#Dev\#TestSciClaimEvalFull747917Unpaired removed704872Table 4:Dataset statistics\.We evaluate on the following multimodal claim verification dataset: SciClaimEval\([Ho et al\., 2026](https://arxiv.org/html/2610.02841#bib.bib6)\)\. In this dataset, the evidence consists of table or figure images from scientific papers\. We choose to center our analysis on SciClaimEval due to its*paired*structure: for each claim, an evidence imageeeis collected and then modified to construct a refuting counterparte′e^\{\\prime\}, resulting in pairs\(c,e\)\(c,e\)and\(c,e′\)\(c,e^\{\\prime\}\)such thatV⁡\(c,e\)=supportV\(c,e\)=\\textsc\{support\}andV⁡\(c,e′\)=refuteV\(c,e^\{\\prime\}\)=\\textsc\{refute\}\. We discard the small number of unpaired instances arising from slight imbalances in the original dataset\. Table[4](https://arxiv.org/html/2610.02841#S4.T4)summarises the dataset statistics after filtering\.

This paired design is critical for our analysis, which modifies claim text and measures how such modifications affect model predictions\. In standard datasets, the supported and refuted subsets contain different claims with potentially different stylistic properties, confounding any modification\-based analysis\. The paired setting controls for this: because the same claims appear under both verdicts, the distributional properties of the two sets of claims are identical, and differences in outcome can be attributed to the modification itself\.

### 4\.2Models

##### Verification models\.

We select five open\-weight VLM families spanning a wide range of architectures and parameter sizes: Gemma 4\([Google DeepMind, 2026](https://arxiv.org/html/2610.02841#bib.bib5)\), GLM\-4\.6V\([Hong et al\., 2025](https://arxiv.org/html/2610.02841#bib.bib7)\), InternVL3\([Zhu et al\., 2025](https://arxiv.org/html/2610.02841#bib.bib25)\), Kimi\-VL\([Team et al\., 2025](https://arxiv.org/html/2610.02841#bib.bib18)\), and Qwen3\-VL\([Bai et al\., 2025](https://arxiv.org/html/2610.02841#bib.bib1)\)\. Specifically, we evaluate Gemma\-4\-26B\-A4B\-it and Gemma\-4\-31B\-it, GLM\-4\.6V\-Flash, InternVL3\-2B, 8B, 38B, Kimi\-VL\-A3B\-Thinking\-2506 \(abbreviated as Kimi\-VL\-A3B\), and Qwen3\-VL\-2B, 4B, 8B, and 32B\-Instruct, totaling 11 models with parameter counts ranging from 2B to 38B\.

##### Rewriting models\.

Claim rewriting is performed using gpt\-oss\-120b\([OpenAI et al\., 2025](https://arxiv.org/html/2610.02841#bib.bib15)\), with temperatureT=0\.7T=0\.7for natural rewriting andT=0\.25T=0\.25for controlled injection\.

##### Inference settings\.

Chain\-of\-thought reasoning is important for solving complex questions such as multimodal claim verification, and methods that extract logprobs without CoT \(e\.g\., forcing the model to output the answer in the first token\) yield comparably worse performance\. To ensure both strong performance and the ability to extract probability estimates, we employ self\-consistency\([Wang et al\., 2023](https://arxiv.org/html/2610.02841#bib.bib23)\)at inference time\. For each data point and model, we sample 10 reasoning paths with temperatureT=1\.0T=1\.0and use the fraction of runs in which the model predictssupportasp⁡\(support\)p\(\\text\{support\}\)\.

##### Computational budgets\.

All experiments were conducted on 8 NVIDIA H200 GPUs over approximately 1–2 days\.

### 4\.3Metrics\.

##### Paired accuracy and F1\.

We report two levels of performance metrics\. F1 score is computed over individual claim\-evidence pairs in the standard binary classification sense \(support vs\. refute\)\. Paired accuracy, following the evaluation design of SciClaimEval, applies a stricter criterion: a claimccis paired\-correct only if the model produces the correct verdict onboththe supporting and refuting evidence, i\.e\.,V⁡\(c,e\)=supportV\(c,e\)=\\text\{support\}andV⁡\(c,e′\)=refuteV\(c,e^\{\\prime\}\)=\\text\{refute\}\. We report both metrics before and after rewriting for each condition\.

Figure 3:AverageΔ​p\\Delta pfor theQwen3\-VLmodel family\. The x\-axis shows the mean change inp⁡\(support\)p\(\\text\{support\}\), i\.e\., the average predicted probability of thesupportlabel across the whole dataset\. The y\-axis represents model size \(2B, 4B, 8B, and 32B\)\.
##### Probability shifts\.

To characterizehowrewriting affects model behavior, we examine the joint pattern of probability shifts on both pairs\. For each claim, we compute:

Δ​psup\\displaystyle\\Delta p\_\{\\text\{sup\}\}=p⁡\(support∣e,c′\)−p⁡\(support∣e,c\)\\displaystyle=p\(\\text\{support\}\\mid e,c^\{\\prime\}\)\-p\(\\text\{support\}\\mid e,c\)\(1\)Δ​pref\\displaystyle\\Delta p\_\{\\text\{ref\}\}=p⁡\(support∣e′,c′\)−p⁡\(support∣e′,c\)\\displaystyle=p\(\\text\{support\}\\mid e^\{\\prime\},c^\{\\prime\}\)\-p\(\\text\{support\}\\mid e^\{\\prime\},c\)\(2\)The sign pattern of\(Δ​psup,Δ​pref\)\(\\Delta p\_\{\\text\{sup\}\},\\Delta p\_\{\\text\{ref\}\}\)classifies each claim into one of four categories: \(1\)Globally Optimistic\(Δ​psup↑\\Delta p\_\{\\text\{sup\}\}\\uparrow,Δ​pref↑\\Delta p\_\{\\text\{ref\}\}\\uparrow\), where the model trusts rewritten claims more across the board; \(2\)Globally Conservative\(Δ​psup↓\\Delta p\_\{\\text\{sup\}\}\\downarrow,Δ​pref↓\\Delta p\_\{\\text\{ref\}\}\\downarrow\), where the model trusts rewritten claims less across the board; \(3\)Sharper\(Δ​psup↑\\Delta p\_\{\\text\{sup\}\}\\uparrow,Δ​pref↓\\Delta p\_\{\\text\{ref\}\}\\downarrow\), where the model becomes more correct on both pairs; and \(4\)Duller\(Δ​psup↓\\Delta p\_\{\\text\{sup\}\}\\downarrow,Δ​pref↑\\Delta p\_\{\\text\{ref\}\}\\uparrow\), where the model becomes less correct on both pairs\. Figure[2](https://arxiv.org/html/2610.02841#S3.F2)visualises this taxonomy as a two\-dimensional compass, placing each rewriting condition in the quadrant\.

##### Statistical tests\.

We assess the significance of each rewrite condition using a unified paired bootstrap procedure\. For pair accuracy, we compute per\-pair differences \(condition−\-base\) and bootstrap their mean over 1,000 resamples; for F1, we resample claim indices and recomputeF1​\(condition\)−F1​\(base\)\\text\{F1\}\(\\text\{condition\}\)\-\\text\{F1\}\(\\text\{base\}\)on each resample; for probability shiftsΔ​p\\Delta p, we bootstrap the mean per\-pair shift\. All tests are two\-sided: thepp\-value is the fraction of resamples that fall on the minority side of zero, doubled to account for both directions\. Multiple comparisons are corrected by Benjamini–Hochberg FDR\([Benjamini and Hochberg, 1995](https://arxiv.org/html/2610.02841#bib.bib2)\)applied jointly over all \(model, condition, metric\) cells within each table \(m=154m=154tests\), treating the table as the unit of correction\. Significance:∗q<0\.05q<0\.05,∗∗q<0\.01q<0\.01,∗∗∗q<0\.001q<0\.001\.

### 4\.4License\.

The dataset SciClaimEval is released under CC BY 4\.0\. All models used in this work, including those for rewriting and evaluation, are open\-weight and distributed under their respective licenses\.

## 5Experimental Results

##### RQ1: Does rewriting change the verdict?

Table[3](https://arxiv.org/html/2610.02841#footnote3)reports paired accuracy and F1 across all models and rewriting conditions\. We observe few statistically significant changes: the most affected model is gemma\-4\-31B\-it, which shows significant drops under FI, OA, and CT \(e\.g\., 0\.610→\\rightarrow0\.557,q<0\.001q<0\.001under CT\), and InternVL3\-8B under CT \(0\.208→\\rightarrow0\.135,q<0\.001q<0\.001\)\. These drops are concentrated in a small number of models rather than a consistent trend across all models or strategies\.

This finding contrasts with prior work showing that LLM\-based rewriting can manipulate review scores\([Kaneko, 2026](https://arxiv.org/html/2610.02841#bib.bib8);[Lin et al\., 2025](https://arxiv.org/html/2610.02841#bib.bib13)\)\. A key difference is that those methods employ adversarial strategies designed to influence evaluation outcomes, whereas our rewriting is non\-adversarial and does not target any specific metric\. Additionally, claim verification is a more constrained setting where the model must ground a single sentence against concrete visual evidence\. Whether similar robustness holds under adversarial rewriting optimized for verification remains an open question\.

##### RQ2: Does rewriting change confidence?

Despite the largely stable accuracy reported in RQ1, Table[3](https://arxiv.org/html/2610.02841#S3.T3)reveals a much clearer effect at the probability level\. The strongest shifts occur under hedging conditions: CT produces significantΔ​pref\\Delta p\_\{\\text\{ref\}\}in 11 out of 11 models, and C1 in 11 out of 11, both atq<0\.05q<0\.05or below\. In contrast, boosting conditions show a more moderate effect: OA produces significantΔ​psup\\Delta p\_\{\\text\{sup\}\}in 6 out of 11 models, but its effect onΔ​pref\\Delta p\_\{\\text\{ref\}\}is mixed\. One possible explanation is that for refuted claims under boosting conditions, the baseprefp\_\{\\text\{ref\}\}is already low, leaving little room for the probability to decrease further\. We observe the same floor\-ceiling dynamic in reverse for hedging conditions, where the effect on the supported subset is smaller because basepsupp\_\{\\text\{sup\}\}is already near 1\. The remaining conditions \(LP, FI, AS\) do not produce meaningful probability shifts, with most values small and non\-significant\. These shifts persist after removing rewrites that insert hedges before numerical or comparative expressions \(Appendix[B](https://arxiv.org/html/2610.02841#A2)\), and a second rewriter \(Gemma\-4\-31B\-it\) yields consistent trends \(Appendix[A](https://arxiv.org/html/2610.02841#A1)\)\.

We also do not observe a clear relationship between model size and the magnitude of probability shifts\. However, within the Qwen3\-VL family \(Figure[3](https://arxiv.org/html/2610.02841#S4.F3)\), a pattern emerges: larger models show larger drops under OA and larger increases under CT, suggesting that these models become more sensitive to writing style and adjust their predictions accordingly\. This pattern is not universal across all model families, so we refrain from drawing a general conclusion\.

##### RQ3: Why do probability shifts not translate to accuracy changes?

Given the significant probability shifts reported in RQ2, why do accuracy metrics remain largely unchanged? The answer lies in model confidence: most models produce base probabilities clustered near 0 or 1, predicting the same label consistently even across 10 reasoning paths at temperatureT=1\.0T=1\.0\. This means that a shift of 0\.03–0\.05, while statistically significant, is unlikely to cross the 0\.5 decision boundary\. Table[10](https://arxiv.org/html/2610.02841#A9.T10)confirms this: flip rates are mostly below 15%\. The model’s internal belief does shift, but the shift is masked by high confidence\. This suggests that on more difficult or subjective cases, where predicted probabilities are closer to 0\.5, the effect of stylistic changes on final verdicts may be considerably larger\. Representative high\-shift and no\-shift examples are shown in Appendix[I](https://arxiv.org/html/2610.02841#A9)\.

## 6Conclusion

In this work, we study how stylistic changes affect model predictions on the multimodal claim verification task\. We find that models across different sizes and families are robust to stylistic rewriting in terms of accuracy metrics, but consistent probability shifts exist under certain rewriting conditions\. The reason probability shifts occur even when accuracy does not change is that most predictions tend to cluster around decisive points \(0 or 1\), making the probability changes ineffective when discretized\. This is in contrast with previous work on paper review prediction, suggesting that fact\-checking tasks are more difficult to sway by simply changing the style of the claims\. A natural next question is how this robustness generalizes to other scientific writing tasks beyond reviewing and claim verification\. Another interesting direction is to examine how the fact\-checking aspect changes when we apply review\-score optimization using previously established methods\. A better separation of semantic change and stylistic change is also needed to provide more rigorous methodological analysis\.

## 7Limitations

##### Separating style from semantics\.

Certain rewriting conditions, particularly CT, produce uniquely strong probability shifts, suggesting that the rewriting model sometimes introduces changes that go beyond style\. A common failure mode is the injection of hedging language before numerical values \(e\.g\., “30%”→\\rightarrow“roughly 30%”\), which reduces the precision and falsifiability of the claim\. Filtering out such rewrites does not remove the observed shifts \(Appendix[B](https://arxiv.org/html/2610.02841#A2)\), although a rule\-based filter cannot capture all semantic changes\. These cases suggest that fully separating stylistic from semantic modification is inherently difficult: even human annotators often disagree on whether a rewrite preserves meaning \(Appendix[H](https://arxiv.org/html/2610.02841#A8)\), and even with explicit instructions to preserve meaning, some rewriting scenarios are more disruptive than others\.

##### Single rewriting model and dataset\.

Our main experiments use a single model \(gpt\-oss\-120b\) for claim rewriting\. A small\-scale check with a second rewriter \(Gemma\-4\-31B\-it; Appendix[A](https://arxiv.org/html/2610.02841#A1)\) shows consistent trends, but it covers only four conditions, one verifier, and the dev set\. Given a fixed computational budget, we prioritize diversity on the*verification*side \(11 models across five families\), as our primary research question concerns how verifiers respond to rewriting rather than how rewriters differ\. Similarly, we evaluate on a single dataset \(SciClaimEval\), whose paired structure \(identical claims appearing under both supported and refuted evidence\) is helpful for our analysis but limits generalizability to other domains and evidence types\. To our knowledge, SciClaimEval is the only available dataset that provides this paired design\.

##### Pre\-existing LLM polish\.

SciClaimEval was constructed from publicly available papers, some of which were published after the widespread adoption of LLMs\. It is therefore possible that certain claims in the dataset have already been polished by LLMs in some way, which could dampen the observed effects of our rewriting methods\.

##### Confidence estimation\.

We estimatep⁡\(support\)p\(\\text\{support\}\)only through self\-consistency over 10 sampled reasoning paths, which yields probabilities in steps of 0\.1\. We did not compare it against other confidence extraction methods, such as token\-level log\-probabilities or verbalized confidence, and the size of the observed shifts may depend on this choice\. We leave this comparison to future work\.

## 8Ethical Considerations

This work is an empirical analysis of existing models and datasets, focusing on stylistic changes in scientific writing\. All datasets used consist of scientific claims and evidence from published open\-access papers and do not contain personally identifying information or offensive content\. The dataset \(SciClaimEval\) is released under CC BY 4\.0, and all models used are open\-weight and distributed under their respective licenses\. The annotation study involved NLP researchers from the research team who evaluated rewritten claims for semantic consistency; no personal or sensitive data was collected\. The annotators were co\-authors of this paper and were not separately compensated for the annotation task\. AI assistants were used for revising author\-written text, reviewing ideas, identifying errors in the manuscript, and as a coding assistant for refactoring and implementation\. All AI\-generated outputs were reviewed and verified by the authors\. We acknowledge that while our rewriting strategies are designed to study model robustness, similar techniques could in principle be applied to manipulate automated verification systems\. Our code and data are publicly available to support transparency and reproducibility\.

## Acknowledgments

This work was supported by JSPS KAKENHI Grant Number 24K03231\. This work was partially funded by the Deutsche Forschungsgemeinschaft \(DFG, German Research Foundation\) –[554559555](https://gepris.dfg.de/gepris/projekt/554559555)\. This work was partially supported by the French National Research Agency \(ANR\) through the FABULEUX project \(ANR\-26\-CE23\-6589\)\. F\. Boudin was also supported by a visiting professor position at the National Institute of Informatics \(NII\), Japan\. We used ABCI 3\.0\([Takano et al\., 2024](https://arxiv.org/html/2610.02841#bib.bib17)\)provided by AIST and AIST Solutions with support from “ABCI 3\.0 Development Acceleration Use”\.

## References

- Bai et al\. \(2025\)Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, and 45 others\. 2025\.[Qwen3\-vl technical report](https://arxiv.org/abs/2511.21631)\.*Preprint*, arXiv:2511\.21631\.
- Benjamini and Hochberg \(1995\)Yoav Benjamini and Yosef Hochberg\. 1995\.[Controlling the false discovery rate: A practical and powerful approach to multiple testing](http://www.jstor.org/stable/2346101)\.*Journal of the Royal Statistical Society\. Series B \(Methodological\)*, 57\(1\):289–300\.
- Chen et al\. \(2024\)Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang\. 2024\.[Humans or LLMs as the judge? a study on judgement bias](https://doi.org/10.18653/v1/2024.emnlp-main.474)\.In*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, pages 8301–8327, Miami, Florida, USA\. Association for Computational Linguistics\.
- Cunningham et al\. \(2025\)Padraig Cunningham, Padhraic Smyth, and Barry Smyth\. 2025\.[Shifting norms in scholarly publications: trends in readability, objectivity, authorship, and ai use](https://arxiv.org/abs/2510.21725)\.*Preprint*, arXiv:2510\.21725\.
- Google DeepMind \(2026\)Google DeepMind\. 2026\.[Gemma 4](https://deepmind.google/models/gemma/gemma-4/)\.
- Ho et al\. \(2026\)Xanh Ho, Yun\-Ang Wu, Sunisth Kumar, Tian Cheng Xia, Florian Boudin, Andre Greiner\-Petter, and Akiko Aizawa\. 2026\.[Sciclaimeval: Cross\-modal claim verification in scientific papers](https://doi.org/10.63317/4ap9rg2gnwmf)\.In*Proceedings of the Fifteenth Language Resources and Evaluation Conference \(LREC 2026\)*, pages 11060–11071, Palma, Mallorca, Spain\. European Language Resources Association \(ELRA\)\.
- Hong et al\. \(2025\)Wenyi Hong, Wenmeng Yu, Xiaotao Gu, Guo Wang, Guobing Gan, Haomiao Tang, Jiale Cheng, Ji Qi, Junhui Ji, Lihang Pan, and 1 others\. 2025\.[Glm\-4\.5 v and glm\-4\.1 v\-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning](https://arxiv.org/abs/2507.01006)\.*arXiv preprint arXiv:2507\.01006*\.
- Kaneko \(2026\)Masahiro Kaneko\. 2026\.[Paraphrasing adversarial attack on llm\-as\-a\-reviewer](https://arxiv.org/abs/2601.06884)\.*Preprint*, arXiv:2601\.06884\.
- Kobak et al\. \(2025\)Dmitry Kobak, Rita González\-Márquez, Emőke Ágnes Horvát, and Jan Lause\. 2025\.[Delving into llm\-assisted writing in biomedical publications through excess vocabulary](https://doi.org/10.1126/sciadv.adt3813)\.*Science Advances*, 11\(27\):eadt3813\.
- Lal et al\. \(2025\)Yash Kumar Lal, Manikanta Bandham, Mohammad Saqib Hasan, Apoorva Kashi, Mahnaz Koupaee, and Niranjan Balasubramanian\. 2025\.[MuSciClaims: Multimodal scientific claim verification](https://doi.org/10.18653/v1/2025.ijcnlp-long.175)\.In*Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics*, pages 3285–3307, Mumbai, India\. The Asian Federation of Natural Language Processing and The Association for Computational Linguistics\.
- Liang et al\. \(2024a\)Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, Daniel Mcfarland, and James Y\. Zou\. 2024a\.[Monitoring AI\-modified content at scale: A case study on the impact of ChatGPT on AI conference peer reviews](https://proceedings.mlr.press/v235/liang24b.html)\.In*Proceedings of the 41st International Conference on Machine Learning*, volume 235 of*Proceedings of Machine Learning Research*, pages 29575–29620\. PMLR\.
- Liang et al\. \(2024b\)Weixin Liang, Yaohui Zhang, Zhengxuan Wu, Haley Lepp, Wenlong Ji, Xuandong Zhao, Hancheng Cao, Sheng Liu, Siyu He, Zhi Huang, Diyi Yang, Christopher Potts, Christopher D Manning, and James Y\. Zou\. 2024b\.[Mapping the increasing use of LLMs in scientific papers](https://openreview.net/forum?id=YX7QnhxESU)\.In*First Conference on Language Modeling*\.
- Lin et al\. \(2025\)Tzu\-Ling Lin, Wei\-Chih Chen, Teng\-Fang Hsiao, Hou\-I Liu, Ya\-Hsin Yeh, Yu\-Kai Chan, Wen\-Sheng Lien, Po\-Yen Kuo, Philip S\. Yu, and Hong\-Han Shuai\. 2025\.[Breaking the reviewer: Assessing the vulnerability of large language models in automated peer review under textual adversarial attacks](https://doi.org/10.18653/v1/2025.findings-emnlp.259)\.In*Findings of the Association for Computational Linguistics: EMNLP 2025*, pages 4819–4839, Suzhou, China\. Association for Computational Linguistics\.
- Matsui \(2025\)Kentaro Matsui\. 2025\.[Delving into PubMed records: How AI\-influenced vocabulary has transformed medical writing since ChatGPT](https://pmejournal.org/articles/10.5334/pme.1929)\.*Perspect\. Med\. Educ\.*, 14\(1\):882–890\.
- OpenAI et al\. \(2025\)OpenAI, :, Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K\. Arora, Yu Bai, Bowen Baker, Haiming Bao, Boaz Barak, Ally Bennett, Tyler Bertao, Nivedita Brett, Eugene Brevdo, Greg Brockman, Sebastien Bubeck, and 108 others\. 2025\.[gpt\-oss\-120b & gpt\-oss\-20b model card](https://arxiv.org/abs/2508.10925)\.*Preprint*, arXiv:2508\.10925\.
- Sclar et al\. \(2024\)Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr\. 2024\.[Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting](https://arxiv.org/abs/2310.11324)\.In*International Conference on Learning Representations*, volume 2024, pages 25055–25083\.
- Takano et al\. \(2024\)Ryousei Takano, Shinichiro Takizawa, Yusuke Tanimura, Hidemoto Nakada, and Hirotaka Ogawa\. 2024\.[Abci 3\.0: Evolution of the leading ai infrastructure in japan](https://arxiv.org/abs/2411.09134)\.*Preprint*, arXiv:2411\.09134\.
- Team et al\. \(2025\)Kimi Team, Angang Du, Bohong Yin, Bowei Xing, Bowen Qu, Bowen Wang, Cheng Chen, Chenlin Zhang, Chenzhuang Du, Chu Wei, Congcong Wang, Dehao Zhang, Dikang Du, Dongliang Wang, Enming Yuan, Enzhe Lu, Fang Li, Flood Sung, Guangda Wei, and 76 others\. 2025\.[Kimi\-vl technical report](https://arxiv.org/abs/2504.07491)\.*Preprint*, arXiv:2504\.07491\.
- Thakkar et al\. \(2025\)Nitya Thakkar, Mert Yuksekgonul, Jake Silberg, Animesh Garg, Nanyun Peng, Fei Sha, Rose Yu, Carl Vondrick, and James Zou\. 2025\.[Can llm feedback enhance review quality? a randomized study of 20k reviews at iclr 2025](https://arxiv.org/abs/2504.09737)\.*Preprint*, arXiv:2504\.09737\.
- Thorne et al\. \(2018\)James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal\. 2018\.[FEVER: a large\-scale dataset for fact extraction and VERification](https://doi.org/10.18653/v1/N18-1074)\.In*Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long Papers\)*, pages 809–819, New Orleans, Louisiana\. Association for Computational Linguistics\.
- Wadden et al\. \(2020\)David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi\. 2020\.[Fact or fiction: Verifying scientific claims](https://doi.org/10.18653/v1/2020.emnlp-main.609)\.In*Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\)*, pages 7534–7550, Online\. Association for Computational Linguistics\.
- Wang et al\. \(2025\)Chengye Wang, Yifei Shen, Zexi Kuang, Arman Cohan, and Yilun Zhao\. 2025\.[SciVer: Evaluating foundation models for multimodal scientific claim verification](https://doi.org/10.18653/v1/2025.acl-long.420)\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 8562–8579, Vienna, Austria\. Association for Computational Linguistics\.
- Wang et al\. \(2023\)Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou\. 2023\.[Self\-consistency improves chain of thought reasoning in language models](https://arxiv.org/abs/2203.11171)\.In*International Conference on Learning Representations*, volume 2023\.
- Xu et al\. \(2024\)Xilie Xu, Keyi Kong, Ning Liu, Lizhen Cui, Di Wang, Jingfeng Zhang, and Mohan Kankanhalli\. 2024\.[An LLM can fool itself: A prompt\-based adversarial attack](https://openreview.net/forum?id=VVgGbB9TNV)\.In*The Twelfth International Conference on Learning Representations*\.
- Zhu et al\. \(2025\)Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, Zhangwei Gao, Erfei Cui, Xuehui Wang, Yue Cao, Yangzhou Liu, Xingguang Wei, Hongjie Zhang, Haomin Wang, Weiye Xu, and 32 others\. 2025\.[Internvl3: Exploring advanced training and test\-time recipes for open\-source multimodal models](https://arxiv.org/abs/2504.10479)\.*Preprint*, arXiv:2504\.10479\.

## Appendix ARobustness to the Choice of Rewriter

To test whether our results generalize to other rewriters, we ran an additional experiment on the OA, CT, C1, and D1 conditions using a different rewriter model, Gemma\-4\-31B\-it\. For this small\-scale experiment, we used Qwen3\-VL\-8B as the verification model; choosing a model from a different family than the rewriter avoids self\-preference bias\. The experiment was conducted on the dev set only\.

ConditionRewriterΔ​psup\\Delta p\_\{\\text\{sup\}\}Δ​pref\\Delta p\_\{\\text\{ref\}\}CTgpt\-oss\+0\.054\+0\.054\+0\.030\+0\.030\[0\.030,0\.078\]\[0\.030,\\,0\.078\]\[0\.012,0\.050\]\[0\.012,\\,0\.050\]Gemma\+0\.056\+0\.056\+0\.037\+0\.037\[0\.032,0\.080\]\[0\.032,\\,0\.080\]\[0\.020,0\.058\]\[0\.020,\\,0\.058\]OAgpt\-oss−0\.035\-0\.035−0\.021\-0\.021\[−0\.058,−0\.012\]\[\-0\.058,\\,\-0\.012\]\[−0\.039,−0\.002\]\[\-0\.039,\\,\-0\.002\]Gemma−0\.027\-0\.027−0\.003\-0\.003\[−0\.051,−0\.005\]\[\-0\.051,\\,\-0\.005\]\[−0\.022,0\.016\]\[\-0\.022,\\,0\.016\]C1gpt\-oss\+0\.016\+0\.016\+0\.018\+0\.018\[−0\.002,0\.035\]\[\-0\.002,\\,0\.035\]\[0\.003,0\.033\]\[0\.003,\\,0\.033\]Gemma\+0\.017\+0\.017\+0\.023\+0\.023\[−0\.000,0\.035\]\[\-0\.000,\\,0\.035\]\[0\.009,0\.038\]\[0\.009,\\,0\.038\]D1gpt\-oss−0\.003\-0\.003−0\.009\-0\.009\[−0\.019,0\.013\]\[\-0\.019,\\,0\.013\]\[−0\.023,0\.005\]\[\-0\.023,\\,0\.005\]Gemma−0\.007\-0\.007−0\.004\-0\.004\[−0\.026,0\.012\]\[\-0\.026,\\,0\.012\]\[−0\.020,0\.014\]\[\-0\.020,\\,0\.014\]Table 5:Δ​psup\\Delta p\_\{\\text\{sup\}\}andΔ​pref\\Delta p\_\{\\text\{ref\}\}for rewrites generated by gpt\-oss and Gemma\-4\-31B\-it on the dev set, verified by Qwen3\-VL\-8B\. Brackets show 95% bootstrap CIs \(1,000 resamples\)\. For Gemma, C1/D1 rewrites that do not insert exactly one pool word are excluded\.As shown in Table[5](https://arxiv.org/html/2610.02841#A1.T5), the direction and magnitude of the confidence drift in CT, C1, and D1 are broadly similar to those observed on the gpt\-oss rewrites\. On OA, the direction holds but the magnitude is smaller\. The difference is clearest on refuted claims, where the shift essentially vanishes \(−0\.021→−0\.003\-0\.021\\rightarrow\-0\.003\)\. Overall, this experiment is consistent with our main finding that hedging\-style conditions produce significant shifts, while boosting conditions show a weaker effect\.

## Appendix BSemantic Change of Claims after Rewriting

As noted in the Limitations section, the model sometimes produces rewrites that alter the semantic meaning of the claim\. To verify that such rewrites do not affect our conclusions, we conducted the following analysis\. A rewrite is flagged if it introduces a hedge word \(Table[11](https://arxiv.org/html/2610.02841#A9.T11)\) within the two tokens immediately preceding one of the following elements: a number, %, app\-value, a unit, or a comparison/ranking word\. This criterion flags 26\.3% of CT rewrites and 22\.9% of C1 rewrites, and fewer than 10% of rewrites in all other conditions\.

FullUnflaggedFlaggedCTΔ​psup\\Delta p\_\{\\text\{sup\}\}\+0\.062\+0\.062\+0\.055\+0\.055\+0\.080\+0\.080\[0\.044,0\.078\]\[0\.044,\\,0\.078\]\[0\.037,0\.074\]\[0\.037,\\,0\.074\]\[0\.051,0\.109\]\[0\.051,\\,0\.109\]CTΔ​pref\\Delta p\_\{\\text\{ref\}\}\+0\.045\+0\.045\+0\.044\+0\.044\+0\.046\+0\.046\[0\.030,0\.059\]\[0\.030,\\,0\.059\]\[0\.027,0\.061\]\[0\.027,\\,0\.061\]\[0\.020,0\.075\]\[0\.020,\\,0\.075\]C1Δ​psup\\Delta p\_\{\\text\{sup\}\}\+0\.017\+0\.017\+0\.016\+0\.016\+0\.020\+0\.020\[0\.006,0\.028\]\[0\.006,\\,0\.028\]\[0\.003,0\.029\]\[0\.003,\\,0\.029\]\[−0\.002,0\.040\]\[\-0\.002,\\,0\.040\]C1Δ​pref\\Delta p\_\{\\text\{ref\}\}\+0\.024\+0\.024\+0\.020\+0\.020\+0\.040\+0\.040\[0\.014,0\.034\]\[0\.014,\\,0\.034\]\[0\.008,0\.031\]\[0\.008,\\,0\.031\]\[0\.016,0\.063\]\[0\.016,\\,0\.063\]Table 6:Δ​psup\\Delta p\_\{\\text\{sup\}\}andΔ​pref\\Delta p\_\{\\text\{ref\}\}on the full, unflagged, and flagged subsets \(Qwen3\-VL\-8B\)\. Brackets show 95% bootstrap CIs \(1,000 resamples\)\.As shown in Table[6](https://arxiv.org/html/2610.02841#A2.T6),Δ​psup\\Delta p\_\{\\text\{sup\}\}andΔ​pref\\Delta p\_\{\\text\{ref\}\}on the unflagged set are similar to or slightly lower than those on the full set, while the flagged set tends to show larger deltas in most conditions\. Because the unflagged set preserves the trend observed on the full set, our conclusions hold even after removing potentially semantics\-altering rewrites\.

## Appendix CPrompts

We provide the full prompt templates used for natural rewriting \(Figure[4](https://arxiv.org/html/2610.02841#A3.F4)\), controlled rewriting \(Figure[5](https://arxiv.org/html/2610.02841#A4.F5)\), and the five prompt variants per rewriting category \(Table[9](https://arxiv.org/html/2610.02841#A9.T9)\)\.

Prompt Used for Natural Rewriting\[Rewriting Prompt\]During rewriting, you must not alter the factual content of the text\. All changes should be purely stylistic and decorative\.Figure 4:Prompt template used for natural rewriting\. The\[Rewriting Prompt\]placeholder is instantiated with one of five rewriting categories: Language Polishing, Fluency Improvement, Academic Style Alignment, Outcome\-Aware, or Cautious Tone\. See Table[9](https://arxiv.org/html/2610.02841#A9.T9)for details\.
## Appendix DFlip Rates Across All Models

Table[10](https://arxiv.org/html/2610.02841#A9.T10)reports flip rates across all models and rewriting conditions, complementing the probability shift analysis in Section[5](https://arxiv.org/html/2610.02841#S5)\.

Prompt Used for Controlled RewritingYou are a linguistic editor\. Your task is to inject exactly 1 word from the provided word list into the given claim\.Rules:•You MUST insert exactly 1 word from the word list below\. This is mandatory\.•Do not delete, replace, or reorder any existing words\. Every original word must remain\.•The injection should be stylistic only—it must not change the factual content\.•The injected word must be grammatically natural in its position\.Word list:
\{word\_list\}Original claim:
\{claim\}Output ONLY the modified claim with exactly one word inserted\. No explanation, no preamble\.Figure 5:Prompt template used for controlled rewriting\. The\{word\_list\}placeholder is instantiated with a randomly sampled subset of candidate words from either the decorative or conservative word pool, and\{claim\}is instantiated with the original claim\.
## Appendix EAnnotation Guidelines

Figure[6](https://arxiv.org/html/2610.02841#A9.F6)presents the annotation guidelines provided to annotators for evaluating semantic consistency between original and rewritten claims, as described in Section[3\.2\.1](https://arxiv.org/html/2610.02841#S3.SS2.SSS1.Px1)\.

## Appendix FWord Lists

Table[11](https://arxiv.org/html/2610.02841#A9.T11)lists the decorative and conservative words used in controlled rewriting, drawn from[Kobak et al\. \(2025\)](https://arxiv.org/html/2610.02841#bib.bib9)and[Matsui \(2025\)](https://arxiv.org/html/2610.02841#bib.bib14)\.

## Appendix GData Exclusions

A small number of claims are excluded from the analysis for two reasons\. First, the inject rewriting pipeline \(D1, C1\) occasionally produces an empty or malformed output; when this affects either member of a claim pair, both members are dropped to preserve the paired evaluation structure \(e\.g\., 29 pairs excluded for D1\)\. Natural rewrite conditions produce no such failures\. Second, for a small number of claims, all 10 self\-consistency responses from a given model are unparseable; these claims are excluded from that model’s evaluation\. Crucially, the affected claims are consistent across the baseline and all rewrite conditions for a given model, so no differential bias is introduced into the comparison\. In both cases the number of excluded items is small relative to the full evaluation set of 788 pairs\.

## Appendix HHuman Annotation

Cond\.Judg\.Pres\. \(%\)n2n\_\{2\}PoP\_\{o\}κ\\kappaAC1CT9678\.1270\.6670\.0880\.475C19784\.5260\.8460\.2460\.807OA9874\.5270\.7410\.2910\.591D110090\.0290\.8970\.3430\.877All39181\.81090\.7890\.2510\.706Table 7:Human annotation results\. Judg\.: number of judgments\. Pres\.: percentage of judgments finding the meaning preserved\.n2n\_\{2\}: claims with two judgments;PoP\_\{o\}\(observed agreement\), Fleiss’κ\\kappa, and Gwet’s AC1 are computed on these only\. CT and C1 are sampled from the unflagged subset\.To verify that rewrites preserve the meaning of the original claims, four annotators, all NLP researchers, answered a single binary question: does the rewrite keep the same meaning as the original? Annotators saw both claims side by side with the changed words highlighted, and the rewrite condition was hidden\. The annotators followed the annotation guidelines \(Figure[6](https://arxiv.org/html/2610.02841#A9.F6)\)\. We sampled 80 claims each from OA, CT, C1, and D1\. For CT and C1, we sampled only from the unflagged subset \(Section[B](https://arxiv.org/html/2610.02841#A2)\), since this is the subset our analysis relies on\. OA and D1 were sampled at random, since the filter targets hedging and these are boosting conditions\. Each claim was assigned to two annotators; as not all assignments were completed, the returned data contain 391 judgments over 282 claims \(109 with two judgments and 173 with one\)\.

Overall, annotators judged the meaning preserved in 81\.8% of judgments, ranging from 74\.5% \(OA\) to 90\.0% \(D1\), with 78\.1% for CT and 84\.5% for C1 \(Table[7](https://arxiv.org/html/2610.02841#A8.T7)\)\. On the 109 doubly annotated claims, observed agreement is 78\.9%, Fleiss’κ\\kappais 0\.251, and Gwet’s AC1 is 0\.706\. Since each condition contributes only 26–29 doubly annotated claims, per\-condition agreement estimates should be interpreted with caution\. We use Fleiss’κ\\kapparather than Cohen’sκ\\kappabecause different annotator pairs judged different claims, whereas Cohen’sκ\\kappaassumes a single fixed pair of annotators\. The lowκ\\kappareflects label skew: 83\.0% of these judgments mark the meaning as preserved, soκ\\kappaattributes most of the observed agreement to chance, whereas AC1 is robust to such skew\. The remaining disagreement is largely systematic\. The rate at which individual annotators judged a meaning change ranges from 4\.7% to 32\.7%, suggesting that assigning meaning\-preservation labels to scientific claims is inherently difficult\.

## Appendix IQualitative Examples

To complement the aggregate results, Table[8](https://arxiv.org/html/2610.02841#A9.T8)presents representative examples of rewritten claims together with the predictions of Qwen3\-VL\-8B before and after rewriting\. We include both high\-shift cases, where rewriting substantially changesp⁡\(support\)p\(\\text\{support\}\), and no\-shift cases, where the prediction remains stable despite the rewrite\.

Case 1: high shiftval\_fig\_0285Condition: CT Gold: Supportedpp: 0\.0→\\rightarrow0\.9 \(Δ​p=\+0\.9\\Delta p=\+0\.9\) Flipped: yesOriginalThe results indicated that there were more bacterial categories in clinical samples than cultures by media\.RewrittenThe resultssuggestedthata greater number ofbacterial categorieswere observedin clinical samples thanwere identified through culturemedia\.Evidence![[Uncaptioned image]](https://arxiv.org/html/2610.02841v1/case_examples/val_fig_0285.png)Case 2: high shiftval\_fig\_0052Condition: OA Gold: Supportedpp: 0\.9→\\rightarrow0\.0 \(Δ​p=−0\.9\\Delta p=\-0\.9\) Flipped: yesOriginalIt is also intriguing to observe from Figure 5 that Vanilla LLaMA\-2\-70B excelled in Writing, but full\-parameter fine\-tuning led to a decline in these areas, a phenomenon known as the “Alignment Tax” \(Ouyang et al\., 2022 \) \.RewrittenFigure 5compellingly demonstratesthatthevanilla LLaMA\-2\-70Bmodel attains peak performance on writing tasks, yet once the model is subjected tofull\-parameter fine\-tuningits scores markedly decline—a clear manifestation ofthewell\-documented“Alignment Tax” \(Ouyang et al\.,2022\)\.Evidence![[Uncaptioned image]](https://arxiv.org/html/2610.02841v1/case_examples/val_fig_0052.png)Case 3: no shiftval\_tab\_0177Condition: LP Gold: Supportedpp: 0\.6→\\rightarrow0\.6 \(Δ​p=0\.0\\Delta p=0\.0\) Flipped: noOriginalSecond, pairwise comparisons with the baseline show both CG and finetuning are effective in formality control, where CG has slightly higher win ratio than FT against the baseline\.RewrittenSecond, pairwise comparisons with the baseline showthatboth CG andfine\-tuningare effectiveforformality control,withCGachieving aslightly higher win ratio than FT against the baseline\.Evidence![[Uncaptioned image]](https://arxiv.org/html/2610.02841v1/case_examples/val_tab_0177.png)Case 4: no shift despite hedgingval\_tab\_0029Condition: C1 Gold: Supportedpp: 0\.6→\\rightarrow0\.6 \(Δ​p=0\.0\\Delta p=0\.0\) Flipped: noOriginalResults in Table 1 show that AD\-Drop achieves consistent improvement, boosting the average scores of BERT base and RoBERTa base by 0\.87 and 0\.62, respectively\.RewrittenResults in Table 1 show that AD\-Drop achievesseeminglyconsistent improvement, boosting the average scores of BERT base and RoBERTa base by 0\.87 and 0\.62, respectively\.Evidence![[Uncaptioned image]](https://arxiv.org/html/2610.02841v1/case_examples/val_tab_0029.png)Table 8:Example cases for Qwen3\-VL\-8B\-Instruct\.ppis the probability of*Supported*estimated from 10 sampled answers, before and after rewriting\. Words changed by the rewrite are in bold\. Evidence images are taken from SciClaimEval\([Ho et al\., 2026](https://arxiv.org/html/2610.02841#bib.bib6)\)\.Annotation Guidelines for Semantic ConsistencyI/O:
You will be shown pairs of scientific claims: an original and a rewritten version\.Input:original claim, rewritten claim, evidence image path \(for reference\)
Output:a binary label \(TRUE/FALSE\), and an optional note fieldTask:
Does the rewritten claim preserve the same factual meaning as the original?•TRUE:The core factual content is the same\. Differences in word choice, tone, emphasis, or formality are fine\.•FALSE:The rewrite adds, removes, or changes a factual detail \(e\.g\., a number, a comparison direction, or which method/dataset is referenced\)\.•If you can’t decide:Imagine verifying the claim against some evidence \(table, figure, etc\.\)\. Ask: does there exist any possible evidence that would support the original claim but not the rewritten one, or vice versa? If yes→\\rightarrowFALSE\. If no→\\rightarrowTRUE\.Examples:•“our method is 20% better”→\\rightarrow“our method is 25% better”→\\rightarrowFALSE\(changed number\)•“our method is 20% better”→\\rightarrow“our method is 20\.0001% better”→\\rightarrowFALSE\(even a trivial numerical change could be invalidated by evidence showing exactly 20%\)•“A is 20% better than B”→\\rightarrow“A is approximately 20% better than B”→\\rightarrowFALSE\(“approximately” widens the range of valid values\)•“our method is 20% better”→\\rightarrow“our fantastic method is 20% better”→\\rightarrowTRUE\(subjective—no evidence can distinguish “fantastic” from “not fantastic”\)•“consistently outperforms”→\\rightarrow“arguably outperforms”→\\rightarrowTRUE\(tone shift, but the factual claim is the same\)•“outperforms method A”→\\rightarrow“outperforms method B”→\\rightarrowFALSE\(different reference\)Figure 6:Annotation guidelines provided to annotators for evaluating semantic consistency between original and rewritten claims\.CategoryVariantPromptLanguagePolishing \(LP\)LP\-1Fix any grammar mistakes in the following text\.LP\-2Proofread this text and correct all grammatical errors\.LP\-3Can you fix the grammar of this text?LP\-4Please correct any grammatical issues in the text below\.LP\-5Review this text for grammar problems and rewrite it with corrections\.FluencyImprovement \(FI\)FI\-1Rewrite this text to make it more fluent and readable\.FI\-2Make this paragraph flow more naturally\.FI\-3Improve the readability of the following text\.FI\-4Smooth out the writing in this text so it reads better\.FI\-5Rephrase this text to improve its clarity and flow\.Academic StyleAlignment \(AS\)AS\-1Rewrite this to match the writing style expected at ACL\.AS\-2Adapt this text to fit the conventions of a AAAI paper\.AS\-3I am submitting this paper to NeurIPS\. Rewrite it to match the style and quality expected there\.AS\-4Revise this text so it reads like a typical ICLR submission\.AS\-5Adjust the tone and style of this text to align with papers published at ICML\.Outcome\-Aware \(OA\)OA\-1Rewrite this to be more impressive and compelling to reviewers\.OA\-2Polish this so it sounds more convincing to reviewers\.OA\-3Rewrite this to enhance the review score and maximize acceptance chances\.OA\-4Make this text as persuasive and impactful as possible for peer review\.OA\-5Strengthen the writing so that reviewers would strongly favor this paper\.CautiousTone \(CT\)CT\-1Rewrite this text to adopt a more cautious and measured tone, softening strong claims and avoiding overgeneralization\.CT\-2Rewrite this text to use a more careful and qualified tone, tempering strong assertions and introducing appropriate nuance\.CT\-3Revise this passage to emphasize balance and restraint, moderating claims and avoiding overly definitive language\.CT\-4Refine this passage to adopt a more reserved academic tone, softening conclusions and avoiding broad or sweeping statements\.CT\-5Rewrite this to sound more analytically cautious, reducing emphasis on certainty and acknowledging potential variability in interpretation\.Table 9:Prompt strategies for natural rewriting experiments\. Each category contains five prompt variants to control for variance introduced by prompt phrasing\.Natural RewritingControlled RewritingModelLPFIASOACTD1C1gemma\-4\-26B\-A4B\-itSupported4\.9%5\.2%7\.9%8\.9%7\.9%4\.5%5\.1%Refuted6\.7%7\.0%8\.5%8\.9%11\.5%5\.7%5\.5%gemma\-4\-31B\-itSupported2\.9%3\.2%3\.9%6\.1%4\.2%3\.0%2\.4%Refuted5\.7%6\.5%7\.9%9\.6%11\.8%5\.1%7\.2%GLM\-4\.6V\-FlashSupported6\.7%7\.5%8\.5%11\.0%10\.8%4\.9%6\.5%Refuted7\.9%9\.5%11\.4%10\.8%15\.5%5\.5%6\.4%InternVL3\-2BSupported5\.5%5\.7%5\.6%6\.0%4\.7%5\.1%4\.7%Refuted7\.4%8\.0%6\.5%8\.6%7\.7%7\.0%5\.9%InternVL3\-8BSupported5\.2%4\.9%5\.3%5\.2%5\.8%4\.3%3\.2%Refuted11\.9%10\.8%11\.3%13\.6%14\.7%11\.2%11\.5%InternVL3\-38BSupported7\.8%7\.1%7\.8%8\.7%6\.7%6\.7%6\.0%Refuted11\.5%11\.2%12\.2%14\.7%15\.1%10\.0%11\.1%Kimi\-VL\-A3BSupported6\.1%9\.8%8\.9%10\.3%8\.0%7\.0%6\.7%Refuted11\.3%14\.8%13\.5%14\.3%16\.1%9\.4%11\.6%Qwen3\-VL\-2B\-InstructSupported10\.9%10\.9%11\.3%13\.7%15\.4%12\.8%11\.4%Refuted12\.3%16\.8%15\.9%16\.7%15\.3%11\.7%12\.7%Qwen3\-VL\-4B\-InstructSupported11\.3%11\.6%12\.1%14\.8%15\.1%7\.9%9\.7%Refuted10\.8%9\.9%12\.1%12\.7%12\.7%8\.6%8\.8%Qwen3\-VL\-8B\-InstructSupported9\.3%10\.4%12\.1%14\.8%13\.5%7\.8%8\.7%Refuted7\.6%8\.3%8\.8%10\.2%11\.8%8\.2%7\.2%Qwen3\-VL\-32B\-InstructSupported6\.0%6\.2%8\.9%10\.7%9\.0%6\.3%6\.9%Refuted7\.0%6\.6%8\.0%9\.8%10\.7%5\.9%7\.2%

Table 10:Flip rate: percentage of claims whose predicted label \(threshold0\.50\.5onp⁡\(support\)p\(\\text\{support\}\)\) changes after rewriting, in either direction\. Rows group claims by gold label:Supported\(claim paired with supporting evidence\) andRefuted\(claim paired with refuting evidence\)\. Colour intensity scales with flip rate\.CategoryWordsConservative \(C\)acknowledges, allegedly, apparently, approximately, arguably, basically, broadly, comparatively, conceivably, essentially, generally, hinting, inadequately, largely, mainly, marginally, moderately, mostly, often, ostensibly, partially, perhaps, plausibly, possibly, potential, presumably, primarily, purportedly, relatively, roughly, seemingly, slightly, somewhat, supposedly, tentatively, typically, underexplored, unexplored, usuallyDecorative \(D\)accentuates, adept, advocates, affirming, aptly, attains, augmenting, avenue, boast, bolster, burgeoning, capitalizing, catalyze, combating, commendable, compelling, compellingly, comprehensive, consolidates, crafted, critical, crucial, culminating, deeper, dependability, dependable, detrimentally, disrupts, distinctive, effectively, effortlessly, elevate, emphasises, emphasising, emphasize, empowers, enduring, enhance, essential, excel, excellently, exceptional, exceptionally, exhaustive, expansive, formidable, fortify, foster, foundational, fresh, fundamental, groundbreaking, illuminate, imperative, impressive, impressively, ingenious, innovative, intriguing, invaluable, leverage, lucidly, meticulous, meticulously, multifaceted, notable, notably, noteworthy, optimizing, orchestrating, outperform, paving, pinpoint, pioneering, pivotal, poised, potent, pressing, promise, pronounced, propelling, remarkable, renowned, revolutionize, seamless, seamlessly, showcase, significant, solidify, strategically, streamline, substantiated, surged, surmount, surpass, swift, swiftly, thorough, transcend, transform, transformative, uncharted, uncovering, underscore, unlocking, unparalleled, unraveling, unveil, uphold, valuable, versatile, versatility, warranting, well\-roundedTable 11:Word pools used in inject rewriting\. These words are frequently included by LLMs during writing and serve as recognizable markers of LLM\-generated text\. Decorative words are sourced from[Kobak et al\. \(2025\)](https://arxiv.org/html/2610.02841#bib.bib9)and[Matsui \(2025\)](https://arxiv.org/html/2610.02841#bib.bib14); conservative words are drawn from the same sources and supplemented with LLM\-assisted hedging terms\. Category C \(Conservative\) contains 39 words; Category D \(Decorative\) contains 114 words\.

Similar Articles

Robust Critics: Defending LLMs Against Multi-Turn Attacks

arXiv cs.AI

This paper proposes Dialogue Critic Guided Sampling (DCGS), a framework that defends LLMs against multi-turn adversarial attacks by inferring user intent from conversation history and using value/regret-based critics to score responses, achieving improved robustness without fine-tuning.