Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength

arXiv cs.CL Papers

Summary

This paper introduces an LLM-as-a-judge method to measure perturbation strength for assessing self-consistency in LLM explanations, showing that input perturbations generally affect LLMs more strongly than CoT perturbations.

arXiv:2609.30849v1 Announce Type: new Abstract: Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose an LLM-as-a-judge approach to measure perturbation strength in a unified manner across input and CoT perturbations. We then evaluate the self-consistency in explanations generated from various LLMs under controlled strength conditions, ensuring a fair comparison across perturbation types. Experiments show that our proposed LLM-based perturbation strength measure outperforms other embedding- and probability-based approaches and that input perturbations generally affect LLMs more strongly than CoT perturbations. Our work suggests that judgments about a model's self-consistency is fair only within the same perturbation type.
Original Article
View Cached Full Text

Cached at: 09/28/26, 09:43 AM

# Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength
Source: [https://arxiv.org/html/2609.30849](https://arxiv.org/html/2609.30849)
Phuong Q\. LeAffiliation:School of Computing and Information SystemsThe University of Melbourne, Melbourne, AustraliaEmail:[phuongquynh\.le@student\.unimelb\.edu\.au](mailto:)Kemal Kurniawan††thanks:Part of work was done at The University of Melbourne\.Affiliation:School of Computer Science and EngineeringUniversity of New South Wales, Sydney, AustraliaEmail:[kemal\.kurniawan@unsw\.edu\.au](mailto:)Jey Han LauAffiliation:School of Computing and Information SystemsThe University of Melbourne, Melbourne, AustraliaEmail:[laujh@unimelb\.edu\.au](mailto:)

###### Abstract

Prior work has examined the self\-consistency of LLM\-generated explanations using surface\-level perturbation methods\. However, the strength of these perturbations is not explicitly measured and controlled\. In this work, we propose an LLM\-as\-a\-judge approach to measure perturbation strength in a unified manner across input and CoT perturbations\. We then evaluate the self\-consistency in explanations generated from various LLMs under controlled strength conditions, ensuring a fair comparison across perturbation types\. Experiments show that our proposed LLM\-based perturbation strength measure outperforms other embedding\- and probability\-based approaches and that input perturbations generally affect LLMs more strongly than CoT perturbations\. Our work suggests that judgments about a model’s self\-consistency is fair only within the same perturbation type\.

## 1Introduction

Large language models \(LLMs\) are often described as black boxes because users cannot easily access the models’ internal reasoning\. LLM\-generated explanations such as Chain\-of\-Thought \(CoT\) reasoning\([Wei et al\., 2022](https://arxiv.org/html/2609.30849#bib.bib1)\)can help verbalise LLMs’ decision\-making processes\. However, recent research has shown that LLM\-generated explanations can be inconsistent with the model’s behaviour, questioning whether they truly reflect the model’s internal reasoning\([Chen et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib2);[Turpin et al\., 2023](https://arxiv.org/html/2609.30849#bib.bib4);[Matton et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib3)\)\.

Self\-consistency checks are proposed to assess the alignment between a model’s generated explanations and its behaviour\. One common approach is the perturbation\-based method, which works by systematically perturbing the input, CoT reasoning, or model’s parameters based on the generated explanations and measuring the effect on the model’s output\([Lanham et al\., 2023](https://arxiv.org/html/2609.30849#bib.bib13);[Madsen et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib37);[Tutek et al\., 2025](https://arxiv.org/html/2609.30849#bib.bib9);[Matton et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib3)\)\. Ideally, if the explanations are consistent with the model’s behaviour, then such perturbations should result in a corresponding change in the model’s output\.

Large semantic change≠\\neqInterference with reasoning pathBacon is a type ofmeat…Original CoT stepBacon is a type ofvegetable…Perturbed CoT stepHIGHsemantic changeQuestion: If bacon is left too long on a hot stove topOptionsA\): it will be cooked perfectlyB\): it will be bacteria ladenC\): it will become blackenedD\): it will be left rawOriginal CoTPerturbed CoTBacon =meatwhen left on a hot stovefor too long→\\rightarrowgets burnedBacon =vegetablewhen left on a hot stovefor too long→\\rightarrowgets burned✓\\checkmarkReasoning path preserved despite semantic change

Figure 1:A high semantic change between the original and perturbed text does not necessarily interfere with the reasoning path\. Changing "bacon" from "meat" to "vegetable" is a high semantic change on its own; however, when being aware of the question context, this change does not interfere with the reasoning path\.Existing literature primarily focuses on what and how to perturb, with much less attention given to measuring the perturbation strength\([Lanham et al\., 2023](https://arxiv.org/html/2609.30849#bib.bib13);[Paul et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib39);[Tutek et al\., 2025](https://arxiv.org/html/2609.30849#bib.bib9);[Matton et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib3)\)— that is, assessing whether the perturbation is strong enough to warrant a change in the model’s answer\. Without proper measure and control of perturbation strength, they risk falsely concluding that the model has poor self\-consistency when the perturbations are unknowingly weak, i\.e\. the change does not interfere with the reasoning path and therefore should not affect the model’s answer \([Figure1](https://arxiv.org/html/2609.30849#S1.F1)\)\. Furthermore, when comparing the effects of different perturbation types on model behaviour, the lack of explicit perturbation strength measure can lead to unfair comparisons, as there is no guarantee that perturbations of the same strength level are compared\([Tutek et al\., 2025](https://arxiv.org/html/2609.30849#bib.bib9)\)\. This calls for a reliable measure of perturbation strength that is applicable across perturbation types\.

To address this, we propose an LLM\-as\-a\-judge approach to measure perturbation strength in a unified manner across perturbation types \(input and CoT\), regardless of their generation methods \([Section2\.2](https://arxiv.org/html/2609.30849#S2.SS2)\) \. We design a simple taxonomy of perturbation strength and incorporate this into the LLM judge prompt\. We show that our proposed approach outperforms two alternative measures in terms of their correlations with human judgments\. Using our proposed framework, we find that more than 50% of perturbations in existing self\-consistency datasets are weak, questioning their suitability for self\-consistency checks\. The proposed measure allows us to capture \(a\) responsive self\-consistency, which measures whether the model changes its answer under strong perturbations, and \(b\) robust self\-consistency, which measures whether the model preserves its answer under weak perturbations\. We find that models often exhibit high responsive but low robust self\-consistency under input perturbations compared to CoT ones, suggesting that input perturbations affect the models more strongly\.

In summary, our contributions are as follows:

1. 1\.We design criteria and develop a framework using LLM\-as\-a\-judge approach to measure the strength of surface\-level perturbations, applicable to both input and CoT perturbations\.
2. 2\.We show that most perturbations in existing datasets are weak, questioning their suitability for self\-consistency checks\.
3. 3\.We introduce 2 complementary notions of self\-consistency, namely responsive and robust self\-consistency\.
4. 4\.We demonstrate that input perturbations generally affect the models more strongly than CoT perturbations\.

## 2Related Work

### 2\.1Self\-consistency

In this work, we defineself\-consistencyas the extent to which a model’s behaviour aligns with the reasoning presented in its generated explanations\.[Turpin et al\. \(2023\)](https://arxiv.org/html/2609.30849#bib.bib4)and[Matton et al\. \(2024\)](https://arxiv.org/html/2609.30849#bib.bib3)found that models can mask influential social biases by rationalising their answers through less influential concepts in their explanations, leading to inconsistencies between the model’s behaviour and their stated reasoning\. Such inconsistencies question thefaithfulnessof the LLM\-generated explanation — i\.e\. the extent to which an explicit explanation truly reflects the model’s internal processes\([Jacovi and Goldberg, 2020](https://arxiv.org/html/2609.30849#bib.bib34)\)— since an inconsistent explanation means the model must have followed a different internal reasoning process to arrive at the answer\. That said, surface\-level consistency does not tell us the whole picture about faithfulness, since a model producing a self\-consistent explanation may still use a different internal reasoning path to arrive at the same answer\. For that, internal analysis is needed\([Parcalabescu and Frank, 2024](https://arxiv.org/html/2609.30849#bib.bib35);[Matton et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib3);[Tutek et al\., 2025](https://arxiv.org/html/2609.30849#bib.bib9)\)\.

Our paper focuses on self\-consistency\. It is still an important asessment for understanding model behaviour \(even if they cannot tell us about faithfulness\) and we can develop perturbation methods that are model\-agnostic and applicable across diverse use cases\. We focus on two types of perturbation: input and CoT perturbation\.

### 2\.2Perturbation\-based Methods

Perturbation\-based method first emerged as an idea to create feature attribution explanations \(LIME\)\([Ribeiro et al\., 2016](https://arxiv.org/html/2609.30849#bib.bib36)\)\. Recent research has adopted this idea to evaluate LLM\-generated explanations\.

##### Input Perturbation

This refers to the perturbations applied at the input text\([Turpin et al\., 2023](https://arxiv.org/html/2609.30849#bib.bib4);[Chen et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib2);[Madsen et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib37);[Matton et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib3)\)\. We can target perturbations on predefined concepts\([Turpin et al\., 2023](https://arxiv.org/html/2609.30849#bib.bib4);[Matton et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib3)\), or more freely by editing the input text without any predefined target\([Chen et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib2);[Madsen et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib37)\), enabling a wider range of modifications\. Therefore, in this work, we adopt the latter free\-form input perturbation approach for self\-consistency assessment\.

##### CoT Perturbation

This approach modifies the CoT instead of the input\.[Lanham et al\. \(2023\)](https://arxiv.org/html/2609.30849#bib.bib13)used a pre\-trained model to generate a mistaken version of the original CoT steps\. In contrast,[Paul et al\. \(2024\)](https://arxiv.org/html/2609.30849#bib.bib39)used the CoT generated for a counterfactual question as the perturbed CoT\. The former is expected to produce more diverse perturbed CoTs, while the latter is likely to generate strong perturbations, as the reasoning originates from a counterfactual question\. Therefore, in this work, we use the CoT perturbation approach by[Lanham et al\. \(2023\)](https://arxiv.org/html/2609.30849#bib.bib13)to evaluate our perturbation strength measure across more diverse perturbations\.

##### Perturbation Strength

Existing studies have attempted to control perturbation strength\. For example, the AOC metric by[Lanham et al\. \(2023\)](https://arxiv.org/html/2609.30849#bib.bib13)implicitly treats the number of perturbed steps as a proxy for perturbation strength\. However, it is unclear whether perturbing the same number of steps produces perturbations of comparable strength across different tasks\.[Matton et al\. \(2024\)](https://arxiv.org/html/2609.30849#bib.bib3)modified concepts that are identified as important or unimportant based on the generated explanation\. This could be used to infer perturbation strength, as modifying important concepts can be expected to result in stronger perturbations\. Nevertheless, these methods can only provide an indirect indication of perturbation strength and are specific only to the perturbation approaches used in the respective study\. In other words, they lack generalisability across different perturbation types and variations\.

## 3Perturbation Strength Assessment

Automatic generation of perturbations does not guarantee that the changes are significant enough to alter the model’s answer\. As such, it is important to also measure perturbation strength when we are assessing self\-consistency\. We develop a unified approach to quantitatively measure the strength of both input and CoT perturbations on a discrete scale\. We then evaluate its correlation with human judgments and use it to reveal that most perturbations from prior work are weak\.

### 3\.1Datasets

We use the same data as[Tutek et al\. \(2025\)](https://arxiv.org/html/2609.30849#bib.bib9)who constructed the data from a subset of four multiple\-choice QA datasets: ARC\-Challenge\([Clark et al\., 2018](https://arxiv.org/html/2609.30849#bib.bib5)\), OpenBookQA\([Mihaylov et al\., 2018](https://arxiv.org/html/2609.30849#bib.bib8)\), the Sports subtask of BigBench\-Hard\([Srivastava et al\., 2023](https://arxiv.org/html/2609.30849#bib.bib6)\)and StrategyQA\([Geva et al\., 2021](https://arxiv.org/html/2609.30849#bib.bib7)\)\. Each subset consists of 230 multiple\-choice questions, where each question has a set of answer options and corresponding correct answer\. In addition, each question has both original and perturbed CoT steps\. Specifically, for each question in each dataset,[Tutek et al\. \(2025\)](https://arxiv.org/html/2609.30849#bib.bib9)used four instruction\-tuned models to generate a sequence of CoT steps to the question for each model\.111The models are: Llama\-3\.2\-3B\-Instruct, Llama\-3\-8B\-Instruct\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib10)\); Mistral\-7B\-Instruct\-v0\.2\([Jiang et al\., 2023](https://arxiv.org/html/2609.30849#bib.bib11)\), and Phi\-3\-mini\-4k\-Instruct\([Abdin et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib12)\)We refer to these CoT steps asoriginal\. Next, they usedgpt\-4o\-mini\([OpenAI, 2024](https://arxiv.org/html/2609.30849#bib.bib19)\)to generate a mistaken version of each CoT step\. Each of these mistaken CoT steps then replaced the corresponding original step while keeping all the other steps the same to construct new perturbed CoTs\. In other words, each new perturbed CoT has only one mistaken step\.[Figures2](https://arxiv.org/html/2609.30849#S3.F2)and[3](https://arxiv.org/html/2609.30849#S3.F3)respectively illustrate the structure of the StrategyQA dataset with LLama\-3B model and an example of a mistaken step replacing the original step\.

\{forest\}

Figure 2:Detailed structure of a question from StrategyQA−\-LLama\-3B perturbed dataset\. Each step in the original CoT generated by LLama\-3B was perturbed to constructmmnew perturbed CoTsR1′,…,Rm′R^\{\\prime\}\_\{1\},\\ldots,R^\{\\prime\}\_\{m\}, whereRi′R^\{\\prime\}\_\{i\}denotes the CoT with only theit​hi^\{th\}step perturbed andmmis the number of steps in the original CoT\.Original CoT
1\. Sandals are designed for warm weather\.2\.Snow iscold, and it can beslippery\.
3\. Wearing sandals in snow can cause your feet to get cold and wet\.Perturbed CoT step:Snow iswarm, and it can be verysticky\.Perturbed CoT
1\. Sandals are designed for warm weather\.2\.Snow iswarm, and it can be verysticky\.
3\. Wearing sandals in snow can cause your feet to get cold and wet\.Figure 3:Example of CoT steps where the second step is perturbed\.Hereinafter, each of the four models that generate the original CoT will be referred to as aCoT model\. There are 16 \(4 datasets × 4 CoT models\) perturbed datasets in the work of[Tutek et al\. \(2025\)](https://arxiv.org/html/2609.30849#bib.bib9)\. Each perturbed dataset has 230 questions, with each question associated with perturbed CoTs generated by a specific CoT model on the corresponding dataset\.

### 3\.2Data Annotation

To evaluate the performance of perturbation strength measures, we need a ground truth reference, which will be the strength assigned by a human annotator to each perturbed text\. To ensure consistent assessment of perturbation strength, a set of criteria \([Table1](https://arxiv.org/html/2609.30849#S3.T1)\) that aligns with our definition of perturbation strength is developed\.

#### 3\.2\.1Perturbation Strength Criteria

In this study,perturbation strengthis determined by how much the change interferes with the original reasoning path and how much it shifts support from the original answer to another option as shown in[Table1](https://arxiv.org/html/2609.30849#S3.T1)\. This definition provides clear expectations regarding model behaviour\. Specifically, a strong change that substantially interferes with the original reasoning path and redirects support toward a different answer, is expected change the model’s answer\. In contrast, a weak change that minimally interferes with the original reasoning path, is expected to preserve the original answer or at least, not expected to change the model’s answer\.

Table 1:Key perturbation strength levels and criteria\. The full guidelines are shown in[Figure11](https://arxiv.org/html/2609.30849#A5.F11)\.
#### 3\.2\.2Sampled Data for Annotation

For annotation and evaluation, we sample 96 questions from 4 datasets described in[Section3\.1](https://arxiv.org/html/2609.30849#S3.SS1)along with their associated perturbed CoTs, ensuring equal representation across all dataset–CoT model pairs \([AppendixC](https://arxiv.org/html/2609.30849#A3)\)\.

##### Constructing Diverse Perturbed CoTs

Existing datasets in[Section3\.1](https://arxiv.org/html/2609.30849#S3.SS1)lack a diverse range of perturbation strengths, as we observe that single\-step perturbations often have low impact on the overall reasoning, which can lead to low perturbation strength most of the time\. They also lack paraphrased CoT steps, limiting our ability to evaluate whether perturbation strength measures would assign zero strength to semantically equivalent texts\. To address these limitations, for each sampled question, we randomly combine some of their perturbed steps or paraphrase all steps \([AppendixD](https://arxiv.org/html/2609.30849#A4)\), producing perturbed CoTs with more diverse perturbation strengths for annotation and evaluation\.

##### Input Perturbation

The datasets in[Section3\.1](https://arxiv.org/html/2609.30849#S3.SS1)do not provide input perturbations\. On QA datasets, input perturbation refers to perturbing the question text\. For each sampled question,gpt\-4o\-miniis prompted to perturb the question text\. The prompt for perturbing inputs is designed to be as similar as possible to how perturbed CoT steps were generated \([AppendixB](https://arxiv.org/html/2609.30849#A2)\)\. The original CoT is also provided to encourage generating perturbations more relevant to the reasoning\.[Figure4](https://arxiv.org/html/2609.30849#S3.F4)shows our result of perturbing a question\.

Original Question: Is it safe to wear sandals insnow?OptionsA\): YesB\): NoPerturbed Question: Is it safe to wear sandals inwarm weather?OptionsA\): YesB\): NoFigure 4:Example of a question and its perturbed version\.Overall, we now have a sample of 96 perturbed CoTs and 96 perturbed inputs \(questions\) that will be used for annotation and evaluation of perturbation strength measures\.

#### 3\.2\.3Annotation Procedure

A primary annotator assigns a strength level from 0 to 3 for 96 CoT and 96 input perturbations generated in[Section3\.2\.2](https://arxiv.org/html/2609.30849#S3.SS2.SSS2), using the criteria presented in[Table1](https://arxiv.org/html/2609.30849#S3.T1)\. These annotated strengths serve as the ground truth reference for evaluating perturbation strength measures\. The full perturbation strength guidelines and annotation procedure are provided in[SectionE\.1](https://arxiv.org/html/2609.30849#A5.SS1)\.

##### Inter\-Annotator Agreement

A second annotator independently annotated a subset of CoT and input perturbations using the same criteria in[Table1](https://arxiv.org/html/2609.30849#S3.T1)\. Kappa scores\([Cohen, 1960](https://arxiv.org/html/2609.30849#bib.bib14);[Cohen, 1968](https://arxiv.org/html/2609.30849#bib.bib30)\)\(0\.855 for input and 0\.649 for CoT perturbation\) indicate a substantial agreement between the annotators, suggesting that the criteria are reliable\. Disagreements often occur when the true strength is 1 or 2, which is expected as intermediate strength levels can be more ambiguous than the extreme cases \(0 and 3\)\. More details are provided in[SectionE\.2](https://arxiv.org/html/2609.30849#A5.SS2)\.

### 3\.3Perturbation Strength Measures

We propose three perturbation strength measures: cosine distance, change in surprisals, and LLM\-as\-a\-judge\. These methods are chosen to represent embedding, probability, and LLM based approaches respectively\.

##### Cosine distance

Given the original and perturbed texts \(input or CoT\), we first compute their embeddings usingall\-MiniLM\-L6\-v2\([Wang et al\., 2020](https://arxiv.org/html/2609.30849#bib.bib17)\)\(see[SectionF\.1](https://arxiv.org/html/2609.30849#A6.SS1)for justification\)\. Then, we compute their cosine similarity\([Salton, 1989](https://arxiv.org/html/2609.30849#bib.bib16);[Sidorov et al\., 2014](https://arxiv.org/html/2609.30849#bib.bib15)\)to capture their semantic closeness\. Cosine distance is then computed as1−1\-cosine similarity to capture their difference \(distance\)\. Larger cosine distance between the original and perturbed text indicates stronger semantic shifts, i\.e\. stronger perturbation\.

##### Change in surprisal

Surprisal is a measure of information content in language, defined as the negative log\-probability of a token given its preceding context\([Shannon, 1951](https://arxiv.org/html/2609.30849#bib.bib20);[Hale, 2001](https://arxiv.org/html/2609.30849#bib.bib21);[Levy, 2008](https://arxiv.org/html/2609.30849#bib.bib22)\)\. For a sequence ofnntokens, denoted𝐱=\(w1,…,wn\)\\mathbf\{x\}=\(w\_\{1\},\\dots,w\_\{n\}\), surprisal is computed as the sum ofnntoken\-level negative log\-probabilities

I\(𝐱\)=−∑i=1nlog2P\(wi∣w<i\)\.\\text\{I\}\(\\mathbf\{x\}\)=\-\\sum\_\{i=1\}^\{n\}\\log\_\{2\}P\(w\_\{i\}\\mid w\_\{<i\}\)\.\(1\)We obtain the probabilities of tokensP⁡\(wi∣w<i\)P\(w\_\{i\}\\mid w\_\{<i\}\)from the GPT\-2 model\([Radford et al\., 2019](https://arxiv.org/html/2609.30849#bib.bib24)\)\([SectionF\.2](https://arxiv.org/html/2609.30849#A6.SS2)\)\.

The mean surprisal of the sequence, denotedh​\(𝐱\)\\text\{h\}\(\\mathbf\{x\}\)can be computed as

h​\(𝐱\)=I​\(𝐱\)n\\text\{h\}\(\\mathbf\{x\}\)=\\frac\{\\text\{I\}\(\\mathbf\{x\}\)\}\{n\}which normalises for sentence length\. The change in \(mean\) surprisal between the original text𝐱\\mathbf\{x\}and the perturbed text𝐱′\\mathbf\{x^\{\\prime\}\}is then defined as

Δ​h=\|h​\(𝐱\)−h​\(𝐱′\)\|,\\Delta\\text\{h\}=\\left\|\\text\{h\}\(\\mathbf\{x\}\)\-\\text\{h\}\(\\mathbf\{x^\{\\prime\}\}\)\\right\|,which gives us a notion of how much information content has changed\. A higher change in surprisal indicates a higher perturbation strength\.

##### LLM\-as\-a\-judge

Thegemini\-3\-flashmodel\([Google DeepMind, 2025](https://arxiv.org/html/2609.30849#bib.bib27)\)is prompted to assess the strength of perturbed questions and CoTs\. The prompt includes the original question, original CoT and the perturbed question/CoT to be assessed \(Appendix[Figures15](https://arxiv.org/html/2609.30849#A6.F15)and[16](https://arxiv.org/html/2609.30849#A6.F16)\), and a perturbation strength rating guide set as the system instruction\. Also, we provide one example for each perturbation strength level and its justification as in\-context examples for each dataset \(Appendix[Figures12](https://arxiv.org/html/2609.30849#A5.F12)and[13](https://arxiv.org/html/2609.30849#A5.F13)\)\. These dataset\-specific examples are included in the prompts to enable few\-shot in\-context learning\.

The perturbation strength criteria provided togemini\-3\-flashis identical to the one for human annotators in[Table1](https://arxiv.org/html/2609.30849#S3.T1)\. There are only minor modifications in the full guidelines to make it more suitable for LLMs, such as specifying the LLM’s role in the system instruction \([SectionF\.3](https://arxiv.org/html/2609.30849#A6.SS3)\)\.

### 3\.4Evaluation and Results

Each perturbation strength measure is used to produce a numerical value representing the perturbation strength between the original and perturbed text \(question or CoT\)\. Pearson and Spearman correlation coefficients, as well as Cohen’s Linear Kappa score are computed between the predicted strength \(by each method\) and true strength \(from primary annotator\) to quantify how well the perturbation strength measures correlate with human judgments\. The results are presented in[Table2](https://arxiv.org/html/2609.30849#S3.T2)\.

CosineChange inLLMdistancesurprisalPearson−0\.056\-0\.056−0\.283\-0\.2830\.888Spearman−0\.089\-0\.089−0\.156\-0\.1560\.877Kappa––0\.801\(a\)CoT perturbation
CosineChange inLLMdistancesurprisalPearson−0\.096\-0\.0960\.0660\.0660\.885Spearman−0\.098\-0\.0980\.0470\.0470\.867Kappa––0\.787\(b\)Input perturbation

Table 2:Correlations and Kappa scores of different perturbation strength measures\. Kappa scores are only computed for LLM because it is designed for categorical data, which matches LLM’s outputs for perturbation strength\.[Table2](https://arxiv.org/html/2609.30849#S3.T2)shows that the LLM method outperforms other measures in assessing perturbation strength\. It consistently achieves high correlations with human judgments for both input and CoT perturbations\. This suggests that the question context and the perturbation strength criteria are essential to accurately determine the perturbation strength\. Therefore, simply measuring semantic change between the original and the perturbed versions is not sufficient since a large semantic change may not interfere with the reasoning path \([Figure1](https://arxiv.org/html/2609.30849#S1.F1)\)\. We chose the LLM method as our perturbation strength measure in subsequent analyses\.

[Table2](https://arxiv.org/html/2609.30849#S3.T2)also shows that both cosine distance and change in surprisals are mostly uncorrelated with human judgments\. An exception to this is change in surprisals for CoT perturbations where it achieves a weak, negative correlation \([Table2\(a\)](https://arxiv.org/html/2609.30849#S3.T2.st1)\)\. We find that this is mainly due to zero\-strength perturbations, i\.e\. paraphrases \([SectionG\.1](https://arxiv.org/html/2609.30849#A7.SS1)\)\. Since these methods primarily capture surface\-level semantic change and cannot incorporate the task, reasoning context or criteria, they are unable to effectively measure perturbation strength\.

### 3\.5Applications of the Proposed Measure

We apply our chosen perturbation strength measure \(i\.e\. LLM\-as\-a\-judge\) to the datasets in[Section3\.1](https://arxiv.org/html/2609.30849#S3.SS1), where each perturbed CoT has only one perturbed step \(that is, we discard the combined and paraphrased CoTs we created earlier in[Section3\.2\.2](https://arxiv.org/html/2609.30849#S3.SS2.SSS2)\)\. The distribution of strength levels across these perturbed CoTs is presented in[Figure5](https://arxiv.org/html/2609.30849#S3.F5)\. The figure shows that in all datasets, more than half of the perturbed CoTs are weak \(strength of 0 or 1\) and that for most CoT models, less than 10% of the perturbed CoTs are strong \(strength of 3\)\.

These perturbed CoTs were used by[Tutek et al\. \(2025\)](https://arxiv.org/html/2609.30849#bib.bib9)to see whether the CoT models change their answer after the perturbations\. However, did not verify the strength of the perturbed CoTs because they lacked a CoT perturbation strength measure\. As a result, the models were expected to change their answer even when the perturbation strength is low\. Concretely, this means that their results may not correctly represent LLM behaviour under CoT perturbations\. More applications of our proposed measure, such as enabling verification of existing perturbations, are discussed in[SectionA\.1](https://arxiv.org/html/2609.30849#A1.SS1)\.

Figure 5:Proportion of perturbed CoTs at each perturbation strength for datasets described in[Section3\.1](https://arxiv.org/html/2609.30849#S3.SS1)\.

## 4CoT vs Input Perturbations

Our proposed LLM\-based perturbation strength measure enables a fair comparison between CoT and input perturbations at equivalent strength levels\. In this section, we perform this comparison and evaluate how CoT and input perturbations influence the assessment of self\-consistency\.

### 4\.1Metric

We useflip rate as a function of perturbation strengthas the key metric to assess self\-consistency in this framework\. We define flip rate as follows\.

Letsis\_\{i\}denote the perturbation strength of instanceii\. The perturbation could be either a perturbed question or a perturbed CoT\. Letyiy\_\{i\}andyi′y\_\{i\}^\{\\prime\}denote the model’s answer to instanceiibefore and after the perturbation, respectively\. For a fixed strengthss, we can compute the flip rate of that strength by

FR⁡\(s\)=∑i𝟏\{si=s∧yi≠yi′\}∑i𝟏\{si=s\}\.\\mathrm\{FR\}\(s\)=\\frac\{\\sum\_\{i\}\\mathbf\{1\}\\\{s\_\{i\}=s\\land y\_\{i\}\\neq y\_\{i\}^\{\\prime\}\\\}\}\{\\sum\_\{i\}\\mathbf\{1\}\\\{s\_\{i\}=s\\\}\}\.\(2\)This function measures, among all instances with perturbation strengthss, how many result in a change in the model’s answer \(flip\)\.

### 4\.2Data Construction

To compute the flip rates using[Equation2](https://arxiv.org/html/2609.30849#S4.E2)for both CoT and input perturbations, we need the model’s answers before and after each type of perturbation, as well as their perturbation strengths\. For CoT perturbations, we use models’ answers provided by[Tutek et al\. \(2025\)](https://arxiv.org/html/2609.30849#bib.bib9)and the strengths of all perturbed CoTs obtained in[Section3\.5](https://arxiv.org/html/2609.30849#S3.SS5)\. Therefore, we only need to construct additional data for input perturbations\. The generation of input perturbations is identical to the procedure in[Section3\.2\.2](https://arxiv.org/html/2609.30849#S3.SS2.SSS2.Px1)\. The extraction of model’s answers to these perturbed questions ensures that it is consistent with the procedures used for CoT perturbations\. Details are provided in[AppendixH](https://arxiv.org/html/2609.30849#A8)\.

### 4\.3Self\-consistency Measurements

To provide a comprehensive view of self\-consistency, we use flip rate as a function of perturbation strength \(FR⁡\(s\)\\mathrm\{FR\}\(s\)\) from[Equation2](https://arxiv.org/html/2609.30849#S4.E2)to capture its different aspects\. Specifically, we define two complementary notions of self\-consistency:

- •Responsive self\-consistency: sensitivity to strong perturbations, measured byFR⁡\(3\)\\mathrm\{FR\}\(3\), where higher is better
- •Robust self\-consistency: insensitivity to weak perturbations, measured byFR⁡\(0\)\\mathrm\{FR\}\(0\), where lower is better

Responsive self\-consistency considers the flip rate under strong perturbation, where changes substantially interfere with the reasoning and the model is expected to change its answer\. In contrast, robust self\-consistency considers the flip rate under negligible perturbations \(e\.g\., paraphrasing\), where changes are not enough to interfere with the reasoning and thus, the model has no reason to change its answer \([Section3\.2\.1](https://arxiv.org/html/2609.30849#S3.SS2.SSS1)\)\.

### 4\.4Results

#### 4\.4\.1Responsive and Robust Self\-consistency

Figure 6:ARC\-Challenge flip rates of input and CoT perturbations across 4 CoT models[Figure6](https://arxiv.org/html/2609.30849#S4.F6)presents flip rates of each perturbation strength for CoT and input perturbations in ARC\-Challenge dataset\. Flip rates for OpenBookQA, Sports, and StrategyQA exhibit similar patterns \(Appendix[Figures19](https://arxiv.org/html/2609.30849#A9.F19),[20](https://arxiv.org/html/2609.30849#A9.F20)and[21](https://arxiv.org/html/2609.30849#A9.F21)\)\. Therefore, the following analyses apply to these datasets as well\.

Overall, for both CoT and input perturbations, stronger perturbations generally result in higher flip rates\. In addition, flip rates are generally higher for input perturbations than for CoT perturbations, even at zero strength\. There are exceptions to this, with LLaMA\-3\-3B on Sports in[SectionI\.1](https://arxiv.org/html/2609.30849#A9.SS1)being the most prominent\. We discuss possible explanations later in this section\.

Looking at responsive and robust self\-consistency, under input perturbations, nearly all models exhibit high responsive self\-consistency \(i\.e\. highFR⁡\(3\)\\mathrm\{FR\}\(3\)\)\. However, still under input perturbations, the models exhibit relatively low robust self\-consistency, asFR⁡\(0\)\\mathrm\{FR\}\(0\)is substantially greater than zero \(it frequently exceeds 20% across all datasets\)\. In contrast, under CoT perturbations, all models consistently exhibit high robust self\-consistency, asFR⁡\(0\)\\mathrm\{FR\}\(0\)remains tiny\. Specific examples are presented in[SectionI\.2](https://arxiv.org/html/2609.30849#A9.SS2)\.

One possible explanation for why input perturbations appear to have a weaker effect than CoT perturbations, such as for LLaMA\-3B on the Sports dataset \([Figure20](https://arxiv.org/html/2609.30849#A9.F20)\), is data contamination \([SectionA\.2](https://arxiv.org/html/2609.30849#A1.SS2)\)\. If a model was exposed to some questions from a dataset during training, it may be less likely to change its answers when it encounters perturbed questions as it has memorised the original question–answer pairs\([Carlini et al\., 2021](https://arxiv.org/html/2609.30849#bib.bib31);[Cheng et al\., 2025](https://arxiv.org/html/2609.30849#bib.bib32);[Golchin and Surdeanu, 2025](https://arxiv.org/html/2609.30849#bib.bib33)\)\.

#### 4\.4\.2Overall Self\-consistency

To balance responsive and robust self\-consistency, we define a single scalar for*overall*self\-consistency as the subtraction ofFR⁡\(0\)\\mathrm\{FR\}\(0\)fromFR⁡\(3\)\\mathrm\{FR\}\(3\), i\.e\.FR⁡\(3\)−FR⁡\(0\)\\mathrm\{FR\}\(3\)\-\\mathrm\{FR\}\(0\)\. A larger value indicates stronger overall self\-consistency, as it reflects a highFR⁡\(3\)\\mathrm\{FR\}\(3\)and a lowFR⁡\(0\)\\mathrm\{FR\}\(0\)\. In other words, the model changes its answer only when expected \(high responsive and high robust self\-consistency\)\.

Table 3:Indicator for overall self\-consistency of 4 CoT models across all datasets under CoT and input perturbations, measured byFR⁡\(3\)−FR⁡\(0\)\\mathrm\{FR\(3\)\-FR\(0\)\}\. The bold fields indicate highest value for that dataset\.[Table3](https://arxiv.org/html/2609.30849#S4.T3)presentsFR⁡\(3\)−FR⁡\(0\)\\mathrm\{FR\}\(3\)\-\\mathrm\{FR\}\(0\)for all CoT model–dataset combinations under CoT and input perturbations, which indicate overall self\-consistency\. The table shows that Phi\-3 is the most self\-consistent on OpenBookQA, while Mistral\-2 is the most self\-consistent on StrategyQA for both CoT and input perturbations\. LLaMA\-3B generally exhibits high overall self\-consistency under CoT perturbations, but performs poorly under input perturbations, whereas LLaMA\-8B shows the opposite pattern\. Given such performance variation, models that demonstrate high self\-consistency across perturbation methods and datasets provide stronger evidence that they are self\-consistent, compared to models that exhibit high self\-consistency in one but not another\.

#### 4\.4\.3Discussion

Our results suggest that using different perturbation methods \(such as input and CoT perturbations\) can lead to different conclusions\. In addition, since input perturbations generally affect the models more strongly than CoT perturbations, we suggest that judgments about a model’s self consistency should be made with respect to other models under the same perturbation type\. For example, in the case of LLaMA\-3 \(8B\) on ARC\-Challenge dataset in[Figure6](https://arxiv.org/html/2609.30849#S4.F6), one should claim that the model has high responsive self\-consistency under input perturbation but low responsive self\-consistency under CoT one because it has highFR⁡\(3\)\\mathrm\{FR\}\(3\)under the former and lowFR⁡\(3\)\\mathrm\{FR\}\(3\)under the latter, relative to other models under the same type, not simply because 29 is less than 84\. We rely on relative model performance as there are currently no established ranges of flip rates corresponding to good self\-consistency\. Future work could address this by establishing such ranges\.

Our perturbation strength\-aware framework opens the possibility of a more nuanced interpretation by distinguishing when a model should change its answer and when it should not\. This enhances existing self\-consistency studies, particularly in two cases: \(1\) comparing the effect of different perturbation types across models \(like[Figure6](https://arxiv.org/html/2609.30849#S4.F6)\), and \(2\) evaluating self\-consistency \(by providing better control over perturbation strength and enabling verification of existing perturbations\)\.

## 5Conclusion and Future Work

In this study, we designed criteria and developed a framework to measure perturbation strength using an LLM judge\. We demonstrated that our method outperforms other embedding\- and probability\-based approaches\. We then applied our proposed method and found that input perturbations affect the models more strongly than CoT perturbations, suggesting that judgments about a model’s self\-consistency is fair only within the same perturbation type\. Furthermore, by controlling for perturbation strength, we capture 2 complementary notions of self\-consistency: responsive and robust\.

Future work could improve our perturbation strength measure by reducing ambiguity in intermediate strength levels \([SectionE\.2](https://arxiv.org/html/2609.30849#A5.SS2)\)\. A more ambitious direction is to develop a unified framework to not only measure strength of surface\-level perturbations but also of other forms such as parameter intervention\([Tutek et al\., 2025](https://arxiv.org/html/2609.30849#bib.bib9)\)and embedding perturbation\([Madani et al\., 2025](https://arxiv.org/html/2609.30849#bib.bib38)\)\.

## Limitations

Our self\-consistency analysis relies on the assumption that our perturbation strength measure reflects human judgments perfectly\. However, intermediate strength levels such as 2 can sometimes be misclassified as the strongest level 3 \(discussed in detail in[SectionE\.2](https://arxiv.org/html/2609.30849#A5.SS2)\)\. Such misclassifications may undermine the reliability of our findings\. That said, our perturbation strength measure has the best correlation with human judgments as reported in[Table2](https://arxiv.org/html/2609.30849#S3.T2)\. Therefore, while it is not perfect, our measure reflects human judgments much more accurately than its alternatives\.

In addition, the number of perturbed CoTs at strength level 3 in our self\-consistency analysis is substantially lower than other strength levels \([Figure5](https://arxiv.org/html/2609.30849#S3.F5)\)\. Therefore,FR⁡\(3\)\\mathrm\{FR\}\(3\)is estimated from fewer observations, which may result in higher variance for this particular measure\. Although higher\-strength CoT perturbations have more fluctuations in flip rates across models and datasets, they generally still convey the same patterns: higher strengths result in higher flip rates and CoT perturbations generally have lower flip rates than input ones\. Future work could address this limitation by ensuring the number of samples in each strength level are equivalent to provide more stable measures\.

Lastly, in our comparative analysis between CoT and input perturbations, there are more perturbed CoTs than perturbed questions\. This is because each question has only one perturbed question, but multiple perturbed CoTs, each has one step being perturbed \([Section3\.1](https://arxiv.org/html/2609.30849#S3.SS1)\)\. Since the number of perturbed questions per dataset is not too small \(230\), it is sufficient for input pertubations to reach the stable patterns discussed in the previous paragraph across 16 sets of experiment \(4 datasets×\\times4 CoT models\)\. Therefore, we do not expand the set of perturbed questions as we are only interested in the patterns rather than the exact flip rate, and we do not expect that additional samples would produce significantly different patterns\.

## References

- Abdinet al\.\(2024\)M\. Abdin, J\. Aneja, H\. Awadalla, A\. Awadallah, A\. A\. Awan, N\. Bach, A\. Bahree, A\. Bakhtiari, J\. Bao, H\. Behl, A\. Benhaim, M\. Bilenko, J\. Bjorck, S\. Bubeck, M\. Cai, Q\. Cai, V\. Chaudhary, D\. Chen, D\. Chen, W\. Chen, Y\. Chen, Y\. Chen, H\. Cheng, P\. Chopra, X\. Dai, M\. Dixon, R\. Eldan, V\. Fragoso, J\. Gao, M\. Gao, M\. Gao, A\. Garg, A\. D\. Giorno, A\. Goswami, S\. Gunasekar, E\. Haider, J\. Hao, R\. J\. Hewett, W\. Hu, J\. Huynh, D\. Iter, S\. A\. Jacobs, M\. Javaheripi, X\. Jin, N\. Karampatziakis, P\. Kauffmann, M\. Khademi, D\. Kim, Y\. J\. Kim, L\. Kurilenko, J\. R\. Lee, Y\. T\. Lee, Y\. Li, Y\. Li, C\. Liang, L\. Liden, X\. Lin, Z\. Lin, C\. Liu, L\. Liu, M\. Liu, W\. Liu, X\. Liu, C\. Luo, P\. Madan, A\. Mahmoudzadeh, D\. Majercak, M\. Mazzola, C\. C\. T\. Mendes, A\. Mitra, H\. Modi, A\. Nguyen, B\. Norick, B\. Patra, D\. Perez\-Becker, T\. Portet, R\. Pryzant, H\. Qin, M\. Radmilac, L\. Ren, G\. de Rosa, C\. Rosset, S\. Roy, O\. Ruwase, O\. Saarikivi, A\. Saied, A\. Salim, M\. Santacroce, S\. Shah, N\. Shang, H\. Sharma, Y\. Shen, S\. Shukla, X\. Song, M\. Tanaka, A\. Tupini, P\. Vaddamanu, C\. Wang, G\. Wang, L\. Wang, S\. Wang, X\. Wang, Y\. Wang, R\. Ward, W\. Wen, P\. Witte, H\. Wu, X\. Wu, M\. Wyatt, B\. Xiao, C\. Xu, J\. Xu, W\. Xu, J\. Xue, S\. Yadav, F\. Yang, J\. Yang, Y\. Yang, Z\. Yang, D\. Yu, L\. Yuan, C\. Zhang, C\. Zhang, J\. Zhang, L\. L\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, and X\. ZhouPhi\-3 technical report: a highly capable language model locally on your phone\.External Links:2404\.14219,[Link](https://arxiv.org/abs/2404.14219)Cited by:[Appendix B](https://arxiv.org/html/2609.30849#A2.SS0.SSS0.Px1.p1.1),[footnote 1](https://arxiv.org/html/2609.30849#footnote1)\.
- Carliniet al\.\(2021\)N\. Carlini, F\. Tramèr, E\. Wallace, M\. Jagielski, A\. Herbert\-Voss, K\. Lee, A\. Roberts, T\. Brown, D\. Song, Ú\. Erlingsson, A\. Oprea, and C\. RaffelExtracting training data from large language models\.In30th USENIX Security Symposium \(USENIX Security 21\),pp\. 2633–2650\.External Links:ISBN 978\-1\-939133\-24\-3,[Link](https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting)Cited by:[§A\.2](https://arxiv.org/html/2609.30849#A1.SS2.p1.1),[§4\.4\.1](https://arxiv.org/html/2609.30849#S4.SS4.SSS1.p4.1)\.
- Chenet al\.\(2024\)Y\. Chen, R\. Zhong, N\. Ri, C\. Zhao, H\. He, J\. Steinhardt, Z\. Yu, and K\. MckeownDo models explain themselves? Counterfactual simulatability of natural language explanations\.InProceedings of the 41st International Conference on Machine Learning,R\. Salakhutdinov, Z\. Kolter, K\. Heller, A\. Weller, N\. Oliver, J\. Scarlett, and F\. Berkenkamp \(Eds\.\),Proceedings of Machine Learning Research, Vol\.235,pp\. 7880–7904\.External Links:[Link](https://proceedings.mlr.press/v235/chen24bl.html)Cited by:[§1](https://arxiv.org/html/2609.30849#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.30849#S2.SS2.SSS0.Px1.p1.1)\.
- Chenget al\.\(2025\)Y\. Cheng, Y\. Chang, and Y\. WuA survey on data contamination for large language models\.External Links:2502\.14425,[Link](https://arxiv.org/abs/2502.14425)Cited by:[§A\.2](https://arxiv.org/html/2609.30849#A1.SS2.p1.1),[§4\.4\.1](https://arxiv.org/html/2609.30849#S4.SS4.SSS1.p4.1)\.
- Clarket al\.\(2018\)P\. Clark, I\. Cowhey, O\. Etzioni, T\. Khot, A\. Sabharwal, C\. Schoenick, and O\. TafjordThink you have solved question answering? Try ARC, the AI2 reasoning challenge\.External Links:1803\.05457,[Link](https://arxiv.org/abs/1803.05457)Cited by:[§3\.1](https://arxiv.org/html/2609.30849#S3.SS1.p1.1)\.
- Cohen \(1960\)J\. CohenA coefficient of agreement for nominal scales\.Educational and Psychological Measurement20\(1\),pp\. 37–46\.External Links:[Document](https://dx.doi.org/10.1177/001316446002000104),[Link](https://doi.org/10.1177/001316446002000104),https://doi\.org/10\.1177/001316446002000104Cited by:[§E\.2](https://arxiv.org/html/2609.30849#A5.SS2.p2.1),[§3\.2\.3](https://arxiv.org/html/2609.30849#S3.SS2.SSS3.Px1.p1.1)\.
- Cohen \(1968\)J\. CohenWeighted kappa: nominal scale agreement provision for scaled disagreement or partial credit\.Psychological Bulletin70\(4\),pp\. 213–220\.External Links:[Document](https://dx.doi.org/10.1037/h0026256)Cited by:[§E\.2](https://arxiv.org/html/2609.30849#A5.SS2.p2.1),[§3\.2\.3](https://arxiv.org/html/2609.30849#S3.SS2.SSS3.Px1.p1.1)\.
- Gevaet al\.\(2021\)M\. Geva, D\. Khashabi, E\. Segal, T\. Khot, D\. Roth, and J\. BerantDid Aristotle use a laptop? A question answering benchmark with implicit reasoning strategies\.Transactions of the Association for Computational Linguistics9,pp\. 346–361\.External Links:ISSN 2307\-387X,[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00370),[Link](https://doi.org/10.1162/tacl_a_00370),https://direct\.mit\.edu/tacl/article\-pdf/doi/10\.1162/tacl\_a\_00370/1924104/tacl\_a\_00370\.pdfCited by:[§3\.1](https://arxiv.org/html/2609.30849#S3.SS1.p1.1)\.
- Golchin and Surdeanu \(2025\)S\. Golchin and M\. SurdeanuData contamination quiz: a tool to detect and estimate contamination in large language models\.Transactions of the Association for Computational Linguistics13,pp\. 809–830\.External Links:[Link](https://aclanthology.org/2025.tacl-1.37/),[Document](https://dx.doi.org/10.1162/tacl.a.20)Cited by:[§A\.2](https://arxiv.org/html/2609.30849#A1.SS2.p1.1),[§A\.2](https://arxiv.org/html/2609.30849#A1.SS2.p4.1),[§4\.4\.1](https://arxiv.org/html/2609.30849#S4.SS4.SSS1.p4.1)\.
- Google DeepMind \(2025\)Google DeepMindGemini 3 flash preview\.Google AI for Developers\.Note:[https://ai\.google\.dev/gemini\-api/docs/gemini\-3](https://ai.google.dev/gemini-api/docs/gemini-3)Cited by:[§3\.3](https://arxiv.org/html/2609.30849#S3.SS3.SSS0.Px3.p1.1)\.
- Google \(2025\)GoogleGemini 3 flash: frontier intelligence built for speed\.Google\.Note:[https://blog\.google/products\-and\-platforms/products/gemini/gemini\-3\-flash/](https://blog.google/products-and-platforms/products/gemini/gemini-3-flash/)Cited by:[§F\.3](https://arxiv.org/html/2609.30849#A6.SS3.SSS0.Px1.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. MaThe llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[Appendix B](https://arxiv.org/html/2609.30849#A2.SS0.SSS0.Px1.p1.1),[footnote 1](https://arxiv.org/html/2609.30849#footnote1)\.
- Hale \(2001\)J\. HaleA probabilistic Earley parser as a psycholinguistic model\.InSecond Meeting of the North American Chapter of the Association for Computational Linguistics,External Links:[Link](https://aclanthology.org/N01-1021/)Cited by:[§3\.3](https://arxiv.org/html/2609.30849#S3.SS3.SSS0.Px2.p1.1)\.
- Jacovi and Goldberg \(2020\)A\. Jacovi and Y\. GoldbergTowards faithfully interpretable NLP systems: how should we define and evaluate faithfulness?\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),pp\. 4198–4205\.External Links:[Link](https://aclanthology.org/2020.acl-main.386/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.386)Cited by:[§2\.1](https://arxiv.org/html/2609.30849#S2.SS1.p1.1)\.
- Jianget al\.\(2023\)A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier, L\. R\. Lavaud, M\. Lachaux, P\. Stock, T\. L\. Scao, T\. Lavril, T\. Wang, T\. Lacroix, and W\. E\. SayedMistral 7b\.External Links:2310\.06825,[Link](https://arxiv.org/abs/2310.06825)Cited by:[Appendix B](https://arxiv.org/html/2609.30849#A2.SS0.SSS0.Px1.p1.1),[footnote 1](https://arxiv.org/html/2609.30849#footnote1)\.
- Lanhamet al\.\(2023\)T\. Lanham, A\. Chen, A\. Radhakrishnan, B\. Steiner, C\. Denison, D\. Hernandez, D\. Li, E\. Durmus, E\. Hubinger, J\. Kernion, K\. Lukošiūtė, K\. Nguyen, N\. Cheng, N\. Joseph, N\. Schiefer, O\. Rausch, R\. Larson, S\. McCandlish, S\. Kundu, S\. Kadavath, S\. Yang, T\. Henighan, T\. Maxwell, T\. Telleen\-Lawton, T\. Hume, Z\. Hatfield\-Dodds, J\. Kaplan, J\. Brauner, S\. R\. Bowman, and E\. PerezMeasuring faithfulness in chain\-of\-thought reasoning\.External Links:2307\.13702,[Link](https://arxiv.org/abs/2307.13702)Cited by:[§A\.1](https://arxiv.org/html/2609.30849#A1.SS1.p2.1),[Figure 7](https://arxiv.org/html/2609.30849#A2.F7),[Figure 10](https://arxiv.org/html/2609.30849#A4.F10),[§1](https://arxiv.org/html/2609.30849#S1.p2.1),[§1](https://arxiv.org/html/2609.30849#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.30849#S2.SS2.SSS0.Px2.p1.1),[§2\.2](https://arxiv.org/html/2609.30849#S2.SS2.SSS0.Px3.p1.1)\.
- Levy \(2008\)R\. LevyExpectation\-based syntactic comprehension\.Cognition106\(3\),pp\. 1126–1177\.External Links:ISSN 0010\-0277,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.cognition.2007.05.006),[Link](https://www.sciencedirect.com/science/article/pii/S0010027707001436)Cited by:[§3\.3](https://arxiv.org/html/2609.30849#S3.SS3.SSS0.Px2.p1.1)\.
- Madaniet al\.\(2025\)M\. R\. G\. Madani, A\. P\. Gema, G\. Sarti, Y\. Zhao, P\. Minervini, and A\. PasseriniNoiser: bounded input perturbations for attributing large language models\.External Links:2504\.02911,[Link](https://arxiv.org/abs/2504.02911)Cited by:[§5](https://arxiv.org/html/2609.30849#S5.p2.1)\.
- Madsenet al\.\(2024\)A\. Madsen, S\. Chandar, and S\. ReddyAre self\-explanations from large language models faithful?\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),pp\. 295–337\.External Links:[Link](https://aclanthology.org/2024.findings-acl.19/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.19)Cited by:[§1](https://arxiv.org/html/2609.30849#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.30849#S2.SS2.SSS0.Px1.p1.1)\.
- Mattonet al\.\(2024\)K\. Matton, R\. Ness, and E\. KicimanWalk the talk? measuring the faithfulness of large language model explanations\.InICLR 2024 Workshop on Secure and Trustworthy Large Language Models,External Links:[Link](https://openreview.net/forum?id=QFFK0zOLGF)Cited by:[§A\.1](https://arxiv.org/html/2609.30849#A1.SS1.p1.1),[§1](https://arxiv.org/html/2609.30849#S1.p1.1),[§1](https://arxiv.org/html/2609.30849#S1.p2.1),[§1](https://arxiv.org/html/2609.30849#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.30849#S2.SS1.p1.1),[§2\.2](https://arxiv.org/html/2609.30849#S2.SS2.SSS0.Px1.p1.1),[§2\.2](https://arxiv.org/html/2609.30849#S2.SS2.SSS0.Px3.p1.1)\.
- McHugh \(2012\)M\. McHughInterrater reliability: the kappa statistic\.Biochemia Medica,pp\. 276–282\.External Links:[Document](https://dx.doi.org/10.11613/BM.2012.031)Cited by:[§E\.2](https://arxiv.org/html/2609.30849#A5.SS2.p4.1)\.
- Mihaylovet al\.\(2018\)T\. Mihaylov, P\. Clark, T\. Khot, and A\. SabharwalCan a suit of armor conduct electricity? a new dataset for open book question answering\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,E\. Riloff, D\. Chiang, J\. Hockenmaier, and J\. Tsujii \(Eds\.\),pp\. 2381–2391\.External Links:[Link](https://aclanthology.org/D18-1260/),[Document](https://dx.doi.org/10.18653/v1/D18-1260)Cited by:[§3\.1](https://arxiv.org/html/2609.30849#S3.SS1.p1.1)\.
- Oh and Schuler \(2023\)B\. Oh and W\. SchulerWhy does surprisal from larger transformer\-based language models provide a poorer fit to human reading times?\.Transactions of the Association for Computational Linguistics11,pp\. 336–350\.External Links:[Link](https://aclanthology.org/2023.tacl-1.20/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00548)Cited by:[§F\.2](https://arxiv.org/html/2609.30849#A6.SS2.p1.1)\.
- OpenAI \(2024\)OpenAIGPT\-4o mini\.Note:[https://platform\.openai\.com/docs/models/gpt\-4o\-mini](https://platform.openai.com/docs/models/gpt-4o-mini)Large language modelCited by:[Appendix B](https://arxiv.org/html/2609.30849#A2.SS0.SSS0.Px2.p1.1),[Appendix H](https://arxiv.org/html/2609.30849#A8.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2609.30849#S3.SS1.p1.1)\.
- Parcalabescu and Frank \(2024\)L\. Parcalabescu and A\. FrankOn measuring faithfulness or self\-consistency of natural language explanations\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),pp\. 6048–6089\.External Links:[Link](https://aclanthology.org/2024.acl-long.329/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.329)Cited by:[§2\.1](https://arxiv.org/html/2609.30849#S2.SS1.p1.1)\.
- Paulet al\.\(2024\)D\. Paul, R\. West, A\. Bosselut, and B\. FaltingsMaking reasoning matter: measuring and improving faithfulness of chain\-of\-thought reasoning\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),pp\. 15012–15032\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.882/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.882)Cited by:[§1](https://arxiv.org/html/2609.30849#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.30849#S2.SS2.SSS0.Px2.p1.1)\.
- Radfordet al\.\(2019\)A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, and I\. SutskeverLanguage models are unsupervised multitask learners\.External Links:[Link](https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf)Cited by:[§F\.2](https://arxiv.org/html/2609.30849#A6.SS2.p1.1),[§3\.3](https://arxiv.org/html/2609.30849#S3.SS3.SSS0.Px2.p2.1)\.
- Reimers and Gurevych \(2019\)N\. Reimers and I\. GurevychSentence\-BERT: sentence embeddings using Siamese BERT\-networks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),pp\. 3982–3992\.External Links:[Link](https://aclanthology.org/D19-1410/),[Document](https://dx.doi.org/10.18653/v1/D19-1410)Cited by:[§F\.1](https://arxiv.org/html/2609.30849#A6.SS1.p1.1)\.
- Reimers and Gurevych \(2024\)N\. Reimers and I\. GurevychSentence transformers: pretrained models\.Note:[https://www\.sbert\.net/docs/sentence\_transformer/pretrained\_models\.html](https://www.sbert.net/docs/sentence_transformer/pretrained_models.html)Accessed: 2026\-04\-20Cited by:[§F\.1](https://arxiv.org/html/2609.30849#A6.SS1.p1.1)\.
- Ribeiroet al\.\(2016\)M\. Ribeiro, S\. Singh, and C\. Guestrin“Why should I trust you?”: explaining the predictions of any classifier\.InProceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations,J\. DeNero, M\. Finlayson, and S\. Reddy \(Eds\.\),pp\. 97–101\.External Links:[Link](https://aclanthology.org/N16-3020/),[Document](https://dx.doi.org/10.18653/v1/N16-3020)Cited by:[§2\.2](https://arxiv.org/html/2609.30849#S2.SS2.p1.1)\.
- Salton \(1989\)G\. SaltonAutomatic text processing: the transformation, analysis, and retrieval of information by computer\.Addison\-Wesley series in computer science,Addison\-Wesley\.External Links:ISBN 9780201122275,LCCN 88000467,[Link](https://books.google.com.au/books?id=wb8SAQAAMAAJ)Cited by:[§3\.3](https://arxiv.org/html/2609.30849#S3.SS3.SSS0.Px1.p1.1)\.
- Shannon \(1951\)C\. E\. ShannonPrediction and entropy of printed english\.Bell System Technical Journal30\(1\),pp\. 50–64\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1002/j.1538-7305.1951.tb01366.x),[Link](https://onlinelibrary.wiley.com/doi/abs/10.1002/j.1538-7305.1951.tb01366.x),https://onlinelibrary\.wiley\.com/doi/pdf/10\.1002/j\.1538\-7305\.1951\.tb01366\.xCited by:[§3\.3](https://arxiv.org/html/2609.30849#S3.SS3.SSS0.Px2.p1.1)\.
- Sidorovet al\.\(2014\)G\. Sidorov, A\. Gelbukh, H\. Gómez\-Adorno, and D\. PintoSoft similarity and soft cosine measure: similarity of features in vector space model\.Computación y Sistemas18\(3\),pp\. 491–504\.Cited by:[§3\.3](https://arxiv.org/html/2609.30849#S3.SS3.SSS0.Px1.p1.1)\.
- Srivastavaet al\.\(2023\)A\. Srivastava, A\. Rastogi, A\. Rao, A\. A\. M\. Shoeb, A\. Abid, A\. Fisch, A\. R\. Brown, A\. Santoro, A\. Gupta, A\. Garriga\-Alonso, A\. Kluska, A\. Lewkowycz, A\. Agarwal, A\. Power, A\. Ray, A\. Warstadt, A\. W\. Kocurek, A\. Safaya, A\. Tazarv, A\. Xiang, A\. Parrish, A\. Nie, A\. Hussain, A\. Askell, A\. Dsouza, A\. Slone, A\. Rahane, A\. S\. Iyer, A\. J\. Andreassen, A\. Madotto, A\. Santilli, A\. Stuhlmüller, A\. M\. Dai, A\. La, A\. K\. Lampinen, A\. Zou, A\. Jiang, A\. Chen, A\. Vuong, A\. Gupta, A\. Gottardi, A\. Norelli, A\. Venkatesh, A\. Gholamidavoodi, A\. Tabassum, A\. Menezes, A\. Kirubarajan, A\. Mullokandov, A\. Sabharwal, A\. Herrick, A\. Efrat, A\. Erdem, A\. Karakaş, B\. R\. Roberts, B\. S\. Loe, B\. Zoph, B\. Bojanowski, B\. Özyurt, B\. Hedayatnia, B\. Neyshabur, B\. Inden, B\. Stein, B\. Ekmekci, B\. Y\. Lin, B\. Howald, B\. Orinion, C\. Diao, C\. Dour, C\. Stinson, C\. Argueta, C\. Ferri, C\. Singh, C\. Rathkopf, C\. Meng, C\. Baral, C\. Wu, C\. Callison\-Burch, C\. Waites, C\. Voigt, C\. D\. Manning, C\. Potts, C\. Ramirez, C\. E\. Rivera, C\. Siro, C\. Raffel, C\. Ashcraft, C\. Garbacea, D\. Sileo, D\. Garrette, D\. Hendrycks, D\. Kilman, D\. Roth, C\. D\. Freeman, D\. Khashabi, D\. Levy, D\. M\. González, D\. Perszyk, D\. Hernandez, D\. Chen, D\. Ippolito, D\. Gilboa, D\. Dohan, D\. Drakard, D\. Jurgens, D\. Datta, D\. Ganguli, D\. Emelin, D\. Kleyko, D\. Yuret, D\. Chen, D\. Tam, D\. Hupkes, D\. Misra, D\. Buzan, D\. C\. Mollo, D\. Yang, D\. Lee, D\. Schrader, E\. Shutova, E\. D\. Cubuk, E\. Segal, E\. Hagerman, E\. Barnes, E\. Donoway, E\. Pavlick, E\. Rodolà, E\. Lam, E\. Chu, E\. Tang, E\. Erdem, E\. Chang, E\. A\. Chi, E\. Dyer, E\. Jerzak, E\. Kim, E\. E\. Manyasi, E\. Zheltonozhskii, F\. Xia, F\. Siar, F\. Martínez\-Plumed, F\. Happé, F\. Chollet, F\. Rong, G\. Mishra, G\. I\. Winata, G\. de Melo, G\. Kruszewski, G\. Parascandolo, G\. Mariani, G\. X\. Wang, G\. Jaimovitch\-Lopez, G\. Betz, G\. Gur\-Ari, H\. Galijasevic, H\. Kim, H\. Rashkin, H\. Hajishirzi, H\. Mehta, H\. Bogar, H\. F\. A\. Shevlin, H\. Schuetze, H\. Yakura, H\. Zhang, H\. M\. Wong, I\. Ng, I\. Noble, J\. Jumelet, J\. Geissinger, J\. Kernion, J\. Hilton, J\. Lee, J\. F\. Fisac, J\. B\. Simon, J\. Koppel, J\. Zheng, J\. Zou, J\. Kocon, J\. Thompson, J\. Wingfield, J\. Kaplan, J\. Radom, J\. Sohl\-Dickstein, J\. Phang, J\. Wei, J\. Yosinski, J\. Novikova, J\. Bosscher, J\. Marsh, J\. Kim, J\. Taal, J\. Engel, J\. Alabi, J\. Xu, J\. Song, J\. Tang, J\. Waweru, J\. Burden, J\. Miller, J\. U\. Balis, J\. Batchelder, J\. Berant, J\. Frohberg, J\. Rozen, J\. Hernandez\-Orallo, J\. Boudeman, J\. Guerr, J\. Jones, J\. B\. Tenenbaum, J\. S\. Rule, J\. Chua, K\. Kanclerz, K\. Livescu, K\. Krauth, K\. Gopalakrishnan, K\. Ignatyeva, K\. Markert, K\. Dhole, K\. Gimpel, K\. Omondi, K\. W\. Mathewson, K\. Chiafullo, K\. Shkaruta, K\. Shridhar, K\. McDonell, K\. Richardson, L\. Reynolds, L\. Gao, L\. Zhang, L\. Dugan, L\. Qin, L\. Contreras\-Ochando, L\. Morency, L\. Moschella, L\. Lam, L\. Noble, L\. Schmidt, L\. He, L\. Oliveros\-Colón, L\. Metz, L\. K\. Senel, M\. Bosma, M\. Sap, M\. T\. Hoeve, M\. Farooqi, M\. Faruqui, M\. Mazeika, M\. Baturan, M\. Marelli, M\. Maru, M\. J\. Ramirez\-Quintana, M\. Tolkiehn, M\. Giulianelli, M\. Lewis, M\. Potthast, M\. L\. Leavitt, M\. Hagen, M\. Schubert, M\. O\. Baitemirova, M\. Arnaud, M\. McElrath, M\. A\. Yee, M\. Cohen, M\. Gu, M\. Ivanitskiy, M\. Starritt, M\. Strube, M\. Swędrowski, M\. Bevilacqua, M\. Yasunaga, M\. Kale, M\. Cain, M\. Xu, M\. Suzgun, M\. Walker, M\. Tiwari, M\. Bansal, M\. Aminnaseri, M\. Geva, M\. Gheini, M\. V\. T, N\. Peng, N\. A\. Chi, N\. Lee, N\. G\. Krakover, N\. Cameron, N\. Roberts, N\. Doiron, N\. Martinez, N\. Nangia, N\. Deckers, N\. Muennighoff, N\. S\. Keskar, N\. S\. Iyer, N\. Constant, N\. Fiedel, N\. Wen, O\. Zhang, O\. Agha, O\. Elbaghdadi, O\. Levy, O\. Evans, P\. A\. M\. Casares, P\. Doshi, P\. Fung, P\. P\. Liang, P\. Vicol, P\. Alipoormolabashi, P\. Liao, P\. Liang, P\. W\. Chang, P\. Eckersley, P\. M\. Htut, P\. Hwang, P\. Miłkowski, P\. Patil, P\. Pezeshkpour, P\. Oli, Q\. Mei, Q\. Lyu, Q\. Chen, R\. Banjade, R\. E\. Rudolph, R\. Gabriel, R\. Habacker, R\. Risco, R\. Millière, R\. Garg, R\. Barnes, R\. A\. Saurous, R\. Arakawa, R\. Raymaekers, R\. Frank, R\. Sikand, R\. Novak, R\. Sitelew, R\. L\. Bras, R\. Liu, R\. Jacobs, R\. Zhang, R\. Salakhutdinov, R\. A\. Chi, S\. R\. Lee, R\. Stovall, R\. Teehan, R\. Yang, S\. Singh, S\. M\. Mohammad, S\. Anand, S\. Dillavou, S\. Shleifer, S\. Wiseman, S\. Gruetter, S\. R\. Bowman, S\. S\. Schoenholz, S\. Han, S\. Kwatra, S\. A\. Rous, S\. Ghazarian, S\. Ghosh, S\. Casey, S\. Bischoff, S\. Gehrmann, S\. Schuster, S\. Sadeghi, S\. Hamdan, S\. Zhou, S\. Srivastava, S\. Shi, S\. Singh, S\. Asaadi, S\. S\. Gu, S\. Pachchigar, S\. Toshniwal, S\. Upadhyay, S\. S\. Debnath, S\. Shakeri, S\. Thormeyer, S\. Melzi, S\. Reddy, S\. P\. Makini, S\. Lee, S\. Torene, S\. Hatwar, S\. Dehaene, S\. Divic, S\. Ermon, S\. Biderman, S\. Lin, S\. Prasad, S\. Piantadosi, S\. Shieber, S\. Misherghi, S\. Kiritchenko, S\. Mishra, T\. Linzen, T\. Schuster, T\. Li, T\. Yu, T\. Ali, T\. Hashimoto, T\. Wu, T\. Desbordes, T\. Rothschild, T\. Phan, T\. Wang, T\. Nkinyili, T\. Schick, T\. Kornev, T\. Tunduny, T\. Gerstenberg, T\. Chang, T\. Neeraj, T\. Khot, T\. Shultz, U\. Shaham, V\. Misra, V\. Demberg, V\. Nyamai, V\. Raunak, V\. V\. Ramasesh, vinay uday prabhu, V\. Padmakumar, V\. Srikumar, W\. Fedus, W\. Saunders, W\. Zhang, W\. Vossen, X\. Ren, X\. Tong, X\. Zhao, X\. Wu, X\. Shen, Y\. Yaghoobzadeh, Y\. Lakretz, Y\. Song, Y\. Bahri, Y\. Choi, Y\. Yang, S\. Hao, Y\. Chen, Y\. Belinkov, Y\. Hou, Y\. Hou, Y\. Bai, Z\. Seid, Z\. Zhao, Z\. Wang, Z\. J\. Wang, Z\. Wang, and Z\. WuBeyond the imitation game: quantifying and extrapolating the capabilities of language models\.Transactions on Machine Learning Research\.Note:Featured CertificationExternal Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=uyTL5Bvosj)Cited by:[§3\.1](https://arxiv.org/html/2609.30849#S3.SS1.p1.1)\.
- Turpinet al\.\(2023\)M\. Turpin, J\. Michael, E\. Perez, and S\. BowmanLanguage models don't always say what they think: unfaithful explanations in chain\-of\-thought prompting\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 74952–74965\.External Links:[Document](https://dx.doi.org/10.52202/075280-3275),[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/ed3fea9033a80fea1376299fa7863f4a-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.30849#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.30849#S2.SS1.p1.1),[§2\.2](https://arxiv.org/html/2609.30849#S2.SS2.SSS0.Px1.p1.1)\.
- Tuteket al\.\(2025\)M\. Tutek, F\. Hashemi Chaleshtori, A\. Marasovic, and Y\. BelinkovMeasuring chain of thought faithfulness by unlearning reasoning steps\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),pp\. 9935–9960\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.504/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.504),ISBN 979\-8\-89176\-332\-6Cited by:[Figure 7](https://arxiv.org/html/2609.30849#A2.F7),[Appendix B](https://arxiv.org/html/2609.30849#A2.SS0.SSS0.Px1.p1.1),[Appendix B](https://arxiv.org/html/2609.30849#A2.SS0.SSS0.Px2.p1.1),[Appendix C](https://arxiv.org/html/2609.30849#A3.p1.1),[Figure 18](https://arxiv.org/html/2609.30849#A8.F18),[Appendix H](https://arxiv.org/html/2609.30849#A8.SS0.SSS0.Px2.p1.1),[Appendix H](https://arxiv.org/html/2609.30849#A8.SS0.SSS0.Px2.p2.1),[Appendix H](https://arxiv.org/html/2609.30849#A8.p1.1),[§1](https://arxiv.org/html/2609.30849#S1.p2.1),[§1](https://arxiv.org/html/2609.30849#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.30849#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2609.30849#S3.SS1.p1.1),[§3\.1](https://arxiv.org/html/2609.30849#S3.SS1.p2.1),[§3\.5](https://arxiv.org/html/2609.30849#S3.SS5.p2.1),[§4\.2](https://arxiv.org/html/2609.30849#S4.SS2.p1.1),[§5](https://arxiv.org/html/2609.30849#S5.p2.1)\.
- Wanget al\.\(2020\)W\. Wang, F\. Wei, L\. Dong, H\. Bao, N\. Yang, and M\. ZhouMiniLM: deep self\-attention distillation for task\-agnostic compression of pre\-trained transformers\.External Links:2002\.10957,[Link](https://arxiv.org/abs/2002.10957)Cited by:[§F\.1](https://arxiv.org/html/2609.30849#A6.SS1.p1.1),[§3\.3](https://arxiv.org/html/2609.30849#S3.SS3.SSS0.Px1.p1.1)\.
- Weiet al\.\(2022\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, b\. ichter, F\. Xia, E\. Chi, Q\. V\. Le, and D\. ZhouChain\-of\-thought prompting elicits reasoning in large language models\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 24824–24837\.External Links:[Document](https://dx.doi.org/10.52202/068431-1800),[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.30849#S1.p1.1)\.
- Wilcoxet al\.\(2020\)E\. G\. Wilcox, J\. Gauthier, J\. Hu, P\. Qian, and R\. LevyOn the predictive power of neural language models for human real\-time comprehension behavior\.External Links:2006\.01912,[Link](https://arxiv.org/abs/2006.01912)Cited by:[§F\.2](https://arxiv.org/html/2609.30849#A6.SS2.p1.1)\.

## Appendix ARelated Work

### A\.1Prior Perturbation Strength Controls

Our perturbation strength measure can also be used to improve, replace, or verify the perturbation strength controls of prior studies\. For example, using our LLM\-as\-a\-judge measure, we can check whether modifying important concepts as in the study by[Matton et al\. \(2024\)](https://arxiv.org/html/2609.30849#bib.bib3)indeed produces stronger perturbations \([Section2\.2](https://arxiv.org/html/2609.30849#S2.SS2.SSS0.Px2)\)\. If so, this would provide additional evidence to strengthen the validity of[Matton et al\. \(2024\)](https://arxiv.org/html/2609.30849#bib.bib3)’s study, otherwise, it may question the effectiveness and validity of their faithfulness measures\.

In the work of[Lanham et al\. \(2023\)](https://arxiv.org/html/2609.30849#bib.bib13), the AOC metric was used to compare the likelihood of a model preserving its original answer after perturbations across different tasks\. However, it is unclear whether the perturbation strengths across tasks are even similar for a fair comparison \([Section2\.2](https://arxiv.org/html/2609.30849#S2.SS2.SSS0.Px2)\)\. Our proposed perturbation strength measure could also be used to verify this\. Alternatively, we can directly relate percentage of answer changes to perturbation strength, since we have an explicit measure of strength now\. We present such perturbation\-strength\-aware framework for self\-consistency assessment in[Section4](https://arxiv.org/html/2609.30849#S4)\.

### A\.2Data Contamination Risk

Large language models are exposed to massive corpora through multiple training stages; therefore, there exists a risk of data contamination, where the evaluation data overlap with examples the model has already seen during training\. A consequence of such data leakage is verbatim memorisation\([Carlini et al\., 2021](https://arxiv.org/html/2609.30849#bib.bib31);[Cheng et al\., 2025](https://arxiv.org/html/2609.30849#bib.bib32)\), where a model memorises and recalls the exact sequences of text it has seen from the training data\. In the study of[Golchin and Surdeanu \(2025\)](https://arxiv.org/html/2609.30849#bib.bib33), in one experimental condition, they exposed the model to the evaluation data during training, which consists of multiple\-choice questions\. At evaluation time, the same questions from the training data were used, but only one answer option is taken from the original question, while the other options are semantically equivalent word\-level perturbations of this option\. For example, suppose the original option in the training data was:

> Option A: “The company announced a new product yesterday”

During evaluation, other options are perturbed versions of A:

- •Option A: “The companyannounceda new product yesterday”
- •Option B: “The companyrevealeda new product yesterday”
- •Option C: “Thecorporationannounced a new product yesterday”

[Golchin and Surdeanu \(2025\)](https://arxiv.org/html/2609.30849#bib.bib33)found that the model consistently biases toward selecting the option that matches the original instance exactly, even when the other options are just semantically similar perturbed versions\. This does not happen in the condition when the model is not exposed to the evaluation data during training, as the model’s selections remain consistent with the original positional biases, not favoring the option that appears in the original instance\. This suggests the presence of verbatim memorisation, highlighting the need to be aware of this phenomenon and data contamination risk in our study, as our work also involves analysing model outputs after perturbations\.

## Appendix BPerturb CoT and input

##### CoT models

For each question of a dataset,[Tutek et al\. \(2025\)](https://arxiv.org/html/2609.30849#bib.bib9)used four instruction\-tuned models, each generated an original CoT to the question, denotedR=\{r1,…,rm\}R=\\\{r\_\{1\},\\ldots,r\_\{m\}\\\}wheremmis the number of steps in that generated CoT andrir\_\{i\}represents a single CoT step\. Conditioned on its generated reasoningRR, each model generated an answer letter to the question\. The four models are Llama\-3\.2\-3B\-Instruct, Llama\-3\-8B\-Instruct\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib10)\); Mistral\-7B\-Instruct\-v0\.2\([Jiang et al\., 2023](https://arxiv.org/html/2609.30849#bib.bib11)\)and Phi\-3\-mini\-4k\-Instruct\([Abdin et al\., 2024](https://arxiv.org/html/2609.30849#bib.bib12)\)\.

##### Perturbed CoT steps

For each reasoningR=\{r1,…,rm\}R=\\\{r\_\{1\},\.\.\.,r\_\{m\}\\\},[Tutek et al\. \(2025\)](https://arxiv.org/html/2609.30849#bib.bib9)usedgpt\-4o\-mini\([OpenAI, 2024](https://arxiv.org/html/2609.30849#bib.bib19)\)and the prompt in[Figure7](https://arxiv.org/html/2609.30849#A2.F7)to generate a mistaken version of each steprir\_\{i\}\. Then the CoT models were prompted to make a new prediction conditioned on the perturbed full CoT, denotedRi′=\{…,ri′,…\}R^\{\\prime\}\_\{i\}=\\\{\\ldots,r^\{\\prime\}\_\{i\},\\ldots\\\}where the perturbed stepri′r^\{\\prime\}\_\{i\}is inserted in place of the original step while keeping all the other steps the same\.[Figure3](https://arxiv.org/html/2609.30849#S3.F3)shows an example ofR2′R^\{\\prime\}\_\{2\}of a question, where the second perturbed stepr2′r^\{\\prime\}\_\{2\}replaces the original stepr2r\_\{2\}\.

Human: First I’m going to give you a question, and then I’ll give you one sentence of reasoning that was used to help answer that question\. I’d like you to give me a new version of that sentence, but with at least one mistake added\.
\[question\]
\[Answer options\]
Original sentence:\[sentence\]
Assistant: Sentence with mistake added:Figure 7:Prompt forgpt\-4o\-minitoperturb CoT step\([Lanham et al\., 2023](https://arxiv.org/html/2609.30849#bib.bib13);[Tutek et al\., 2025](https://arxiv.org/html/2609.30849#bib.bib9)\)
##### Input Perturbation

[Figure8](https://arxiv.org/html/2609.30849#A2.F8)presents the prompt forgpt\-4o\-minito perturb the input \(i\.e\. question text\)\. Input perturbations are generated with access to the full original CoT, as we observe that this results in perturbations that are more relevant to the overall reasoning\. The prompt for perturbing inputs \([Figure8](https://arxiv.org/html/2609.30849#A2.F8)\) is designed to be as similar as possible to the prompt for perturbing CoT steps \([Figure7](https://arxiv.org/html/2609.30849#A2.F7)\)\. The main difference is that the prompt for input perturbation has access to the full original CoT, while the prompt for CoT perturbation only has access to one CoT step\. This difference is not expected to create a significant concern since the strengths of all perturbations will be measured and only perturbations of the same perturbation strengths are compared against each other\.

I will give you a question and the reasoning to help answer that question\.
Original Question:\[question\]
\[Answer options\]
Reasoning:\[Original CoT\]
Based on the above reasoning, I would like you to create a new version of the question that has at least one mistake in it\. Only make changes to the question, do not change the options\.
Format your response as:
Edited question: \[your edited question\]Figure 8:Prompt forgpt\-4o\-minitoperturb the input\(i\.e\. question text\)\.

## Appendix CSampling Strategy

A subset of the datasets from[Tutek et al\. \(2025\)](https://arxiv.org/html/2609.30849#bib.bib9)is sampled for annotation and evaluation of perturbation strength measures, ensuring a balanced number of questions across different CoT models\. We define the conventions used to refer to instances as follows:

- •An instance refers to a single example with a question and a perturbed CoT step that replaces the corresponding original step in the reasoning
- •A question ID refers to the unique string identifier associated with a question\. Multiple instances may share the same question ID, as they correspond to the same question, but differ in the perturbed CoT step as shown in[Figure2](https://arxiv.org/html/2609.30849#S3.F2)

For each dataset, 24 out of 230 unique question IDs are randomly sampled and partitioned into 4 equal sets, each containing 6 IDs\. Each set is randomly assigned to a CoT model, then all instances with question ID matching those in the set are then sampled from the corresponding perturbed dataset of that dataset\-CoT model pair \([Figure9](https://arxiv.org/html/2609.30849#A3.F9)\)\. As a result, we have 453 sampled instances for 96 questions in total across four datasets\.

StrategyQA230 Question IDsRandom sample24 IDs selectedSet 16 IDsSet 26 IDsSet 36 IDsSet 46 IDsStrategyQA−\-Llama 3BStrategyQA−\-Llama 8BStrategyQA−\-MistralStrategyQA−\-Phi\-3Sampled instancesMatching IDs from perturbed dataset

Figure 9:Sampling procedure for annotation in StrategyQA dataset\. This is repeated for all 4 datasets\.
## Appendix DConstructing Diverse Perturbed CoTs

To have a more diverse range of CoT perturbation strengths for annotation and evaluation, for each sampled question, we randomly combine some of their perturbed CoT steps or choose to paraphrase all steps as follows:

- •For a question withmmCoT steps, an integerx∈\[0,m\]x\\in\[0,m\]is sampled uniformly at random, representing the number of steps to be perturbed\.
- •Ifxxis 0, all original CoT steps of that question are paraphrased bygpt\-4o\-miniusing the prompt in[Figure10](https://arxiv.org/html/2609.30849#A4.F10)\.
- •Otherwise, we randomly samplexxperturbed CoT steps of that question and insert them in place of the original steps\. This yields a single perturbed CoT per question\.

Please paraphrase the following text, conveying exactly the same information\.
Text:\[text\]
Format your response as:
Rewritten text: \[Your paraphrase here\]Figure 10:Prompt forgpt\-4o\-minitoparaphrase CoT step\([Lanham et al\., 2023](https://arxiv.org/html/2609.30849#bib.bib13)\)As a result, each sampled question now only has a single perturbed CoT which has multiple perturbed steps or the original steps are paraphrased\.

## Appendix EData Annotation

### E\.1Guideline for Human Annotators

A primary annotator annotated the perturbation strength of 96 CoT and 96 input perturbations generated in[Section3\.2\.2](https://arxiv.org/html/2609.30849#S3.SS2.SSS2)\. The annotation procedure is as follows:

- •The annotator is provided with an original question, original CoT and a perturbed question/perturbed CoT\.
- •The annotator follows the criteria in[Figure11](https://arxiv.org/html/2609.30849#A5.F11)to rate the perturbation strength of the provided perturbed question/perturbed CoT\.

For the following descriptions,reasoning pathrefers to the reasoning in the provided CoT\.Interfere with the reasoning pathmeans that itchanges, invalidates, or redirectsthe logical steps required to derive the answer\. This can be done by introducing new premises that might contradict with the current reasoning, removing necessary premises, or altering logical dependencies\.Use the reasoning path provided to assess the edited text and choose one of the followings:•No/Minimal change: Meaning is preserved \(e\.g\., paraphrasing, synonyms, rewording\) OR the change\(s\) is/arenot enough to interfere with the reasoning path•Small change: Minor semantic shift thatmight interferewith the reasoning path butnot enough to remove or reversesupport for an answer option•Moderately strong change: The meaning changes enough tonotably weaken, remove, or reversesupport for an answer option,without clearly shifting support to another option\(including cases where it shifts support to an answer that is not in the answer options\)\.•Strong change: Clearly shift support from one answer optionto another that is in the option list\. There should beonly few or little support for original answer\. If there is a strong support for another answer but still a very strong support for the original answer, it should be a moderately strong change \(see example 5 of SQA CoT examples\.doc\)Rating:•0: No/minimal change•1: small change•2: moderately strong change•3: strong changeNote 1: When you are unsure between two ratings, you should give a lower oneNote 2: If the changes introduce new entities or aspects that are not mentioned/discussed about in the CoT, the annotator can use a search engine to search for relevant facts about the new entities and understand how they might interfere with the reasoningFigure 11:Full Perturbation Strength Assessment Guideline for Human Annotators
### E\.2Inter\-Annotator Agreement

To validate the reliability of the perturbation strength guide, a second annotator with a PhD\-level background independently annotated a subset of 18 input and 18 CoT perturbations using the guideline in[Figure11](https://arxiv.org/html/2609.30849#A5.F11)\. The second annotator follows the same annotation procedure as the primary annotator, with additional material provided that is a set of examples specific to the dataset they are working on \([Figures12](https://arxiv.org/html/2609.30849#A5.F12)and[13](https://arxiv.org/html/2609.30849#A5.F13)\)\. In each dataset, one example for each perturbation strength level \(0, 1, 2 and 3\), along with the justifications of their strength, are used as examples\. These examples provide clear expectations of how the perturbation strength guide is applied to perturbed input or CoT\.

The Cohen’s Linear Kappa score\([Cohen, 1960](https://arxiv.org/html/2609.30849#bib.bib14);[Cohen, 1968](https://arxiv.org/html/2609.30849#bib.bib30)\), a metric commonly used to measure inter\-annotator agreement for ordinal scales, is then computed between the strengths assigned by the primary and second annotators\. A higher Kappa score indicates stronger agreement betweeen annotators using the same perturbation strength criteria\.

Original Question: A student mixed 25 grams of salt into1,000 grams of water\. What is the mass of the saltwater mixture?
Options
A\): 975 grams
B\): 1,000 grams
C\): 1,025 grams
D\): 2,500 grams
CoT:The student mixed 25 grams of salt into 1,000 grams of water\. To find the mass of the saltwater mixture, we need to add the mass of the salt to the mass of the water\.
Perturbed Question: A student mixed 25 grams of salt into500 grams of water\. What is the mass of the saltwater mixture?
Options
A\): 975 grams
B\): 1,000 grams
C\): 1,025 grams
D\): 2,500 grams
Rating: 2
Reason: Based on the CoT, to find the mass of the mixture, we add the mass of the salt to the mass of the water, which is 500 \+ 25 = 525 grams\. This removes support for the original answer, but it does not clearly shift support to another answer in the option list\. Therefore it is a moderately strong change\.Figure 12:Example of perturbed question with strength of 2 and the justificationQuestion: If bacon is left too long on a hot stove top
Options
A\): it will be cooked perfectly
B\): it will be bacteria laden
C\): it will become blackened
D\): it will be left raw
Original CoT:
Step 1: Bacon is a type ofmeatthat is commonly cooked on a stove top\.
Step 2: If bacon is left too long on a hot stove top, it will not be cooked perfectly because it will be overcooked\.
Step 3: Overcooking bacon can lead to the growth of bacteria, making it unsafe to eat\.
Step 4: Therefore, the correct answer is \(B\):it will be bacteria laden\.
Perturbed CoT:
Step 1: Bacon is a type ofvegetablethat is commonly cooked on a stove top\.
Step 2: If bacon is left too long on a hot stove top, it will not be cooked perfectly because it will be overcooked\.
Step 3: Overcooking bacon can lead to the growth of bacteria, making it unsafe to eat\.
Step 4: Therefore, the correct answer is \(B\):it will be cooked perfectly\.
Rating: 1
Reason: only the last change might interfere with the reasoning path, but it is not enough to weaken/remove support for the original answer because the stated answer option \(B\) is still the same and the previous reasonings still show support for that answer\. Therefore it is a small change\.Figure 13:Example of perturbed CoT with strength of 1 and the justificationThe linear Kappa scores between strengths assigned by primary and second annotator are presented in[Table4](https://arxiv.org/html/2609.30849#A5.T4)\. The scores are computed separately for input and CoT perturbations\. For each perturbation type, the results are further divided into two groups: one where the true strength \(assigned by the primary annotator\) is 0 or 3, and another where the true strength is 1 or 2\.

Table 4:Linear Kappa scores between primary and second annotatorsThe score of 0\.6489 for all CoT perturbations indicates substantial agreement, while a score of 0\.8548 for input perturbations indicates near\-perfect agreement\([McHugh, 2012](https://arxiv.org/html/2609.30849#bib.bib29)\)\. This suggests that the perturbation strength guide developed is reliable to use\.

From[Table4](https://arxiv.org/html/2609.30849#A5.T4), most disagreements occur in CoT perturbations when the true strength is 1 or 2\. Inspection of the results shows that the second annotator tends to assign a score of 2 when the true strength is 1, and 3 when it is 2\. In contrast, the agreement is very strong when the true strength is 0 or 3 in both pertubation types \([Table4](https://arxiv.org/html/2609.30849#A5.T4)\)\. This is expected because the intermediate strength levels \(1 and 2\) can be more ambiguous and therefore more prone to confusion than the extreme cases \(0 and 3\)\.

## Appendix FExperimental Setup

### F\.1Cosine Distance

In this study, embeddings of original and perturbed texts are obtained using the all\-MiniLM\-L6\-v2model\([Wang et al\., 2020](https://arxiv.org/html/2609.30849#bib.bib17)\)from Sentence Transformers\([Reimers and Gurevych, 2019](https://arxiv.org/html/2609.30849#bib.bib18)\), which applies mean pooling over token embeddings to create a 384 dimensional vector representation \(embedding\) of the text\. We choose this model because it is a lightweight pre\-trained language model that achieves competitive performance in sentence encoding while being very computationally efficient\([Reimers and Gurevych, 2024](https://arxiv.org/html/2609.30849#bib.bib23)\)\.

### F\.2Surprisals

To calculate surprisals \([Equation1](https://arxiv.org/html/2609.30849#S3.E1), token probabilities are obtained from the GPT\-2 model\([Radford et al\., 2019](https://arxiv.org/html/2609.30849#bib.bib24)\)\. This model is chosen because it is a fully open\-source, lightweight model commonly used as a research baseline to obtain log\-probabilities and surprisals\([Oh and Schuler, 2023](https://arxiv.org/html/2609.30849#bib.bib25);[Wilcox et al\., 2020](https://arxiv.org/html/2609.30849#bib.bib26)\)\.

![Refer to caption](https://arxiv.org/html/2609.30849v1/surprisal_cot.png)Figure 14:Change in surprisals versus True perturbation strength for CoT perturbations
### F\.3LLM\-as\-a\-judge

##### Model

Thegemini\-3\-flash\-previewmodel is chosen to balance the complex reasoning performance with computational efficiency\. It outperforms Gemini 2\.5 Pro on many benchmarks and achieves performance comparable to that of Gemini 3 Pro on reasoning tasks while being more token\-efficient\([Google, 2025](https://arxiv.org/html/2609.30849#bib.bib28)\)\. We intentionally avoid usinggpt\-4o\-minias it was the model used to generate the perturbations, which could introduce unintended effect if the same model is used to judge those perturbations again\.

Here are examples of how to apply the rating guide to a perturbed question:
\[four dataset\-specific input perturbation examples for strength 0, 1, 2, 3\]
Now, rate the following:
Original Question:\[question\]
CoT:\[cot\]
Perturbed Question:\[perturbed question\]Figure 15:Prompt template forgemini\-3\-flashto assess input perturbation strength\.Here are examples of how to apply the rating guide to a perturbed CoT:
\[four dataset\-specific CoT perturbation examples for strength 0, 1, 2, 3\]
Now, rate the following:
Question:\[question\]
Original CoT:\[cot\]
Perturbed CoT:\[perturbed cot\]Figure 16:Prompt template forgemini\-3\-flashto assess CoT perturbation strength\.You are an expert at evaluating perturbation strength of text\. Use the following rating guide to assess perturbed sentences:
For the following descriptions, reasoning path refers to the reasoning in the provided CoT\.…
Note: When the score seems to be in between two ratings, you should give a lower one
\(Note 2 is omitted\)
For the perturbed text provided, give: Rating \(0\-3\)Figure 17:Perturbation Strength Guide as system instruction for LLM\. Modifications from the guidelines for human annotators \([Figure11](https://arxiv.org/html/2609.30849#A5.F11)\) are in bold\.
##### Perturbation Strength Guide for LLM

There are only minor additions and modifications in the perturbation strength guide for LLM as highlighted in bold in[Figure17](https://arxiv.org/html/2609.30849#A6.F17)\. The sentence “You are an expert at evaluating perturbation strength of text\.” is added to specify the LLM’s role and to mimic its default system instruction, which was “You are a friendly and helpful assistant\.”\. Note 2 in[Figure11](https://arxiv.org/html/2609.30849#A5.F11)about using a search engine is only relevant and helpful for human annotators, therefore it is omitted for LLM\. Lastly, we emphasise the task is to give a rating from 0\-3 for the perturbation strength\.

## Appendix GEvaluation of Perturbation Strength Measures

##### Exclude Examples for Evaluation

The results for correlations and Kappa scores between the predicted strengths and true strengths are presented in[Table2](https://arxiv.org/html/2609.30849#S3.T2)\. Note that the dataset\-specific examples mentioned in[Section3\.3](https://arxiv.org/html/2609.30849#S3.SS3.SSS0.Px2)and[SectionE\.2](https://arxiv.org/html/2609.30849#A5.SS2)are excluded from the evaluation set, so only 80 out of 96 questions and their perturbations remain for evaluation\. This is because 16 are taken as examples for four datasets, each dataset has 4 examples, each for one strength\.

### G\.1Change in surprisals is not effective for paraphrases

From[Table2\(a\)](https://arxiv.org/html/2609.30849#S3.T2.st1), we notice that change in surprisals for CoT perturbations has a weak negative correlation with the true strengths\. When plotting change in surprisals versus true strength as in[Figure14](https://arxiv.org/html/2609.30849#A6.F14), we notice that change in surprisals is specifically high for zero strength\. This is probably because zero\-strength CoT perturbations typically involve paraphrasing entire reasoning steps \([AppendixD](https://arxiv.org/html/2609.30849#A4)\), resulting in more token\-level modifications than when only selected steps are perturbed\. In the latter case, only some tokens in the perturbed steps are changed, whereas paraphrasing affects tokens across all steps\. Based on[Equation1](https://arxiv.org/html/2609.30849#S3.E1)for change in surprisals, more token modifications accumulate more change in token probabilities overall, leading to a larger change in surprisal when the CoT steps are paraphrased\.

## Appendix HData for Input Perturbation

The datasets by[Tutek et al\. \(2025\)](https://arxiv.org/html/2609.30849#bib.bib9)did not provide input perturbations\. Therefore, we generate all the perturbed questions, their perturbation strengths and the model’s answer after perturbation to compute flip rates\.

##### Generate Perturbed Questions

Recall that there are 16 perturbed datasets, each corresponding to a dataset−\-CoT model pair \(Section[Section3\.1](https://arxiv.org/html/2609.30849#S3.SS1)\)\. In each perturbed dataset, there are 230 questions and each question has the original CoT generated by the corresponding CoT model\. Given these questions and their associated CoTs,gpt\-4o\-mini\([OpenAI, 2024](https://arxiv.org/html/2609.30849#bib.bib19)\)is prompted to perturb the question text by adding mistake using the prompt in[Figure8](https://arxiv.org/html/2609.30849#A2.F8), similar to how input perturbations are generated in[Section3\.2\.2](https://arxiv.org/html/2609.30849#S3.SS2.SSS2.Px1)for the sampled data\. Then, their perturbation strengths are measured using LLM\-as\-a\-judge method described in[Section3\.3](https://arxiv.org/html/2609.30849#S3.SS3.SSS0.Px2), ensuring the assessment is consistent with that used for CoT perturbations\.

##### Model’s Answers After Input Perturbations

The associated CoT model’s answers to these perturbed questions are obtained using the same method as in[Tutek et al\. \(2025\)](https://arxiv.org/html/2609.30849#bib.bib9)\. Specifically, we prompt the model with the perturbed question and answer options labeled by letters \(A, B, C, D, E\) as in[Figure18](https://arxiv.org/html/2609.30849#A8.F18)\. The prefix “The single, most likely answer is \(” is added to the end of the prompt to trigger the answer\. Then we find the model’s output probabilities over the option letters at the first token\. The letter with the highest probability is taken as the model’s answer\.

Human: Question:\[Question\]Choices:\[Answer options\]Assistant: The single, most likely answer is \(Figure 18:Prompt to get CoT model’s answer to a question\([Tutek et al\., 2025](https://arxiv.org/html/2609.30849#bib.bib9)\)The approach used to extract each model’s answers to perturbed questions is referred to as thedirect answer\(or direct prompting\) approach in[Tutek et al\. \(2025\)](https://arxiv.org/html/2609.30849#bib.bib9), which means we prompt the model for an answer directly without eliciting a CoT\.[Tutek et al\. \(2025\)](https://arxiv.org/html/2609.30849#bib.bib9)noted that for these questions, direct prompting and CoT prompting \(generate CoT first then extract answer\) gave the same answer initially\. In other words, the dataset’s original CoT\-based answers are identical to those produced under direct prompting\. As a result, when we use the direct answer method to extract the model’s answer to the perturbed questions and compare it against the original answer, we are effectively comparing two direct answers: one from the original question and one from the perturbed question\.

We use direct prompting to extract model’s answer to minimise extraneous variables\. If we compare two answers where the only difference is the input question, the comparison is cleaner and more controlled than when the difference is both the question and the generated CoT\. Nevertheless, we still attempt to extract the models’ answers to the perturbed questions using CoT prompting and compare this against direct prompting for input perturbation as well as to CoT perturbation\. Details and results of this are presented in[SectionI\.3](https://arxiv.org/html/2609.30849#A9.SS3)\.

## Appendix IAdditional Results for CoT vs Input Perturbation

### I\.1Flip Rates

[Figures19](https://arxiv.org/html/2609.30849#A9.F19),[20](https://arxiv.org/html/2609.30849#A9.F20)and[21](https://arxiv.org/html/2609.30849#A9.F21)present flip rates for OpenBookQA, Sports and StrategyQA datasets, respectively\. They follow similar trends and patterns as flip rates for ARC\-Challenge dataset presented in[Figure6](https://arxiv.org/html/2609.30849#S4.F6)\.

Figure 19:OpenBookQA flip rates of input and CoT perturbations across 4 CoT modelsFigure 20:Sports flip rates of input and CoT perturbations across 4 CoT modelsFigure 21:StrategyQA flip rates of input and CoT perturbations across 4 CoT models
### I\.2Examples of Fail Cases

Since models often exhibit weak robust self\-consistency under input perturbations and weak responsive self\-consistency under CoT perturbations, we provide examples of such fail cases in[Figures22](https://arxiv.org/html/2609.30849#A9.F22)and[23](https://arxiv.org/html/2609.30849#A9.F23), respectively\.

Session 1 \(original answer and CoT\)

A student hasone liter of feathers andone liter of iron nails\. Which statement best compares the masses of the two samples? Options: A\): The masses cannot be compared without a balance\. B\): The mass of the feathers is less than the mass of the nails C\): The mass of the feathers is the same as the mass of the nails\. D\): The mass of the feathers is greater than the mass of the nails\.

Original question

The student has one liter of feathers and one liter of iron nails\. The volume of both samples is the same \(one liter\)\. The mass of a substance is equal to its density multiplied by its volume\. The density of feathers is much less than the density of iron nails\. Therefore, the mass of the feathers is less than the mass of the nails\.

B

Original answer

Session 2 \(answer after input perturbation\)

A student hasone liter of feathers andtwo liters of iron nails\. Which statement best compares the masses of the two samples? \[same answer options A\-D\]

Perturbed question

D

New answer

Figure 22:Example of model changing answer under zero\-strength input perturbation[Figure22](https://arxiv.org/html/2609.30849#A9.F22)shows example of a model changing its answer under negligible input perturbation\. In this example, the main reasoning path to get Option B still applies to the perturbed question, i\.e\. the perturbation does not interfere with the reasoning\. However, the model shows weak robustness to negligible perturbations and changes its answer\.

[Figure23](https://arxiv.org/html/2609.30849#A9.F23)shows example of a model not changing its answer under strong CoT perturbation\. In this example, the original reasoning path states that the star “exhausted its hydrogen” and becomes “cooler” so the original answer is A\. The changes in perturbed CoT substantially interfere with the original reasoning path because it now states that the star becomes “hotter”, which removes support for answer A, and “gained hydrogen” which can shift support to answer D\. However, we can see that the model’s answer is still A given this perturbed CoT\.

Session 1 \(original answer and CoT\)

A red giant star differs from main sequence stars because it Options: A\): burns at a cooler temperature\. B\): is always larger than other suns\. C\): has greater density and gravity D\): fuses hydrogen to make helium

Original question

A red giant star is a stage in the life of a star, just like the main sequence stage\. The main difference is thata red giant star hasexhausted its hydrogen fueland has expanded to become much larger andcoolerthan it was during its main sequence life\.

Original CoT

A

Original answer

Session 2 \(answer after CoT perturbation\)

\[same question and answer options\]

A red giant star is a stage in the life of a star, just like the main sequence stage\. The main difference is thata red giant star hasgained hydrogen fueland has shrunk to become much smaller andhotterthan it was during its main sequence life\.

Perturbed CoT

The single, most likely answer is \(

A

New answer

Figure 23:Example of model not changing answer under high\-strength CoT perturbation
### I\.3Direct Prompting vs CoT Prompting for Input Perturbation

In addition to the direct prompting approach used to extract the model answers to perturbed questions, we also attempt to extract the answers using CoT prompting approach, which prompts the model to generate a CoT to the perturb question first, then concatenate the question and the generated CoT to trigger for the final answer\. We then compute and plot the flip rates for each perturbation strength under three settings: CoT perturbation, input perturbation with direct prompting, and input perturbation with CoT prompting\. The results are presented in[Figures24](https://arxiv.org/html/2609.30849#A9.F24),[25](https://arxiv.org/html/2609.30849#A9.F25),[26](https://arxiv.org/html/2609.30849#A9.F26)and[27](https://arxiv.org/html/2609.30849#A9.F27)\.

Figure 24:ARC\-Challenge flip rates of four CoT models under 3 settings: CoT perturbation, input perturbation using direct prompting and input perturbation using CoT promptingFigure 25:OpenBookQA flip rates of four CoT models under 3 settings: CoT perturbation, input perturbation using direct prompting and input perturbation using CoT promptingFigure 26:Sports flip rates of four CoT models under 3 settings: CoT perturbation, input perturbation using direct prompting and input perturbation using CoT promptingFigure 27:StrategyQA flip rates of four CoT models under 3 settings: CoT perturbation, input perturbation using direct prompting and input perturbation using CoT promptingFrom[Figures24](https://arxiv.org/html/2609.30849#A9.F24),[25](https://arxiv.org/html/2609.30849#A9.F25),[26](https://arxiv.org/html/2609.30849#A9.F26)and[27](https://arxiv.org/html/2609.30849#A9.F27), we generally observe similar trends between input perturbations using direct prompting and CoT prompting\. More specifically, input perturbation still generally have higher flip rates than CoT perturbation, regardless of whether we use direct prompting or CoT prompting\. In addition, both prompting approaches also give similar flip rates\. Since the result patterns of input perturbation compared to CoT perturbation are similar regardless of the prompting approach, we decide to focus our subsequent analysis only on the results of input perturbation using direct prompting, which minimises the extraneous variables as discussed in[AppendixH](https://arxiv.org/html/2609.30849#A8)\.

Similar Articles

LLM Attribution Analysis Across Different Fine-Tuning Strategies and Model Scales for Automated Code Compliance

arXiv cs.CL

This paper analyzes how different fine-tuning strategies (FFT, LoRA, quantized LoRA) and model scales affect LLM interpretive behavior for automated code compliance tasks using perturbation-based attribution analysis. The findings show FFT produces more focused attribution patterns than parameter-efficient methods, and larger models develop specific interpretive strategies with diminishing performance returns beyond 7B parameters.