Uncovering Uncontrolled Repetition through Residual Stream Dynamics
Summary
This paper introduces Tokenwise Residual Comparison (TRC), a method that localizes repetition anomalies from residual stream dynamics during generation in LVLMs/LLMs and selectively suppresses them, reducing loop rates by 57% on average, while showing that repetition semantics emerge in shallow layers before propagating deeper.
View Cached Full Text
Cached at: 10/01/26, 09:45 AM
# Uncovering Uncontrolled Repetition through Residual Stream Dynamics Source: [https://arxiv.org/html/2609.38802](https://arxiv.org/html/2609.38802) Yuanhe ZhangXinyao ZhouAffiliation:Beijing University of Posts and TelecommunicationsEmail:[susen@bupt\.edu\.cn;](mailto:[email protected];)Haoran GaoAffiliation:JIUTIAN ResearchYuyao ZhangAffiliation:JIUTIAN ResearchZhenhong ZhouAffiliation:Nanyang Technological UniversityFanyu MengAffiliation:JIUTIAN ResearchLi SunAffiliation:Beijing University of Posts and TelecommunicationsSen SuAffiliation:Beijing University of Posts and TelecommunicationsAffiliation:Chongqing University of Posts and Telecommunications ###### Abstract Uncontrolled repetition can prolong autoregressive generation in large language models \(LLMs\) and enable resource consumption attacks\. Prior analyses of repetitive generation have identified strongly activated features in intermediate and late layers\. However, how uncontrolled repetition activity emerges and develops before becoming prominent in these layers remains insufficiently understood\. In this paper, we investigate this question primarily in large vision\-language models \(LVLMs\), which support a richer set of uncontrolled repetitions through both visual and textual inputs\. We proposeTokenwise Residual Comparison\(TRC\), a method that identifies and localizes anomalies associated with repetition from residual dynamics during generation\. TRC compares attention and multilayer perceptron writes to the residual stream across generated tokens to identify patterns associated with repetition\. It then selectively suppresses coordinates in the residual stream at the identified layer\. Experiments show that TRC effectively mitigates uncontrolled repetition, reducing loop rates by 57% on average\. Our analysis further shows that repetition semantics emerge in shallow layers and propagate through the residual stream, disrupting normal representations\. TRC also generalizes to large language models \(LLMs\) and large reasoning models \(LRMs\), where it consistently captures analogous repetition dynamics and achieves effective mitigation\. Our work broadens the study of repetitive generation from its prominent internal representations to earlier opportunities for intervention, providing insights for mitigating resource consumption attacks\. ††footnotemark:††footnotetext:†\\daggerindicates corresponding author\.## 1Introduction Language models can become trapped in repetitive generation, where recurring content disrupts the normal development of the response\([Hiraoka and Inui, 2025](https://arxiv.org/html/2609.38802#bib.bib9)\)\. Attackers can deliberately induce such repetition through textual or visual inputs to increase inference cost and latency, threatening service availability\([Li et al\., 2026](https://arxiv.org/html/2609.38802#bib.bib41);[Gao et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib40)\)\. Prior studies have primarily analyzed repetition using manually constructed repetitive sequences\([Yao et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib10);[Hiraoka and Inui, 2025](https://arxiv.org/html/2609.38802#bib.bib9)\)\. These analyses identify repetition\-related features concentrated in intermediate and final layers\([Hiraoka and Inui, 2025](https://arxiv.org/html/2609.38802#bib.bib9);[Yao et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib10)\)\. We argue that understanding these prominent representations also requires examining how uncontrolled repetition emerges across layers during model generation\. LVLMs provide a convenient setting for this analysis, as the continuous space of visual inputs facilitates the construction of examples that induce uncontrolled generation\([Gao et al\., 2024a](https://arxiv.org/html/2609.38802#bib.bib8);[Gao et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib40)\)\. By varying visual and textual triggers, we can study a broader range of model\-generated failure cases and examine repetition activity before it becomes prominent in deeper layers\. Existing mitigation strategies address repetition through decoding controls and internal activation interventions\([Xu et al\., 2022](https://arxiv.org/html/2609.38802#bib.bib22);[Yao et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib10);[Zhang et al\., 2026](https://arxiv.org/html/2609.38802#bib.bib58)\)\. At the output level, nucleus sampling and repetition penalties counter degenerate continuations by changing token selection\([Holtzman et al\., 2019](https://arxiv.org/html/2609.38802#bib.bib35);[Zhu et al\., 2023](https://arxiv.org/html/2609.38802#bib.bib36)\)\. Output\-budget constraints require balancing token cost against answer accuracy, while aggressive repetition penalties can impair legitimate generation\([Han et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib20);[Gao et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib40)\)\. At the activation level, prior studies identify repetition related neurons or features from constructed repetitive sequences and show that suppressing their activations can mitigate repetition\([Hiraoka and Inui, 2025](https://arxiv.org/html/2609.38802#bib.bib9);[Yao et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib10)\)\. However, these approaches primarily identify representations after the repetitive semantics have become strongly activated, leading to their localization mainly in intermediate or later layers\([Hiraoka and Inui, 2025](https://arxiv.org/html/2609.38802#bib.bib9);[Yao et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib10)\)\. Our question is whether early layers already exhibit signals of repetition that can support analysis\. In this paper, we investigate this question using repetition that naturally emerges from LVLM generation under visual and textual input triggers, rather than artificially constructed repetition in model outputs\. We proposeTokenwise Residual Comparison\(TRC\), which characterizes uncontrolled repetition changes in residual stream contributions\. TRC compares residual stream contributions across generated sequence to derive tokenwise repetition signals at each layer\. It then identifies candidate layers based on the magnitude of these signals relative to benign references and their local consistency across layers\. TRC scores at the selected layer are then used to construct a mask that selectively suppresses anomalous residual stream contributions\. By comparing residual stream contributions along the generated sequence, TRC exposes fine grained changes associated with repetition and localizes them to specific layers and coordinates\. Experiments show that suppressing residual contributions at the layer with the minimum TRC score reduces repetition by 57% on average\. The same criterion consistently localizes repetition related changes to shallow layers, suggesting that the signals captured by TRC emerge well before repetition becomes prominent in later representations\. TRC further generalizes to large language models \(LLMs\) and large reasoning models \(LRMs\), where it similarly mitigates repetitive generation and identifies corresponding changes at early layers\. These findings indicate that fine grained residual changes in shallow layers provide effective targets for observing and suppressing uncontrolled repetition before its representations become dominant\. In summary, we introduce TRC, which uses features from residual connections to characterize tokenwise repetition signals across layers\. TRC reveals that uncontrolled repetition already leaves distinguishable traces in shallow layers, where targeted suppression substantially mitigates repetitive generation under both visual and textual triggers in LVLMs\. This behavior further generalizes to LLMs and LRMs, indicating that early repetition signals and their intervention effects are shared across different model families\. Figure 1:Overview ofTRC, showing uncontrolled repetition across model families \(Left\), tokenwise residual comparison and selective intervention \(Middle\), and mitigation outcomes\(Right\)\. ## 2Related Work ### 2\.1Input Modalities and Adversarial Example Construction Different input modalities provide distinct perturbation spaces for adversarial example construction\. Text attacks operate over discrete tokens through efficient gradient guided modifications or direct contextual replacements, with continuous relaxations enabling optimization over discrete sequences\([Ebrahimi et al\., 2018](https://arxiv.org/html/2609.38802#bib.bib11);[Li et al\., 2020a](https://arxiv.org/html/2609.38802#bib.bib12);[Garg and Ramakrishnan, 2020](https://arxiv.org/html/2609.38802#bib.bib23);[Guo et al\., 2021](https://arxiv.org/html/2609.38802#bib.bib24)\)\. Visual and speech attacks instead optimize continuous pixels or waveforms under perceptual constraints\([Goodfellow et al\., 2014](https://arxiv.org/html/2609.38802#bib.bib25);[Carlini and Wagner, 2017](https://arxiv.org/html/2609.38802#bib.bib26);[Carlini and Wagner, 2018](https://arxiv.org/html/2609.38802#bib.bib27);[Qin et al\., 2019](https://arxiv.org/html/2609.38802#bib.bib13)\)\. Multimodal attacks further exploit interactions across channels through coordinated perturbations and cross modal alignment\([Zhang et al\., 2022](https://arxiv.org/html/2609.38802#bib.bib29);[Yin et al\., 2023](https://arxiv.org/html/2609.38802#bib.bib15);[Lu et al\., 2023](https://arxiv.org/html/2609.38802#bib.bib28)\)\. In LVLMs, optimized pixels and typographic visual prompts can bypass textual safety alignment, while visual optimization can also prolong generation\([Qi et al\., 2024](https://arxiv.org/html/2609.38802#bib.bib14);[Gong et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib30);[Gao et al\., 2024a](https://arxiv.org/html/2609.38802#bib.bib8)\)\. This flexibility facilitates the construction of diverse visual and textual triggers for studying model generated repetition\. ### 2\.2Resource Consumption Attacks and Defenses Resource consumption attacks arise through different mechanisms across model types\. Sponge examples increase energy use and latency, while DeepSloth and SlowBERT undermine early exit efficiency\([Shumailov et al\., 2021](https://arxiv.org/html/2609.38802#bib.bib16);[Hong et al\., 2020](https://arxiv.org/html/2609.38802#bib.bib33);[Zhang et al\., 2023](https://arxiv.org/html/2609.38802#bib.bib17)\)\. In autoregressive models, optimized prompts, natural instructions, adversarial images, poisoned training data, and injected reasoning decoys can all prolong computation or induce excessive outputs\([Dong et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib31);[Chen et al\., 2026](https://arxiv.org/html/2609.38802#bib.bib18);[Gao et al\., 2024a](https://arxiv.org/html/2609.38802#bib.bib8);[Gao et al\., 2024b](https://arxiv.org/html/2609.38802#bib.bib32);[Kumar et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib19);[Zhang et al\., 2025b](https://arxiv.org/html/2609.38802#bib.bib57)\)\. Existing mitigation includes token budget control, filtering or paraphrasing external context, decoding strategies, and training objectives that discourage degeneration or repetition\([Han et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib20);[Kumar et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib19);[Holtzman et al\., 2019](https://arxiv.org/html/2609.38802#bib.bib35);[Zhu et al\., 2023](https://arxiv.org/html/2609.38802#bib.bib36);[Su et al\., 2022](https://arxiv.org/html/2609.38802#bib.bib37);[Li et al\., 2023](https://arxiv.org/html/2609.38802#bib.bib21);[Welleck et al\., 2019](https://arxiv.org/html/2609.38802#bib.bib34);[Li et al\., 2020b](https://arxiv.org/html/2609.38802#bib.bib38);[Xu et al\., 2022](https://arxiv.org/html/2609.38802#bib.bib22)\)\. Existing internal analyses primarily capture prominent repetition activations, providing limited insight into when repetition first emerges during generation\([Hiraoka and Inui, 2025](https://arxiv.org/html/2609.38802#bib.bib9);[Yao et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib10)\)\. TRC instead tracks evolving semantic changes in residual stream contributions across the generated sequence to identify the emergence of repetition signals\. ## 3Tokenwise Residual Comparison Tokenwise Residual Comparison \(TRC\) identifies repetition signals by comparing differences in residual stream contributions, then selects residual coordinates for targeted suppression\. After specifying the calibration setting in Section[3\.1](https://arxiv.org/html/2609.38802#S3.SS1), we derive TRC scores and layer localization scores \(TRC\-l\) in Section[3\.2](https://arxiv.org/html/2609.38802#S3.SS2)\. Section[3\.3](https://arxiv.org/html/2609.38802#S3.SS3)converts the selected TRC scores into a fixed soft mask applied at the selected layer in residual addition\. ### 3\.1Preliminaries and Problem Setup We consider an LVLM whose autoregressive generation can be driven into uncontrolled repetition\. The defender has access to the transformer backbone and records the contributions transmitted through its residual connections during generation\. Calibration uses a set of attack\-induced repetition outputs𝒟A\\mathcal\{D\}\_\{\\mathrm\{A\}\}and a separate set of benign reference outputs𝒟N\\mathcal\{D\}\_\{\\mathrm\{N\}\}\. Attention and multilayer perceptron \(MLP\) modules are analyzed independently within each model, so we omit the module index throughout\. The formulation also applies separately to text\-only LLMs and LRMs\. Given an input request, the model generates one token at a time, conditioned on that request and its previously generated tokens\. We denote a completed output by𝐲=\(y1,…,yT\)\\mathbf\{y\}=\(y\_\{1\},\\ldots,y\_\{T\}\), whereyty\_\{t\}is the token at output positionttandT≥2T\\geq 2is the output length\. The backbone hasLLlayers indexed byl∈\{0,…,L−1\}l\\in\\\{0,\\ldots,L\-1\\\}, each with hidden dimensiondd\. At generation positiontt,𝐑l,t∈ℝd×St\\mathbf\{R\}\_\{l,t\}\\in\\mathbb\{R\}^\{d\\times S\_\{t\}\}denotes the recorded contribution of the selected module in residual addition, whereStS\_\{t\}is the total number of tokens from the input through generation positiontt\. Our goal is to characterize changes in the residual stream associated with uncontrolled repetition and use them to determine a localized intervention during generation\. Given the recorded module contributions, we seek to identify layers and hidden coordinates that exhibit distinctive changes under repetition, and selectively suppress these contributions while preserving benign generation\. ### 3\.2TRC Scores and Layer Localization We measure changes in the contributions transmitted through residual connections across generation positions\. To identify the dominant repeated pattern, we search overnn\-grams withn∈\{1,…,⌊T/2⌋\}n\\in\\\{1,\\ldots,\\lfloor T/2\\rfloor\\\}\. For eachnn, we consider onlynn\-grams occurring at least twice and measure the fraction of output tokens covered by their occurrences\. We then select the smallestnnwhose most\-covered repeatednn\-gram accounts for more than half of the output: n⋆=\\displaystyle n^\{\\star\}=min\{n:maxg:\|𝒪n\(g\)\|≥2\|⋃t∈𝒪n\(g\)\{t,…,t\+n−1\}\|T\>0\.5\},\\displaystyle\\min\\left\\\{n:\\max\_\{g:\\,\|\\mathcal\{O\}\_\{n\}\(g\)\|\\geq 2\}\\frac\{\\left\|\\bigcup\_\{t\\in\\mathcal\{O\}\_\{n\}\(g\)\}\\\{t,\\ldots,t\+n\-1\\\}\\right\|\}\{T\}\>0\.5\\right\\\},\(1\)g⋆=\\displaystyle g^\{\\star\}=argmaxg:\|𝒪n⋆\(g\)\|≥2\|⋃t∈𝒪n⋆\(g\)\{t,…,t\+n⋆−1\}\|,𝒫⋆=𝒪n⋆\(g⋆\),\\displaystyle\\operatorname\{arg\\,max\}\_\{g:\\,\|\\mathcal\{O\}\_\{n^\{\\star\}\}\(g\)\|\\geq 2\}\\left\|\\bigcup\_\{t\\in\\mathcal\{O\}\_\{n^\{\\star\}\}\(g\)\}\\\{t,\\ldots,t\+n^\{\\star\}\-1\\\}\\right\|,\\qquad\\mathcal\{P\}^\{\\star\}=\\mathcal\{O\}\_\{n^\{\\star\}\}\(g^\{\\star\}\),whereg=\(g1,…,gn\)g=\(g\_\{1\},\\ldots,g\_\{n\}\)denotes annn\-gram in𝐲\\mathbf\{y\}, and𝒪n\(g\)=\{t∈\{1,…,T−n\+1\}:\(yt,…,yt\+n−1\)=g\}\\mathcal\{O\}\_\{n\}\(g\)=\\\{t\\in\\\{1,\\ldots,T\-n\+1\\\}:\(y\_\{t\},\\ldots,y\_\{t\+n\-1\}\)=g\\\}denotes the set of starting positions at whichggoccurs in the output\. Thus,n⋆n^\{\\star\}gives the minimumnnfor which a repeatednn\-gram covers more than50%50\\%of the output, while𝒫⋆\\mathcal\{P\}^\{\\star\}records all starting positions of the selected repeatednn\-gram\. We then determine the comparison startss, token spacingqq, and comparison positions𝒬\\mathcal\{Q\}as: \(s,q\)=\{\(min𝒫⋆,n⋆\),ifn⋆is defined,\(1,1\),otherwise,𝒬=\{s\+rq\|r∈ℕ0,s\+\(r\+1\)q≤T\}\.\(s,q\)=\\begin\{cases\}\(\\min\\mathcal\{P\}^\{\\star\},n^\{\\star\}\),&\\text\{if \}n^\{\\star\}\\text\{ is defined\},\\\\ \(1,1\),&\\text\{otherwise\},\\end\{cases\}\\qquad\\mathcal\{Q\}=\\\{s\+rq\|r\\in\\mathbb\{N\}\_\{0\},\\ s\+\(r\+1\)q\\leq T\\\}\.\(2\)The setℕ0\\mathbb\{N\}\_\{0\}contains the nonnegative integers, and eacht∈𝒬t\\in\\mathcal\{Q\}serves as the earlier position in a pairwise comparison with\(t,t\+q\)\(t,t\+q\)\. SinceSt\+q=St\+qS\_\{t\+q\}=S\_\{t\}\+q, we align each comparison pair by retaining the trailing columns corresponding to the newly generated tokens in𝐑l,t\\mathbf\{R\}\{l,t\}and𝐑l,t\+q\\mathbf\{R\}\{l,t\+q\}\. We then compute the element\-wise absolute difference between these aligned module contributions: 𝐯l=1\|𝒬\|∑t∈𝒬Mean\(\|𝐑l,t\+q\[:,St\+q−1:\]−𝐑l,t\[:,St−1:\]\|\)\.\\mathbf\{v\}\_\{l\}=\\frac\{1\}\{\|\\mathcal\{Q\}\|\}\\sum\_\{t\\in\\mathcal\{Q\}\}\\operatorname\{Mean\}\\left\(\\left\|\\mathbf\{R\}\_\{l,t\+q\}\[:,S\_\{t\+q\-1\}:\]\-\\mathbf\{R\}\_\{l,t\}\[:,S\_\{t\-1\}:\]\\right\|\\right\)\.\(3\)Here,Mean\(⋅\)\\operatorname\{Mean\}\(\\cdot\)denotes averaging over the column dimension corresponding to token positions, and𝐯l∈ℝd\\mathbf\{v\}\_\{l\}\\in\\mathbb\{R\}^\{d\}represents the contribution change at layerllaveraged across all comparison pairs\. We then aggregate the sample\-level vectors across all attack calibration samples under the same request condition to obtain the attack TRC score vector: 𝒯lA=1\|𝒟A\|∑z∈𝒟A𝐯l\(z\),\\mathscr\{T\}^\{\\mathrm\{A\}\}\_\{l\}=\\frac\{1\}\{\|\\mathcal\{D\}\_\{\\mathrm\{A\}\}\|\}\\sum\_\{z\\in\\mathcal\{D\}\_\{\\mathrm\{A\}\}\}\\mathbf\{v\}\_\{l\}\(z\),\(4\)𝐯l\(z\)\\mathbf\{v\}\_\{l\}\(z\)is the contribution\-change vector computed for samplezz\. The resulting𝒯lA∈ℝd\\mathscr\{T\}^\{\\mathrm\{A\}\}\_\{l\}\\in\\mathbb\{R\}^\{d\}represents the attack TRC scores at layerll, with each entry corresponding to one hidden coordinate\. We analogously aggregate the contribution\-change vectors over the benign calibration set to obtain𝒯lN\\mathscr\{T\}^\{\\mathrm\{N\}\}\_\{l\}\. For layer\-wise analysis, we summarize the attack and benign TRC scores over hidden coordinates\. For a controllable window spanh∈\{1,…,L−1\}h\\in\\\{1,\\ldots,L\-1\\\}, we define the normal\-relative magnitude and the cross\-layer variation as: al=Mean\(𝒯lA\)Mean\(𝒯lN\),elA=\(Mean\(𝒯l\+hA\)−Mean\(𝒯lA\)h\)2\.a\_\{l\}=\\frac\{\\operatorname\{Mean\}\(\\mathscr\{T\}^\{\\mathrm\{A\}\}\_\{l\}\)\}\{\\operatorname\{Mean\}\(\\mathscr\{T\}^\{\\mathrm\{N\}\}\_\{l\}\)\},\\qquad e\_\{l\}^\{A\}=\\left\(\\frac\{\\operatorname\{Mean\}\(\\mathscr\{T\}^\{\\mathrm\{A\}\}\_\{l\+h\}\)\-\\operatorname\{Mean\}\(\\mathscr\{T\}^\{\\mathrm\{A\}\}\_\{l\}\)\}\{h\}\\right\)^\{2\}\.\(5\)Where,ala\_\{l\}measures the magnitude of attack\-induced contribution changes relative to benign generation, whileelAe\_\{l\}^\{A\}measures the net variation of the attack TRC scores over a local layer interval of spanhh\. Attention and MLP modules are calibrated separately within each model\. To normalize the cross\-layer variation using benign observations only, we construct a benign layer\-wise reference curve from𝒟N\\mathcal\{D\}\_\{\\mathrm\{N\}\}using the same aggregation procedure as for the attack samples\. We collect its positive cross\-layer variation values and define the TRC\-l score as ℒl=al\(1\+elAexp\(Meane∈ℰNloge\)\),ℰN=\{elN∣0≤l≤L−h−1,elN\>0\}\.\\mathscr\{L\}\_\{l\}=a\_\{l\}\\left\(1\+\\frac\{e\_\{l\}^\{\\mathrm\{A\}\}\}\{\\exp\\left\(\\operatorname\{Mean\}\_\{e\\in\\mathcal\{E\}\_\{\\mathrm\{N\}\}\}\\log e\\right\)\}\\right\),\\qquad\\mathcal\{E\}\_\{\\mathrm\{N\}\}=\\left\\\{e\_\{l\}^\{\\mathrm\{N\}\}\\mid 0\\leq l\\leq L\-h\-1,\\ e\_\{l\}^\{\\mathrm\{N\}\}\>0\\right\\\}\.\(6\)elNe\_\{l\}^\{\\mathrm\{N\}\}denotes the cross\-layer variation of the benign reference curve starting at layerll, computed using the same formulation aselAe\_\{l\}^\{\\mathrm\{A\}\}\. The geometric mean overℰN\\mathcal\{E\}\_\{\\mathrm\{N\}\}provides a benign reference scale for the variation term, such thatℒl\\mathscr\{L\}\_\{l\}jointly captures the normal\-relative magnitude and the normalized cross\-layer variation of the attack curve\. We useℒl\\mathscr\{L\}\_\{l\}as the criterion for localizing the intervention layer, favoring candidate starting layers with both a low normal\-relative magnitude and limited variation over the followinghhlayers\. The starting layer is selected by minimizingℒl\\mathscr\{L\}\_\{l\}, with ties broken in favor of the shallowest layerl⋆=min\(argmin0≤l≤L−h−1ℒl\)l^\{\\star\}=\\min\\left\(\\operatorname\*\{argmin\}\_\{0\\leq l\\leq L\-h\-1\}\\mathscr\{L\}\_\{l\}\\right\)\. The selected localization window is represented by\(l⋆,h\)\(l^\{\\star\},h\), wherel⋆l^\{\\star\}specifies the starting layer andhhis the controllable window span\. ### 3\.3Selective Residual\-Stream Intervention The localization stage returnsl⋆l^\{\\star\}as the selected intervention layer\. Atl⋆l^\{\\star\}, we use the TRC scores computed from attack samplesℒl⋆A\\mathscr\{L\}^\{\\mathrm\{A\}\}\_\{l^\{\\star\}\}, to identify the residual\-stream contributions most associated with repetition\. Given a direction\-selection ratioρ∈\(0,1\]\\rho\\in\(0,1\], we selectΩl⋆=Top⌊ρd⌋\(𝒯l⋆A\)\\Omega\_\{l^\{\\star\}\}=\\operatorname\{Top\}\_\{\\lfloor\\rho d\\rfloor\}\\left\(\\mathscr\{T\}^\{\\mathrm\{A\}\}\_\{l^\{\\star\}\}\\right\), whereΩl⋆\\Omega\_\{l^\{\\star\}\}contains the⌊ρd⌋\\lfloor\\rho d\\rfloorcoordinates with the largest attack TRC scores at the selected layer\. We further determine the suppression strength directly from the layer\-level TRC\-l score: αl⋆=max\(0,lnℒl⋆N−lnℒl⋆A\)1\+max\(0,lnℒl⋆N−lnℒl⋆A\)\.\\alpha\_\{l^\{\\star\}\}=\\frac\{\\max\\left\(0,\\ln\\mathscr\{L\}^\{\\mathrm\{N\}\}\_\{l^\{\\star\}\}\-\\ln\\mathscr\{L\}^\{\\mathrm\{A\}\}\_\{l^\{\\star\}\}\\right\)\}\{1\+\\max\\left\(0,\\ln\\mathscr\{L\}^\{\\mathrm\{N\}\}\_\{l^\{\\star\}\}\-\\ln\\mathscr\{L\}^\{\\mathrm\{A\}\}\_\{l^\{\\star\}\}\\right\)\}\.\(7\)The resultingαl⋆∈\[0,1\)\\alpha\_\{l^\{\\star\}\}\\in\[0,1\)increases as the TRC\-l score decreases, assigning stronger suppression to layers exhibiting a smaller normal\-relative magnitude and more stable cross\-layer variation\. The direction\-selection ratioρ\\rhocontrols the intervention sparsity, whileαl⋆\\alpha\_\{l^\{\\star\}\}controls the suppression strength\. During inference, let𝐁l⋆∈ℝd×S\\mathbf\{B\}\_\{l^\{\\star\}\}\\in\\mathbb\{R\}^\{d\\times S\}denote the current module contribution overSSsequence positions and𝐔l⋆∈ℝd×S\\mathbf\{U\}\_\{l^\{\\star\}\}\\in\\mathbb\{R\}^\{d\\times S\}the residual stream entering the corresponding residual addition\. We construct a diagonal soft mask𝐌l⋆∈\[0,1\]d×d\\mathbf\{M\}\_\{l^\{\\star\}\}\\in\[0,1\]^\{d\\times d\}and apply it in residual addition: \[𝐌l⋆\]u,u=\{1−αl⋆,u∈Ωl⋆,𝐁~l⋆=𝐁l⋆\+𝐌l⋆𝐔l⋆\.\[\\mathbf\{M\}\_\{l^\{\\star\}\}\]\_\{u,u\}=\\begin\{cases\}1\-\\alpha\_\{l^\{\\star\}\},&u\\in\\Omega\_\{l^\{\\star\}\},\\\\ \\end\{cases\}\\qquad\\widetilde\{\\mathbf\{B\}\}\_\{l^\{\\star\}\}=\\mathbf\{B\}\_\{l^\{\\star\}\}\+\\mathbf\{M\}\_\{l^\{\\star\}\}\\mathbf\{U\}\_\{l^\{\\star\}\}\.\(8\)Left\-multiplying𝐔l⋆\\mathbf\{U\}\_\{l^\{\\star\}\}by𝐌l⋆\\mathbf\{M\}\_\{l^\{\\star\}\}scales each selected row by1−αl⋆1\-\\alpha\_\{l^\{\\star\}\}in residual addition\. The mask is shared across sequence positions and is applied only to the selected target layer, leaving all other layers unchanged\. For every layerl≠l⋆l\\neq l^\{\\star\}, the module contribution is added without modification\. The selected layer, coordinate set, and soft mask are fixed after calibration and directly applied to subsequent evaluation requests\. ## 4Experiments ### 4\.1Experimental Setup #### Models and Evaluation Scope Our evaluation suite comprises three groups\. The LVLM group includes InstructBLIP\-Vicuna\-7B\([Dai et al\., 2023](https://arxiv.org/html/2609.38802#bib.bib1)\), Qwen2\.5\-VL\-3B\-Instruct\([Team, 2025](https://arxiv.org/html/2609.38802#bib.bib3)\), and LLaVA\-1\.5\-7B\([Liu et al\., 2024a](https://arxiv.org/html/2609.38802#bib.bib2)\)\. The text\-only LLM group includes Llama\-3\.2\-3B222Official Llama 3\.2 model card:[https://github\.com/meta\-llama/llama\-models/blob/main/models/llama3\_2/MODEL\_CARD\.md](https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md)\.and Qwen2\.5\-3B\([Team, 2024](https://arxiv.org/html/2609.38802#bib.bib4)\)\. The reasoning\-model \(LRM\) group includes DeepSeek\-Llama\-8B\([Guo et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib5)\), Qwen3\.6\-27B\([Team, 2026](https://arxiv.org/html/2609.38802#bib.bib6)\), and GLM\-4\.7\-Flash\([Zeng et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib7)\)\. #### Attack and Benign Evaluation Data The attack suite separates visual perturbations from textual induction\. RECITE optimizes image perturbations to elicit repeated output\([Gao et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib40)\)\. Textual conditions comprise GCG\-based prompt optimization\([Zou et al\., 2023](https://arxiv.org/html/2609.38802#bib.bib44)\), LoopLLM’s repetition\-inducing attacks\([Li et al\., 2026](https://arxiv.org/html/2609.38802#bib.bib41)\), and direct repetition instructions \(Direct\)\. Benign evaluation uses ScienceQA for multimodal science question answering\([Lu et al\., 2022](https://arxiv.org/html/2609.38802#bib.bib45)\)and TextVQA for answering questions requiring scene\-text understanding\([Singh et al\., 2019](https://arxiv.org/html/2609.38802#bib.bib46)\)\. MMLU assesses knowledge and reasoning across text\-based subjects for the LLMs and LRMs\([Hendrycks et al\., 2020](https://arxiv.org/html/2609.38802#bib.bib47)\)\. Table[6](https://arxiv.org/html/2609.38802#A1.T6)in the Appendix specifies the attack coverage and benign tasks for each model\. #### Baselines We compare TRC with three defense baselines, alongside the unmodified model\.Fixed\-length truncation \(Fixed\-length\)stops decoding at a preset output\-token limit\([Zhang et al\., 2025a](https://arxiv.org/html/2609.38802#bib.bib39);[Gao et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib40)\)\.No\-repeatuses no\-repeatnn\-gram blocking to prevent any next token from reproducing a previously generated tokennn\-gram\([Fu et al\., 2026](https://arxiv.org/html/2609.38802#bib.bib43);[Zhu et al\., 2023](https://arxiv.org/html/2609.38802#bib.bib36)\)\.Interpretability\-based intervention \(AUSteer\)adapts AUSteer’s selection and steering of individual activation dimensions to repetition suppression\([Feng et al\., 2026](https://arxiv.org/html/2609.38802#bib.bib42)\)\. #### Evaluation Metrics We evaluate generation length, task accuracy, and loop rate\. Following Hiraoka and Inui\([Hiraoka and Inui, 2025](https://arxiv.org/html/2609.38802#bib.bib9)\), we identify repetitive outputs by detecting recurring token subsequences in the generated sequence\. The loop rate is defined asNrep/NN\_\{\\mathrm\{rep\}\}/N, whereNrepN\_\{\\mathrm\{rep\}\}is the number of samples exhibiting repetitive loops andNNis the total number of evaluated samples\. ### 4\.2Mitigation Effectiveness and Generalization #### Defense Effectiveness\. Table 1:Defense results on three LVLMs\. Lower is better for both metrics\.Table 2:Benign task accuracy \(%\)\.Table[1](https://arxiv.org/html/2609.38802#S4.T1)shows that TRC consistently suppresses uncontrolled repetition across visual and textual attacks, reducing both generation length and loop rate in most settings\. Its numerical efficiency is weaker than direct decoding constraints such as no\-repeat, since TRC intervenes through localized internal signals rather than explicitly blocking repeated output patterns\. Importantly, Table[2](https://arxiv.org/html/2609.38802#S4.T2)shows that this intervention has little effect on benign task accuracy\. These results indicate that TRC can mitigate repetition while preserving normal generation, supporting the identified residual changes as meaningful intervention targets\. #### Generalization\. We next examine whether TRC generalizes beyond LVLMs to text only LLMs and LRMs\. As shown in Figure[2](https://arxiv.org/html/2609.38802#S4.F2), TRC consistently reduces generation length and loop rate across model families whenever the underlying attack succeeds, with especially strong effects on LRMs\. Table[3](https://arxiv.org/html/2609.38802#S4.T3)shows that these gains come with little change in benign MMLU accuracy\. We further observe more stable intervention effects on newer and larger models\. A possible explanation is that advances in model scale and architecture lead to better separation between repetition related residual changes and normal semantic representations, allowing targeted suppression to preserve task relevant behavior more effectively\. The cross model results indicate that TRC captures internal repetition signals that transfer beyond the original LVLM setting\. Figure 2:Generalization of TRC to LLMs and LRMs under Direct, GCG, and LoopLLM attacks\. Bars show generation length and lines show loop rate; lower is better for both\.Table 3:Benign task accuracy on LLMs and LRMs\. TRC causes minor changes in performance\. ### 4\.3Mechanistic Analysis We use Qwen2\.5\-VL\-3B\-Instruct as the primary model for the detailed mechanistic analyses in this section; experiments comparing model families include the other specified models\. \(a\) Top 1 TRC localization across model families\. \(b\) Cross layer ranking of TRC Figure 3:Layer wise localization of uncontrolled repetition\. \(a\) Top 1 TRC localization across LVLM, LLM, and LRM families, showing consistently shallow attack locations\. \(b\) Cross layer TRC rankings under uncontrolled repetition, benign repetition, and normal requests, revealing a distinct early layer profile for uncontrolled repetition\.#### Shallow Localization of Repetition Signals\. Figure[3](https://arxiv.org/html/2609.38802#S4.F3)\(a\) compares the Top\-1 layers identified by TRC for an LVLM, a text\-only LLM, and an LRM\. Across all evaluated attack conditions, both Attention and MLP branches consistently localize repetition signals to shallow layers within 0–3, whereas benign references are localized substantially later, between layers 9 and 18\. This separation is preserved across three models, suggesting that uncontrolled repetition residual changes emerge early rather than at a model\-specific depth\. Attention provides the more stable localization signal, with nine of ten conditions concentrated at layer 1, while MLP locations vary across layers 0, 1, and 3\. These results therefore identify the residual stream entering shallow Attention layers as a particularly consistent location of repetition related changes\. Notably, its concentration at layer 1 corresponds to the residual representation immediately after the layer 0 MLP, indicating that repetition related changes are already prominent after the first MLP transformation\. #### Specificity to Uncontrolled Repetition\. To test whether TRC merely responds to repeated tokens or repeated semantics, we compare layer\-wise localization under three conditions: uncontrolled\-repetition failures, benign requests containing legitimate repetition, and ordinary normal requests\. Figure[3](https://arxiv.org/html/2609.38802#S4.F3)\(b\) shows that benign repetition closely follows the localization profile of normal requests across layers\. In both cases, the highest\-ranked region shifts toward the middle layers before weakening in later layers\. Uncontrolled repetition follows a different trajectory, with candidate locations ranked substantially higher in the early layers and progressively lower at greater depths\. This separation rules out the simple explanation that TRC detects repetition semantics or repeated token identity alone\. Instead, the localized signal is associated with uncontrolled repetition collapse\. The result therefore supports the specificity of TRC to failure\-related repetition dynamics, while the causal role of the localized components is evaluated separately in subsequent chapters\. The absolute scores and their depth\-dependent trends are examined in Appendix[D](https://arxiv.org/html/2609.38802#A4)\. #### Attention and MLP Branch Contributions\. Figure[4](https://arxiv.org/html/2609.38802#S4.F4)\(a\) shows that suppressing the full Attention write provides a better mitigation–utility tradeoff than suppressing individual heads, suggesting that repetition is not yet concentrated in a specific head at shallow layers\. In contrast, suppressing the MLP write or both branches causes substantially larger degradation on benign tasks\. Since the intervened residual stream already contains the output of the preceding block, these results suggest that repetition related features are formed early in MLP computations and then propagated through subsequent residual updates\. This interpretation is consistent with prior studies characterizing MLPs as key value memories and linking them to concept and knowledge representations\([Geva et al\., 2021](https://arxiv.org/html/2609.38802#bib.bib48);[Geva et al\., 2022](https://arxiv.org/html/2609.38802#bib.bib49);[Dai et al\., 2022](https://arxiv.org/html/2609.38802#bib.bib50);[Meng et al\., 2022](https://arxiv.org/html/2609.38802#bib.bib51)\)\. \(a\) Branch selection \(b\) Layer selection Figure 4:Mechanistic evidence for branch and layer selection\. \(a\) Branch\-wise suppression on attack and normal examples\. Attention\-only suppression preserves normal accuracy, whereas MLP\-only and joint suppression reduce attack length at substantially higher utility cost\. \(b\) Effect of the suppression start layer on Recite repetition suppression and ScienceQA accuracy\. The star denotes the original model without suppression\. #### Intervention Effects of Layer and Direction Selection\. To isolate the effect of layer choice, we keep the suppression rule fixed and vary only the intervention layer\. As shown in Figure[4](https://arxiv.org/html/2609.38802#S4.F4)\(b\), the shallow layer selected by TRC achieves the best mitigation–utility trade\-off: layer 1 suppresses all of repetitive failures while retaining 81\.5% ScienceQA accuracy, close to the original model\. Applying the same intervention at deeper layers can maintain high suppression but causes substantially larger accuracy degradation\. These results show that effective repetition mitigation depends critically on where suppression is applied, supporting TRC’s shallow\-layer localization rather than depth\-agnostic intervention\. #### Attention Update Magnitude and Residual Propagation\. Table 4:Attention\-update magnitudes in the first two blocks\. Gap/τl<1\\tau\_\{l\}<1indicates that the condition difference remains within the normal token\-level fluctuation bound\.We further examine whether shallow attention blocks repeatedly amplify the repetition signal, or whether the signal is mainly retained through residual propagation\. For the first two attention blocks, we measure the update magnitude𝐃l=𝐔lout−𝐔lin\\mathbf\{D\}\_\{l\}=\\mathbf\{U\}^\{\\mathrm\{out\}\}\_\{l\}\-\\mathbf\{U\}^\{\\mathrm\{in\}\}\_\{l\}using both L2 norm and RMS, and normalize the condition gap by the corresponding normal fluctuation thresholdτl\\tau\_\{l\}\. As shown in Table[4](https://arxiv.org/html/2609.38802#S4.T4), the repetition\-induced gaps remain within normal fluctuation ranges for both metrics, reaching only 23\.4%/17\.2% of the threshold in Block 0 and 73\.0%/69\.2% in Block 1 for L2/RMS, respectively\. This indicates that repetitive generation is not accompanied by an abnormal increase in shallow attention\-update magnitude\. Combined with the earlier localization results, this is more consistent with repetition\-related features emerging through shallow attention transformations and then being retained through the residual stream, rather than being repeatedly amplified by subsequent attention transformations\. This early emergence precedes the progressive strengthening of repetition related semantics in intermediate layers, as observed in prior activation level studies\([Hiraoka and Inui, 2025](https://arxiv.org/html/2609.38802#bib.bib9);[Yao et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib10)\)\. #### Propagation and Activation Restoration\. We test whether the coordinates localized by TRC capture an early attack\-associated feature using GCG examples that differ only by the optimized suffix\. At the first layer, we perform unit\-strength cross\-interventions by either restoring localized attack coordinates with values from the paired normal trajectory or injecting the corresponding attack\-minus\-normal difference into the normal trajectory, with equal\-size random coordinates as controls\. As shown in Table[5](https://arxiv.org/html/2609.38802#S4.T5), restoring the TRC\-localized coordinates converts 80% repetitive attack trajectories to non\-repetitive outputs and reduces average generation length from 4096 to 828 tokens, substantially outperforming random restoration\. This indicates that the localized coordinates already encode behaviorally relevant attack\-associated information at the first layer and that restoring these early activations can propagate to downstream recovery\. Conversely, injecting the attack\-associated difference into normal trajectories disrupts generation but does not reproduce repetition, suggesting that these coordinates are important for the failure state but are not independently sufficient to induce it\. Together, the results support TRC as identifying early residual features that contribute to repetition and can be causally restored to mitigate the failure\. Appendix[E](https://arxiv.org/html/2609.38802#A5)presents representative outputs for both intervention directions\. Table 5:Cross intervention results with unit strength on GCG examples\.TRCdenotes coordinates selected by TRC, whileRandomuses an equal size coordinate set\. Transition success measures repetition to nonrepetition for attack inputs and nonrepetition to repetition for normal inputs\.Δ\\DeltaLen is computed relative to the corresponding no intervention baseline\.Appendix[F](https://arxiv.org/html/2609.38802#A6)further reports sensitivity analyses on the layer window span, the number of attack training examples, the suppression ratio, and the maximum repetition distance used by the repeatednngram detector\. These results examine the robustness of TRC to its main design choices\. ## 5Conclusion We presentedTokenwise Residual Comparison\(TRC\), a framework for identifying and mitigating uncontrolled repetition by comparing residual stream contributions across generation positions\. Rather than focusing only on prominent repetition representations in intermediate or later layers, TRC tracks how residual contributions evolve along the generated sequence and localizes fine grained repetition signals to specific layers and coordinates\. Across LVLMs, text only LLMs, and LRMs, TRC consistently identifies shallow repetition signals and enables targeted suppression that reduces repetitive generation while largely preserving benign task performance\. Mechanistic analyses further show that these signals emerge early, remain distinguishable from benign and legitimate repetition, and can be partially restored through localized activation replacement, supporting residual propagation as a plausible mechanism by which repetition related features persist through the network\. Together, these findings shift the analysis of uncontrolled repetition from where strong repetition representations are observed to how they emerge early and become actionable, providing a finer grained perspective for understanding and mitigating resource consumption failures in autoregressive models\. ## AI use statement Generative AI tools were used to assist with literature retrieval, review manuscript formatting, and suggest caption and editorial revisions\. The authors reviewed all AI\-assisted outputs and suggestions and take full responsibility for the final text, data, results, and claims\. ## Ethics statement This work studies repetition\-inducing attacks and a defense against them using existing model and benchmark data\. Attack procedures are reported to support evaluation and defense research; they may also be misused to increase inference costs or disrupt model services\. We report defensive results alongside benign\-task performance to make this trade\-off visible\. ### Reproducibility statement The method and evaluation metrics are described in Sections[3](https://arxiv.org/html/2609.38802#S3)and[4\.1](https://arxiv.org/html/2609.38802#S4.SS1); the models, attack conditions, and benign tasks are listed in Appendix[A](https://arxiv.org/html/2609.38802#A1)\. The appendix also documents the construction and validation of the legitimate\-repetition controls\. ## References - Buenoet al\.\(2024\)M\. C\. Bueno, R\. Lotufo, and R\. F\. NogueiraMLissard: multilingual long and simple sequential reasoning benchmarks\.InProceedings of the 2nd GenBench Workshop on Generalisation \(Benchmarking\) in NLP,pp\. 86–95\.Cited by:[Appendix C](https://arxiv.org/html/2609.38802#A3.SS0.SSS0.Px1.p2.1)\. - Carlini and Wagner \(2017\)N\. Carlini and D\. WagnerTowards evaluating the robustness of neural networks\.In2017 ieee symposium on security and privacy \(sp\),pp\. 39–57\.Cited by:[§2\.1](https://arxiv.org/html/2609.38802#S2.SS1.p1.1)\. - Carlini and Wagner \(2018\)N\. Carlini and D\. WagnerAudio adversarial examples: targeted attacks on speech\-to\-text\.In2018 IEEE security and privacy workshops \(SPW\),pp\. 1–7\.Cited by:[§2\.1](https://arxiv.org/html/2609.38802#S2.SS1.p1.1)\. - Chenet al\.\(2026\)Y\. Chen, Z\. Li, X\. Yue, R\. T\. Tan, and H\. LiNaturalSloth: revisiting denial\-of\-service attacks on large language models\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 19685–19702\.Cited by:[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Daiet al\.\(2022\)D\. Dai, L\. Dong, Y\. Hao, Z\. Sui, B\. Chang, and F\. WeiKnowledge neurons in pretrained transformers\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 8493–8502\.Cited by:[§4\.3](https://arxiv.org/html/2609.38802#S4.SS3.SSS0.Px3.p1.1)\. - Daiet al\.\(2023\)W\. Dai, J\. Li, D\. Li, A\. Tiong, J\. Zhao, W\. Wang, B\. Li, P\. N\. Fung, and S\. HoiInstructblip: towards general\-purpose vision\-language models with instruction tuning\.Vol\.36\.Cited by:[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px1.p1.1)\. - Donget al\.\(2025\)J\. Dong, Z\. Zhang, Q\. Zhang, T\. Zhang, H\. Wang, H\. Li, Q\. Li, C\. Zhang, K\. Xu, and H\. QiuAn engorgio prompt makes large language model babble on\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 67280–67307\.Cited by:[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Ebrahimiet al\.\(2018\)J\. Ebrahimi, A\. Rao, D\. Lowd, and D\. DouHotflip: white\-box adversarial examples for text classification\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),pp\. 31–36\.Cited by:[§2\.1](https://arxiv.org/html/2609.38802#S2.SS1.p1.1)\. - Fenget al\.\(2026\)Z\. Feng, T\. Li, Z\. Zhu, H\. Zhou, J\. Qian, L\. Zhang, C\. Deryl, L\. Mak, G\. Ng, and K\. MaoFine\-grained activation steering: steering less, achieving more\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 39421–39443\.Cited by:[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px3.p1.1)\. - Fuet al\.\(2026\)J\. Fu, K\. Jiang, L\. Hong, J\. Li, H\. Guo, D\. Yang, Z\. Chen, and W\. ZhangLingoloop attack: trapping mllms via linguistic context and state entrapment into endless loops\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 87860–87893\.Cited by:[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px3.p1.1)\. - Gaoet al\.\(2025\)H\. Gao, Y\. Zhang, Z\. Zhou, L\. Jiang, F\. Meng, Y\. Xiao, L\. Sun, K\. Wang, Y\. Liu, and J\. FengResource consumption red\-teaming for large vision\-language models\.Cited by:[Appendix A](https://arxiv.org/html/2609.38802#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38802#S1.p1.1),[§1](https://arxiv.org/html/2609.38802#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px3.p1.1)\. - Gaoet al\.\(2024a\)K\. Gao, Y\. Bai, J\. Gu, S\. Xia, P\. Torr, Z\. Li, and W\. LiuInducing high energy\-latency of large vision\-language models with verbose images\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 17156–17182\.Cited by:[§1](https://arxiv.org/html/2609.38802#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.38802#S2.SS1.p1.1),[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Gaoet al\.\(2024b\)K\. Gao, T\. Pang, C\. Du, Y\. Yang, S\. Xia, and M\. LinDenial\-of\-service poisoning attacks against large language models\.arXiv preprint arXiv:2410\.10760\.Cited by:[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Garg and Ramakrishnan \(2020\)S\. Garg and G\. RamakrishnanBAE: bert\-based adversarial examples for text classification\.InProceedings of the 2020 conference on empirical methods in natural language processing \(EMNLP\),pp\. 6174–6181\.Cited by:[§2\.1](https://arxiv.org/html/2609.38802#S2.SS1.p1.1)\. - Gevaet al\.\(2022\)M\. Geva, A\. Caciularu, K\. Wang, and Y\. GoldbergTransformer feed\-forward layers build predictions by promoting concepts in the vocabulary space\.InProceedings of the 2022 conference on empirical methods in natural language processing,pp\. 30–45\.Cited by:[Table 7](https://arxiv.org/html/2609.38802#A2.T7.2.2.1.1.1),[Appendix D](https://arxiv.org/html/2609.38802#A4.p4.1),[§4\.3](https://arxiv.org/html/2609.38802#S4.SS3.SSS0.Px3.p1.1)\. - Gevaet al\.\(2021\)M\. Geva, R\. Schuster, J\. Berant, and O\. LevyTransformer feed\-forward layers are key\-value memories\.InProceedings of the 2021 conference on empirical methods in natural language processing,pp\. 5484–5495\.Cited by:[Table 7](https://arxiv.org/html/2609.38802#A2.T7.2.2.1.1.1),[Appendix D](https://arxiv.org/html/2609.38802#A4.p4.1),[Appendix D](https://arxiv.org/html/2609.38802#A4.p5.1),[§4\.3](https://arxiv.org/html/2609.38802#S4.SS3.SSS0.Px3.p1.1)\. - Gonget al\.\(2025\)Y\. Gong, D\. Ran, J\. Liu, C\. Wang, T\. Cong, A\. Wang, S\. Duan, and X\. WangFigstep: jailbreaking large vision\-language models via typographic visual prompts\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 23951–23959\.Cited by:[§2\.1](https://arxiv.org/html/2609.38802#S2.SS1.p1.1)\. - Goodfellowet al\.\(2014\)I\. J\. Goodfellow, J\. Shlens, and C\. SzegedyExplaining and harnessing adversarial examples\.arXiv preprint arXiv:1412\.6572\.Cited by:[§2\.1](https://arxiv.org/html/2609.38802#S2.SS1.p1.1)\. - Guoet al\.\(2021\)C\. Guo, A\. Sablayrolles, H\. Jégou, and D\. KielaGradient\-based adversarial attacks against text transformers\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 5747–5757\.Cited by:[§2\.1](https://arxiv.org/html/2609.38802#S2.SS1.p1.1)\. - Guoet al\.\(2025\)D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi,et al\.Deepseek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.Cited by:[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px1.p1.1)\. - Hanet al\.\(2025\)T\. Han, Z\. Wang, C\. Fang, S\. Zhao, S\. Ma, and Z\. ChenToken\-budget\-aware llm reasoning\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 24842–24855\.Cited by:[§1](https://arxiv.org/html/2609.38802#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Hendryckset al\.\(2020\)D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. SteinhardtMeasuring massive multitask language understanding\.arXiv preprint arXiv:2009\.03300\.Cited by:[Appendix A](https://arxiv.org/html/2609.38802#A1.SS0.SSS0.Px6.p1.1),[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px2.p1.1)\. - Hiraoka and Inui \(2025\)T\. Hiraoka and K\. InuiRepetition neurons: how do language models produce repetitions?\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 2: Short Papers\),pp\. 483–495\.Cited by:[Table 7](https://arxiv.org/html/2609.38802#A2.T7.2.4.1.1.1),[§1](https://arxiv.org/html/2609.38802#S1.p1.1),[§1](https://arxiv.org/html/2609.38802#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1),[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px4.p1.1),[§4\.3](https://arxiv.org/html/2609.38802#S4.SS3.SSS0.Px5.p1.1)\. - Holtzmanet al\.\(2019\)A\. Holtzman, J\. Buys, L\. Du, M\. Forbes, and Y\. ChoiThe curious case of neural text degeneration\.Cited by:[§1](https://arxiv.org/html/2609.38802#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Honget al\.\(2020\)S\. Hong, Y\. Kaya, I\. Modoranu, and T\. DumitraşA panda? no, it’s a sloth: slowdown attacks on adaptive multi\-exit neural network inference\.arXiv preprint arXiv:2010\.02432\.Cited by:[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Kumaret al\.\(2025\)A\. Kumar, J\. Roh, A\. Naseh, M\. Karpinska, M\. Iyyer, A\. Houmansadr, and E\. BagdasarianOverthink: slowdown attacks on reasoning llms\.Cited by:[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Liet al\.\(2020a\)L\. Li, R\. Ma, Q\. Guo, X\. Xue, and X\. QiuBert\-attack: adversarial attack against bert using bert\.InProceedings of the 2020 conference on empirical methods in natural language processing \(EMNLP\),pp\. 6193–6202\.Cited by:[§2\.1](https://arxiv.org/html/2609.38802#S2.SS1.p1.1)\. - Liet al\.\(2020b\)M\. Li, S\. Roller, I\. Kulikov, S\. Welleck, Y\. Boureau, K\. Cho, and J\. WestonDon’t say that\! making inconsistent dialogue unlikely with unlikelihood training\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 4715–4728\.Cited by:[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Liet al\.\(2023\)X\. L\. Li, A\. Holtzman, D\. Fried, P\. Liang, J\. Eisner, T\. B\. Hashimoto, L\. Zettlemoyer, and M\. LewisContrastive decoding: open\-ended text generation as optimization\.InProceedings of the 61st annual meeting of the association for computational linguistics \(volume 1: Long papers\),pp\. 12286–12312\.Cited by:[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Liet al\.\(2026\)X\. Li, X\. Liu, C\. Liu, Y\. Xu, K\. Ding, B\. Xin, and J\. YinLoopllm: transferable energy\-latency attacks in llms via repetitive generation\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 31770–31777\.Cited by:[Appendix A](https://arxiv.org/html/2609.38802#A1.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2609.38802#S1.p1.1),[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px2.p1.1)\. - Liuet al\.\(2024a\)H\. Liu, C\. Li, Y\. Li, and Y\. J\. LeeImproved baselines with visual instruction tuning\.In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 26286–26296\.Cited by:[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px1.p1.1)\. - Liuet al\.\(2024b\)Z\. Liu, C\. Kong, Y\. Liu, and M\. SunFantastic semantics and where to find them: investigating which layers of generative llms reflect lexical semantics\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 14551–14558\.Cited by:[Table 7](https://arxiv.org/html/2609.38802#A2.T7.2.3.1.1.1),[Appendix D](https://arxiv.org/html/2609.38802#A4.p4.1)\. - Luet al\.\(2023\)D\. Lu, Z\. Wang, T\. Wang, W\. Guan, H\. Gao, and F\. ZhengSet\-level guidance attack: boosting adversarial transferability of vision\-language pre\-training models\.In2023 IEEE/CVF International Conference on Computer Vision \(ICCV\),pp\. 102–111\.Cited by:[§2\.1](https://arxiv.org/html/2609.38802#S2.SS1.p1.1)\. - Luet al\.\(2022\)P\. Lu, S\. Mishra, T\. Xia, L\. Qiu, K\. Chang, S\. Zhu, O\. Tafjord, P\. Clark, and A\. KalyanLearn to explain: multimodal reasoning via thought chains for science question answering\.Advances in neural information processing systems35,pp\. 2507–2521\.Cited by:[Appendix A](https://arxiv.org/html/2609.38802#A1.SS0.SSS0.Px6.p1.1),[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px2.p1.1)\. - Menget al\.\(2022\)K\. Meng, D\. Bau, A\. Andonian, and Y\. BelinkovLocating and editing factual associations in gpt\.Advances in neural information processing systems35,pp\. 17359–17372\.Cited by:[§4\.3](https://arxiv.org/html/2609.38802#S4.SS3.SSS0.Px3.p1.1)\. - Qiet al\.\(2024\)X\. Qi, K\. Huang, A\. Panda, P\. Henderson, M\. Wang, and P\. MittalVisual adversarial examples jailbreak aligned large language models\.InProceedings of the AAAI conference on artificial intelligence,Vol\.38,pp\. 21527–21536\.Cited by:[§2\.1](https://arxiv.org/html/2609.38802#S2.SS1.p1.1)\. - Qinet al\.\(2019\)Y\. Qin, N\. Carlini, G\. Cottrell, I\. Goodfellow, and C\. RaffelImperceptible, robust, and targeted adversarial examples for automatic speech recognition\.InInternational conference on machine learning,pp\. 5231–5240\.Cited by:[§2\.1](https://arxiv.org/html/2609.38802#S2.SS1.p1.1)\. - Shumailovet al\.\(2021\)I\. Shumailov, Y\. Zhao, D\. Bates, N\. Papernot, R\. Mullins, and R\. AndersonSponge examples: energy\-latency attacks on neural networks\.In2021 IEEE European symposium on security and privacy \(EuroS&P\),pp\. 212–231\.Cited by:[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Singhet al\.\(2019\)A\. Singh, V\. Natarajan, M\. Shah, Y\. Jiang, X\. Chen, D\. Batra, D\. Parikh, and M\. RohrbachTowards vqa models that can read\.In2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 8309–8318\.Cited by:[Appendix A](https://arxiv.org/html/2609.38802#A1.SS0.SSS0.Px6.p1.1),[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px2.p1.1)\. - Srivastavaet al\.\(2022\)A\. Srivastava, A\. Rastogi, A\. Rao, A\. A\. M\. Shoeb, A\. Abid, A\. Fisch, A\. R\. Brown, A\. Santoro, A\. Gupta, A\. Garriga\-Alonso,et al\.Beyond the imitation game: quantifying and extrapolating the capabilities of language models\.arXiv preprint arXiv:2206\.04615\.Cited by:[Appendix C](https://arxiv.org/html/2609.38802#A3.SS0.SSS0.Px1.p2.1)\. - Suet al\.\(2022\)Y\. Su, T\. Lan, Y\. Wang, D\. Yogatama, L\. Kong, and N\. CollierA contrastive framework for neural text generation\.Advances in neural information processing systems35,pp\. 21548–21561\.Cited by:[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Team \(2024\)Q\. TeamQwen2\.5: a party of foundation models\.External Links:[Link](https://qwenlm.github.io/blog/qwen2.5/)Cited by:[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px1.p1.1)\. - Team \(2025\)Q\. TeamQwen2\.5\-vl\.External Links:[Link](https://qwenlm.github.io/blog/qwen2.5-vl/)Cited by:[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px1.p1.1)\. - Team \(2026\)Q\. TeamQwen3\. 6\-27b: flagship\-level coding in a 27b dense model\.April\.Cited by:[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px1.p1.1)\. - Wellecket al\.\(2019\)S\. Welleck, I\. Kulikov, S\. Roller, E\. Dinan, K\. Cho, and J\. WestonNeural text generation with unlikelihood training\.Cited by:[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Xuet al\.\(2022\)J\. Xu, X\. Liu, J\. Yan, D\. Cai, H\. Li, and J\. LiLearning to break the loop: analyzing and mitigating repetitions for neural text generation\.Advances in Neural Information Processing Systems35,pp\. 3082–3095\.Cited by:[§1](https://arxiv.org/html/2609.38802#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Yaoet al\.\(2025\)J\. Yao, S\. Yang, J\. Xu, L\. Hu, M\. Li, and D\. WangUnderstanding the repeat curse in large language models from a feature perspective\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 7787–7815\.Cited by:[Table 7](https://arxiv.org/html/2609.38802#A2.T7.2.5.1.1.1),[§1](https://arxiv.org/html/2609.38802#S1.p1.1),[§1](https://arxiv.org/html/2609.38802#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1),[§4\.3](https://arxiv.org/html/2609.38802#S4.SS3.SSS0.Px5.p1.1)\. - Yinet al\.\(2023\)Z\. Yin, M\. Ye, T\. Zhang, T\. Du, J\. Zhu, H\. Liu, J\. Chen, T\. Wang, and F\. MaVlattack: multimodal adversarial attacks on vision\-language tasks via pre\-trained models\.Advances in Neural Information Processing Systems36,pp\. 52936–52956\.Cited by:[§2\.1](https://arxiv.org/html/2609.38802#S2.SS1.p1.1)\. - Yonaet al\.\(2025\)I\. Yona, I\. Shumailov, J\. Hayes, F\. Barbero, and Y\. GandelsmanInterpreting the repeated token phenomenon in large language models\.arXiv preprint arXiv:2503\.08908\.Cited by:[Appendix A](https://arxiv.org/html/2609.38802#A1.SS0.SSS0.Px5.p1.1)\. - Zenget al\.\(2025\)A\. Zeng, X\. Lv, Q\. Zheng, Z\. Hou, B\. Chen, C\. Xie, C\. Wang, D\. Yin, H\. Zeng, J\. Zhang,et al\.Glm\-4\.5: agentic, reasoning, and coding \(arc\) foundation models\.Cited by:[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px1.p1.1)\. - Zhanget al\.\(2022\)J\. Zhang, Q\. Yi, and J\. SangTowards adversarial attack on vision\-language pre\-training models\.InProceedings of the 30th ACM international conference on multimedia,pp\. 5005–5013\.Cited by:[§2\.1](https://arxiv.org/html/2609.38802#S2.SS1.p1.1)\. - Zhanget al\.\(2023\)S\. Zhang, X\. Pan, M\. Zhang, and M\. YangSlowBERT: slow\-down attacks on input\-adaptive multi\-exit bert\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 9992–10007\.Cited by:[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Zhanget al\.\(2026\)Y\. Zhang, X\. Wang, Z\. Chen, W\. Wang, Z\. Zhang, Z\. Gong, Z\. Zhou, K\. Wang, L\. Sun, Y\. Liu,et al\.Resource consumption threats in large language models\.arXiv preprint arXiv:2603\.16068\.Cited by:[§1](https://arxiv.org/html/2609.38802#S1.p2.1)\. - Zhanget al\.\(2025a\)Y\. Zhang, X\. Wang, H\. Gao, Z\. Zhou, F\. Meng, Y\. Zhang, and S\. SuPD3\{\}^\{3\}f: a pluggable and dynamic dos\-defense framework against resource consumption attacks targeting large language models\.\.InEMNLP \(Findings\),pp\. 3641–3671\.Cited by:[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px3.p1.1)\. - Zhanget al\.\(2025b\)Y\. Zhang, Z\. Zhou, W\. Zhang, X\. Wang, X\. Jia, Y\. Liu, and S\. SuCrabs: consuming resource via auto\-generation for llm\-dos attack under black\-box settings\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 11128–11150\.Cited by:[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1)\. - Zhanget al\.\(2025c\)Z\. Zhang, S\. Yadav, F\. Han, and E\. ShutovaCross\-modal information flow in multimodal large language models\.In2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 19781–19791\.Cited by:[Table 7](https://arxiv.org/html/2609.38802#A2.T7.2.6.1.1.1)\. - Zhuet al\.\(2023\)W\. Zhu, H\. Hao, and R\. WangPenalty decoding: well suppress the self\-reinforcement effect in open\-ended text generation\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 1218–1228\.Cited by:[§1](https://arxiv.org/html/2609.38802#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.38802#S2.SS2.p1.1),[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px3.p1.1)\. - Zouet al\.\(2023\)A\. Zou, Z\. Wang, N\. Carlini, M\. Nasr, J\. Z\. Kolter, and M\. FredriksonUniversal and transferable adversarial attacks on aligned language models\.Cited by:[Appendix A](https://arxiv.org/html/2609.38802#A1.SS0.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2609.38802#S4.SS1.SSS0.Px2.p1.1)\. ## Appendix AModel and Attack Coverage Table[6](https://arxiv.org/html/2609.38802#A1.T6)maps the attack pools and benign evaluation tasks to the models used in our experiments\. We construct attack pools separately for each checkpoint so that model\-specific tokenization, chat templates, and multimodal preprocessing are preserved\. A candidate is retained only when greedy decoding produces an uncontrolled one\- or two\-token cycle for at least ten consecutive repetitions\. A dash in the table therefore denotes a condition outside the evaluated scope, rather than a failed attack\. #### Training–test separation and default configuration\. The attack examples used to localize TRC are disjoint from all examples used for evaluation\. For each target model, we perform localization once using a designated training attack: RECITE for LVLMs and GCG for LLMs and LRMs\. We do not retrain or retune TRC separately for every evaluation attack or dataset; instead, the resulting model\-specific configuration is applied directly to the held\-out attack pools summarized in Table[6](https://arxiv.org/html/2609.38802#A1.T6)\. In the main experiments, the adaptive rule yields a suppression strength of approximatelyαl⋆=0\.73\\alpha\_\{l^\{\\star\}\}=0\.73\(73%\)\. We use a direction\-selection ratio ofρ=5%\\rho=5\\%and a layer\-window span ofh=3h=3\. Table 6:Model\-specific attack coverage and benign evaluation tasks\. Y marks an included attack condition; – denotes a condition outside the evaluated scope\.ModelRECITEGCGLoopLLMDirectBenign tasksLVLMsInstructBLIP\-Vicuna\-7BYYY–ScienceQA, TextVQAQwen2\.5\-VL\-3B\-InstructYYY–ScienceQA, TextVQALLaVA\-1\.5\-7BYY––ScienceQA, TextVQALLMsLlama\-3\.2\-3B–YYYMMLUQwen2\.5\-3B–YYYMMLULRMsDeepSeek\-Llama\-8B–YYYMMLUQwen3\.6\-27B–YYYMMLUGLM\-4\.7\-Flash–YYYMMLU #### RECITE\. For the multimodal models, we follow the original RECITE construction procedure\([Gao et al\., 2025](https://arxiv.org/html/2609.38802#bib.bib40)\)\. Each source item contains an image, its associated request, and a repeated\-token target\. A targeted projected\-gradient attack modifies the image while leaving the text request fixed\. We then run the corresponding LVLM on the adversarial image and retain only candidates whose generated continuation satisfies the repetition criterion above\. This produces paired clean and adversarial images for the same request and avoids changing the linguistic content of the input\. #### GCG\. We implement GCG\([Zou et al\., 2023](https://arxiv.org/html/2609.38802#bib.bib44)\)with[NanoGCG](https://github.com/GraySwanAI/nanoGCG)and adapt its optimization to the tokenizer, chat template, and cache representation of each model family\. Starting from the clean request pool associated with a target model, NanoGCG optimizes a discrete suffix toward a repeated continuation\. The resulting suffix is transferred into the corresponding target input format and screened again on the actual target checkpoint\. Thus, the GCG datasets contain model\-specific request–suffix pairs rather than one universal suffix reused across all models\. #### LoopLLM\. We construct the LoopLLM pools with the same model\-wise adaptation and transfer protocol used for GCG, but replace the GCG target loss with LoopLLM’s repetition\-inducing objective\([Li et al\., 2026](https://arxiv.org/html/2609.38802#bib.bib41)\)\. For each clean request, coordinate optimization searches for a suffix that concentrates probability mass on a short cyclic token pattern\. The suffix is appended to the request under the target model’s own chat template, and the generated continuation is retained only after target\-model screening\. This keeps the base request intact while making the optimized suffix specific to the model and source pool\. #### Direct\. Direct follows the repeated\-token attack setting studied by[Yona et al\. \(2025\)](https://arxiv.org/html/2609.38802#bib.bib56)\. We first serialize the original request with the model’s chat template, then concatenate a short repeated token sequence to the assistant\-side generation prefix\. Autoregressive decoding resumes from this prefilled partial output, so the model receives the repetition as its own unfinished continuation rather than as a new user instruction\. We retain examples only when the newly generated tokens continue into an uncontrolled repetition cycle\. #### Benign task sets\. For utility evaluation, LVLMs use ScienceQA and TextVQA, which test multimodal science reasoning and question answering over scene text, respectively\([Lu et al\., 2022](https://arxiv.org/html/2609.38802#bib.bib45);[Singh et al\., 2019](https://arxiv.org/html/2609.38802#bib.bib46)\)\. Text\-only LLMs and LRMs use MMLU across its subject domains\([Hendrycks et al\., 2020](https://arxiv.org/html/2609.38802#bib.bib47)\)\. These benign inputs preserve the original images, questions, and answer choices and receive no attack suffix or assistant\-side repetition prefix\. For each dataset, we use a subset of 200 samples for evaluation\. ## Appendix BDistinction from Existing Interpretability Methods Table[7](https://arxiv.org/html/2609.38802#A2.T7)compares TRC with representative interpretability approaches along the three axes most relevant to our design: the quantity being observed, the internal site being localized, and the model modalities on which the method has been demonstrated\. The comparison concerns methodological scope rather than a shared numerical benchmark; the listed approaches answer complementary questions and were not all designed to mitigate uncontrolled repetition\. Table 7:Comparison with representative interpretability approaches\. “Scope” records the model families demonstrated in the cited work, rather than a theoretical restriction\. TRC differs primarily in observing changes between cycle\-aligned generation positions, localizing the corresponding pre\-addition residual writes, and applying the same formulation across LVLMs, LLMs, and LRMs\.#### Observation dimension\. Most neuron, probe, and feature based approaches inspect activation magnitude, prediction attribution, recoverable content, or a learned feature basis at a position or stage\. TRC instead treats repetition as a temporal change pattern: it aligns positions separated by the detected repetition cycle and measures how the native attention and MLP contributions change across those positions\. The attack curve is then interpreted relative to benign generation and its local cross\-layer variation\. This observation axis is suited to distinguishing a continuation that becomes dynamically stationary from one that continues to advance semantically\. #### Localization position\. TRC localizes both a layer and residual coordinates at the module output before residual addition, with attention and MLP analyzed separately\. Our experiments consistently place the repetition\-associated low\-change regime in shallow layers, whereas normal generation reaches its minimum later\. This gives TRC an early and directly actionable intervention site, rather than requiring a probe, an SAE dictionary, or a search over late output activations\. The comparison does not imply that shallow layers are universally optimal for every behavior; it identifies the location supported for uncontrolled repetition under our evaluated settings\. #### Cross\-modal applicability\. TRC observes the Transformer backbone’s residual pathway after modality\-specific inputs have entered the autoregressive decoder\. Consequently, the same score, layer\-selection rule, and soft\-mask construction can be used for visually induced repetition in LVLMs and textually induced repetition in LLMs and LRMs\. The resulting advantage is not merely support for multimodal inputs: it is a shared internal analysis and intervention interface across visual and textual attack constructions, as demonstrated by the model coverage and held\-out evaluations in this paper\. ## Appendix CConstruction of Legitimate\-Repetition Controls #### Models\. For generation, we cap the number of newly generated tokens at 2,048 for InstructBLIP\-Vicuna\-7B \(its default\), 4,096 for the other LVLMs, 8,192 for all text\-only LLMs, and 16,384 for all LRMs\. The specificity analysis in Section[4\.3](https://arxiv.org/html/2609.38802#S4.SS3.SSS0.Px2)requires a control condition that contains intentional repetition without an uncontrolled generation failure\. We construct this condition using two design principles rather than copying benchmark instances\. MLissard controls sequential complexity through repeated applications of simple rules, while BIG\-bench emphasizes explicitly specified, auditable tasks that probe distinct model capabilities\([Bueno et al\., 2024](https://arxiv.org/html/2609.38802#bib.bib53);[Srivastava et al\., 2022](https://arxiv.org/html/2609.38802#bib.bib54)\)\. Following these principles, each of our image\-grounded items requires the model to first produce a normal semantic description and then execute a finite repetition instruction\. Each prompt contains two ordered requirements for Qwen2\.5\-VL\-3B\-Instruct\. The model must first write one concise English sentence grounded in the input image and must then repeat a designated common English token exactly ten times\. We select repeat units that map to one token both in isolation and with a leading space under the model tokenizer\. The repeated token is excluded from the semantic description so that the descriptive and repetitive portions can be validated separately\. An item is accepted only when the output begins with an image\-grounded description, contains an expected visual keyword before the repeated segment, omits the repeat unit from that semantic prefix, and ends with ten contiguous copies of the same token ID\. All five constructed items pass these criteria\. Table[8](https://arxiv.org/html/2609.38802#A3.T8)shows two representative examples\. They preserve a bounded, task\-compliant repetition after a meaningful visual response, in contrast to an uncontrolled loop that displaces the task answer or fails to terminate\. Table 8:Two representative legitimate\-repetition controls for Qwen2\.5\-VL\-3B\-Instruct\. Each prompt requires an image\-grounded description followed by exactly ten copies of a designated token\. Output text is reproduced verbatim; line breaks are normalized for table layout\. ## Appendix DLayer\-wise TRC\-l Profiles under Repetition This analysis expands the shallow\-localization and specificity results in Sections[4\.3](https://arxiv.org/html/2609.38802#S4.SS3.SSS0.Px1)and[4\.3](https://arxiv.org/html/2609.38802#S4.SS3.SSS0.Px2)\. Figure[5](https://arxiv.org/html/2609.38802#A4.F5)plots the TRC\-l score at every eligible candidate start layer for Direct, GCG, RECITE, and normal requests on Qwen2\.5\-VL\-3B\. Attention and MLP are evaluated separately at window spansh∈\{2,3,4\}h\\in\\\{2,3,4\\\}\. Lower scores indicate a smaller attack\-to\-normal contribution change ratio and/or less net change across the layer window, as defined in Section[3\.2](https://arxiv.org/html/2609.38802#S3.SS2)\. This does not affect layer ranking, and the score should not be interpreted as a direct measure of semantic similarity\. Figure 5:Layer\-wise TRC\-l profiles for Qwen2\.5\-VL\-3B Attention \(top\) and MLP \(bottom\), with window spansh=2,3,4h=2,3,4\. Curves show Direct, GCG, and RECITE failures and the mean of normal\-reference splits; gray shading denotes the normal range\.Across the plotted candidate layers, the failure curves generally lie below the normal curve\. This scale difference is compatible with the construction of TRC: when successive comparison positions revisit a similar repetitive content state, their aligned residual contributions can differ less than those in a normal continuation that advances the answer\. The smaller tokenwise difference lowers the magnitude term of TRC\-l\. Because TRC\-l also contains a cross\-layer variation term, and because comparison positions depend on the detected repetition unit, the plot alone does not establish semantic similarity as the unique cause of the score gap\. The more diagnostic observation is where each curve reaches its minimum within its own condition\. For normal requests, the minimum remains in the middle portion of the network: Attention selects layers 19, 11, and 10 forh=2,3,4h=2,3,4, respectively, while MLP selects layers 12, 11, and 14\. Ordinary generation can therefore exhibit a relatively stable local contribution profile even while the model continues to process a coherent task\. This pattern is compatible with intermediate layers integrating task\-relevant information before later output decisions; it does not establish a universal reasoning layer or imply that normal representations are semantically unchanged between tokens\. The repetitive failures show a different location for their most stable, low\-change state\. All three attacks select Attention layer 1 for every plotted span\. In MLP, Direct and RECITE select layer 1, whereas GCG selects layer 0\. The separation from normal requests persists for each testedhh\. Prior layer\-wise studies offer a more specific basis for interpreting this shallow location: lower feed\-forward layers tend to recognize relatively simple textual patterns, whereas later layers capture more semantic patterns\([Geva et al\., 2021](https://arxiv.org/html/2609.38802#bib.bib48)\); probing a generative LLM likewise finds lexical semantics in lower\-layer representations\([Liu et al\., 2024b](https://arxiv.org/html/2609.38802#bib.bib55)\)\. Feed\-forward updates can also promote human\-interpretable concepts in the model’s vocabulary space\([Geva et al\., 2022](https://arxiv.org/html/2609.38802#bib.bib49)\)\. Together, these findings make it plausible that basic lexical and local\-pattern features in shallow layers provide a substrate from which a repetition\-associated representation can emerge\. Our TRC\-l curves locate a low\-change regime consistent with this interpretation; they do not identify a dedicated repetition neuron or prove that any cited model implements the same failure mechanism\. Once established, such a feature can remain available to later blocks through the residual stream, consistent with the account of feed\-forward outputs being refined across layers via residual connections\([Geva et al\., 2021](https://arxiv.org/html/2609.38802#bib.bib48)\)\. This interpretation also aligns with Section[4\.3](https://arxiv.org/html/2609.38802#S4.SS3.SSS0.Px5): shallow attention\-update magnitudes are nonzero, yet their attack–normal gaps remain within normal token\-level fluctuation bounds\. ## Appendix EQualitative Cross\-Intervention Examples We provide one representative paired example for each cross\-intervention direction discussed in Section[4\.3](https://arxiv.org/html/2609.38802#S4.SS3.SSS0.Px6)\. These examples illustrate the aggregate patterns in Table[5](https://arxiv.org/html/2609.38802#S4.T5)\. In each pair, the normal and attack inputs share the same base request, and the attack input differs only by the optimized GCG suffix\. The ellipsis below marks an output truncated for presentation\. Example A: Restoring a normal activation on an attack trajectory\.Attack baselineHello, hello, hello, hello, hello, \.\.\.*Outcome:*2048 newly generated tokens, repetitive, stopped at the generation cap\.Localized coordinates, normal residual replacementHello, hello, hello, hello, hello\.*Outcome:*11 newly generated tokens and non\-repetitive generation\.Matched random coordinates, normal residual replacementHello, hello, hello, hello, hello, \.\.\.*Outcome:*2048 newly generated tokens, repetitive, stopped at the generation cap\.Observed contrast\.Restoring the localized coordinates recovers the short normal response, whereas the equal\-size random replacement leaves this attack trajectory in the repetitive state\. Example B: Injecting the attack difference into a normal trajectory\.Normal baselineHello, hello, hello, hello, hello\!*Outcome:*11 newly generated tokens and non\-repetitive generation\.Localized coordinates, attack\-minus\-normal addition\.\.*Outcome:*2 newly generated tokens and non\-repetitive but erroneous termination\.Matched random coordinates, attack\-minus\-normal additionHello, hello, hello, hello, hello\!*Outcome:*11 newly generated tokens and non\-repetitive generation, matching the normal baseline\.Observed contrast\.The localized perturbation severely disrupts the normal response but does not recreate repetition; the matched random perturbation leaves the response unchanged\. ## Appendix FAblation Studies ### F\.1Sensitivity to the Layer\-Window Span The TRC\-l score measures cross\-layer variation over a local spanhh\(Section[3\.2](https://arxiv.org/html/2609.38802#S3.SS2)\)\. We varyh∈\{2,3,4\}h\\in\\\{2,3,4\\\}while keeping the remaining localization procedure fixed\. Table[9](https://arxiv.org/html/2609.38802#A6.T9)reports the resulting Top\-1 start layer for the attention and MLP branches\. Table 9:Sensitivity of the localized Top\-1 start layer to the layer\-window spanhh\. Each entry is Attention / MLP, and layers are zero\-indexed\.The selected locations are invariant over the three evaluated spans\. In particular, the branch\-specific difference for GCG \(attention layer 1 versus MLP layer 0\) is preserved, whereas both branches select layer 1 for RECITE and LoopLLM\. ### F\.2Sensitivity to the Suppression Ratio on Attacks This ablation varies the suppression ratioρ∈\{1,5,10,15,20\}%\\rho\\in\\\{1,5,10,15,20\\\}\\%while keeping the attack training data and other intervention settings fixed\. Table[10](https://arxiv.org/html/2609.38802#A6.T10)reports attack behavior; each attack cell gives mean newly generated tokens followed by repetition rate\. The final column reports the overlap between the coordinates selected at each ratio and the coordinates selected atρ=20%\\rho=20\\%\. Table 10:Sensitivity to the suppression ratio on attacks\. Cells under each attack and the macro average report mean newly generated tokens / repetition rate \(%\)\. Coordinate overlap is measured against the mask selected atρ=5%\\rho=5\\%\. Lower is better for both reported attack metrics\.Every evaluated suppression ratio lowers the macro\-average generation length and repetition rate relative to no defense, but the improvement is not monotonic\. Atρ=20%\\rho=20\\%, the macro\-average output is shortest \(620\.56 tokens\), whileρ=1%\\rho=1\\%andρ=20%\\rho=20\\%tie for the lowest macro\-average repetition rate \(21\.3%\)\. The attack\-level results expose important heterogeneity: RECITE is strongly mitigated at every evaluated ratio, GCG benefits most atρ=20%\\rho=20\\%\. ### F\.3Effect of the Suppression Ratio on Benign Utility Table[11](https://arxiv.org/html/2609.38802#A6.T11)is a separate benign\-utility study\. Here, the learned direction ranking, the training\-set size, and all other training settings are fixed, and we vary only the suppression ratioρ\\rho, i\.e\., the fraction of hidden coordinates included in the soft mask\. Table 11:Effect of the suppression ratioρ\\rhoon benign task\. Accuracy is reported in percent\.Across the evaluated range, task accuracy decreases gradually as the suppression ratio increases, with a maximum drop of only 1\.5 percentage points relative to the unsuppressed model\. Meanwhile, the average response length varies by no more than 0\.1 tokens\. ### F\.4Sensitivity to the Maximum Repetition Distance Finally, we vary the maximum token distanceqmaxq\_\{\\max\}considered by the repeated\-nn\-gram detector\. This control determines how far apart matching positions may be when constructing the tokenwise comparisons used by TRC\. Table 12:Top\-1 start layer under different maximum repetition distancesqmaxq\_\{\\max\}\. Layers are zero\-indexed\.RECITE remains localized at layer 1 throughout, while GCG shifts from layer 12 atqmax=1q\_\{\\max\}=1to layer 1 for everyqmax≥2q\_\{\\max\}\\geq 2\. A one\-token limit can be too restrictive because a repeated semantic unit may be segmented into several subword tokens; comparisons restricted to adjacent single\-token matches can therefore miss the actual cycle alignment\. Allowing even a short multi\-token distance is sufficient to recover the stable shallow location in this experiment\. This result motivates using a small but non\-unit search range rather than assuming that one semantic repetition always corresponds to one token\. ## Appendix GOnline Efficiency TRC does not append network layers, auxiliary detectors, or additional model forward passes at deployment\. After offline localization, the fixed intervention is applied directly to the residual stream at the selected target layers by suppressing the identified residual coordinates during the original residual update\. All other layers and residual coordinates remain unchanged\. TRC therefore operates within the model’s existing residual pathway, without introducing a separate inference module, additional forward computation, or changes to the depth of the computational graph\. Table[13](https://arxiv.org/html/2609.38802#A7.T13)reports attack\-generation time for two LLMs\. TRC shortens generation in every listed model–attack condition\. The reductions range from 90\.9% to 99\.6%, because suppressing uncontrolled repetition allows generation to terminate much earlier instead of continuing the attack\-induced loop\. Table 13:Attack\-generation time on two LLMs\. Reduction is computed as\(Torig−TTRC\)/Torig\(T\_\{\\mathrm\{orig\}\}\-T\_\{\\mathrm\{TRC\}\}\)/T\_\{\\mathrm\{orig\}\}\. Lower is better\.Table[14](https://arxiv.org/html/2609.38802#A7.T14)separately summarizes mean throughput\. Across the two LLMs, throughput changes only slightly after intervention, decreasing by 3\.00 tokens/s for Qwen2\.5\-3B and increasing by 1\.05 tokens/s for Llama\-3\.2\-3B\. These small variations indicate that the intervention has negligible impact on generation throughput\. Table 14:Mean throughput on two LLMs\. Difference is TRC minus Original\. Throughput is measured in tokens per second\.
Similar Articles
PARTREP: Learning What to Repeat for Decoder-only LLMs
PartRep proposes a selective prompt repetition method for decoder-only LLMs that appends only the most informative tokens (selected via NLL) instead of the full prompt, reducing KV cache and prefill FLOPs while retaining most of the accuracy gains across multiple benchmarks.
Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition Probability
The paper introduces Telescope Perplexity, a metric that measures token repetition probability to detect LLM-generated text in a zero-shot manner, achieving state-of-the-art or competitive performance across diverse datasets.
@Raman_bansal_: If you’ve trained an LLM, you’ve may have seen doom loop, in which a LLM endlessly repeats the same token or sentence, …
A Substack article explains the 'doom loop' problem in LLMs where models repeat tokens endlessly, and introduces Final Token Preference Optimization (FTPO) from Liquid AI as a method to detect and fix such loops during fine-tuning.
What Makes Recurrence Effective in Looped Language Models?
This paper systematically studies when recurrence helps in looped language models (LoopLMs), finding that extra recurrence can improve reasoning beyond the training horizon but degrade knowledge retention, and proposes channel-wise history-state injection with timestep conditioning as a more robust design for variable inference budgets.
Why Limit the Residual Stream to Layers and Not Tokens? Persistent Memory for Continuous Latent Reasoning
This paper identifies a 'concept bottleneck' in the CoCoNuT latent reasoning paradigm where hidden states are overwritten across passes, and proposes AGCLR, which adds a gated persistent memory stream to retain intermediate facts. Evaluations on GSM8K, HotpotQA, and ProsQA using GPT-2 show consistent improvements, especially on multi-hop tasks.