Multimodal Resource-Exhaustion Attacks on Vision-Language Models via Joint Pixel-Prompt Optimization

arXiv cs.AI Papers

Summary

The paper introduces Joint Pixel-Prompt Optimization (JPPO), a novel adversarial framework that jointly optimizes pixel perturbations and visible prompts to exhaust resources in autoregressive vision-language models, achieving significant latency and energy amplification compared to existing methods.

arXiv:2609.05889v1 Announce Type: new Abstract: Resource-exhaustion attacks against autoregressive vision-language models (VLMs) typically assume unimodal threat models, treating the image branch as the primary optimization surface while holding user-visible prompts fixed. Even recent loop-centric variants remain confined to this single-channel paradigm, leaving the exploitation of availability unexplored as a cross-modal optimization problem over jointly controllable input surfaces. We introduce Joint Pixel-Prompt Optimization (JPPO), the first compound adversarial framework elevating the visible prompt to a first-class adversarial variable alongside image perturbations. Under a restricted joint-input threat model, JPPO performs coupled, stagewise optimization over both the pixel and prompt surfaces. This produces synergistic cost amplification, mechanistically distinct from loop-dependent failures, exhibiting negligible loop incidence in our experiments. Evaluating five open-source VLM families on MS COCO and ImageNet under an 8/255 infinity-norm budget, JPPO achieves over 4.6x latency and 5.3x energy amplification on Qwen2.5-VL-7B, and over 36.6x latency with 32.7x energy amplification on BLIP-2. This represents the strongest cost amplification among directly compared baselines while requiring substantially fewer optimization iterations. Ablations confirm this amplification arises from multimodal coordination rather than prompt length or isolated modalities. These findings reveal structural blind spots in current VLM serving defenses, motivating cost-aware robustness evaluation as a first-class security requirement for multimodal deployments.
Original Article
View Cached Full Text

Cached at: 09/10/26, 08:44 AM

# Multimodal Resource-Exhaustion Attacks on Vision-Language Models via Joint Pixel-Prompt Optimization
Source: [https://arxiv.org/html/2609.05889](https://arxiv.org/html/2609.05889)
Zhaoxiong NiAffiliation:Guangzhou University,School of Computer Science and Cyber Engineering,Guangzhou,Chinaemail:[zhaoxiong\.ni@e\.gzhu\.edu\.cn](mailto:[email protected])Yatie XiaoNote:Corresponding author\.Affiliation:Guangzhou University,School of Computer Science and Cyber Engineering,Guangzhou,Chinaemail:[ytxiao21@gzhu\.edu\.cn](mailto:[email protected]),Chi\-Man PunAffiliation:University of Macau,Department of Computer and Information Science,Macau,Chinaemail:[cmpun@umac\.mo](mailto:[email protected]),Fei PengAffiliation:Guangzhou University,School of Artificial Intelligence,Guangzhou,Chinaemail:[eepengf@gmail\.com](mailto:[email protected]),Qingxiao GuanAffiliation:Guangzhou University,School of Artificial Intelligence,Guangzhou,Chinaemail:[258817567@qq\.com](mailto:[email protected])andKeke TangAffiliation:Guangzhou University,Cyberspace Institute of Advanced Technology,Guangzhou,Chinaemail:[tangbohutbh@gmail\.com](mailto:[email protected])

###### Abstract\.

Resource\-exhaustion attacks against autoregressive vision\-language models \(VLMs\) typically assume unimodal threat models, treating the image branch as the primary optimization surface while holding user\-visible prompts fixed\. Even recent loop\-centric variants remain confined to this single\-channel paradigm, leaving the exploitation of availability unexplored as a cross\-modal optimization problem over jointly controllable input surfaces\.

We introduce Joint Pixel\-Prompt Optimization \(JPPO\), the first compound adversarial framework elevating the visible prompt to a first\-class adversarial variable alongside image perturbations\. Under a restricted joint\-input threat model, JPPO performs coupled, stagewise optimization over both the pixel and prompt surfaces\. This produces synergistic cost amplification, mechanistically distinct from loop\-dependent failures, exhibiting negligible loop incidence in our experiments\.

Evaluating five open\-source VLM families on MS COCO and ImageNet under an 8/255 infinity\-norm budget, JPPO achieves over 4\.6× latency and 5\.3× energy amplification on Qwen2\.5\-VL\-7B, and over 36\.6× latency with 32\.7× energy amplification on BLIP\-2\. This represents the strongest cost amplification among directly compared baselines while requiring substantially fewer optimization iterations\. Ablations confirm this amplification arises from multimodal coordination rather than prompt length or isolated modalities\. These findings reveal structural blind spots in current VLM serving defenses, motivating cost\-aware robustness evaluation as a first\-class security requirement for multimodal deployments\.

###### Keywords:

autoregressive vision\-language models, resource exhaustion attacks, joint pixel\-prompt optimization, multimodal availability, decoding dynamics

## 1\.Introduction

![Benign versus JPPO inference comparison](https://arxiv.org/html/2609.05889v1/1.png)Figure 1\.Comparison between a benign inference \(top\) and the proposed JPPO attack \(bottom\), along with the resulting cost curve \(right\)\.While the clean input yields a concise response, our joint optimization of bounded pixel perturbations and a budget\-constrained interface\-visible prompt context drives the model toward prolonged decoding, resulting in substantially higher cumulative latency and estimated energy under the same serving configuration\.Benign versus JPPO inference comparisonTwo example inference cases are shown side by side\. The benign input produces a short response, while the JPPO\-adversarial input produces a much longer continuation\. A cost plot on the right indicates substantially larger cumulative latency and estimated energy for the attacked case under the same serving configuration\.Vision\-language models \(VLMs\) increasingly serve as autoregressive multimodal assistants, agents, and content\-understanding systems\. Unlike discriminative vision models, their per\-request cost depends on the realized decoding trajectory: longer generations require more decoding steps, larger effective context, and higher latency and energy overhead\. This makes availability a natural security concern: inputs that appear benign at the interface may still induce disproportionately expensive decoding trajectories, as illustrated by the benign\-versus\-JPPO comparison in Figure[1](https://arxiv.org/html/2609.05889#acmlabel1)\.

Prior VLM security work has mainly studied semantic failures, including jailbreaks and adversarial image attacks\([Qi et al\., 2024](https://arxiv.org/html/2609.05889#bib.bib22);[Bailey et al\., 2024](https://arxiv.org/html/2609.05889#bib.bib24)\)\. Recent availability\-oriented studies show that image\-only perturbations can delay termination and increase serving costs\([Gao et al\., 2024a](https://arxiv.org/html/2609.05889#bib.bib18)\)\. More recently, LingoLoop demonstrates that strong VLM resource exhaustion can be achieved through repetition\-centric and loop\-inducing mechanisms\([Fu et al\., 2026](https://arxiv.org/html/2609.05889#bib.bib41)\)\. These results establish image\-only availability attacks as a realistic threat\. Still, they also leave open a distinct question: can substantial cost amplification arise without relying primarily on explicit repetitive loops, when a bounded pixel perturbation and the user\-visible prompt are jointly optimized as attack\-controlled inputs? We study this question under a restricted joint\-input threat model\.

We presentJoint Pixel\-Prompt Optimization \(JPPO\), a restricted joint\-input resource\-exhaustion attack that jointly optimizes bounded pixel perturbations and a budget\-constrained visible prompt\. JPPO is not an unrestricted multimodal jailbreak or prompt\-injection method; it targets availability under interface\-compatible constraints\. Its core mechanism is multimodal coordination: pixel perturbations reshape the initial multimodal state, while optimized visible prompts exert continuation pressure during decoding\. This threat model is distinct from fixed\-prompt image\-only attacks and from loop\-centric failures that primarily rely on repetitive decoding\.

Our evaluation on MS COCO and ImageNet shows a substantial gap between the attacker’s bounded input budget and the resulting serving cost\. Under anℓ∞=8/255\\ell\_\{\\infty\}=8/255budget, JPPO achieves over4\.6×4\.6\\timeslatency and5\.3×5\.3\\timesestimated\-energy amplification on Qwen2\.5\-VL, and over36\.6×36\.6\\timeslatency and32\.7×32\.7\\timesestimated\-energy amplification on BLIP\-2 \(Table[1](https://arxiv.org/html/2609.05889#S6.T1)\)\. These gains indicate a high\-leverage per\-request availability risk rather than a marginal slowdown\.

The success of JPPO suggests a broader weakness in the availability of multimodal serving among autoregressive VLMs\. Existing image\-only attacks reveal that costly decoding can be induced through visual manipulation alone, including recent loop\-centric failures\. Our results further show that when the visible prompt is also optimized as part of the attack, the attack surface expands: multimodal coordination can sustain a distinct high\-cost decoding regime that cannot be fully explained by loop\-centric mechanisms alone\. We therefore position JPPO not as a replacement for recent image\-only attacks, but as a complementary threat model and attack mechanism for cost\-aware security evaluation in multimodal systems\.

Our main contributions are as follows:

- •We identify a restricted joint\-input availability threat model for autoregressive VLMs, distinguishing offline construction from deployment\-time submission\.
- •We present JPPO, a two\-stage multimodal attack that combines pixel\-space optimization with budgeted, interface\-compatible prompt construction to steer VLMs into high\-cost decoding regimes\.
- •Through matched pixel\-only, prompt\-only, and joint ablations, together with neutral\-prompt controls, we show that JPPO’s amplification is explained by multimodal coordination rather than by prompt length alone or either branch in isolation\.
- •Across five VLM families, we evaluate effectiveness, transferability, costs, output quality, defenses, shared\-worker impact, and local API constraints\.

## 2\.Related Work

### 2\.1\.Autoregressive VLMs and Inference Cost

Modern vision\-language models \(VLMs\) increasingly rely on LLM\-centric architectures for autoregressive text generation\. Representative examples include Flamingo\([Alayrac et al\., 2022](https://arxiv.org/html/2609.05889#bib.bib2)\), BLIP\-2\([Li et al\., 2023](https://arxiv.org/html/2609.05889#bib.bib3)\), LLaVA\([Liu et al\., 2023](https://arxiv.org/html/2609.05889#bib.bib4)\), MiniGPT\-4\([Zhu et al\., 2024](https://arxiv.org/html/2609.05889#bib.bib6)\), InstructBLIP\([Dai et al\., 2023](https://arxiv.org/html/2609.05889#bib.bib5)\), and Qwen2\.5\-VL\([Bai et al\., 2025](https://arxiv.org/html/2609.05889#bib.bib10)\)\. More recent open multimodal systems further broaden this design space\([Li et al\., 2025](https://arxiv.org/html/2609.05889#bib.bib11);[Lu et al\., 2024](https://arxiv.org/html/2609.05889#bib.bib8);[Wu et al\., 2025](https://arxiv.org/html/2609.05889#bib.bib9);[Dai et al\., 2024](https://arxiv.org/html/2609.05889#bib.bib12)\)\. System\-level optimizations such as PagedAttention\([Kwon et al\., 2023](https://arxiv.org/html/2609.05889#bib.bib13)\), VL\-Cache\([Tu et al\., 2025](https://arxiv.org/html/2609.05889#bib.bib14)\), and token/KV\-cache compression\([Tan et al\., 2025](https://arxiv.org/html/2609.05889#bib.bib15)\)aim to reduce the overhead of long\-context inference\. Despite these optimizations, serving costs remain tightly coupled to input\-triggered decoding length\. This creates a natural availability attack surface in autoregressive VLMs: adversarial inputs that prolong decoding or delay termination can directly exhaust latency, memory, and energy resources\.

### 2\.2\.Multimodal Adversarial and Jailbreak Attacks

A large body of prior work studies the adversarial robustness of multimodal models and how they can be manipulated through visual and textual channels\([Zhao et al\., 2023](https://arxiv.org/html/2609.05889#bib.bib23)\)\. Existing attacks show that adversarial images can directly elicit unsafe outputs from aligned multimodal models\([Qi et al\., 2024](https://arxiv.org/html/2609.05889#bib.bib22);[Bailey et al\., 2024](https://arxiv.org/html/2609.05889#bib.bib24)\)\. Typographic visual prompts can bypass text\-centric safeguards through the visual stream\([Gong et al\., 2025](https://arxiv.org/html/2609.05889#bib.bib25)\)\. More recent coordinated image\-text jailbreaks further show that multimodal alignment can be undermined through cross\-modal composition and concealment\([Shayegani et al\., 2024](https://arxiv.org/html/2609.05889#bib.bib26);[Wang et al\., 2025a](https://arxiv.org/html/2609.05889#bib.bib27);[Wang et al\., 2025d](https://arxiv.org/html/2609.05889#bib.bib28);[Jiang et al\., 2025](https://arxiv.org/html/2609.05889#bib.bib29)\)\. Existing research primarily targets semantic integrity, forcing models into policy violations or typographic bypasses\. While effective at altering what a model says, these attacks overlook the temporal persistence of the inference process\. JPPO shifts the focus from semantic corruption to resource amplification, leveraging coordinated cross\-modal manipulation to increase the computational leverage of a single malicious request\.

### 2\.3\.Availability and Resource\-Exhaustion Attacks

Availability\-oriented attacks on machine learning systems have received growing attention in recent years\. Sponge Examples show that carefully crafted test\-time inputs can increase latency and energy usage\([Shumailov et al\., 2021](https://arxiv.org/html/2609.05889#bib.bib16)\), while Sponge Poisoning extends this efficiency\-security perspective to training\-time manipulation\([Cinà et al\., 2025](https://arxiv.org/html/2609.05889#bib.bib17)\)\. More recent studies broaden the discussion to generative systems, including efficiency degradation in neural machine translation\([Chen et al\., 2022a](https://arxiv.org/html/2609.05889#bib.bib43)\), denial\-of\-service poisoning against large language models\([Gao et al\., 2024b](https://arxiv.org/html/2609.05889#bib.bib42)\), repetitive\-generation attacks on LLM decoding\([Li et al\., 2026](https://arxiv.org/html/2609.05889#bib.bib19)\), and context\-poisoning attacks on RAG\-based code generation\([Wang et al\., 2025c](https://arxiv.org/html/2609.05889#bib.bib20)\)\.

Within the VLM setting, NICGSlowDown first studies the robustness of efficiency for image caption generation\([Chen et al\., 2022b](https://arxiv.org/html/2609.05889#bib.bib35)\)\. Verbose Images and VLMInferSlow extend this line to modern large VLMs and service\-oriented settings\([Gao et al\., 2024a](https://arxiv.org/html/2609.05889#bib.bib18);[Wang et al\., 2025b](https://arxiv.org/html/2609.05889#bib.bib32)\), while Hidden Tail further studies stealthy resource consumption under constrained pixel perturbations by inducing continuations that can contain user\-invisible special tokens\([Zhang et al\., 2026](https://arxiv.org/html/2609.05889#bib.bib33)\)\. These works primarily assume an image\-only threat surface and show that pixel\-level perturbations alone can already amplify decoding cost\.

Recent work further shows that image\-only attacks can be significantly strengthened through explicit loop\-centric mechanisms\. Here, we use*loop\-centric*to denote attacks whose cost amplification relies primarily on explicit cyclic or highly repetitive token generation\. In particular, LingoLoop demonstrates that strong VLM/MLLM resource exhaustion can arise by manipulating token prediction and inducing loops, pushing models toward excessively verbose trajectories\([Fu et al\., 2026](https://arxiv.org/html/2609.05889#bib.bib41)\)\. This result indicates that powerful image\-only availability attacks need not remain weak once decoding becomes increasingly language\-dominated\.

Unlike recent image\-only and loop\-centric attacks, JPPO treats the interface\-visible prompt context as part of the attack optimization process, enabling coordinated manipulation of pixel initialization and prompt\-driven continuation pressure\.

## 3\.Problem Formulation and Threat Model

### 3\.1\.Autoregressive VLM Inference and Resource Cost

We consider an autoregressive vision\-language model \(VLM\)fθf\_\{\\theta\}that takes an imagex∈𝒳x\\in\\mathcal\{X\}and a textual promptp∈𝒫p\\in\\mathcal\{P\}as input, and produces an output sequence

\(1\)y1:G=fθ\(x,p\),y\_\{1:G\}=f\_\{\\theta\}\(x,p\),where generation terminates either when the model\-specific end\-of\-sequence \(EOS\) token is emitted or when a system\-imposed maximum decoding token budgetTmaxT\_\{\\max\}is reached\.

Inference in such systems consists of two stages: \(i\) multimodal prefilling, which encodes the image and prompt into a joint context, and \(ii\) autoregressive decoding, in which the output sequence is generated conditioned on\(x,p,y<t\)\(x,p,y\_\{<t\}\)\. Unlike fixed\-cost discriminative models, the resource consumption of autoregressive VLMs depends on the realized decoding trajectory\. In particular, longer generations increase the number of decoding steps, enlarge the effective textual context and KV\-cache footprint, and typically lead to higher inference\-time overhead\.

To capture this availability\-relevant behavior, we associate each inference request with a cost vector

\(2\)𝐜⁡\(x,p\)=\(W⁡\(x,p\),τ⁡\(x,p\),E^​\(x,p\)\),\\mathbf\{c\}\(x,p\)=\\big\(W\(x,p\),\\tau\(x,p\),\\widehat\{E\}\(x,p\)\\big\),whereWWdenotes the reported output length measured in words after detokenization,τ\\taudenotes the software\-observed inference latency under a fixed serving configuration, andE^\\widehat\{E\}denotes an NVML\-based estimated\-energy proxy under our evaluation protocol\. In our evaluation, the decoding cap is enforced in tokens, whereasWWis reported in words for readability\. The dominant serving overhead of autoregressive VLMs is driven by token\-level decoding and KV\-cache growth; accordingly,τ\\tauandE^\\widehat\{E\}are treated as the primary realized\-cost metrics, whileWWserves as an output\-length indicator empirically associated with prolonged decoding\.

### 3\.2\.Attack Objective

Our goal is not targeted semantic failure but increased cost per single VLM request while preserving input plausibility\. Let\(x,p0\)\(x,p\_\{0\}\)be a benign image and prompt pair and\(x⋆,p⋆\)\(x^\{\\star\},p^\{\\star\}\)the selected adversarial state\. We evaluate attack strength using generation length, latency, and estimated energy amplification\.

The attack satisfies a shared pixel budget,

\(3\)‖δ‖∞≤ϵ,\\\|\\delta\\\|\_\{\\infty\}\\leq\\epsilon,withϵ=8/255\\epsilon=8/255by default\. Pixel updates are projected under this budget, andx\+δx\+\\deltadenotes the clipped image evaluated by the model\.

Under this formulation, the attacker aims to solve

\(4\)max‖δ‖∞≤ϵ,p∈𝒯⁡\(p0\)⁡𝒥eval​\(x\+δ,p\),𝒥eval=E^\\max\_\{\\begin\{subarray\}\{c\}\\\|\\delta\\\|\_\{\\infty\}\\leq\\epsilon,\\\\ p\\in\\mathcal\{T\}\(p\_\{0\}\)\\end\{subarray\}\}\\mathcal\{J\}\_\{\\mathrm\{eval\}\}\(x\+\\delta,p\),\\qquad\\mathcal\{J\}\_\{\\mathrm\{eval\}\}=\\widehat\{E\}where𝒯⁡\(p0\)\\mathcal\{T\}\(p\_\{0\}\)denotes the restricted set of prompts reachable from the fixed templatep0p\_\{0\}under the budget\-constrained construction procedure described in Section[4](https://arxiv.org/html/2609.05889#S4)\. We use estimated energy as the primary evaluation objective because it jointly reflects decoding duration and GPU power draw under a fixed serving stack\.

This objective defines the evaluation goal rather than a directly differentiable optimization loss\. Since latency and estimated energy are noisy, non\-differentiable, and serving\-stack dependent, JPPO uses the surrogate optimization objectives described in Section[4\.4](https://arxiv.org/html/2609.05889#S4.SS4)during attack construction\.

### 3\.3\.Threat Model

#### Adversarial goal\.

We consider an input\-space availability attacker against an autoregressive VLM service\. The attacker’s goal is to increase the cost of accepted inference requests, measured by output length, latency, and estimated energy, and thereby create localized serving\-cost and availability pressure under repeated submissions when malicious and benign requests share workers or queues\. We focus on per\-request cost amplification and its resulting effect on shared serving resources\.

#### Attacker\-controlled and protected components\.

The attacker controls a user\-provided image and an interface\-visible prompt\.

The attacker does not modify the target model parameters, tokenizer, hidden system instructions, decoding implementation, server\-side scheduling policy, admission controller, or serving infrastructure, and requires no privileged access to other users’ requests or data\.

#### Offline construction and deployment\-time submission\.

During offline construction, the attacker uses either the exact open\-source model or a locally available surrogate\. Gradients are required only in this phase to update the bounded pixel perturbation; the prompt branch uses budget\-constrained text\-space operators\. Exact\-model construction provides the white\-box reference, while cross\-model and external evaluations assess black\-box transfer\.

At deployment time, the attacker submits a previously constructed image–prompt pair through the ordinary user interface as a black\-box request, without target\-side gradients, parameters, hidden prompts, or serving signals\. When construction and target models differ, deployment relies on cross\-model transferability\.

#### Cost allocation and repeated use\.

JPPO is most relevant to flat\-rate, free\-quota, fixed\-price, subscription, internal, or multi\-tenant services in which an expensive request’s marginal cost is not fully charged to its submitter\. In fully usage\-metered services, it may primarily increase the attacker’s own expense\.

Offline construction is attacker\-borne and separate from target\-side serving cost; it can be amortized only through reuse or repeated submission of related variants\.

#### Relation to image\-only threat models\.

JPPO studies a broader joint\-input setting than fixed\-prompt image\-only attacks\. Many VLM interfaces expose both image upload and visible textual instructions, allowing the availability risk of jointly controllable input modalities to be evaluated\.

### 3\.4\.Scope, Assumptions, and Deployment Considerations

JPPO applies to autoregressive VLM services exposing both image and visible\-prompt inputs and whose serving cost depends on the decoding trajectory\. It does not directly cover single\-modality interfaces, largely fixed\-cost non\-autoregressive systems, or services preventing attacker\-controlled prompt context\.

Practical feasibility depends on cross\-model transfer, resource sharing, and platform controls\. Section[6\.3](https://arxiv.org/html/2609.05889#S6.SS3)evaluates cache\-hit sensitivity, while Section[6\.8](https://arxiv.org/html/2609.05889#S6.SS8)evaluates a local HTTP deployment with request\-rate limiting, safety admission, and runtime early stopping\. Other controls such as output caps, timeouts, and production scheduling remain deployment\-dependent\. For black\-box services, attack practicality therefore depends on surrogate transferability and the extent to which platform\-level controls admit or mitigate high\-cost requests\.

## 4\.Methodology

### 4\.1\.Overview of JPPO

![JPPO pipeline overview](https://arxiv.org/html/2609.05889v1/2.png)Figure 2\.Overview of JPPO\.JPPO is a two\-stage multimodal availability attack\. In Stage I, the attacker maintains three prompt paths, namely visual, spatial, and semantic, and iteratively merges them into a diverse warm\-up prompt while jointly updating a bounded pixel perturbation\. The strongest measured Stage\-I state, consisting of its paired warm\-up prompt and perturbation, initializes Stage II, in which JPPO performs generation\-driven prompt refinement through selective accumulation, together with further perturbation updates\. Across both stages, the attack is guided by three availability\-oriented objectives, and the final adversarial state is selected based on the realized resource amplification\.JPPO pipeline overviewA flow diagram illustrates the two\-stage JPPO pipeline\. Stage I builds a warm\-up prompt from visual, spatial, and semantic path memories while jointly updating a bounded pixel perturbation\. Stage II starts from the warm\-up prompt and the strongest Stage I perturbation, then further refines the running prompt and pixel perturbation using model\-generated feedback\. The final adversarial state is selected based on the realized resource amplification\.JPPO is a two\-stage attack framework designed to enhance the per\-request inference cost of autoregressive vision\-language models \(VLMs\) by coordinating bounded pixel perturbations with a budget\-constrained visible prompt context\. Given a benign image and prompt pair\(x,p0\)\(x,p\_\{0\}\), JPPO first constructs Stage\-I candidate states and then refines them further in Stage II\. The final selected adversarial state is denoted by\(x⋆,p⋆\)\(x^\{\\star\},p^\{\\star\}\), wherex⋆=clip⁡\(x\+δ⋆\)x^\{\\star\}=\\mathrm\{clip\}\(x\+\\delta^\{\\star\}\)and‖δ⋆‖∞≤ϵ\\\|\\delta^\{\\star\}\\\|\_\{\\infty\}\\leq\\epsilon\. In figures, we also write\(xadv,padv\)≡\(x⋆,p⋆\)\(x\_\{\\mathrm\{adv\}\},p\_\{\\mathrm\{adv\}\}\)\\equiv\(x^\{\\star\},p^\{\\star\}\)to emphasize the final adversarial image and prompt\. This state is selected from the strongest stage\-wise candidates retained across Stage I and Stage II, inducing prolonged decoding and higher realized inference cost while remaining within bounded multimodal input constraints\. As discussed in Section[3](https://arxiv.org/html/2609.05889#S3), our objective is availability degradation rather than semantic task failure\.

Figure[2](https://arxiv.org/html/2609.05889#acmlabel2)provides a schematic view of the JPPO pipeline\. A key distinction between the two stages lies in the structure of prompt evolution\. Stage I maintains three complementary path memories, namely visual, spatial, and semantic, and uses aspect\-aware merging and compression to construct a diverse warm\-up prompt together with an initial bounded perturbation\. Stage II abandons explicit multi\-path separation and instead performs single\-context refinement through an iterative selective\-accumulation procedure initialized from the strongest measured Stage\-I state, preserving the pairing between the warm\-up prompt and the perturbation that produced the largest realized estimated energy\. This design reflects the different roles of the two stages: broad contextual diversification in Stage I and aggressive resource\-seeking refinement in Stage II\.

A key design choice is that JPPO does not directly optimize discrete prompt tokens via gradient descent\. Instead, the prompt branch evolves through attacker\-directed, budget\-constrained text\-space operators, including aspect\-aware path merging, selective phrase accumulation based on lexical novelty, compression, and light de\-duplication\. In contrast, the pixel branch is optimized using projected gradient updates\([Madry et al\., 2018](https://arxiv.org/html/2609.05889#bib.bib21)\)\. This design aligns with realistic user\-facing interfaces, in which an attacker can control visible prompt content but not hidden system prompts or the tokenizer’s internals\.

Across both stages, three availability\-oriented objectives guide JPPO, and the final adversarial example is selected from the strongest state observed during optimization\. Concretely, within each stage, JPPO maintains a best\-state buffer of measured image and prompt pairs, where each entry stores the perturbation, the visible prompt, and the realized metrics\(W,τ,E^\)\(W,\\tau,\\widehat\{E\}\)\. The buffer is ranked by the estimated energy of the realized candidate state and is updated only when a newly observed paired state yields a strictly higher estimated energy; ties retain their first occurrence\. After Stage I and Stage II are complete, JPPO compares the stage\-wise buffers using the same estimated\-energy ranking criterion and returns the strongest paired state\.

JPPO targets high\-cost decoding through coordinated pixel–prompt steering rather than explicit loop induction\. Algorithm[1](https://arxiv.org/html/2609.05889#alg1)summarizes the complete two\-stage construction procedure, including stage\-wise prompt evolution, projected pixel updates, realized\-cost measurement, and final paired\-state selection\.

Algorithm 1Joint Pixel\-Prompt Optimization \(JPPO\)0:benign input

\(x,p0\)\(x,p\_\{0\}\); perturbation budget

ϵ\\epsilon; step sizes

\(α1,α2\)\(\\alpha\_\{1\},\\alpha\_\{2\}\); iteration budgets

\(K,T\)\(K,T\); prompt budgets

\(B1,B2\)\(B\_\{1\},B\_\{2\}\)
0:final selected adversarial state

\(x⋆,p⋆\)\(x^\{\\star\},p^\{\\star\}\)
1:Initialize path memories

Svis,Sspa,SsemS\_\{\\mathrm\{vis\}\},S\_\{\\mathrm\{spa\}\},S\_\{\\mathrm\{sem\}\}
2:Initialize perturbation

δ←0\\delta\\leftarrow 0
3:Initialize best records

ℬ\(1\)←∅\\mathcal\{B\}^\{\(1\)\}\\leftarrow\\emptyset,

ℬ\(2\)←∅\\mathcal\{B\}^\{\(2\)\}\\leftarrow\\emptyset
4:Stage I: multi\-path contextual warm\-up

5:for

k=1k=1to

KKdo

6:

ak←AspectSchedule⁡\(k,K\)a\_\{k\}\\leftarrow\\operatorname\{AspectSchedule\}\(k,K\)
7:form prompt

pw\(k\)←ℳ⁡\(Svis,Sspa,Ssem,ak,B1\)p\_\{w\}^\{\(k\)\}\\leftarrow\\mathcal\{M\}\(S\_\{\\mathrm\{vis\}\},S\_\{\\mathrm\{spa\}\},S\_\{\\mathrm\{sem\}\};a\_\{k\},B\_\{1\}\)
8:prefix forward on

\(x\+δ,pw\(k\)\)\(x\+\\delta,p\_\{w\}^\{\(k\)\}\); extract logits/states

9:compute

ℒ\(1\)=λeos\(1\)​ℒeos\+λalign\(1\)​ℒalign\+λbot\(1\)​ℒbot\\mathcal\{L\}^\{\(1\)\}=\\lambda\_\{\\mathrm\{eos\}\}^\{\(1\)\}\\mathcal\{L\}\_\{\\mathrm\{eos\}\}\+\\lambda\_\{\\mathrm\{align\}\}^\{\(1\)\}\\mathcal\{L\}\_\{\\mathrm\{align\}\}\+\\lambda\_\{\\mathrm\{bot\}\}^\{\(1\)\}\\mathcal\{L\}\_\{\\mathrm\{bot\}\}
10:update

δ\\deltausing Eq\. \([5](https://arxiv.org/html/2609.05889#S4.E5)\) with

\(ϵ,α1\)\(\\epsilon,\\alpha\_\{1\}\)
11:generate

y^k←fθ​\(x\+δ,pw\(k\)\)\\hat\{y\}\_\{k\}\\leftarrow f\_\{\\theta\}\(x\+\\delta,p\_\{w\}^\{\(k\)\}\)and measure

\(W,τ,E^\)\(W,\\tau,\\widehat\{E\}\)
12:update

ℬ\(1\)\\mathcal\{B\}^\{\(1\)\}with

\(δ,pw\(k\),W,τ,E^\)\(\\delta,p\_\{w\}^\{\(k\)\},W,\\tau,\\widehat\{E\}\)
13:update active memory

Sak←𝒰⁡\(Sak,y^k\)S\_\{a\_\{k\}\}\\leftarrow\\mathcal\{U\}\(S\_\{a\_\{k\}\},\\hat\{y\}\_\{k\}\)
14:endfor

15:

\(δwarm,pw\)←arg⁡max\(δ,p,W,τ,E^\)∈ℬ\(1\)⁡E^\(\\delta\_\{\\mathrm\{warm\}\},p\_\{w\}\)\\leftarrow\\arg\\max\_\{\(\\delta,p,W,\\tau,\\widehat\{E\}\)\\in\\mathcal\{B\}^\{\(1\)\}\}\\widehat\{E\}
16:

δ←δwarm,pc←pw\\delta\\leftarrow\\delta\_\{\\mathrm\{warm\}\},\\quad p\_\{c\}\\leftarrow p\_\{w\}
17:Stage II: joint pixel\-prompt refinement

18:for

t=1t=1to

TTdo

19:prefix forward on

\(x\+δ,pc\)\(x\+\\delta,p\_\{c\}\); extract logits/states

20:compute

ℒ\(2\)=λeos\(2\)​ℒeos\+λalign\(2\)​ℒalign\+λbot\(2\)​ℒbot\\mathcal\{L\}^\{\(2\)\}=\\lambda\_\{\\mathrm\{eos\}\}^\{\(2\)\}\\mathcal\{L\}\_\{\\mathrm\{eos\}\}\+\\lambda\_\{\\mathrm\{align\}\}^\{\(2\)\}\\mathcal\{L\}\_\{\\mathrm\{align\}\}\+\\lambda\_\{\\mathrm\{bot\}\}^\{\(2\)\}\\mathcal\{L\}\_\{\\mathrm\{bot\}\}
21:update

δ\\deltausing Eq\. \([5](https://arxiv.org/html/2609.05889#S4.E5)\) with

\(ϵ,α2\)\(\\epsilon,\\alpha\_\{2\}\)
22:generate

y^t←fθ​\(x\+δ,pc\)\\hat\{y\}\_\{t\}\\leftarrow f\_\{\\theta\}\(x\+\\delta,p\_\{c\}\)and measure

\(W,τ,E^\)\(W,\\tau,\\widehat\{E\}\)
23:update

ℬ\(2\)\\mathcal\{B\}^\{\(2\)\}with

\(δ,pc,W,τ,E^\)\(\\delta,p\_\{c\},W,\\tau,\\widehat\{E\}\)
24:refine prompt

pc←ℛ⁡\(pc,y^t,B2\)p\_\{c\}\\leftarrow\\mathcal\{R\}\(p\_\{c\},\\hat\{y\}\_\{t\},B\_\{2\}\)
25:endfor

26:

\(δ⋆,p⋆,W⋆,τ⋆,E^⋆\)←Best⁡\(ℬ\(1\)∪ℬ\(2\)\)\(\\delta^\{\\star\},p^\{\\star\},W^\{\\star\},\\tau^\{\\star\},\\widehat\{E\}^\{\\star\}\)\\leftarrow\\operatorname\{Best\}\(\\mathcal\{B\}^\{\(1\)\}\\cup\\mathcal\{B\}^\{\(2\)\}\)
27:return

\(x⋆,p⋆\)\(x^\{\\star\},p^\{\\star\}\), where

x⋆=clip⁡\(x\+δ⋆\)x^\{\\star\}=\\mathrm\{clip\}\(x\+\\delta^\{\\star\}\)

For stages∈\{1,2\}s\\in\\\{1,2\\\}, the pixel perturbation is updated by

\(5\)δ←Πϵ​\(δ−αs​sign​\(∇δℒ\(s\)\)\),s∈\{1,2\},\\delta\\leftarrow\\Pi\_\{\\epsilon\}\\\!\\left\(\\delta\-\\alpha\_\{s\}\\,\\mathrm\{sign\}\(\\nabla\_\{\\delta\}\\mathcal\{L\}^\{\(s\)\}\)\\right\),\\quad s\\in\\\{1,2\\\},whereΠϵ\\Pi\_\{\\epsilon\}denotes projection onto the sharedℓ∞\\ell\_\{\\infty\}ball of radiusϵ\\epsilon\. For notational simplicity, we writex\+δx\+\\deltain the algorithm and stage\-specific descriptions below, with valid\-range clipping applied before model evaluation\. Within each iteration, prefix losses are computed before free\-form generation, and realized\-cost measurement is performed after the perturbation update using the current prompt state\.

### 4\.2\.Stage I: Multi\-path Contextual Warm\-up

The goal of Stage I is to construct a class\-agnostic, diverse prompt scaffold that biases the optimization process toward prompts empirically associated with longer continuations\. Starting from a benign imagexxand a fixed short prompt templatep0p\_\{0\}, JPPO maintains three textual path memories,SvisS\_\{\\mathrm\{vis\}\},SspaS\_\{\\mathrm\{spa\}\}, andSsemS\_\{\\mathrm\{sem\}\}, corresponding to visual attributes, spatial relations, and scene\-level semantics, respectively\. Although their initialization may differ, all three paths are subsequently evolved under the same stage\-wise prompt\-construction and refinement framework\. These memories are not generated through unrestricted free\-form prompt writing\. Instead, Stage I uses three predefined class\-agnostic prompt sets as aspect\-specific guidance during prompt construction\.

For readability, Stage I updates the current path memories in place and suppresses iteration superscripts when they are not needed\. At iterationkk, JPPO determines the active aspect according to the deterministic scheduleAspectSchedule⁡\(k,K\)\\operatorname\{AspectSchedule\}\(k,K\)\. Specifically, the visual, spatial, and semantic aspects are activated during the first, middle, and final thirds of Stage I, respectively\. JPPO then merges the current path memories into an enriched prompt

\(6\)pw\(k\)=ℳ⁡\(Svis,Sspa,Ssem,ak,B1\)\.p\_\{w\}^\{\(k\)\}=\\mathcal\{M\}\\\!\\left\(S\_\{\\mathrm\{vis\}\},S\_\{\\mathrm\{spa\}\},S\_\{\\mathrm\{sem\}\};a\_\{k\},B\_\{1\}\\right\)\.In our implementation,ℳ⁡\(⋅\)\\mathcal\{M\}\(\\cdot\)is an aspect\-aware merge operator with dynamic path\-specific quotas rather than fixed mixing ratios\. Given the current active aspect, it allocates different word budgets to the visual, spatial, and semantic paths, retains the most recent content from each path memory by tail truncation at the word level, concatenates the retained segments in a fixed order \(visual, then spatial, then semantic\), and then appends the corresponding aspect\-specific hint text\. The resulting Stage\-I prompt is finally truncated at the word level to satisfy the Stage\-I budgetB1B\_\{1\}\.

Givenpw\(k\)p\_\{w\}^\{\(k\)\}, JPPO first performs a differentiable prefix forward pass on\(x\+δ,pw\(k\)\)\(x\+\\delta,p\_\{w\}^\{\(k\)\}\)to compute the availability\-oriented losses defined in Section[4\.4](https://arxiv.org/html/2609.05889#S4.SS4)\. These losses are computed on the optimization prefix before free\-form decoding begins, and they are used only to update the pixel perturbation through the projected\-sign rule in Eq\. \([5](https://arxiv.org/html/2609.05889#S4.E5)\)\. After this perturbation update, JPPO queries the target VLM with the current perturbed image and prompt, obtains a generated responsey^k\\hat\{y\}\_\{k\}, and records the realized cost of that evaluated state\.

The response is not copied verbatim into the prompt\. Instead, it is treated as a feedback signal that updates the active path memory through a refinement operator

\(7\)Sak←𝒰⁡\(Sak,y^k\),S\_\{a\_\{k\}\}\\leftarrow\\mathcal\{U\}\\\!\\left\(S\_\{a\_\{k\}\},\\hat\{y\}\_\{k\}\\right\),where𝒰⁡\(⋅\)\\mathcal\{U\}\(\\cdot\)denotes phrase\-level selective accumulation with novelty filtering and light compression\. Concretely, the generated text is segmented into short phrase\-like fragments in their original order, and a fragment is retained only if it contains a sufficient proportion of words not already present in the existing context\. If the accumulated text exceeds the update budget, JPPO applies lightweight compression by retaining a subset of informative fragments under the budget while preserving their original order\. In our implementation, this operator performs whitespace normalization but does not rely on a separate heavy de\-duplication stage\.

Thus, Stage I performs two coupled functions\. First, it initializes a resource\-seeking pixel perturbation via prefix\-level differentiable losses\. Second, it constructs a broad, feedback\-adaptive prompt scaffold using realized generation feedback, which is refined more aggressively in Stage II\.

Stage I uses no success threshold or restart criterion: it always runs forKKiterations and then selects the strongest measured state from the Stage\-I buffer:

\(8\)\(δwarm,pw\)=arg⁡max\(δ,p,W,τ,E^\)∈ℬ\(1\)⁡E^\.\(\\delta\_\{\\mathrm\{warm\}\},p\_\{w\}\)=\\arg\\max\_\{\(\\delta,p,W,\\tau,\\widehat\{E\}\)\\in\\mathcal\{B\}^\{\(1\)\}\}\\widehat\{E\}\.This selection preserves the pairing between the warm\-up prompt and the perturbation that jointly produced the largest realized estimated energy during Stage I\. The selected pair\(δwarm,pw\)\(\\delta\_\{\\mathrm\{warm\}\},p\_\{w\}\), rather than a separately merged final prompt, defines the initialization for Stage II\.

### 4\.3\.Stage II: Joint Pixel\-Prompt Refinement

Stage II starts from the highest\-energy paired state retained in the Stage\-I buffer and transforms it into a stronger resource\-amplifying trajectory\. For readability, Stage II uses in\-place updates and initializes the running prompt and perturbation aspc←pwp\_\{c\}\\leftarrow p\_\{w\}andδ←δwarm\\delta\\leftarrow\\delta\_\{\\mathrm\{warm\}\}\.

At iterationtt, JPPO first performs a differentiable prefix forward pass on\(x\+δ,pc\)\(x\+\\delta,p\_\{c\}\)and computes the Stage\-II availability\-oriented loss\. The pixel perturbation is then updated with the Stage\-II step sizeα2\\alpha\_\{2\}under the shared perturbation budgetϵ\\epsilon\. After this update, the target VLM produces a generationy^t\\hat\{y\}\_\{t\}from the current state\(x\+δ,pc\)\(x\+\\delta,p\_\{c\}\); this generation is used for realized\-cost measurement and prompt feedback\.

JPPO then refines the prompt context according to

\(9\)pc←ℛ⁡\(pc,y^t,B2\),p\_\{c\}\\leftarrow\\mathcal\{R\}\\\!\\left\(p\_\{c\},\\hat\{y\}\_\{t\},B\_\{2\}\\right\),whereℛ⁡\(⋅\)\\mathcal\{R\}\(\\cdot\)is a single\-context prompt\-refinement operator that reuses the same selective\-accumulation principle as Stage I, but no longer maintains separate visual, spatial, and semantic memories\. Instead, Stage II updates a single running prompt by appending newly retained phrase\-level fragments, then applying the same novelty filtering and lightweight compression procedure under the Stage II budgetB2B\_\{2\}\. Thus, Stage II preserves the same accumulation rule as Stage I while replacing multi\-path contextual merging with continuation\-oriented refinement over a single running context\.

The text branch remains interface\-compatible and attacker\-directed through restricted prompt\-construction operators\. Only the pixel perturbation is updated by gradients; generated text affects subsequent iterations only through the prompt\-refinement operator and is not used as a differentiable optimization target\.

Both stages operate under a fixed cap on the number of newly generated tokens during realized\-cost evaluation\. After each perturbation update, JPPO performs free\-form generation with the current prompt state, records generation length, latency, and estimated energy, and updates the corresponding best\-state buffer\. The final adversarial example is selected from the strongest measured state across both stages rather than from the final surrogate\-loss iterate alone\.

By default, Stage II continues from the highest\-energy Stage\-I candidate using one refinement trajectory\. We evaluate random and multi\-candidate alternatives in Section[6\.3](https://arxiv.org/html/2609.05889#S6.SS3)\.

### 4\.4\.Availability\-Oriented Objectives

JPPO uses three availability\-oriented losses during pixel\-space optimization\. All losses are computed on the optimization prefix available before free\-form generation\. In particular, the stop\-related loss is applied on the full optimization prefix used during the differentiable attack forward pass, rather than on tokens generated after decoding has begun\. The generated responsesy^k\\hat\{y\}\_\{k\}andy^t\\hat\{y\}\_\{t\}are used only for prompt feedback and realized\-cost measurement, not as differentiable targets for the pixel update\.

#### Prefix Stop\-Suppression Objective\.

Letℐprefix\\mathcal\{I\}\_\{\\mathrm\{prefix\}\}denote the full optimization\-prefix region used during the attack forward pass\. For eachi∈ℐprefixi\\in\\mathcal\{I\}\_\{\\mathrm\{prefix\}\}, letqiq\_\{i\}denote the model probability assigned to the model\-specific EOS token at positionii\. We define

\(10\)ℒeos=1\|ℐprefix\|​∑i∈ℐprefixwi​qi,\\mathcal\{L\}\_\{\\mathrm\{eos\}\}=\\frac\{1\}\{\|\\mathcal\{I\}\_\{\\mathrm\{prefix\}\}\|\}\\sum\_\{i\\in\\mathcal\{I\}\_\{\\mathrm\{prefix\}\}\}w\_\{i\}q\_\{i\},wherewiw\_\{i\}follows a linearly decaying schedule from 10 to 2 over the optimization\-prefix positions, i\.e\.,

wi=\{10−8⋅i−1\|ℐprefix\|−1,\|ℐprefix\|\>1,10,\|ℐprefix\|=1\.i=1,…,\|ℐprefix\|\.w\_\{i\}=\\begin\{cases\}10\-8\\cdot\\dfrac\{i\-1\}\{\|\\mathcal\{I\}\_\{\\mathrm\{prefix\}\}\|\-1\},&\|\\mathcal\{I\}\_\{\\mathrm\{prefix\}\}\|\>1,\\\\ 10,&\|\\mathcal\{I\}\_\{\\mathrm\{prefix\}\}\|=1\.\\end\{cases\}\\qquad i=1,\\dots,\|\\mathcal\{I\}\_\{\\mathrm\{prefix\}\}\|\.Thus, earlier prefix positions receive larger weights than later ones\. Minimizingℒeos\\mathcal\{L\}\_\{\\mathrm\{eos\}\}suppresses premature preference for emitting the model\-specific EOS token on the optimization prefix and encourages the model to enter continuation from a less termination\-prone initial state\.

#### Cross\-Modal Prompt\-Image Misalignment Objective\.

The second objective operates on model\-dependent image and prompt representations extracted from the optimization prefix\. LetV∈ℝM×dV\\in\\mathbb\{R\}^\{M\\times d\}denote the image sequence andP∈ℝN×dP\\in\\mathbb\{R\}^\{N\\times d\}denote the prompt sequence extracted on the optimization prefix\. The exact extraction points are architecture\-dependent and follow each model’s multimodal fusion design\. For Q\-Former\-style VLMs, the image sequence is taken after the query\-to\-language projection\. In contrast, for direct or early fusion VLMs, it is taken from the final multimodal hidden sequence at image\-token positions\. The prompt sequence is taken from the corresponding prompt span in the final hidden sequence used by the attack loss\. Because the two sequences may have different lengths, we resize the prompt sequence to lengthMMusing adaptive average pooling, yieldingP~∈ℝM×d\\tilde\{P\}\\in\\mathbb\{R\}^\{M\\times d\}\. We then compute position\-wise cosine similarity

\(11\)ri=cos\(Vi,P~i\),i=1,…,M,r\_\{i\}=\\cos\(V\_\{i\},\\tilde\{P\}\_\{i\}\),\\qquad i=1,\\dots,M,and define

\(12\)ℒalign=1M​∑i=1Mri\.\\mathcal\{L\}\_\{\\mathrm\{align\}\}=\\frac\{1\}\{M\}\\sum\_\{i=1\}^\{M\}r\_\{i\}\.This term, therefore, measures the mean position\-wise alignment between the image sequence and the pooled prompt sequence on the optimization prefix\. Although the concrete extraction points may differ across models such as BLIP\-2 and Qwen2\.5\-VL, we treat these extraction choices as architecture\-specific implementations of the same alignment surrogate and evaluate their effects empirically through ablations\.

#### Bottleneck Regularization Objective\.

LetH∈ℝS×DH\\in\\mathbb\{R\}^\{S\\times D\}denote the final multimodal hidden sequence extracted on the optimization prefix, and letZ=ϕ⁡\(H\)∈ℝS×d′Z=\\phi\(H\)\\in\\mathbb\{R\}^\{S\\times d^\{\\prime\}\}be its low\-dimensional bottleneck representation, where

\(13\)d′=max⁡\(⌊ρ​D⌋,1\)\.d^\{\\prime\}=\\max\(\\lfloor\\rho D\\rfloor,1\)\.In our implementation,ϕ\\phiandψ\\psiare fixed random linear projections instantiated once per attack and reused across both Stage I and Stage II; they are not learned model parameters\. We define

\(14\)ℒbot=MSE⁡\(ψ⁡\(Z\),H\)\+β​Var​\(Z\),\\mathcal\{L\}\_\{\\mathrm\{bot\}\}=\\mathrm\{MSE\}\(\\psi\(Z\),H\)\+\\beta\\,\\mathrm\{Var\}\(Z\),whereβ=0\.1\\beta=0\.1,MSE⁡\(⋅,⋅\)\\mathrm\{MSE\}\(\\cdot,\\cdot\)denotes elementwise mean\-squared reconstruction loss, andVar⁡\(Z\)\\mathrm\{Var\}\(Z\)is computed by first taking the variance ofZZalong the bottleneck feature dimension and then averaging over positions\. This term serves as a fixed\-projection bottleneck regularizer for the optimization\-prefix hidden states, penalizing large reconstruction residuals and excessive variance in the compressed state\. We use this term as an empirical regularizer and do not claim that it alone causes resource amplification\. The same objective is used across all evaluated VLMs, while the concrete hidden\-state extraction follows each model’s multimodal architecture\. We useρ=0\.10\\rho=0\.10by default; sensitivity to the bottleneck ratio is reported in Appendix Figure[8](https://arxiv.org/html/2609.05889#acmlabel8)and Table[22](https://arxiv.org/html/2609.05889#A4.T22)\.

The full optimization objective is

\(15\)ℒ\(s\)=λeos\(s\)​ℒeos\+λalign\(s\)​ℒalign\+λbot\(s\)​ℒbot,s∈\{1,2\}\.\\mathcal\{L\}^\{\(s\)\}=\\lambda\_\{\\mathrm\{eos\}\}^\{\(s\)\}\\mathcal\{L\}\_\{\\mathrm\{eos\}\}\+\\lambda\_\{\\mathrm\{align\}\}^\{\(s\)\}\\mathcal\{L\}\_\{\\mathrm\{align\}\}\+\\lambda\_\{\\mathrm\{bot\}\}^\{\(s\)\}\\mathcal\{L\}\_\{\\mathrm\{bot\}\},\\qquad s\\in\\\{1,2\\\}\.

### 4\.5\.Discussion of Design Choices

JPPO is built around three design choices\. First, the two\-stage decomposition separates broad contextual exploration from aggressive refinement\. Second, the prompt branch is deliberately updated via attacker\-directed, budget\-constrained text\-space operators rather than via discrete\-token gradient optimization, thereby improving realism and reducing methodological brittleness\. Third, JPPO separates differentiable prefix\-level optimization from realized\-cost evaluation\. The final adversarial example is selected from the strongest realized\-cost state observed across both stages, since realized serving cost is the threat of interest and may not perfectly correlate with the final surrogate\-loss iterate in the presence of stochastic decoding\.

## 5\.Experimental Methodology

### 5\.1\.Setup

#### Models and datasets\.

We evaluate JPPO on five autoregressive VLMs: the official LLaVA\-NeXT\-Mistral\-7B release\([Liu et al\., 2024](https://arxiv.org/html/2609.05889#bib.bib34);[Jiang et al\., 2023](https://arxiv.org/html/2609.05889#bib.bib40)\), Qwen2\.5\-VL\-7B\-Instruct\([Bai et al\., 2025](https://arxiv.org/html/2609.05889#bib.bib10)\), MiniGPT\-4 with Vicuna\-7B\([Zhu et al\., 2024](https://arxiv.org/html/2609.05889#bib.bib6);[Chiang et al\., 2023](https://arxiv.org/html/2609.05889#bib.bib39)\), BLIP\-2 with OPT\-2\.7B\([Li et al\., 2023](https://arxiv.org/html/2609.05889#bib.bib3);[Zhang et al\., 2022](https://arxiv.org/html/2609.05889#bib.bib7)\), and InstructBLIP with Vicuna\-7B\([Dai et al\., 2023](https://arxiv.org/html/2609.05889#bib.bib5);[Chiang et al\., 2023](https://arxiv.org/html/2609.05889#bib.bib39)\)\. Experiments are conducted on the MS COCO val2017 split and the ImageNet validation split\([Lin et al\., 2014](https://arxiv.org/html/2609.05889#bib.bib30);[Deng et al\., 2009](https://arxiv.org/html/2609.05889#bib.bib31)\)\. For each dataset, we randomly sample 1,000 images and repeat the evaluation with three random seeds\. For a given victim model, all compared methods are evaluated on the same sampled subset under each seed\. We report the mean results over the three seeds\.

#### Attack setting\.

The initial promptp0p\_\{0\}is a fixed model\-specific template, as specified in Appendix[C\.1](https://arxiv.org/html/2609.05889#A3.SS1)\. JPPO constructs subsequent prompts through the stage\-wise text\-space procedure described in Section[4](https://arxiv.org/html/2609.05889#S4), using the fixed initial templatep0p\_\{0\}, three predefined class\-agnostic prompt sets for visual, spatial, and semantic aspects, and model\-generated feedback as prompt\-building signals\. Unless otherwise stated, JPPO uses anℓ∞\\ell\_\{\\infty\}perturbation budget ofϵ=8/255\\epsilon=8/255, step size1/2551/255, 100 iterations in Stage I and 100 iterations in Stage II, and prompt\-length budgets of 150 and 200 words for the two stages, respectively\. Throughout the paper, the 100/100 and 150/200 configurations are treated as the default JPPO settings unless a table explicitly reports a separate ablation on optimization schedules or prompt budgets\. In our implementation, prompt growth is controlled by word\-level truncation under these stage\-wise budgets\. We set the availability\-objective weights forℒeos\\mathcal\{L\}\_\{\\mathrm\{eos\}\},ℒalign\\mathcal\{L\}\_\{\\mathrm\{align\}\}, andℒbot\\mathcal\{L\}\_\{\\mathrm\{bot\}\}to 2\.0, 1\.0, and 1\.0 in Stage I, and to 2\.0, 3\.0, and 1\.0 in Stage II\. All default settings are fixed across models and datasets\. We useϵ=8/255\\epsilon=8/255following the perturbation setting of Verbose Images\([Gao et al\., 2024a](https://arxiv.org/html/2609.05889#bib.bib18)\), which maintains a consistent low\-perturbation budget for the main comparison\. The remaining stage\-wise settings are supported by the BLIP\-2 sensitivity analyses: the150/200150/200\-word prompt budgets perform best across both datasets, increasing the iteration schedule to200/200200/200improves estimated energy by only 11\.2–12\.4% while doubling the iteration count, and the selected loss weights yield the highest estimated energy in the tested sweep\. Complete sensitivity results appear in the appendix\.

#### Inference configuration\.

All methods are evaluated under the same decoding setting for a given model\. We cap decoding at 512 newly generated tokens and employ nucleus sampling \(top\-pp\) with a temperature of1\.01\.0andp=0\.9p=0\.9\. We report output length in words after detokenization, whereas the decoding cap itself is enforced in tokens\. Model\-specific input wrappers and benign prompt templates are fixed within each model; only explicitly prompt\-involving settings modify the user\-visible prompt content\.

#### Hardware and measurement\.

Except for the shared\-worker and local API evaluations described below, all locally executed experiments are run on a single NVIDIA RTX 3090 GPU \(24 GB\), with CUDA 12\.8, PyTorch 2\.8\.0, and Python 3\.9\. The GPU is exclusively used for each experiment\. Before formal measurement, we perform warm\-up runs to stabilize the runtime environment\. For each final inference request, we record the reported output length, the software\-observed inference latency, and the estimated energy\. Latency is measured from the start of multimodal inference to the end of generation\. Estimated energy is computed as an average\-power\-times\-latency proxy based on NVML readings\. For each repeated inference run, we record one NVML power reading at the end of generation; we then average these power readings across repeats and multiply the result by the average latency over the same repeats\. This quantity is a software\-level estimate, not a hardware\-level time integral of power, and is used only for relative comparison under an identical measurement protocol\. NVML is the NVIDIA Management Library underlying the NVIDIA\-supportednvidia\-smitool\([NVIDIA, 2026](https://arxiv.org/html/2609.05889#bib.bib38)\)\. Our energy\-estimation protocol follows the end\-of\-generation NVML reading strategy used in Verbose Images\([Gao et al\., 2024a](https://arxiv.org/html/2609.05889#bib.bib18)\)\. Reported latency and estimated energy are measured only during the final inference run of the crafted adversarial example and do not include offline attack\-construction cost\. Unless otherwise stated, all amplification claims in this paper refer only to online serving costs\. For each input\-method pair under a fixed dataset seed, inference is repeated 3 times; we first average over the 3 runs, then average across the sampled subset for that seed\. The final reported value is then obtained by averaging over the three dataset seeds\. For stochastic decoding, we use the same set of decoding seeds across clean, baseline, and JPPO runs for each input\. For the latency\-amplification visualization in Figure[3](https://arxiv.org/html/2609.05889#acmlabel3), we further aggregate the latency\-amplification values across MS COCO and ImageNet for each model\-method pair; bars show the mean across the two datasets, and error bars indicate cross\-dataset variation\.

#### Shared\-worker and local API workloads\.

We additionally evaluate shared\-worker and local API workloads on localhost Qwen2\.5\-VL and BLIP\-2 deployments with four target\-model replicas behind a shared FIFO queue; WildGuard\([Han et al\., 2024](https://arxiv.org/html/2609.05889#bib.bib46)\)runs on a separate fifth RTX 3090 GPU\. Both workload evaluations use three independent arrival\-process seeds\. The uncontrolled setting uses 80% benign utilization, while the local API replays fixed JPPO traces under Open, Request\-RL, Safety\+Request\-RL, and EarlyStop\+Request\-RL\. Exact workload and gateway settings appear in Appendix[C\.5](https://arxiv.org/html/2609.05889#A3.SS5)\.

### 5\.2\.Baselines and Fairness Controls

#### Baselines\.

We compare JPPO against: \(i\)*Clean*, the benign image and prompt pair; \(ii\)*Noise*, which applies a single random bounded perturbation within the sameℓ∞\\ell\_\{\\infty\}budget, followed by clipping to the valid input range, without any optimization; and \(iii\) prior image\-only resource\-exhaustion attacks, including NICGSlowDown\([Chen et al\., 2022b](https://arxiv.org/html/2609.05889#bib.bib35)\)and Verbose Images\([Gao et al\., 2024a](https://arxiv.org/html/2609.05889#bib.bib18)\)\. We report Hidden Tail\([Zhang et al\., 2026](https://arxiv.org/html/2609.05889#bib.bib33)\)separately in Appendix[D\.3](https://arxiv.org/html/2609.05889#A4.SS3)\. Its original evaluation regime is not directly aligned with our low\-perturbation main comparison; under our unified protocol, it yields nonzero but comparatively weak amplification\. We discuss recent loop\-centric image\-only attacks\([Fu et al\., 2026](https://arxiv.org/html/2609.05889#bib.bib41)\)separately under matched greedy decoding rather than including them in the main baseline table, because they differ from JPPO in attack surface, optimization target, and default decoding protocol\.

#### Fairness controls\.

All methods use the same sampled images, model checkpoints, hardware, decoding cap, sampling policy, and measurement protocol\. Prior baselines are reproduced from their open\-source implementations with reported hyperparameters: NICGSlowDown and Verbose Images use 1000 optimization iterations, while Hidden Tail is reproduced under its 5000\-iteration setting\. JPPO uses 100 Stage\-I and 100 Stage\-II iterations by default\. Explicit word budgets additionally constrain prompt\-involving methods\. The Neutral\-long\-prompt control uses a fixed 200\-word neutral descriptive instruction matching JPPO’s default Stage\-II visible\-prompt cap without using JPPO’s stage\-wise feedback, pixel perturbation, or prompt\-refinement procedure\. Claims of superiority in the main table are restricted to the baselines directly evaluated there\. In the cross\-model defense evaluation, all attacks use the same top\-pp/512\-token serving protocol and each defended result is paired with its corresponding undefended attack\.

#### Evaluation protocol\.

For JPPO, we select the final adversarial state from the measured image–prompt pairs visited during optimization using estimated energy as the primary ranking criterion, and report generation length and latency for the same selected pair\. Each measured state is evaluated after the corresponding perturbation update through free\-form generation under the fixed serving protocol\. After selection, the chosen adversarial state is re\-evaluated from scratch under the fixed serving protocol and repeated decoding seeds; all reported latency and estimated\-energy values come from these final evaluation runs, not from the construction\-time ranking measurements\. This candidate\-selection procedure is specific to JPPO and should be interpreted as part of the attack construction pipeline rather than as an online serving\-time advantage\.

### 5\.3\.Metrics

#### Generation length\.

We measure the number of newly generated output words before termination or the decoding cap is reached\.

#### Latency\.

We measure software\-observed inference time, including multimodal prefilling and autoregressive decoding\.

#### Estimated energy\.

We report the NVML\-based estimated energy proxy, defined above, in joules\. In all tables,E^\\widehat\{E\}\(J\) denotes this software\-level proxy rather than an integrated hardware energy measurement\.

#### Amplification factor\.

When reporting amplification, we compute it relative to the corresponding clean baseline for the same model, dataset, decoding protocol, and measurement setup:

\(16\)Ampm=m⁡\(x⋆,p⋆\)m⁡\(x,p0\),m∈\{W,τ,E^\}\.\\mathrm\{Amp\}\_\{m\}=\\frac\{m\(x^\{\\star\},p^\{\\star\}\)\}\{m\(x,p\_\{0\}\)\},\\qquad m\\in\\\{W,\\tau,\\widehat\{E\}\\\}\.Absolute tables report the aggregated measurements above, and defense degradation is computed relative to the paired undefended attack\. Table[7](https://arxiv.org/html/2609.05889#S6.T7)reports equal\-weight means over the two datasets and three serving\-cost metrics by model, grouping the nine non\-EarlyStop defenses and reporting EarlyStop separately\. Figure[6](https://arxiv.org/html/2609.05889#acmlabel6)further reports equal\-weight means over the five models, two datasets, and three serving\-cost metrics by defense\.

#### API and queueing metrics\.

For the uncontrolled FIFO study, we report last\-stable benign P95 inflation, the implied overload bracket, and additional target\-model GPU\-hours per 1,000 attack requests\. For the fixed\-trace API study, we report JPPO admission, attacker worker share, normalized loadρ\\rho, and benign P95 end\-to\-end latency inflation; the safety\-admission evaluation reports rejection rate and end\-to\-end latency amplification\. Detailed timing and aggregation definitions appear in Appendix[C\.5](https://arxiv.org/html/2609.05889#A3.SS5)\.

#### Transfer\-gain retention\.

For metricmm, letAs→t\(m,d\)A\_\{s\\rightarrow t\}^\{\(m,d\)\}denote amplification for an input constructed on sourcessand evaluated on targetttfor datasetdd\. For each off\-diagonal pair, retained exact\-model gain is

\(17\)Rs→t\(m,d\)=As→t\(m,d\)−1At→t\(m,d\)−1,s≠t\.R\_\{s\\rightarrow t\}^\{\(m,d\)\}=\\frac\{A\_\{s\\rightarrow t\}^\{\(m,d\)\}\-1\}\{A\_\{t\\rightarrow t\}^\{\(m,d\)\}\-1\},\\qquad s\\neq t\.We average over the 20 directed off\-diagonal pairs for each dataset and then across MS COCO and ImageNet\.

## 6\.Evaluation

### 6\.1\.Main Results

Table 1\.Full comparison on MS COCO and ImageNet\. We report absolute output length, estimated\-energy proxyE^\\widehat\{E\}, and latency\. Means are over three repeated inference runs and three dataset seeds; image\-level standard deviations are reported in the appendix\. Bold marks the largest non\-Clean value in each setting\.ModelMethodMS COCOImageNetLen\.E^\\widehat\{E\}\(J\)Lat\. \(s\)Len\.E^\\widehat\{E\}\(J\)Lat\. \(s\)BLIP\-2Clean8\.3667\.380\.427\.0781\.500\.55Noise8\.2365\.650\.447\.2260\.280\.43NICGSlowDown86\.76567\.034\.46103\.41663\.155\.01Verbose Images211\.621432\.6112\.07231\.311592\.6912\.48JPPO373\.232324\.2919\.58361\.602665\.9620\.16Qwen2\.5\-VLClean80\.24945\.974\.6181\.52912\.434\.53Noise83\.18934\.214\.7279\.86940\.804\.47NICGSlowDown146\.121561\.607\.28150\.641700\.148\.01Verbose Images252\.462898\.4110\.88247\.182765\.4910\.41JPPO441\.235458\.5924\.03422\.974867\.4421\.16LLaVA\-NeXTClean85\.531577\.856\.7476\.671339\.236\.39Noise88\.431479\.396\.6272\.631312\.715\.96NICGSlowDown187\.452382\.0510\.02146\.402053\.489\.82Verbose Images233\.433219\.7111\.78235\.063208\.5912\.04JPPO369\.106520\.1227\.53340\.035706\.2127\.18InstructBLIPClean50\.12618\.403\.6853\.95741\.224\.47Noise52\.34606\.913\.7955\.61756\.044\.39NICGSlowDown83\.081243\.187\.0189\.011194\.387\.02Verbose Images129\.841530\.9310\.72125\.631720\.4511\.11JPPO250\.923476\.1817\.58262\.113968\.3421\.43MiniGPT\-4Clean52\.60790\.883\.8069\.06983\.264\.37Noise58\.16864\.513\.9769\.30943\.824\.56NICGSlowDown196\.592906\.5214\.07183\.133139\.3615\.43Verbose Images314\.304251\.7721\.57307\.624057\.5218\.45JPPO385\.355378\.7723\.64370\.535204\.6823\.10Figure 3\.Latency amplification across victim models, aggregated over MS COCO and ImageNet\.Bars show mean amplification over the two datasets, with error bars indicating cross\-dataset variation\.Latency amplification bar chartA grouped bar chart compares latency amplification across BLIP\-2, Qwen2\.5\-VL, LLaVA\-NeXT, InstructBLIP, and MiniGPT\-4\. JPPO is the tallest bar for each model among the directly evaluated baselines\. BLIP\-2 shows the largest latency amplification, while LLaVA\-NeXT shows comparatively smaller amplification\. Error bars show variation between the two datasets\.Table[1](https://arxiv.org/html/2609.05889#S6.T1)summarizes the main comparison on MS COCO and ImageNet across representative autoregressive VLMs, while Figure[3](https://arxiv.org/html/2609.05889#acmlabel3)provides a latency\-focused view aggregated over the two datasets\. Under the fixed top\-pp/512\-token protocol, JPPO yields the largest latency and estimated\-energy amplification among the directly evaluated baselines\.

Random noise yields little consistent amplification\. Across the ten model–dataset settings, Neutral\-long\-prompt reaches at most2\.04×2\.04\\timesestimated\-energy and2\.40×2\.40\\timeslatency amplification, whereas JPPO’s minimum corresponding amplifications are4\.13×4\.13\\timesand4\.08×4\.08\\times\. Thus, matching JPPO’s maximum visible\-prompt budget alone does not reproduce its serving\-cost increase\.

### 6\.2\.Cross\-Model Transferability

To evaluate deployment\-time transfer, we construct JPPO inputs on each source model and submit them unchanged to every target model\. We refer to the off\-diagonal setting as surrogate\-based black\-box target transfer: construction uses a white\-box source model, while the target is used only for final inference, without target\-side gradients, refinement, or target\-informed selection\. Diagonal entries provide exact\-model white\-box references\.

Figure 4\.White\-box\-to\-black\-box latency survivability\.For each target, light\-blue circles denote the four off\-diagonal transfers, dark\-blue circles their arithmetic mean, and orange diamonds the exact\-model white\-box result\. Panels show MS COCO and ImageNet on a shared logarithmic scale; the dashed line marks the matched Clean baseline \(1×1\\times\)\.White\-box\-to\-black\-box latency survivabilityTwo side\-by\-side point plots compare JPPO latency amplification on MS COCO and ImageNet\. For each of five target models, four light\-blue points show transfer from the other source models, a dark\-blue point shows their mean, and an orange diamond shows the exact\-model white\-box result\. All transfer means remain above the Clean baseline but below their exact\-model references, showing partial and asymmetric black\-box target transfer\.Table 2\.JPPO latency across source–target pairs \(s; MS COCO / ImageNet\)\.Off\-diagonal cells use no target\-side adaptation; shaded diagonals are exact\-model white\-box references\. Bold marks the highest latency for each target model and dataset\.Source ModelBLIP\-2Qwen2\.5\-VLLLaVA\-NeXTInstructBLIPMiniGPT\-4BLIP\-219\.58/20\.1616\.98 / 18\.7422\.49 / 24\.2912\.33 / 13\.0217\.71 / 8\.94Qwen2\.5\-VL4\.40 / 7\.3624\.03/ 21\.1618\.92 / 21\.6116\.32 / 16\.5614\.61 / 12\.35LLaVA\-NeXT3\.73 / 3\.3823\.68 /21\.3027\.53/27\.1814\.73 / 12\.8017\.97 / 19\.19InstructBLIP0\.92 / 1\.6515\.97 / 17\.0418\.45 / 18\.4617\.58/21\.4312\.22 / 12\.86MiniGPT\-43\.64 / 3\.3110\.69 / 8\.8316\.82 / 11\.457\.52 / 8\.0223\.64/23\.10Figure[4](https://arxiv.org/html/2609.05889#acmlabel4)contrasts exact\-model white\-box amplification with surrogate\-based black\-box target transfer, while Table[2](https://arxiv.org/html/2609.05889#S6.T2)reports the source–target latency matrix\. All 40 dataset\-specific off\-diagonal directions exceed the matched Clean baseline, ranging from1\.79×1\.79\\timesto13\.38×13\.38\\times\. Across both datasets, off\-diagonal transfer averages5\.15×5\.15\\times,3\.97×3\.97\\times, and4\.12×4\.12\\timesoutput\-length, energy, and latency amplification, retaining 48\.7% and 50\.7% of the exact\-model energy and latency gains, respectively\. Full length and energy matrices appear in Appendix[D\.9](https://arxiv.org/html/2609.05889#A4.SS9)\.

#### External black\-box API transfer\.

We further evaluate external black\-box transfer on Gemini 3\.1 Pro Preview and Qwen2\.5\-VL\-72B\. After within\-case averaging over 500 paired Clean/JPPO cases per endpoint, 18\.8% and 24\.2% of cases, respectively, exceed1\.2×1\.2\\timesoutput\-length amplification, with median amplifications of1\.07×1\.07\\timesand1\.12×1\.12\\times\. These results indicate limited but nonzero transfer to external black\-box endpoints\.

### 6\.3\.Construction and Serving Costs

Table 3\.Construction and serving costs averaged across datasets\.Values are equal\-weight means over MS COCO and ImageNet\. VI denotes Verbose Images\. For JPPO, construction cost includes Stage I and one Stage\-II trajectory initialized from the highest\-energy Stage\-I candidate; serving cost measures only the final selected request\.ModelMethodConstructionServingTime \(s\)E^\\widehat\{E\}\(kJ\)E^\\widehat\{E\}\(J\)Lat\. \(s\)BLIP\-2VI316\.0441\.951512\.6512\.28JPPO582\.2653\.042495\.1319\.87Qwen2\.5\-VLVI4813\.97183\.852831\.9510\.65JPPO3354\.70146\.725163\.0222\.60LLaVA\-NeXTVI8645\.60658\.263214\.1511\.91JPPO7495\.38558\.766113\.1727\.36InstructBLIPVI1433\.08233\.981625\.6910\.92JPPO1658\.18211\.503722\.2619\.51MiniGPT\-4VI2664\.83534\.504154\.6520\.01JPPO2097\.47401\.485291\.7323\.37Table[3](https://arxiv.org/html/2609.05889#S6.T3)separates attacker\-side offline construction from victim\-side online serving\. JPPO’s build time and energy show no uniform ordering relative to Verbose Images, whereas its final requests incur higher online energy and latency for every model on both datasets\. Identical replay is directly cacheable after the first request\. In a cache\-hit sensitivity analysis using 5\-percentage\-point increments, avoiding overload requires 35–50% hits for Qwen2\.5\-VL and 65–85% for BLIP\-2 across the two datasets\. Thus, repeated\-input caching can substantially reduce the impact of identical replay when sufficiently high hit rates are achieved\.

#### Stage\-I schedule and transition\.

Across five models and both datasets, phase\-based scheduling increases average length, estimated\-energy, and latency amplification from4\.46×4\.46\\times/4\.96×4\.96\\times/4\.28×4\.28\\timesunder random scheduling to5\.95×5\.95\\times/6\.06×6\.06\\times/5\.40×5\.40\\times\. Stage I uses a fixed iteration budget without a success threshold or restart rule; 31\.8% of construction runs attain the final selected optimum during Stage I, while 5\.6% of final inference runs reach the 512\-token cap\. These results support the deterministic schedule while retaining a fixed\-budget transition to Stage II\. Detailed schedule results are provided in Appendix[D\.4](https://arxiv.org/html/2609.05889#A4.SS4)\.

#### Stage\-II candidate selection\.

Random top\-3/top\-5 selection is model\-dependent: each strategy slightly improves finalE^\\widehat\{E\}on one model but reduces it on the other\. Refining all top\-3/top\-5 candidates increases finalE^\\widehat\{E\}by 4\.30–4\.65% and 5\.12–6\.17%, respectively, while increasing both build\-time and build\-energy costs to 2\.12–2\.20×\\timesand 3\.29–3\.36×\\times\. We therefore retain the highest\-energy Stage\-I candidate as the default initialization; complete results appear in Appendix[D\.5](https://arxiv.org/html/2609.05889#A4.SS5)\.

#### Judge\-based output quality\.

We use GPT\-5\.5 to score every final\-inference output from Clean, Neutral\-long\-prompt, Prompt\-only, Verbose Images, and JPPO on relevance, informativeness, redundancy, and usefulness\. Neutral\-long\-prompt uses the fixed 200\-word length\-control instruction defined in the evaluation controls\. Prompt\-only reuses the independently constructed prompt\-only setting from the optimization\-space ablation: its output is generated from the original image using the selected optimized prompt, with all pixel\-space availability losses disabled\. Every candidate output is evaluated against the corresponding original image and benign instruction using a common original\-task reference; the complete judge prompt and protocol appear in Appendix[D\.20](https://arxiv.org/html/2609.05889#A4.SS20)\.

Table 4\.Original\-task output quality\.GPT\-5\.5 scores for 450,000 candidate outputs are reported as equal\-weight means across the ten model–dataset settings\. Each output is evaluated against the corresponding original image and benign instruction, irrespective of the method\-specific input used during generation\. Rel\., Info\., Red\., and Use\. denote relevance, informativeness, redundancy, and usefulness; higher is better except for Red\.MethodRel\.Info\.Red\.Use\.Clean4\.493\.661\.434\.05Neutral\-long\-prompt4\.294\.012\.043\.92Prompt\-only3\.723\.073\.452\.98Verbose Images2\.822\.154\.082\.91JPPO3\.593\.443\.323\.38Table[4](https://arxiv.org/html/2609.05889#S6.T4)shows that JPPO has lower relevance, informativeness, and usefulness and higher redundancy than Clean and Neutral\-long\-prompt under the common original\-task reference\. Compared with Verbose Images, JPPO scores higher in relevance, informativeness, and usefulness and lower in redundancy; relative to Prompt\-only, it improves informativeness and usefulness while reducing redundancy\. These results indicate that JPPO’s increased serving cost is not explained solely by a collapse into low\-utility repetitive output: although task quality degrades relative to Clean, the generated responses retain substantially more original\-task utility than Verbose Images under the same judge protocol\. The comparison with Neutral\-long\-prompt further separates resource amplification from simply eliciting longer but otherwise conventional descriptive responses\.

### 6\.4\.Loop\-Oriented Characterization of JPPO

We characterize explicit looping on generated token sequences from MS COCO and ImageNet before detokenization\. For each dataset, an output is loop\-positive if it contains a repeated token block of length at least 16 repeated 3 times, or length at least 32 repeated 2 times\. We report loop incidence \(Inc\.\), longest repeated\-span ratio \(Span\), and tail loop coverage \(Cov\.\), where the last metric measures repeated\-span coverage in the final 128 generated tokens\.

Table 5\.Loop characterization of JPPO on MS COCO and ImageNet\.Entries are reported as MS COCO / ImageNet\. Inc\., Span, and Cov\. denote loop incidence, longest repeated\-span ratio, and tail loop coverage, respectively\. Lower values are better\.ModelLoopInc\. \(%\)LongestSpan \(%\)TailCov\. \(%\)BLIP\-21\.90 / 1\.012\.04 / 1\.261\.31 / 0\.52Qwen2\.5\-VL0\.00 / 0\.000\.00 / 0\.000\.00 / 0\.00LLaVA\-NeXT0\.00 / 0\.000\.47 / 0\.580\.00 / 0\.00InstructBLIP0\.00 / 0\.000\.01 / 0\.130\.00 / 0\.00MiniGPT\-41\.60 / 0\.001\.82 / 1\.580\.93 / 0\.59A complementary token\-level novelty analysis in Appendix Figure[7](https://arxiv.org/html/2609.05889#acmlabel7)further shows that attacked outputs retain moderate novelty despite longer generation, consistent with the low explicit\-loop incidence above\.

### 6\.5\.Optimization\-Space Ablation

On BLIP\-2, this ablation serves as a focused diagnostic of our threat\-model claim\. Rather than asking only whether JPPO is stronger than prior image\-only attacks, we ask whether coordinated multimodal control can create additional availability risk in a representative model\. We compare Pixel\-only, Prompt\-only, and Joint optimization\. Pixel\-only retains the benign prompt and optimizes only the bounded perturbation, whereas Prompt\-only keeps the image unperturbed and independently executes the same stage\-wise prompt\-construction procedure with all pixel\-space availability losses disabled\.

Figure 5\.Optimization\-space ablation on BLIP\-2\.Clean\-normalized output\-length, estimated\-energy, and latency amplification for Pixel\-only, Prompt\-only, and Joint optimization on MS COCO and ImageNet\.Optimization\-space ablation bar chartTwo bar charts compare Pixel\-only, Prompt\-only, and Joint optimization on MS COCO and ImageNet\. Each panel reports amplification over clean for length, energy, and latency\. Prompt\-only is generally stronger than Pixel\-only, while Joint optimization is strongest across all metrics and both datasets\.Figure[5](https://arxiv.org/html/2609.05889#acmlabel5)shows that both single\-space variants increase serving cost, while Joint optimization yields the largest output\-length, estimated\-energy, and latency amplification on both MS COCO and ImageNet\. Across all five models on MS COCO, Joint also yields the largest estimated\-energy amplification, exceeding the stronger unimodal branch by 28\.7–83\.7% \(Table[6](https://arxiv.org/html/2609.05889#S6.T6)\)\. On BLIP\-2, Prompt\-only is generally stronger than Pixel\-only, indicating that the visible prompt provides an effective handle on generation persistence, but neither single\-space variant recovers the full Joint effect\. The gain of Joint over Prompt\-only shows that the bounded pixel branch remains useful even when the prompt is already optimized, while its gain over Pixel\-only shows the additional contribution of prompt\-driven continuation pressure\. As a qualitative supplement, Appendix Figure[9](https://arxiv.org/html/2609.05889#acmlabel9)illustrates changes in visual evidence localization under JPPO\.

### 6\.6\.Ablation on Loss Components

We compare the three individual availability\-oriented objectives,ℒeos\\mathcal\{L\}\_\{\\mathrm\{eos\}\},ℒalign\\mathcal\{L\}\_\{\\mathrm\{align\}\}, andℒbot\\mathcal\{L\}\_\{\\mathrm\{bot\}\}, with the full objective across all five models on MS COCO\. As shown in Table[6](https://arxiv.org/html/2609.05889#S6.T6), the strongest individual loss varies by model, while the full objective yields the largest estimated\-energy amplification in every case, exceeding the strongest single\-loss variant by 14\.9–22\.5%\. Detailed two\-dataset results on BLIP\-2 are reported in Appendix Table[19](https://arxiv.org/html/2609.05889#A4.T19)\.

Table 6\.Cross\-model ablation on MS COCO\.Estimated\-energy amplification over Clean\. Bold marks the largest value per model\.ModelModalitySingle LossJointPixelPrompt𝓛𝐞𝐨𝐬\\boldsymbol\{\\mathcal\{L\}\_\{\\mathrm\{eos\}\}\}𝓛𝐚𝐥𝐢𝐠𝐧\\boldsymbol\{\\mathcal\{L\}\_\{\\mathrm\{align\}\}\}𝓛𝐛𝐨𝐭\\boldsymbol\{\\mathcal\{L\}\_\{\\mathrm\{bot\}\}\}BLIP\-211\.0220\.1929\.6822\.5222\.7234\.50Qwen2\.5\-VL3\.913\.364\.124\.684\.715\.77LLaVA\-NeXT3\.043\.212\.963\.183\.394\.13InstructBLIP2\.373\.064\.894\.314\.575\.62MiniGPT\-44\.274\.335\.115\.465\.726\.80
### 6\.7\.Cross\-Model Robustness under Serving Defenses

Table 7\.Cross\-model defense degradation \(%\)\.Panel \(a\) averages the nine non\-EarlyStop defenses, and Panel \(b\) reports EarlyStop\. Values are equal\-weight means over datasets and serving\-cost metrics; lower means less mitigation\. Bold marks the smallest degradation per model and the smallest overall average\.ModelJPPOVerboseImagesNICGSlowDownLingoLoop\(a\) Nine\-defense meanBLIP\-211\.2513\.057\.9125\.12Qwen2\.5\-VL2\.784\.802\.036\.04LLaVA\-NeXT5\.143\.723\.3214\.28InstructBLIP1\.2414\.924\.2915\.97MiniGPT\-44\.6921\.9510\.9825\.93Average5\.0211\.695\.7117\.47\(b\) EarlyStopBLIP\-28\.408\.019\.016\.98Qwen2\.5\-VL1\.085\.264\.012\.04LLaVA\-NeXT1\.537\.885\.421\.09InstructBLIP7\.1012\.517\.079\.13MiniGPT\-40\.1920\.4116\.009\.08Average3\.6610\.818\.305\.66![Per-defense degradation heatmap](https://arxiv.org/html/2609.05889v1/defense_per_defense_heatmap.png)Figure 6\.Per\-defense degradation across five VLMs \(%\)\.Rows denote defenses and columns denote attacks\. Each cell is an equal\-weight mean over the five models, MS COCO and ImageNet, and the three serving\-cost metrics \(output length, estimated energy, and latency\)\. Darker blue indicates greater degradation; lower means less mitigation, and negative values indicate increased cost\.Per\-defense degradation heatmapA heatmap with ten defenses as rows and four attack methods as columns\. Each cell is annotated with the corresponding degradation percentage\. Lighter blue indicates lower degradation and darker blue indicates higher degradation, with the color bar spanning approximately minus 2 percent to 25 percent\. Negative values are displayed directly in the corresponding cells\.Table[7](https://arxiv.org/html/2609.05889#S6.T7)shows the smallest aggregate reduction for JPPO under both the nine non\-EarlyStop defenses \(5\.02%\) and EarlyStop \(3\.66%\)\. Figure[6](https://arxiv.org/html/2609.05889#acmlabel6)reports per\-defense degradation relative to paired undefended attacks\. Each cell is an equal\-weight mean over the five models, the two datasets \(MS COCO and ImageNet\), and the three serving\-cost metrics, and is annotated to one decimal place\. Across the nine non\-EarlyStop defenses, JPPO’s degradation ranges from−2\.3\-2\.3% to 8\.7%, remaining below the corresponding average degradation of Verbose Images and LingoLoop for every reported defense\. Model\-level and metric\-level breakdowns appear in Appendix[D\.2](https://arxiv.org/html/2609.05889#A4.SS2); implementation details for prompt sanitization and EarlyStop appear in Appendix[C\.4](https://arxiv.org/html/2609.05889#A3.SS4)\.

### 6\.8\.Shared\-Worker Impact and Local API Constraints

We next evaluate benign\-user contention and local API controls\. We first quantify the uncontrolled four\-worker FIFO impact and then evaluate request\-rate limiting, safety admission, and runtime early stopping under a fixed workload\.

Table 8\.Uncontrolled four\-worker FIFO impact at 80% benign utilization\.VI denotes Verbose Images\. Each overload bracket reports the largest stable and smallest overloaded offered attack shares\. Bold marks the stronger effect within each model–dataset setting: higher P95 inflation and GPU\-h/1K, and a lower overload threshold\.ModelMethodP95OverloadGPU\-h/1KMS COCOQwen2\.5\-VLVI1\.82×\\times\[9%,10%\]1\.76JPPO1\.94×\\times\[4%,5%\]6\.60BLIP\-2VI37\.1×\\times\[0\.75%,1\.00%\]3\.05JPPO55\.9×\\times\[0\.50%,0\.75%\]4\.92ImageNetQwen2\.5\-VLVI1\.73×\\times\[9%,10%\]1\.68JPPO2\.37×\\times\[5%,6%\]5\.28BLIP\-2VI35\.2×\\times\[1\.00%,1\.25%\]3\.17JPPO58\.8×\\times\[0\.50%,0\.75%\]6\.11Table[8](https://arxiv.org/html/2609.05889#S6.T8)connects isolated request cost to benign\-user queueing pressure\. At 80% benign utilization, JPPO reaches the queue saturation threshold at 4–5% / 5–6% offered attack share on Qwen2\.5\-VL and 0\.50–0\.75% on BLIP\-2, earlier than Verbose Images; the corresponding benign P95 inflation and additional GPU\-hours are reported in the table\.

To isolate content\-safety admission before evaluating the complete gateway, we apply WildGuard to the interface\-visible Clean and JPPO prompts, pooling 1,000 MS COCO and 1,000 ImageNet prompts for each target model and prompt type\. WildGuard rejects 1\.2% and 0\.3% of JPPO prompts for Qwen2\.5\-VL and BLIP\-2, respectively, with no Clean rejections\. Relative to matched runs without WildGuard, Clean/JPPO end\-to\-end latency amplification is 1\.24×\\times/1\.27×\\timeson Qwen2\.5\-VL and 1\.21×\\times/1\.32×\\timeson BLIP\-2\. WildGuard receives only the visible text and runs on a separate GPU; its latency is included in end\-to\-end timing, while its energy consumption is not reported\.

Without retuning the workload for any control, the fixed\-trace API evaluation reuses the first implied\-overload JPPO operating points from Table[8](https://arxiv.org/html/2609.05889#S6.T8)\. These operating points correspond to mean external JPPO submission rates of 2\.19 / 2\.71 requests/min for Qwen2\.5\-VL and 3\.45 / 2\.64 requests/min for BLIP\-2\. Request\-RL uses a per\-identity token bucket with a 5\-request/min refill rate and burst capacity 2\. Although the mean attack rates are below the refill rate, Poisson bursts can temporarily exhaust the burst capacity and therefore produce nonzero request\-rate rejection\.

Table 9\.Fixed\-trace local API evaluation under serving controls\.Entries are MS COCO / ImageNet\. Admission is measured before target\-model execution, while P95 reports benign end\-to\-end latency inflation\.ModelPolicyJPPOAdmissionAttackerWorker ShareLoad𝝆\\boldsymbol\{\\rho\}Benign API P95E2E InflationQwen2\.5\-VLOpen100\.00% / 100\.00%23\.5% / 25\.3%1\.04 / 1\.0742\.9×\\times/ 47\.6×\\timesRequest\-RL91\.34% / 87\.80%22\.1% / 23\.6%1\.01 / 1\.0213\.1×\\times/ 16\.2×\\timesSafety\+Request\-RL90\.15% / 86\.83%21\.8% / 23\.2%1\.00 / 1\.0214\.5×\\times/ 18\.0×\\timesEarlyStop\+Request\-RL91\.34% / 87\.80%21\.7% / 22\.9%1\.00 / 1\.0112\.4×\\times/ 15\.4×\\timesBLIP\-2Open100\.00% / 100\.00%27\.2% / 24\.1%1\.10 / 1\.05108\.6×\\times/ 94\.2×\\timesRequest\-RL82\.74% / 89\.55%24\.4% / 22\.3%1\.04 / 1\.0180\.7×\\times/ 76\.8×\\timesSafety\+Request\-RL82\.57% / 89\.19%24\.0% / 22\.1%1\.03 / 1\.0186\.4×\\times/ 80\.1×\\timesEarlyStop\+Request\-RL82\.74% / 89\.55%22\.2% / 20\.6%1\.01 / 0\.9972\.5×\\times/ 66\.3×\\timesTable[9](https://arxiv.org/html/2609.05889#S6.T9)shows that Request\-RL reduces JPPO admission and benign P95 inflation on both models, while Safety\+Request\-RL produces only a small additional change\. EarlyStop\+Request\-RL leaves admission unchanged but further reduces attacker worker share and benign P95, particularly on BLIP\-2\. Under the remaining admitted JPPO load, benign requests nevertheless retain elevated P95 latency\.

All requests in Table[9](https://arxiv.org/html/2609.05889#S6.T9)are black\-box submissions of previously constructed image–prompt pairs\. Request\-rate limiting and safety admission operate before target\-model execution, whereas EarlyStop acts during generation; none requires target\-side gradients\.

These controls operate at different stages of the serving pipeline\. Request\-rate limiting constrains admission bursts, while the low rejection rates under WildGuard reflect an important property of JPPO: its interface\-visible prompts contain no overtly harmful content despite inducing substantially increased serving cost\. Runtime EarlyStop targets loop\- or repetition\-dominant generations and therefore provides a complementary mechanism for suppressing this class of decoding behavior\. Together, these results show that JPPO can retain substantial resource\-amplification effects under ordinary content\-safety admission and loop\-oriented runtime controls, while request\-rate limiting further reduces its practical serving impact\.

### 6\.9\.Comparison with Recent Loop\-Centric Attacks

To position JPPO relative to recent loop\-centric image\-only attacks, we compare against LingoLoop under a matched greedy/512\-token protocol using the same sampled images and local measurement procedure\. We report word\-level length, generated tokens, latency, and estimated energy\.

Table 10\.Matched comparison with LingoLoop\.All methods use greedy decoding and a 512\-token cap; we report word and token lengths, estimated energy, and latency\. Bold marks the largest non\-Clean value for each model, dataset, and metric\.ModelMethodMS COCOImageNetLen\.TokensE^\\widehat\{E\}\(J\)Lat \(s\)Len\.TokensE^\\widehat\{E\}\(J\)Lat \(s\)Qwen2\.5\-VL\-7BClean70\.5182\.13825\.213\.1674\.1388\.47893\.153\.68LingoLoop435\.29501\.945056\.1318\.75439\.18500\.185014\.3218\.01JPPO456\.15510\.785528\.4123\.61440\.68502\.325184\.1822\.46InstructBLIPClean53\.8465\.35675\.213\.6159\.1771\.56857\.445\.08LingoLoop417\.61495\.205001\.6234\.23409\.92488\.364751\.6633\.71JPPO318\.89392\.664414\.1828\.13326\.47409\.284684\.4129\.52Table[10](https://arxiv.org/html/2609.05889#S6.T10)shows mixed results: JPPO is competitive with or stronger than LingoLoop on Qwen2\.5\-VL, but weaker on InstructBLIP\. This supports the interpretation that JPPO and loop\-centric image\-only attacks probe different high\-cost regimes rather than forming a single monotone ranking\.

## 7\.Limitations

JPPO is primarily evaluated in an exact\-model white\-box offline\-construction setting using open\-source models, complemented by cross\-model black\-box transfer and limited external API testing\. Black\-box transfer is generally weaker and more variable than exact\-model results, and the external API evaluation is limited to output\-length measurements because provider\-side resource metrics are unavailable\. Its joint\-input threat model is broader than fixed\-prompt image\-only settings, and comparisons to image\-only attacks should therefore be read as complementary rather than substitutive\. Absolute latency and estimated\-energy values depend on the local hardware and software stack, and our NVML\-based energy metric is a proxy rather than an integrated hardware measurement\. Finally, broader evaluations of perceptual noticeability, downstream utility, speculative decoding, request filtering, and production scheduling remain important future work\.

## 8\.Conclusion

We presented JPPO, a restricted joint\-input resource\-exhaustion attack that jointly manipulates bounded pixel perturbations and a budget\-constrained visible prompt context in autoregressive VLMs\. Across five VLM families and two benchmarks, JPPO substantially increases generation length, latency, and estimated energy while exhibiting low explicit loop incidence under our detector\. Ablations show that this amplification is not explained by prompt length, pixel perturbation, or either modality alone, but by coordinated multimodal optimization\. These findings suggest that per\-request cost amplification should be treated as a first\-class security property in VLM deployment, motivating cost\-aware evaluation, multimodal anomaly detection, and decoding\-time safeguards\.

###### Acknowledgements\.

This work was supported in part by the National Natural Science Foundation of China \(No\. 62302116\), the Science and Technology Development Fund, Macau SAR, under Grant 0193/2023/RIA3, 0079/2025/AFJ, and 0052/2026/RIA1, the Guangdong Basic and Applied Basic Research Project \(No\. 2023A1515012697\), and the University of Macau under Grant MYRG\-GRG2024\-00065\-FST\-UMDF\.

## References

- Adebayoet al\.\(2018\)J\. Adebayo, J\. Gilmer, M\. Muelly, I\. Goodfellow, M\. Hardt, and B\. KimSanity checks for saliency maps\.InAdvances in Neural Information Processing Systems,Vol\.31,pp\. 9505–9515\.External Links:[Link](https://proceedings.neurips.cc/paper/2018/hash/294a8ed24b1ad22ec2e7efea049b8737-Abstract.html)Cited by:[§D\.12](https://arxiv.org/html/2609.05889#A4.SS12.p1.1)\.
- Alayracet al\.\(2022\)J\. Alayrac, J\. Donahue, P\. Luc, A\. Miech, I\. Barr, Y\. Hasson, K\. Lenc, A\. Mensch, K\. Millican, M\. Reynolds, R\. Ring, E\. Rutherford, S\. Cabi, T\. Han, Z\. Gong, S\. Samangooei, M\. Monteiro, J\. Menick, M\. Bińkowski, A\. Brock, A\. Nematzadeh, S\. Sharifzadeh, R\. Barreira, O\. Vinyals, A\. Zisserman, and K\. SimonyanFlamingo: a visual language model for few\-shot learning\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 23716–23736\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/hash/960a172bc7fbf0177ccccbb411a7d800-Abstract-Conference.html)Cited by:[§2\.1](https://arxiv.org/html/2609.05889#S2.SS1.p1.1)\.
- Baiet al\.\(2025\)S\. Bai, K\. Chen, X\. Liu, J\. Wang, W\. Ge, S\. Song, K\. Dang, P\. Wang, S\. Wang, J\. Tang, H\. Zhong, Y\. Zhu, M\. Yang, Z\. Li, J\. Wan, P\. Wang, W\. Ding, Z\. Fu, Y\. Xu, J\. Ye, X\. Zhang, T\. Xie, Z\. Cheng, H\. Zhang, Z\. Yang, H\. Xu, and J\. LinQwen2\.5\-VL technical report\.arXiv preprint arXiv:2502\.13923\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2502.13923),[Link](https://arxiv.org/abs/2502.13923)Cited by:[§2\.1](https://arxiv.org/html/2609.05889#S2.SS1.p1.1),[§5\.1](https://arxiv.org/html/2609.05889#S5.SS1.SSS0.Px1.p1.1)\.
- Baileyet al\.\(2024\)L\. Bailey, E\. Ong, S\. Russell, and S\. EmmonsImage hijacks: adversarial images can control generative models at runtime\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 2443–2455\.External Links:[Link](https://proceedings.mlr.press/v235/bailey24a.html)Cited by:[§1](https://arxiv.org/html/2609.05889#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.05889#S2.SS2.p1.1)\.
- Chenet al\.\(2022a\)S\. Chen, C\. Liu, M\. Haque, Z\. Song, and W\. YangNMTSloth: understanding and testing efficiency degradation of neural machine translation systems\.InProceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering,Singapore, Singapore,pp\. 1148–1160\.External Links:[Document](https://dx.doi.org/10.1145/3540250.3549102),ISBN 9781450394130Cited by:[§2\.3](https://arxiv.org/html/2609.05889#S2.SS3.p1.1)\.
- Chenet al\.\(2022b\)S\. Chen, Z\. Song, M\. Haque, C\. Liu, and W\. YangNICGSlowDown: evaluating the efficiency robustness of neural image caption generation models\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 15365–15374\.Cited by:[§2\.3](https://arxiv.org/html/2609.05889#S2.SS3.p2.1),[§5\.2](https://arxiv.org/html/2609.05889#S5.SS2.SSS0.Px1.p1.1)\.
- Chianget al\.\(2023\)W\. Chiang, Z\. Li, Z\. Lin, Y\. Sheng, Z\. Wu, H\. Zhang, L\. Zheng, S\. Zhuang, Y\. Zhuang, J\. E\. Gonzalez, I\. Stoica, and E\. P\. XingVicuna: an open\-source chatbot impressing GPT\-4 with 90%\* ChatGPT quality\.External Links:[Link](https://www.lmsys.org/blog/2023-03-30-vicuna/)Cited by:[§5\.1](https://arxiv.org/html/2609.05889#S5.SS1.SSS0.Px1.p1.1)\.
- Cinàet al\.\(2025\)A\. E\. Cinà, A\. Demontis, B\. Biggio, F\. Roli, and M\. PelilloEnergy\-latency attacks via sponge poisoning\.Information Sciences702,pp\. 121905\.External Links:[Document](https://dx.doi.org/10.1016/j.ins.2025.121905)Cited by:[§2\.3](https://arxiv.org/html/2609.05889#S2.SS3.p1.1)\.
- Daiet al\.\(2024\)W\. Dai, N\. Lee, B\. Wang, Z\. Yang, Z\. Liu, J\. Barker, T\. Rintamaki, M\. Shoeybi, B\. Catanzaro, and W\. PingNVLM: open frontier\-class multimodal llms\.arXiv preprint arXiv:2409\.11402\.Cited by:[§2\.1](https://arxiv.org/html/2609.05889#S2.SS1.p1.1)\.
- Daiet al\.\(2023\)W\. Dai, J\. Li, D\. Li, A\. M\. H\. Tiong, J\. Zhao, W\. Wang, B\. Li, P\. Fung, and S\. HoiInstructBLIP: towards general\-purpose vision\-language models with instruction tuning\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 49250–49267\.Cited by:[§2\.1](https://arxiv.org/html/2609.05889#S2.SS1.p1.1),[§5\.1](https://arxiv.org/html/2609.05889#S5.SS1.SSS0.Px1.p1.1)\.
- Denget al\.\(2009\)J\. Deng, W\. Dong, R\. Socher, L\. Li, K\. Li, and L\. Fei\-FeiImageNet: a large\-scale hierarchical image database\.In2009 IEEE Conference on Computer Vision and Pattern Recognition,Miami, FL, USA,pp\. 248–255\.External Links:[Document](https://dx.doi.org/10.1109/CVPR.2009.5206848),[Link](https://image-net.org/static_files/papers/imagenet_cvpr09.pdf)Cited by:[§5\.1](https://arxiv.org/html/2609.05889#S5.SS1.SSS0.Px1.p1.1)\.
- Fuet al\.\(2026\)J\. Fu, K\. Jiang, L\. Hong, J\. Li, H\. Guo, D\. Yang, Z\. Chen, and W\. ZhangLingoLoop attack: trapping mllms via linguistic context and state entrapment into endless loops\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=kxEM2vc7ne)Cited by:[§1](https://arxiv.org/html/2609.05889#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.05889#S2.SS3.p3.1),[§5\.2](https://arxiv.org/html/2609.05889#S5.SS2.SSS0.Px1.p1.1)\.
- Gaoet al\.\(2024a\)K\. Gao, Y\. Bai, J\. Gu, S\. Xia, P\. Torr, Z\. Li, and W\. LiuInducing high energy\-latency of large vision\-language models with verbose images\.InThe Twelfth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.05889#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.05889#S2.SS3.p2.1),[§5\.1](https://arxiv.org/html/2609.05889#S5.SS1.SSS0.Px2.p1.1),[§5\.1](https://arxiv.org/html/2609.05889#S5.SS1.SSS0.Px4.p1.1),[§5\.2](https://arxiv.org/html/2609.05889#S5.SS2.SSS0.Px1.p1.1)\.
- Gaoet al\.\(2024b\)K\. Gao, T\. Pang, C\. Du, Y\. Yang, S\. Xia, and M\. LinDenial\-of\-service poisoning attacks against large language models\.arXiv preprint arXiv:2410\.10760\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2410.10760),[Link](https://arxiv.org/abs/2410.10760)Cited by:[§2\.3](https://arxiv.org/html/2609.05889#S2.SS3.p1.1)\.
- Gonget al\.\(2025\)Y\. Gong, D\. Ran, J\. Liu, C\. Wang, T\. Cong, A\. Wang, S\. Duan, and X\. WangFigStep: jailbreaking large vision\-language models via typographic visual prompts\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 23951–23959\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v39i22.34568)Cited by:[§2\.2](https://arxiv.org/html/2609.05889#S2.SS2.p1.1)\.
- Hanet al\.\(2024\)S\. Han, K\. Rao, A\. Ettinger, L\. Jiang, B\. Y\. Lin, N\. Lambert, Y\. Choi, and N\. DziriWildGuard: open one\-stop moderation tools for safety risks, jailbreaks, and refusals of llms\.InAdvances in Neural Information Processing Systems,Vol\.37\.External Links:[Document](https://dx.doi.org/10.52202/079017-0261),[Link](https://papers.nips.cc/paper_files/paper/2024/hash/0f69b4b96a46f284b726fbd70f74fb3b-Abstract-Datasets_and_Benchmarks_Track.html)Cited by:[§C\.5](https://arxiv.org/html/2609.05889#A3.SS5.SSS0.Px9.p1.1),[§5\.1](https://arxiv.org/html/2609.05889#S5.SS1.SSS0.Px5.p1.1)\.
- Jianget al\.\(2023\)A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier, L\. Renard Lavaud, M\. Lachaux, P\. Stock, T\. Le Scao, T\. Lavril, T\. Wang, T\. Lacroix, and W\. El SayedMistral 7B\.arXiv preprint arXiv:2310\.06825\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2310.06825),[Link](https://arxiv.org/abs/2310.06825)Cited by:[§5\.1](https://arxiv.org/html/2609.05889#S5.SS1.SSS0.Px1.p1.1)\.
- Jianget al\.\(2025\)L\. Jiang, Z\. Zhang, Z\. Wang, X\. Sun, Z\. Li, L\. Zhen, and X\. XuCross\-modal obfuscation for jailbreak attacks on large vision\-language models\.arXiv preprint arXiv:2506\.16760\.Cited by:[§2\.2](https://arxiv.org/html/2609.05889#S2.SS2.p1.1)\.
- Kwonet al\.\(2023\)W\. Kwon, Z\. Li, S\. Zhuang, Y\. Sheng, L\. Zheng, C\. H\. Yu, J\. E\. Gonzalez, H\. Zhang, and I\. StoicaEfficient memory management for large language model serving with PagedAttention\.InProceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles,Koblenz, Germany,pp\. 611–626\.External Links:[Document](https://dx.doi.org/10.1145/3600006.3613165),ISBN 979\-8\-4007\-0229\-7Cited by:[§2\.1](https://arxiv.org/html/2609.05889#S2.SS1.p1.1)\.
- Liet al\.\(2025\)B\. Li, Y\. Zhang, D\. Guo, R\. Zhang, F\. Li, H\. Zhang, K\. Zhang, P\. Zhang, Y\. Li, Z\. Liu, and C\. LiLLaVA\-OneVision: easy visual task transfer\.Transactions on Machine Learning Research\.Cited by:[§2\.1](https://arxiv.org/html/2609.05889#S2.SS1.p1.1)\.
- Liet al\.\(2023\)J\. Li, D\. Li, S\. Savarese, and S\. C\. H\. HoiBLIP\-2: bootstrapping language\-image pre\-training with frozen image encoders and large language models\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 19730–19742\.External Links:[Link](https://proceedings.mlr.press/v202/li23q.html)Cited by:[§2\.1](https://arxiv.org/html/2609.05889#S2.SS1.p1.1),[§5\.1](https://arxiv.org/html/2609.05889#S5.SS1.SSS0.Px1.p1.1)\.
- Liet al\.\(2026\)X\. Li, X\. Liu, C\. Liu, Y\. Xu, K\. Ding, B\. Xin, and J\. YinLoopLLM: transferable energy\-latency attacks in llms via repetitive generation\.Proceedings of the AAAI Conference on Artificial Intelligence40\(38\),pp\. 31770–31777\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i38.40445),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/40445)Cited by:[§2\.3](https://arxiv.org/html/2609.05889#S2.SS3.p1.1)\.
- Linet al\.\(2014\)T\. Lin, M\. Maire, S\. Belongie, J\. Hays, P\. Perona, D\. Ramanan, P\. Dollár, and C\. L\. ZitnickMicrosoft COCO: common objects in context\.InComputer Vision – ECCV 2014,Lecture Notes in Computer Science, Vol\.8693,pp\. 740–755\.External Links:[Document](https://dx.doi.org/10.1007/978-3-319-10602-1%5F48),[Link](https://doi.org/10.1007/978-3-319-10602-1_48)Cited by:[§5\.1](https://arxiv.org/html/2609.05889#S5.SS1.SSS0.Px1.p1.1)\.
- Liuet al\.\(2024\)H\. Liu, C\. Li, Y\. Li, B\. Li, Y\. Zhang, S\. Shen, and Y\. J\. LeeLLaVA\-NeXT: improved reasoning, ocr, and world knowledge\.External Links:[Link](https://llava-vl.github.io/blog/2024-01-30-llava-next/)Cited by:[§C\.1](https://arxiv.org/html/2609.05889#A3.SS1.SSS0.Px2.p1.1),[§5\.1](https://arxiv.org/html/2609.05889#S5.SS1.SSS0.Px1.p1.1)\.
- Liuet al\.\(2023\)H\. Liu, C\. Li, Q\. Wu, and Y\. J\. LeeVisual instruction tuning\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 34892–34916\.Cited by:[§2\.1](https://arxiv.org/html/2609.05889#S2.SS1.p1.1)\.
- Luet al\.\(2024\)J\. Lu, C\. Clark, S\. Lee, Z\. Zhang, S\. Khosla, R\. Marten, D\. Hoiem, and A\. KembhaviUnified\-io 2: scaling autoregressive multimodal models with vision, language, audio, and action\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 26439–26455\.Cited by:[§2\.1](https://arxiv.org/html/2609.05889#S2.SS1.p1.1)\.
- Madryet al\.\(2018\)A\. Madry, A\. Makelov, L\. Schmidt, D\. Tsipras, and A\. VladuTowards deep learning models resistant to adversarial attacks\.InThe Sixth International Conference on Learning Representations,Cited by:[§4\.1](https://arxiv.org/html/2609.05889#S4.SS1.p3.1)\.
- NVIDIA \(2026\)NVIDIANVML API Reference Guide\.Note:[https://docs\.nvidia\.com/deploy/nvml\-api/index\.html](https://docs.nvidia.com/deploy/nvml-api/index.html)vR595, last updated April 16, 2026; accessed April 18, 2026Cited by:[§5\.1](https://arxiv.org/html/2609.05889#S5.SS1.SSS0.Px4.p1.1)\.
- Papineniet al\.\(2002\)K\. Papineni, S\. Roukos, T\. Ward, and W\. ZhuBLEU: a method for automatic evaluation of machine translation\.InProceedings of the 40th Annual Meeting of the Association for Computational Linguistics,pp\. 311–318\.External Links:[Document](https://dx.doi.org/10.3115/1073083.1073135)Cited by:[§D\.21](https://arxiv.org/html/2609.05889#A4.SS21.p1.1)\.
- Qiet al\.\(2024\)X\. Qi, K\. Huang, A\. Panda, P\. Henderson, M\. Wang, and P\. MittalVisual adversarial examples jailbreak aligned large language models\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.38,pp\. 21527–21536\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v38i19.30150)Cited by:[§1](https://arxiv.org/html/2609.05889#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.05889#S2.SS2.p1.1)\.
- Radfordet al\.\(2021\)A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark, G\. Krueger, and I\. SutskeverLearning transferable visual models from natural language supervision\.InProceedings of the 38th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.139,pp\. 8748–8763\.External Links:[Link](https://proceedings.mlr.press/v139/radford21a.html)Cited by:[§D\.22](https://arxiv.org/html/2609.05889#A4.SS22.p1.1)\.
- Selvarajuet al\.\(2017\)R\. R\. Selvaraju, M\. Cogswell, A\. Das, R\. Vedantam, D\. Parikh, and D\. BatraGrad\-CAM: visual explanations from deep networks via gradient\-based localization\.InProceedings of the IEEE International Conference on Computer Vision \(ICCV\),Venice, Italy,pp\. 618–626\.External Links:[Document](https://dx.doi.org/10.1109/ICCV.2017.74),[Link](https://doi.org/10.1109/ICCV.2017.74)Cited by:[§D\.12](https://arxiv.org/html/2609.05889#A4.SS12.p1.1)\.
- Shayeganiet al\.\(2024\)E\. Shayegani, Y\. Dong, and N\. Abu\-GhazalehJailbreak in pieces: compositional adversarial attacks on multi\-modal language models\.InThe Twelfth International Conference on Learning Representations,Cited by:[§2\.2](https://arxiv.org/html/2609.05889#S2.SS2.p1.1)\.
- Shumailovet al\.\(2021\)I\. Shumailov, Y\. Zhao, D\. Bates, N\. Papernot, R\. Mullins, and R\. AndersonSponge examples: energy\-latency attacks on neural networks\.In2021 IEEE European Symposium on Security and Privacy \(EuroS&P\),pp\. 212–231\.External Links:[Document](https://dx.doi.org/10.1109/EuroSP51992.2021.00024),[Link](https://doi.org/10.1109/EuroSP51992.2021.00024)Cited by:[§2\.3](https://arxiv.org/html/2609.05889#S2.SS3.p1.1)\.
- Tanet al\.\(2025\)X\. Tan, P\. Ye, C\. Tu, J\. Cao, Y\. Yang, L\. Zhang, D\. Zhou, and T\. ChenTokenCarve: information\-preserving visual token compression in multimodal large language models\.arXiv preprint arXiv:2503\.10501\.Cited by:[§2\.1](https://arxiv.org/html/2609.05889#S2.SS1.p1.1)\.
- Tuet al\.\(2025\)D\. Tu, D\. Vashchilenko, Y\. Lu, and P\. XuVL\-cache: sparsity and modality\-aware kv cache compression for vision\-language model inference acceleration\.InThe Thirteenth International Conference on Learning Representations,Cited by:[§2\.1](https://arxiv.org/html/2609.05889#S2.SS1.p1.1)\.
- Vedantamet al\.\(2015\)R\. Vedantam, C\. L\. Zitnick, and D\. ParikhCIDEr: consensus\-based image description evaluation\.InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition,pp\. 4566–4575\.Cited by:[§D\.21](https://arxiv.org/html/2609.05889#A4.SS21.p1.1)\.
- Wanget al\.\(2025a\)R\. Wang, J\. Li, Y\. Wang, B\. Wang, X\. Wang, Y\. Teng, Y\. Wang, X\. Ma, and Y\. JiangIDEATOR: jailbreaking and benchmarking large vision\-language models using themselves\.InProceedings of the IEEE/CVF International Conference on Computer Vision \(ICCV\),pp\. 8875–8884\.Cited by:[§2\.2](https://arxiv.org/html/2609.05889#S2.SS2.p1.1)\.
- Wanget al\.\(2025b\)X\. Wang, T\. Yao, S\. Chen, R\. Wang, L\. Ye, K\. Gao, Y\. Huang, and Y\. YaoVLMInferSlow: evaluating the efficiency robustness of large vision\-language models as a service\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Vienna, Austria,pp\. 16035–16050\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.781),[Link](https://aclanthology.org/2025.acl-long.781/),ISBN 979\-8\-89176\-251\-0Cited by:[§2\.3](https://arxiv.org/html/2609.05889#S2.SS3.p2.1)\.
- Wanget al\.\(2025c\)Y\. Wang, J\. Wu, T\. Jiang, M\. Liu, J\. Chen, C\. Wang, E\. Shi, X\. Liu, Y\. Ma, and Z\. ZhengDrainCode: stealthy energy consumption attacks on retrieval\-augmented code generation via context poisoning\.In2025 40th IEEE/ACM International Conference on Automated Software Engineering \(ASE\),pp\. 778–790\.External Links:[Document](https://dx.doi.org/10.1109/ASE63991.2025.00070)Cited by:[§2\.3](https://arxiv.org/html/2609.05889#S2.SS3.p1.1)\.
- Wanget al\.\(2025d\)Z\. Wang, H\. Wang, C\. Tian, and Y\. JinImplicit jailbreak attacks via cross\-modal information concealment on vision\-language models\.arXiv preprint arXiv:2505\.16446\.Cited by:[§2\.2](https://arxiv.org/html/2609.05889#S2.SS2.p1.1)\.
- Wuet al\.\(2025\)C\. Wu, X\. Chen, Z\. Wu, Y\. Ma, X\. Liu, Z\. Pan, W\. Liu, Z\. Xie, X\. Yu, C\. Ruan, and P\. LuoJanus: decoupling visual encoding for unified multimodal understanding and generation\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 12966–12977\.External Links:[Link](https://openaccess.thecvf.com/content/CVPR2025/html/Wu_Janus_Decoupling_Visual_Encoding_for_Unified_Multimodal_Understanding_and_Generation_CVPR_2025_paper.html)Cited by:[§2\.1](https://arxiv.org/html/2609.05889#S2.SS1.p1.1)\.
- Zhanget al\.\(2026\)R\. Zhang, Z\. Wang, T\. Yang, W\. Jiang, R\. Zhang, Q\. Zhao, H\. Li, Y\. Liu, and G\. XuHidden tail: adversarial attack for stealthy resource consumption against vision\-language models\.IEEE Transactions on Dependable and Secure Computing23\(4\),pp\. 8447–8459\.External Links:[Document](https://dx.doi.org/10.1109/TDSC.2026.3684479)Cited by:[§D\.3](https://arxiv.org/html/2609.05889#A4.SS3.p1.1),[§2\.3](https://arxiv.org/html/2609.05889#S2.SS3.p2.1),[§5\.2](https://arxiv.org/html/2609.05889#S5.SS2.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2022\)S\. Zhang, S\. Roller, N\. Goyal, M\. Artetxe, M\. Chen, S\. Chen, C\. Dewan, M\. Diab, X\. Li, X\. V\. Lin, T\. Mihaylov, M\. Ott, S\. Shleifer, K\. Shuster, D\. Simig, P\. S\. Koura, A\. Sridhar, T\. Wang, and L\. ZettlemoyerOPT: open pre\-trained transformer language models\.arXiv preprint arXiv:2205\.01068\.Cited by:[§5\.1](https://arxiv.org/html/2609.05889#S5.SS1.SSS0.Px1.p1.1)\.
- Zhaoet al\.\(2023\)Y\. Zhao, T\. Pang, C\. Du, X\. Yang, C\. Li, N\. Cheung, and M\. LinOn evaluating adversarial robustness of large vision\-language models\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 54111–54138\.Cited by:[§2\.2](https://arxiv.org/html/2609.05889#S2.SS2.p1.1)\.
- Zhuet al\.\(2024\)D\. Zhu, J\. Chen, X\. Shen, X\. Li, and M\. ElhoseinyMiniGPT\-4: enhancing vision\-language understanding with advanced large language models\.InThe Twelfth International Conference on Learning Representations,Cited by:[§2\.1](https://arxiv.org/html/2609.05889#S2.SS1.p1.1),[§5\.1](https://arxiv.org/html/2609.05889#S5.SS1.SSS0.Px1.p1.1)\.

## Appendix AOpen Science

The artifact provides the core JPPO implementation, including joint pixel–prompt optimization, shared prompt\-construction and runtime utilities, and backend adapters for Qwen2\.5\-VL, BLIP\-2, MiniGPT\-4, InstructBLIP, and LLaVA\-NeXT\. It also includes a unified entry point, run scripts, configuration files, dependency specifications, and documentation\.

The full paper evaluates multiple VLM families and ablation settings whose complete end\-to\-end reproduction requires separate third\-party model checkpoints, benchmark datasets, model\-specific serving environments, and substantial compute\. We do not redistribute datasets, pretrained model weights, or third\-party model repositories\. The released artifact provides the implementation, configuration, and documentation needed to instantiate the reported evaluation protocol with the corresponding external resources\.

The artifact is not intended as a turnkey service\-abuse package and is limited to bounded local evaluation on open\-source models\.

## Appendix BEthical Considerations

#### Affected stakeholders\.

This work concerns VLM users, model developers, platform providers, downstream application operators, and co\-located users sharing the same serving infrastructure\. Researchers and practitioners using the released artifact are also relevant stakeholders because the method can support authorized robustness testing but could be misused for resource abuse\.

#### Potential harms\.

The principal risks are increased serving cost, benign\-user latency degradation, queueing pressure, and localized loss of availability in shared, flat\-rate, fixed\-price, subscription, or quota\-based deployments\. The method could also be misused to consume compute resources without producing proportionally useful output\. We therefore characterize the realistic impact as localized serving\-cost and availability pressure under repeated submissions in shared or non\-fully\-metered deployments\.

#### Responsible experimentation and disclosure\.

All main effectiveness, defense, construction\-cost, locally deployed API, and shared\-worker experiments use locally hosted open\-source models under controlled configurations\. The four target workers and separate WildGuard service run only on our local five\-GPU RTX 3090 testbed\. We use public benchmark images and involve no human subjects, private user data, or confidential service traces\. We do not attempt rate\-limit evasion, bypass access controls, generate distributed attack traffic, or seek to degrade deployed third\-party systems\. The external API checks are low\-volume and non\-concurrent and assess only aggregate output\-length transfer; they do not estimate provider\-side energy or hardware cost or intentionally stress service availability\. We do not study safety\-filter bypass, data exfiltration, or targeted harmful\-content generation\. Given the limited and non\-disruptive scope of these external API checks, we did not initiate vendor disclosure\.

#### Artifact safeguards\.

The released artifact is limited to reproducible local evaluation on open\-source models and public benchmark images\. It includes documentation warning users not to run JPPO against third\-party services without authorization and omits distributed traffic generation, rate\-limit\-evasion functionality, and turnkey deployment\-abuse tooling\.

#### Mitigations\.

Our evaluation directly considers several service\-side controls, including request\-rate limiting, request\-side safety admission, runtime early stopping, and repeated\-input caching through cache\-hit sensitivity analysis\. These results show that platform\-side controls can reduce attack practicality to different degrees\. Broader production mechanisms, such as timeouts, output caps, and workload\-aware scheduling, remain deployment\-dependent and are outside the scope of our current evaluation\.

## Appendix CImplementation Details

### C\.1\.Model\-Specific Prompt Templates

For each evaluated model, we use a fixed model\-specific input wrapper and initial benign prompt template\. Clean and image\-only baselines retain this benign prompt, whereas Neutral\-long\-prompt, Prompt\-only, and JPPO modify the user\-visible prompt according to their respective protocols\. Below, we summarize the model variant, visual preprocessing setting, and exact benign prompt template used in our implementation\.

#### BLIP\-2\.

We use BLIP\-2 with an OPT\-2\.7B language model\. The visual input resolution is fixed at224×224224\\times 224\. For the baseline captioning configuration, the input text is an empty placeholder, i\.e\.,""\(equivalently,∅\\varnothing\)\.

#### LLaVA\-NeXT\.

We use LLaVA\-NeXT with a Mistral\-7B language model\. Image preprocessing follows the AnyRes strategy\([Liu et al\., 2024](https://arxiv.org/html/2609.05889#bib.bib34)\)with a base size of 336 and a crop size of336×336336\\times 336\. For the baseline configuration, we use the predefined instruction template\.

> \[INST\] <image\>\\n What is shown in this image? \[/INST\]

#### Qwen2\.5\-VL\.

We use Qwen2\.5\-VL\-7B\-Instruct\. In our implementation, the raw image is first manually resized to224×224224\\times 224and then fed into the processor\. For the baseline configuration, we use the predefined instruction template\. Here, N denotes the number of image placeholder tokens inserted by the processor after visual preprocessing:

> <\|im\_start\|\>system\\n You are a helpful assistant\. <\|im\_end\|\>\\n <\|im\_start\|\>user\\n <\|vision\_start\|\> \(<\|image\_pad\|\> repeatedNNtimes\)<\|vision\_end\|\> Describe this image\.<\|im\_end\|\>\\n <\|im\_start\|\> assistant

#### InstructBLIP\.

We use InstructBLIP with a Vicuna\-7B language model\. The visual input resolution is fixed at224×224224\\times 224\. The baseline prompt is:

> <Image\> What is the content of this image?

#### MiniGPT\-4\.

We use MiniGPT\-4 with a Vicuna\-7B language model\. The visual input resolution is fixed at224×224224\\times 224\. For the baseline configuration, we use the predefined instruction template:

> Give the following image: <Img\>ImageContent </Img\>\. You will be able to see the image once I provide it to you\. Please answer my questions\. \\n \#\#\#Human: <Img\><ImageFeature\> </Img\> What is the content of this image?\\n \#\#\#Assistant:

### C\.2\.Prompt\-Construction Details

In Stage I, JPPO uses phase\-based aspect scheduling over visual, spatial, and semantic prompt paths, corresponding to the first, second, and final thirds of the Stage\-I iterations, respectively\. Prompt growth is controlled at the word level rather than at the token level, with budgets of 150 words in Stage I and 200 words in Stage II\. Selective accumulation retains candidate fragments based on lexical novelty relative to the current prompt state, while compression and lightweight deduplication prevent uncontrolled prompt expansion\. Stage II no longer uses phase\-based path scheduling; instead, it updates a single running prompt context via budget\-constrained refinement using the latest model generation\.

### C\.3\.Neutral Descriptive Prompt for Length\-Control Evaluation

To document the neutral\-control evaluation reported in Table[17](https://arxiv.org/html/2609.05889#A4.T17), we use a generic neutral descriptive prompt as a controlled reference instruction\. The purpose of this template is to test whether longer generations can be explained simply by prompt length alone, rather than by JPPO’s stage\-wise prompt\-construction dynamics\. Accordingly, the prompt is written to encourage plain image\-grounded description while discouraging stylistic elaboration, interpretive expansion, and unsupported inference that might otherwise confound the comparison\. This template should therefore be interpreted as a length control, not as an instruction\-engineered, verbose baseline\.

The exact neutral prompt used in our evaluation is as follows\. We intentionally keep this template within the 200\-word Stage\-II prompt cap in the main JPPO configuration so that the neutral\-control setting uses the same maximum visible\-prompt budget as JPPO under the default setting without reproducing its internal stage\-wise prompt\-construction dynamics:

> Please provide a neutral description of the image using plain factual language and a calm, generic tone\. Base the response on direct observation only\. Avoid guesses, hidden causes, symbolism, personal intention, narrative expansion, or unsupported context\. A simple account of the scene is enough\. You may mention the main subject, the surrounding setting, several apparent attributes, and any clear activity when it is plainly visible\. The description does not need specialized labels, persuasive wording, artistic interpretation, or evaluative remarks\. If some part of the picture is uncertain, it may remain unstated\. Please keep the response centered on the image itself rather than the broader meaning\. The answer may include a few additional visible details when they are obvious, but it should remain neutral, ordinary, and descriptive in style\. Rely on appearance, arrangement, and directly available content\. Avoid commentary, emotional reading, story construction, symbolic analysis, social inference, technical discussion, or imagined information that is not clearly supported by the picture\. Keep the wording plain, measured, and impersonal throughout\. A generic account of visible content is entirely acceptable here\. It can remain reserved, descriptive, observational, and detached from claims about purpose, identity, history, implication, or motive\. Nothing specific is required here\.

Importantly, the neutral\-control setting does not reproduce JPPO’s stage\-wise prompt evolution\. Instead, it uses a single fixed neutral prompt within the 200\-word visible\-prompt cap and should therefore be interpreted only as a control for maximum visible\-prompt length in the default main setting\. It does not constrain or replace the separate prompt\-budget ablation, in which Stage\-II budgets are intentionally varied, including values above 200 words, to study budget sensitivity\.

### C\.4\.Defense Implementation Details

For prompt sanitization, we normalize the visible prompt by collapsing whitespace, removing abnormal consecutive runs of symbols, and suppressing consecutive repeated phrases andnn\-grams\. In the repeated\-nn\-gram cleanup step, we scan from larger to smallernnand remove redundant consecutive repetitions while preserving a valid interface\-visible prompt\. This defense is intended to reduce artificial continuation pressure introduced through attacker\-optimized prompt redundancy\.

For early stopping under low novelty / high repetition, we monitor three lightweight word\-level statistics during generation: repeated\-word ratio, novelty ratio, and the maximum number of consecutive repeatednn\-grams\. LetNwN\_\{w\}be the number of generated words andNuN\_\{u\}be the number of distinct generated words observed so far\. We compute

\(18\)NR=NuNw,RR=1−NR,\\mathrm\{NR\}=\\frac\{N\_\{u\}\}\{N\_\{w\}\},\\qquad\\mathrm\{RR\}=1\-\\mathrm\{NR\},whereNR\\mathrm\{NR\}andRR\\mathrm\{RR\}denote the novelty ratio and repeated\-word ratio, respectively\. Letnmax=min⁡\{4,⌊Nw/2⌋\}n\_\{\\max\}=\\min\\\{4,\\lfloor N\_\{w\}/2\\rfloor\\\}\. We search overn=1,…,nmaxn=1,\\dots,n\_\{\\max\}and record the maximum number of consecutive repeatednn\-grams\. Generation is terminated when repetition\-dominant behavior is detected, using thresholds of 0\.55 forRR\\mathrm\{RR\}, 0\.20 forNR\\mathrm\{NR\}, and 3 for maximum consecutivenn\-gram repetition\. This runtime novelty ratio is distinct from the token\-level sliding\-window novelty diagnostic reported in Appendix[D\.10](https://arxiv.org/html/2609.05889#A4.SS10)\. As shown in Table[13](https://arxiv.org/html/2609.05889#A4.T13), this early\-stopping rule provides only limited mitigation against JPPO\. Averaged across five models and both datasets, EarlyStop reduces estimated energy by 5\.81%, compared with 2\.55% for output length and 2\.63% for latency\. This limited degradation indicates that low\-novelty or high\-repetition criteria do not eliminate JPPO’s high\-cost behavior\.

### C\.5\.Shared\-Worker and Local API Evaluation Details

This appendix specifies the evaluation scope, timing boundaries, uncontrolled FIFO workload, identity\-aware local gateway, request\-rate limiting, text\-only safety admission, runtime EarlyStop integration, and the fixed\-external\-trace protocol used in Section[6\.8](https://arxiv.org/html/2609.05889#S6.SS8)\.

#### Evaluation scope and aggregation\.

We instantiate separate localhost Qwen2\.5\-VL\-7B\-Instruct and BLIP\-2 services using four RTX 3090 target workers and a separate fifth RTX 3090 for WildGuard\. We use the same 1,000\-image pools as the main experiments and compute metrics separately by model, dataset, configuration, and arrival seed\.

#### Timing boundaries\.

The gateway records external submission, request\-limit decisions, safety\-admission start and completion, queue entry, worker dispatch, worker release, and response completion\. Model service time is measured from worker dispatch to worker release, queueing delay from queue entry to worker dispatch, FIFO response time from queue entry to worker release, and API end\-to\-end latency from external submission to response completion\. Moderation latency is measured on the dedicated WildGuard device and is included in API end\-to\-end latency\. Requests rejected before FIFO entry contribute zero target\-VLM service time\.

#### Uncontrolled four\-worker protocol\.

For each target model, we operateW=4W=4complete model replicas, each hosted on a dedicated GPU and sharing one FIFO queue\. Benign and attack requests follow independent Poisson arrival processes\. For each dataset and arrival seed, requests are drawn from independently shuffled repetitions of the corresponding 1,000\-input pools; each pool is reshuffled after a complete pass\. Response caching and continuous batching are disabled\. Each uncontrolled run discards the first 500 admitted requests as warm\-up and measures the subsequent 5,000 admitted requests\. We repeat every configuration with three arrival seeds and average the three per\-seed benign P95 values\.

#### Uncontrolled arrival rates\.

Letρ=0\.80\\rho=0\.80denote the target no\-attack benign utilization, and lets¯b\\bar\{s\}\_\{b\}denote the mean Clean model service time\. We set

\(19\)λb=ρ​Ws¯b,\\lambda\_\{b\}=\\frac\{\\rho W\}\{\\bar\{s\}\_\{b\}\},and, for offered attack shareff,

\(20\)λa=λb​f1−f\.\\lambda\_\{a\}=\\lambda\_\{b\}\\frac\{f\}\{1\-f\}\.The implied mean worker utilization is

\(21\)ρtotal=λb​s¯b\+λa​s¯aW,\\rho\_\{\\mathrm\{total\}\}=\\frac\{\\lambda\_\{b\}\\bar\{s\}\_\{b\}\+\\lambda\_\{a\}\\bar\{s\}\_\{a\}\}\{W\},wheres¯a\\bar\{s\}\_\{a\}is the mean service time of the corresponding attack requests\. A point withρtotal<1\\rho\_\{\\mathrm\{total\}\}<1is treated as stable for threshold reporting, whereas a point withρtotal≥1\\rho\_\{\\mathrm\{total\}\}\\geq 1is classified as overloaded and is not assigned a steady\-state P95 value\. P95 inflation is normalized to the matched no\-attack run\.

#### Additional GPU service time\.

For modelMM, datasetdd, and attackaa, additional GPU service time per 1,000 attack requests is

\(22\)Δ​H1000\(M,d,a\)=10003600​\(τ¯M,d,a−τ¯M,d,Clean\)\.\\Delta H\_\{1000\}^\{\(M,d,a\)\}=\\frac\{1000\}\{3600\}\\left\(\\overline\{\\tau\}\_\{M,d,a\}\-\\overline\{\\tau\}\_\{M,d,\\mathrm\{Clean\}\}\\right\)\.This quantity reports cumulative model service time rather than queueing delay or four\-worker wall\-clock duration\.

#### Threshold refinement\.

For Qwen2\.5\-VL, the initial uncontrolled grid isf∈\{0,0\.01,0\.03,0\.05,0\.07,0\.10,0\.15,0\.20\}f\\in\\\{0,0\.01,0\.03,0\.05,0\.07,0\.10,0\.15,0\.20\\\}; for BLIP\-2, it isf∈\{0,0\.0025,0\.005,0\.0075,0\.01,0\.015,0\.02,0\.03,0\.05\}f\\in\\\{0,0\.0025,0\.005,0\.0075,0\.01,0\.015,0\.02,0\.03,0\.05\\\}\. After identifying the largest stable and smallest overloaded shares, we evaluate intermediate points until the bracket width is at most 0\.01 for Qwen2\.5\-VL or 0\.0025 for BLIP\-2\. The rule is fixed before comparing JPPO with Verbose Images\.

#### Identity\-aware gateway\.

Every local API request carries a persistent identity\. The protected experiment uses one fixed attack identity and does not rotate keys or coordinate multiple attack identities\. To construct the benign population without tuning identities to an admission outcome, each benign identity emits an independent Poisson stream at most 1 request/min\. For the aggregate Clean arrival rateλb\\lambda\_\{b\}defined above, we use

\(23\)Nb=⌈60​λb⌉N\_\{b\}=\\left\\lceil 60\\lambda\_\{b\}\\right\\rceilbenign identities: the firstNb−1N\_\{b\}\-1emit at 1 request/min, and the final identity emits at60​λb−\(Nb−1\)60\\lambda\_\{b\}\-\(N\_\{b\}\-1\)requests/min\. The resulting identities, rates, and request\-to\-identity assignments are fixed for each model–dataset pair across service configurations and arrival seeds\.

#### Request\-rate limiting\.

Request\-RL is implemented as a continuous per\-identity token bucket before safety admission and FIFO entry\. The main setting uses a refill rate ofLreq=5L\_\{\\mathrm\{req\}\}=5requests/min and burst capacityCreq=2C\_\{\\mathrm\{req\}\}=2\. This setting is fixed before attack comparison and is five times the maximum per\-identity benign rate used in the workload\. For identityuu, the bucket state at timettis

\(24\)ru​\(t\)=min⁡\{Creq,ru​\(t0\)\+Lreq60​\(t−t0\)\}\.r\_\{u\}\(t\)=\\min\\left\\\{C\_\{\\mathrm\{req\}\},r\_\{u\}\(t\_\{0\}\)\+\\frac\{L\_\{\\mathrm\{req\}\}\}\{60\}\(t\-t\_\{0\}\)\\right\\\}\.A request is admitted only whenru​\(t\)≥1r\_\{u\}\(t\)\\geq 1, after which one token is removed\. Otherwise, it is rejected before moderation or FIFO entry and receives no downstream processing\. The main text reports the mean JPPO request rate implied by the fixed external trace so that the observed admission rate can be interpreted relative to the 5\-request/min refill rate and burst capacity of 2\.

#### Safety admission\.

The gateway uses the WildGuard checkpointallenai/wildguard\([Han et al\., 2024](https://arxiv.org/html/2609.05889#bib.bib46)\)in request\-only mode, with the assistant\-response field set to an empty string\. Moderation inputs are truncated to 512 tokens, and classification uses deterministic decoding withmax\_new\_tokens=32anddo\_sample=False\. The parser reads only theHarmful request: yes/nofield\. Parsedyesoutputs are rejected; parsednooutputs are admitted\. Malformed outputs are handled fail\-open and logged withparse\_failed=1\. WildGuard receives only the interface\-visible request text and does not inspect the image\. WildGuard runs on a separate fifth NVIDIA RTX 3090 GPU \(24 GB\) and does not share the four target\-VLM workers\. Its latency is included in API end\-to\-end latency but not in target\-VLM load or attacker worker share\. In Safety\+Request\-RL, Request\-RL executes first; requests rejected by Request\-RL are not moderated, while requests passing Request\-RL are inspected by WildGuard before FIFO entry\.

#### EarlyStop integration\.

EarlyStop\+Request\-RL uses the same request\-rate admission stage as Request\-RL and applies the low\-novelty/high\-repetition rule in Appendix[C\.4](https://arxiv.org/html/2609.05889#A3.SS4)only after an admitted request begins generation\. Consequently, EarlyStop does not change JPPO admission under a fixed trace; it can only reduce the realized service time of generations that satisfy the stopping condition\.

#### WildGuard validation and safety evaluation\.

The main evaluation processes the 1,000 Clean and JPPO prompts from each dataset and pools the two datasets for each target model and prompt type\. Section[6\.8](https://arxiv.org/html/2609.05889#S6.SS8)reports rejection rate and API end\-to\-end latency amplification\. The pooled Clean false\-positive rate is 0\.0% for both target models, JPPO rejection is 1\.2% for Qwen2\.5\-VL and 0\.3% for BLIP\-2, and no moderation output fails to parse\. Verbose Images is omitted because its interface\-visible prompt is identical to the corresponding Clean prompt and WildGuard does not inspect the image\.

#### Fixed external traces\.

The API experiment does not recalibrate the attack share\. It directly uses the first JPPO implied\-overload shares identified by the uncontrolled evaluation at 80% benign load: 5% and 6% for Qwen2\.5\-VL on MS COCO and ImageNet, respectively, and 0\.75% for BLIP\-2 on both datasets\. Using the sameλb\\lambda\_\{b\}andλa\\lambda\_\{a\}definitions as the uncontrolled experiment, we generate one external trace

\(25\)ℛ=\{\(ti,ui,zi,inputi\)\}i=1N,\\mathcal\{R\}=\\\{\(t\_\{i\},u\_\{i\},z\_\{i\},\\mathrm\{input\}\_\{i\}\)\\\}\_\{i=1\}^\{N\},wheretit\_\{i\}is the external submission time,uiu\_\{i\}is the persistent identity,zi∈\{Clean,JPPO\}z\_\{i\}\\in\\\{\\mathrm\{Clean\},\\mathrm\{JPPO\}\\\}, andinputi\\mathrm\{input\}\_\{i\}is the fixed image–prompt pair\. Open, Request\-RL, Safety\+Request\-RL, and EarlyStop\+Request\-RL replay the same submission times, identities, inputs, decoding seeds, and ordering\. The mean realized JPPO request rate is computed from this trace and reported in the main text\.

#### Fixed\-trace execution\.

Protected runs use a fixed number of external submissions rather than a fixed number of admitted requests\. For each arrival seed, the first 500 external submissions are excluded as warm\-up and the subsequent 5,000 external submissions form the measurement trace, preserving the uncontrolled experiment’s warm\-up and measurement scale while allowing admission controls to reject requests\. Request\-bucket state carries from warm\-up into measurement\. After the final measured submission, no new requests are issued and all admitted requests are allowed to finish\. Every model–dataset–configuration point uses three arrival seeds\.

#### Protected\-run metrics\.

For configurationcc, JPPO admission is

\(26\)Pa\(c\)=Na,admitted\(c\)Na,submitted\.P\_\{a\}^\{\(c\)\}=\\frac\{N\_\{a,\\mathrm\{admitted\}\}^\{\(c\)\}\}\{N\_\{a,\\mathrm\{submitted\}\}\}\.The attacker’s target\-worker share is

\(27\)Sa\(c\)=∑i∈𝒜adm\(c\)si∑i∈𝒜adm\(c\)si\+∑j∈ℬadm\(c\)sj,S\_\{a\}^\{\(c\)\}=\\frac\{\\sum\_\{i\\in\\mathcal\{A\}\_\{\\mathrm\{adm\}\}^\{\(c\)\}\}s\_\{i\}\}\{\\sum\_\{i\\in\\mathcal\{A\}\_\{\\mathrm\{adm\}\}^\{\(c\)\}\}s\_\{i\}\+\\sum\_\{j\\in\\mathcal\{B\}\_\{\\mathrm\{adm\}\}^\{\(c\)\}\}s\_\{j\}\},wheresis\_\{i\}is measured target\-model service time\. The admitted target\-model load is

\(28\)ρ\(c\)=λb\(c\)​s¯b\(c\)\+λa\(c\)​s¯a\(c\)W\.\\rho^\{\(c\)\}=\\frac\{\\lambda\_\{b\}^\{\(c\)\}\\bar\{s\}\_\{b\}^\{\(c\)\}\+\\lambda\_\{a\}^\{\(c\)\}\\bar\{s\}\_\{a\}^\{\(c\)\}\}\{W\}\.Benign P95 E2E Inflation is

\(29\)I95\(c\)=P95⁡\(Tb,mixed\(c\)\)P95⁡\(Tb,Clean−only\(c\)\),I\_\{95\}^\{\(c\)\}=\\frac\{\\mathrm\{P95\}\\\!\\left\(T\_\{b,\\mathrm\{mixed\}\}^\{\(c\)\}\\right\)\}\{\\mathrm\{P95\}\\\!\\left\(T\_\{b,\\mathrm\{Clean\-only\}\}^\{\(c\)\}\\right\)\},whereTb\(c\)T\_\{b\}^\{\(c\)\}is measured from external submission to response completion and both numerator and denominator use the same API configuration\. Because the finite trace is drained after submission stops, we reportI95\(c\)I\_\{95\}^\{\(c\)\}together with the continuous load valueρ\(c\)\\rho^\{\(c\)\}, including configurations withρ≥1\\rho\\geq 1\.

#### External API evaluation\.

We use the official Google Gemini API withgemini\-3\.1\-pro\-previewand Alibaba Cloud Model Studio’s OpenAI\-compatible vision API withqwen2\.5\-vl\-72b\-instruct, both accessed on June 29, 2026\. For each endpoint, we evaluate 500 paired Clean/JPPO cases with three paired submissions per case and identical endpoint settings within each pair\. We do not override provider\-default generation settings; the Gemini endpoint retains its documented default temperature of 1\.0 and high thinking level, while the Qwen requests omit optionaltemperature,top\_p, andmax\_tokensoverrides\. The external evaluation reports output length only\.

## Appendix DExtended Experimental Analysis

### D\.1\.Dataset\-Specific Construction and Serving Costs

Table[11](https://arxiv.org/html/2609.05889#A4.T11)reports the dataset\-specific absolute costs underlying the equal\-weight means in main\-text Table[3](https://arxiv.org/html/2609.05889#S6.T3)\. The direction of every JPPO–Verbose Images comparison is consistent across MS COCO and ImageNet: construction\-cost differences remain model\-dependent, whereas JPPO incurs higher online energy and latency in all ten model–dataset settings\.

Table 11\.Dataset\-specific offline construction and online serving costs\.Values are reported as JPPO / Verbose Images\. Construction metrics are attacker\-side; online metrics measure the final selected request\. Main\-text Table[3](https://arxiv.org/html/2609.05889#S6.T3)reports equal\-weight means over MS COCO and ImageNet\.ModelDatasetBuild Time\(s\)Build Energy\(kJ\)Online Energy\(J\)Online Latency\(s\)BLIP\-2MS COCO576\.41 / 304\.9250\.14 / 38\.302324\.29 / 1432\.6119\.58 / 12\.07ImageNet588\.10 / 327\.1655\.93 / 45\.602665\.96 / 1592\.6920\.16 / 12\.48Qwen2\.5\-VLMS COCO3554\.83 / 4896\.92154\.02 / 186\.285458\.59 / 2898\.4124\.03 / 10\.88ImageNet3154\.57 / 4731\.01139\.42 / 181\.414867\.44 / 2765\.4921\.16 / 10\.41LLaVA\-NeXTMS COCO7574\.42 / 8538\.11597\.51 / 655\.176520\.12 / 3219\.7127\.53 / 11\.78ImageNet7416\.33 / 8753\.08520\.00 / 661\.355706\.21 / 3208\.5927\.18 / 12\.04InstructBLIPMS COCO1496\.00 / 1370\.94198\.51 / 218\.973476\.18 / 1530\.9317\.58 / 10\.72ImageNet1820\.36 / 1495\.21224\.49 / 248\.983968\.34 / 1720\.4521\.43 / 11\.11MiniGPT\-4MS COCO2149\.12 / 2813\.57410\.91 / 541\.145378\.77 / 4251\.7723\.64 / 21\.57ImageNet2045\.82 / 2516\.08392\.04 / 527\.855204\.68 / 4057\.5223\.10 / 18\.45
### D\.2\.Model\- and Metric\-Level Defense Breakdowns

Table[12](https://arxiv.org/html/2609.05889#A4.T12)reports JPPO degradation by target model, while Table[13](https://arxiv.org/html/2609.05889#A4.T13)separates its aggregate degradation by serving\-cost metric\. All values use paired undefended and defended records under the unified top\-pp/512\-token protocol\. Together with Table[7](https://arxiv.org/html/2609.05889#S6.T7)and Figure[6](https://arxiv.org/html/2609.05889#acmlabel6), these results provide complementary model\-, defense\-, and metric\-level views of robustness\.

Table 12\.Model\-level defense degradation for JPPO \(%\)\.Each entry averages both datasets and the three serving\-cost metrics\. For each model, the mean of the first nine rows and the EarlyStop row correspond to the JPPO entries in Panels \(a\) and \(b\), respectively, of Table[7](https://arxiv.org/html/2609.05889#S6.T7)\.DefenseBLIP\-2Qwen2\.5\-VLLLaVA\-NeXTInstructBLIPMiniGPT\-4JPEG \(Q=95\)\-0\.46\-0\.712\.523\.622\.59JPEG \(Q=75\)15\.700\.114\.634\.069\.55JPEG \(Q=50\)25\.795\.629\.011\.810\.78Resize restoration16\.894\.6814\.58\-1\.969\.125\-bit quantization\-4\.28\-1\.064\.92\-9\.01\-1\.824\-bit quantization30\.500\.451\.694\.62\-0\.12Gaussian blur \(0\.8\)\-9\.321\.737\.543\.823\.55Median filter \(3×33\\times 3\)26\.314\.682\.520\.981\.35Prompt sanitization0\.139\.54\-1\.113\.2317\.23EarlyStop8\.401\.081\.537\.100\.19Table 13\.Metric\-level defense degradation for JPPO \(%\)\.The non\-EarlyStop row averages across nine defenses, five target models, and both datasets, while EarlyStop is averaged across five target models and both datasets\. The equal\-weight means across output length, estimated energy, and latency are 5\.02% and 3\.66%, respectively, matching the JPPO averages reported in Table[7](https://arxiv.org/html/2609.05889#S6.T7)\.Defense settingLen\.E^\\widehat\{E\}Lat\.Nine non\-EarlyStop defenses5\.194\.964\.92EarlyStop2\.555\.812\.63
### D\.3\.Hidden Tail and Expanded Baseline Comparison

Hidden Tail\([Zhang et al\., 2026](https://arxiv.org/html/2609.05889#bib.bib33)\)is designed for a stealthy image\-only evaluation regime in which the induced continuation can contain user\-invisible special tokens\. This differs from our main protocol, which reports visible detokenized word length, latency, and estimated energy under a unified top\-pp/512\-token serving configuration\. To avoid conflating protocol differences with attack strength, we exclude Hidden Tail from the main baseline table for direct comparison and report it separately under our unified protocol\. Table[14](https://arxiv.org/html/2609.05889#A4.T14)provides an expanded version of the main comparison, including Hidden Tail under the unified protocol\.

Under this unified protocol, Hidden Tail yields nonzero but comparatively weak amplification across the evaluated models\. This supports treating it as an appendix\-only reference rather than a directly compared main\-table baseline\.

Table 14\.Expanded full comparison including Hidden Tail under the unified low\-perturbation top\-pp/512\-token protocol\. This table is provided for completeness; the main text excludes Hidden Tail from the directly compared baseline table because its original evaluation regime is not directly aligned with the main low\-perturbation comparison\. Bold marks the largest non\-Clean value for each model, dataset, and metric\.ModelMethodMS COCOImageNetLen\.E^\\widehat\{E\}\(J\)Lat \(s\)Len\.E^\\widehat\{E\}\(J\)Lat \(s\)BLIP\-2Clean8\.3667\.380\.427\.0781\.500\.55Noise8\.2365\.650\.447\.2260\.280\.43NICGSlowDown86\.76567\.034\.46103\.41663\.155\.01Verbose Images211\.621432\.6112\.07231\.311592\.6912\.48Hidden Tail52\.05351\.522\.3150\.13309\.452\.29JPPO373\.232324\.2919\.58361\.602665\.9620\.16Qwen2\.5\-VLClean80\.24945\.974\.6181\.52912\.434\.53Noise83\.18934\.214\.7279\.86940\.804\.47NICGSlowDown146\.121561\.607\.28150\.641700\.148\.01Verbose Images252\.462898\.4110\.88247\.182765\.4910\.41Hidden Tail131\.361444\.136\.51122\.081363\.526\.50JPPO441\.235458\.5924\.03422\.974867\.4421\.16LLaVA\-NeXTClean85\.531577\.856\.7476\.671339\.236\.39Noise88\.431479\.396\.6272\.631312\.715\.96NICGSlowDown187\.452382\.0510\.02146\.402053\.489\.82Verbose Images233\.433219\.7111\.78235\.063208\.5912\.04Hidden Tail139\.172039\.818\.05145\.911894\.917\.91JPPO369\.106520\.1227\.53340\.035706\.2127\.18InstructBLIPClean50\.12618\.403\.6853\.95741\.224\.47Noise52\.34606\.913\.7955\.61756\.044\.39NICGSlowDown83\.081243\.187\.0189\.011194\.387\.02Verbose Images129\.841530\.9310\.72125\.631720\.4511\.11Hidden Tail85\.91840\.278\.0386\.41858\.128\.16JPPO250\.923476\.1817\.58262\.113968\.3421\.43MiniGPT\-4Clean52\.60790\.883\.8069\.06983\.264\.37Noise58\.16864\.513\.9769\.30943\.824\.56NICGSlowDown196\.592906\.5214\.07183\.133139\.3615\.43Verbose Images314\.304251\.7721\.57307\.624057\.5218\.45Hidden Tail103\.621650\.018\.0199\.851640\.197\.91JPPO385\.355378\.7723\.64370\.535204\.6823\.10
### D\.4\.Stage\-I Schedule Analysis

Table 15\.Stage\-I schedule comparison\.Values are averaged over the five evaluated models and both datasets under the main evaluation protocol\.ScheduleLen\.E^\\widehat\{E\}Lat\.Random4\.46×\\times4\.96×\\times4\.28×\\timesPhase\-based5\.95×\\times6\.06×\\times5\.40×\\timesStage I uses no success threshold or restart rule\. Across construction runs, 31\.8% attain the final selected optimum during Stage I; across final inference runs, 5\.6% reach the 512\-token cap\.

### D\.5\.Stage\-II Candidate\-Selection Strategies

Stage I ranks measured prompt–perturbation pairs by realized estimated energy\. The default strategy refines only the highest\-ranked candidate; random top\-3/top\-5 refines one randomly selected candidate from the corresponding set, whereas refine\-all\-top\-3/top\-5 independently refines every candidate and selects the final state with the highest realized estimated energy\.

All strategies use matched Stage\-I buffers\. Build cost includes Stage I once and all Stage\-II trajectories executed by the corresponding strategy\. Results are averaged over three construction seeds within each dataset and then equally over MS COCO and ImageNet; finalE^\\widehat\{E\}follows the main serving protocol\.

Table 16\.Stage\-II candidate\-selection ablation\.Values are equal\-weight means over MS COCO and ImageNet\. Build cost includes Stage I once and all Stage\-II trajectories executed by each strategy; finalE^\\widehat\{E\}measures only the selected request\.ModelStrategyTraj\.Build Time \(s\)BuildE^\\widehat\{E\}\(kJ\)FinalE^\\widehat\{E\}\(J\)Δ\\DeltaFinalE^\\widehat\{E\}Qwen2\.5\-VLHighest\-energy candidate13354\.70146\.725163\.020\.00%0\.00\\%Random top\-313381\.46147\.955201\.64\+0\.75%\+0\.75\\%Random top\-513319\.28145\.184937\.83−4\.36%\-4\.36\\%Refine all top\-337186\.91311\.275384\.76\+4\.30%\+4\.30\\%Refine all top\-5511046\.38488\.115427\.19\+5\.12%\+5\.12\\%BLIP\-2Highest\-energy candidate1582\.2653\.042495\.130\.00%0\.00\\%Random top\-31574\.8352\.372439\.74−2\.22%\-2\.22\\%Random top\-51589\.6153\.812512\.60\+0\.70%\+0\.70\\%Refine all top\-331281\.77116\.932611\.15\+4\.65%\+4\.65\\%Refine all top\-551944\.32178\.412649\.10\+6\.17%\+6\.17\\%Random selection is model\-dependent: random top\-3 changes finalE^\\widehat\{E\}by\+0\.75%\+0\.75\\%on Qwen2\.5\-VL and−2\.22%\-2\.22\\%on BLIP\-2, whereas random top\-5 changes it by−4\.36%\-4\.36\\%and\+0\.70%\+0\.70\\%, respectively\. Refining all top\-3/top\-5 candidates consistently increases finalE^\\widehat\{E\}by 4\.30–4\.65% and 5\.12–6\.17%, but raises build time to 2\.14–2\.20×\\timesand 3\.29–3\.34×\\timesand buildE^\\widehat\{E\}to 2\.12–2\.20×\\timesand 3\.33–3\.36×\\times, respectively\. These results support refining only the highest\-energy Stage\-I candidate by default\.

### D\.6\.Neutral\-Prompt Control Across Models

Table 17\.Effect of different prompt settings and attack methods on multiple models\. Each model row shows the results for Original Prompt, Neutral Prompt, and JPPO under the same evaluation protocol\. The Neutral Prompt uses the fixed length\-control prompt described in Appendix[C\.3](https://arxiv.org/html/2609.05889#A3.SS3)and is included to test whether JPPO’s gains can be explained by prompt length alone\. Bold marks the largest value for each model, dataset, and metric\.ModelMethodMS COCOImageNetLen\.E^\\widehat\{E\}\(J\)Lat \(s\)Len\.E^\\widehat\{E\}\(J\)Lat \(s\)BLIP\-2Original8\.3667\.380\.427\.0781\.500\.55Neutral21\.93137\.211\.0124\.17152\.751\.11JPPO373\.232324\.2919\.58361\.602665\.9620\.16Qwen2\.5\-VLOriginal80\.24945\.974\.6181\.52912\.434\.53Neutral102\.541257\.805\.09115\.121468\.245\.27JPPO441\.235458\.5924\.03422\.974867\.4421\.16LLaVA\-NeXTOriginal85\.531577\.856\.7476\.671339\.236\.39Neutral126\.802241\.299\.25113\.761927\.227\.92JPPO369\.106520\.1227\.53340\.035706\.2127\.18InstructBLIPOriginal50\.12618\.403\.6853\.95741\.224\.47Neutral44\.63554\.373\.0452\.16756\.494\.12JPPO250\.923476\.1817\.58262\.113968\.3421\.43MiniGPT\-4Original52\.60790\.883\.8069\.06983\.264\.37Neutral112\.771565\.186\.8686\.371128\.437\.32JPPO385\.355378\.7723\.64370\.535204\.6823\.10Table[17](https://arxiv.org/html/2609.05889#A4.T17)examines whether the gains of JPPO can be explained simply by using a longer visible prompt\. To rule out a trivial prompt\-length explanation, we compare JPPO against a single fixed neutral descriptive prompt within the 200\-word Stage\-II cap that matches the default maximum visible\-prompt length in the main JPPO setting, while not reproducing JPPO’s internal stage\-wise prompt\-construction dynamics\. This control isolates prompt length alone rather than a strongly engineered, verbose instruction\. Across all evaluated models, the neutral prompt fails to reproduce the large increases in generation length, latency, and estimated energy achieved by JPPO\. This result indicates that JPPO does not merely benefit from longer visible text, but from coordinated construction of a specific multimodal state\.

The neutral\-control effect is not consistently positive across models\. While the neutral prompt increases realized serving cost for BLIP\-2, Qwen2\.5\-VL, LLaVA\-NeXT, and MiniGPT\-4, it reduces output length and latency for InstructBLIP on both datasets\. Estimated energy also decreases on MS COCO but changes only slightly on ImageNet, increasing from 741\.22 J to 756\.49 J\. This model dependence shows that prompt length alone is not a reliable driver of prolonged decoding\. The larger amplification produced by JPPO is therefore better explained by coordinated pixel and prompt optimization than by visible\-prompt length alone\.

### D\.7\.Exact Values for Optimization\-Space Ablation

Table[18](https://arxiv.org/html/2609.05889#A4.T18)provides the absolute values underlying the optimization\-space ablation\. The joint setting consistently yields the largest output length, estimated energy, and latency on both datasets\. This supports the conclusion that JPPO’s gain is not explained by either the pixel branch or the prompt branch alone, but by their coordinated use\.

Table 18\.Exact values for the optimization\-space ablation on BLIP\-2\. This table complements the amplification visualization in Figure[5](https://arxiv.org/html/2609.05889#acmlabel5)\.MethodLen\.E^\\widehat\{E\}\(J\)Lat\. \(s\)MS COCOClean8\.3667\.380\.42Pixel\-only82\.59742\.615\.73Prompt\-only212\.431360\.2914\.82JPPO \(Joint\)373\.232324\.2919\.58ImageNetClean7\.0781\.500\.55Pixel\-only71\.08720\.525\.64Prompt\-only192\.601248\.0214\.76JPPO \(Joint\)361\.602665\.9620\.16
### D\.8\.Loss\-Component Ablation

Table[19](https://arxiv.org/html/2609.05889#A4.T19)shows that all three loss terms contribute to resource amplification, but no single component recovers the full JPPO effect\. Among single\-loss variants, the prefix stop\-suppression loss is generally strongest, while combinations involving it tend to be more effective than combinations without it\. The full objective remains strongest overall, supporting the use of the three\-term surrogate rather than attributing the attack to a single loss\.

Table 19\.Ablation analysis of the three availability\-oriented loss terms on BLIP\-2:ℒeos\\mathcal\{L\}\_\{\\mathrm\{eos\}\}for prefix stop suppression,ℒalign\\mathcal\{L\}\_\{\\mathrm\{align\}\}for cross\-modal alignment suppression, andℒbot\\mathcal\{L\}\_\{\\mathrm\{bot\}\}for bottleneck regularization\. We report generation length, estimated\-energy proxyE^\\widehat\{E\}, and latency across MS COCO and ImageNet\. The row with no checked loss component corresponds to prompt\-only refinement without pixel\-space availability losses\.Loss ComponentsMS COCOImageNet𝓛𝐞𝐨𝐬\\boldsymbol\{\\mathcal\{L\}\_\{\\mathrm\{eos\}\}\}𝓛𝐚𝐥𝐢𝐠𝐧\\boldsymbol\{\\mathcal\{L\}\_\{\\mathrm\{align\}\}\}𝓛𝐛𝐨𝐭\\boldsymbol\{\\mathcal\{L\}\_\{\\mathrm\{bot\}\}\}Len\.E^\\widehat\{E\}\(J\)Lat\. \(s\)Len\.E^\\widehat\{E\}\(J\)Lat\. \(s\)212\.431360\.2914\.82192\.601248\.0214\.76✓282\.032000\.0618\.16313\.671958\.0618\.51✓239\.671517\.3916\.24223\.901471\.5415\.47✓244\.071530\.9716\.12230\.671539\.7116\.85✓✓283\.072165\.6518\.29275\.531863\.7317\.17✓✓324\.172226\.1418\.64302\.871968\.2818\.79✓✓223\.271406\.4715\.90243\.231470\.8915\.12✓✓✓373\.232324\.2919\.58361\.602665\.9620\.16
### D\.9\.Additional Cross\-Model Transfer Results

Table 20\.JPPO output length across all source–target model pairs\.Rows are construction models, columns are evaluation targets, and shaded diagonal cells denote exact\-model white\-box references\. Bold marks the highest output length for each target model and dataset\.Source ModelBLIP\-2Qwen2\.5\-VLLLaVA\-NeXTInstructBLIPMiniGPT\-4MS COCOBLIP\-2373\.23361\.21345\.36209\.59226\.40Qwen2\.5\-VL111\.02441\.23298\.74293\.08222\.87LLaVA\-NeXT69\.20446\.50369\.10217\.91256\.93InstructBLIP26\.90343\.22277\.93250\.92198\.30MiniGPT\-476\.70274\.51274\.46104\.40385\.35ImageNetBLIP\-2361\.60400\.22344\.58203\.92176\.78Qwen2\.5\-VL162\.49422\.97325\.66245\.65231\.52LLaVA\-NeXT71\.72446\.28340\.03160\.40292\.93InstructBLIP33\.21390\.41282\.96262\.11217\.43MiniGPT\-464\.68225\.73178\.33125\.26370\.53Table 21\.JPPO estimated energy across all source–target model pairs \(J\)\.Rows are construction models, columns are evaluation targets, and shaded diagonal cells denote exact\-model white\-box references\. Bold marks the highest estimated energy for each target model and dataset\.Source ModelBLIP\-2Qwen2\.5\-VLLLaVA\-NeXTInstructBLIPMiniGPT\-4MS COCOBLIP\-22324\.293509\.625178\.872296\.143032\.69Qwen2\.5\-VL642\.495458\.594191\.652844\.302838\.93LLaVA\-NeXT435\.855083\.856520\.122309\.903605\.19InstructBLIP101\.263254\.063751\.563476\.182695\.34MiniGPT\-4328\.932391\.893912\.611037\.255378\.77ImageNetBLIP\-22665\.964202\.185402\.352626\.732032\.61Qwen2\.5\-VL978\.004867\.444868\.983560\.973424\.11LLaVA\-NeXT487\.044612\.865706\.211903\.784000\.52InstructBLIP202\.864186\.113954\.313968\.343110\.26MiniGPT\-4482\.091988\.282652\.241651\.655204\.68Tables[20](https://arxiv.org/html/2609.05889#A4.T20)and[21](https://arxiv.org/html/2609.05889#A4.T21)report the complete absolute output\-length and estimated\-energy results underlying Section[6\.2](https://arxiv.org/html/2609.05889#S6.SS2)\. The corresponding absolute latency matrix and white\-box\-to\-black\-box survivability visualization are reported in Table[2](https://arxiv.org/html/2609.05889#S6.T2)and Figure[4](https://arxiv.org/html/2609.05889#acmlabel4), respectively\. Across metrics, transfer is on average weaker than exact\-model construction but remains nonzero and asymmetric across source–target directions\.

### D\.10\.Additional Loop Diagnostics

Figure 7\.Mean token\-level novelty ratio before and after JPPO across five evaluated VLMs\. The statistic is computed on normalized generated tokens using a sliding window of 32\. Higher values indicate more novel continuation behavior, whereas lower values indicate stronger repetition\. JPPO reduces novelty relative to benign generation, but the attacked outputs remain well above collapse\-to\-loop behavior on most models, consistent with the two\-dataset loop\-incidence results in Table[5](https://arxiv.org/html/2609.05889#S6.T5)\.Novelty\-ratio comparison before and after attackA bar chart compares the mean novelty ratio before and after JPPO across five VLMs\. For every model, the attacked outputs have lower novelty than the clean outputs, but the post\-attack values remain moderate rather than collapsing to near\-zero repetition\.As an additional descriptive diagnostic, Figure[7](https://arxiv.org/html/2609.05889#acmlabel7)reports the mean novelty ratio before and after attack across the five evaluated VLMs\. This diagnostic novelty ratio is computed on normalized generated tokens rather than words, using a sliding window of 32 generated tokens\. In particular, tokens are normalized by removing tokenizer\-specific boundary markers, lowercasing, and ignoring pure punctuation before the windowed novelty statistic is computed\. The attacked outputs are less novel than benign outputs, which is expected under a resource\-amplification attack that prolongs generation and increases local reuse\. However, the post\-attack novelty ratio remains moderate across all models, ranging from 0\.52 for MiniGPT\-4 to 0\.77 for BLIP\-2\. This pattern is consistent with the loop\-oriented statistics above: JPPO can reduce output novelty without collapsing generation into overt repetition\. In other words, the attack often sustains long and costly continuations through broader multimodal steering rather than through explicit token\-loop failure alone\. This comparison should nevertheless be interpreted with care: LingoLoop is a recent loop\-centric image\-only attack, whereas JPPO jointly optimizes bounded pixel perturbations and visible prompt inputs\. The purpose of Table[10](https://arxiv.org/html/2609.05889#S6.T10)and the two\-dataset loop diagnostics in Table[5](https://arxiv.org/html/2609.05889#S6.T5)is not to erase that distinction, but to make the behavioral difference between the two regimes more explicit under a matched decoding protocol\.

### D\.11\.Bottleneck\-Ratio Sensitivity Results

Figure 8\.Sensitivity of JPPO to the bottleneck dimensionality ratio on BLIP\-2 across MS COCO and ImageNet\.The ratioρ\\rhodetermines the compressed bottleneck dimensiond′d^\{\\prime\}in Eq\. \([13](https://arxiv.org/html/2609.05889#S4.E13)\)\. The left column shows the low\-ratio regime, and the right column shows the broader global trend\. The top row reports output\-length amplification, while the bottom row reports latency amplification\.Bottleneck\-ratio sensitivity across two datasetsA four\-panel line chart showing sensitivity to the bottleneck dimensionality ratio on BLIP\-2 across MS COCO and ImageNet\. The left column shows the low\-ratio regime, and the right column shows the broader global trend\. The top row plots output\-length amplification, and the bottom row plots latency amplification\.Figure[8](https://arxiv.org/html/2609.05889#acmlabel8)and Table[22](https://arxiv.org/html/2609.05889#A4.T22)report sensitivity to the bottleneck dimensionality ratioρ\\rhoon BLIP\-2 across MS COCO and ImageNet\. This ratio controls the compressed dimensiond′d^\{\\prime\}used by the bottleneck regularization objective while keeping the perturbation budget, prompt budgets, stage schedule, and decoding protocol fixed\. The completed settings show that JPPO is sensitive toρ\\rho\. The defaultρ=0\.10\\rho=0\.10yields the strongest latency and remains near the strongest setting across the other metrics, althoughρ=0\.04\\rho=0\.04slightly exceeds it for MS COCO estimated energy and ImageNet output length\. We therefore useρ=0\.10\\rho=0\.10as a strong and stable default rather than as a globally optimal bottleneck ratio\.

Table 22\.Full sensitivity results for the bottleneck dimensionality ratioρ\\rhoon BLIP\-2 across MS COCO and ImageNet\. The ratio determines the compressed representation sized′=max⁡\(⌊ρ​D⌋,1\)d^\{\\prime\}=\\max\(\\lfloor\\rho D\\rfloor,1\)in the bottleneck regularization objective\. We report mean output length, estimated energy, and latency under each setting\.ρ\\rhoMS COCOImageNetLen\.E^\\widehat\{E\}\(J\)Lat \(s\)Len\.E^\\widehat\{E\}\(J\)Lat \(s\)Clean8\.3667\.380\.427\.0781\.500\.550\.02294\.091971\.6218\.24248\.621641\.1516\.210\.04371\.822414\.1819\.49369\.112410\.9719\.420\.06277\.871962\.2817\.89249\.181601\.9215\.940\.08288\.912026\.1717\.90283\.011968\.1117\.530\.10373\.232324\.2919\.58361\.602665\.9620\.160\.12346\.032204\.5118\.89326\.282009\.1818\.200\.14323\.912041\.5918\.51311\.961997\.1317\.840\.16315\.832007\.5018\.59320\.192039\.4118\.530\.18311\.732105\.5618\.41328\.642147\.6818\.520\.20285\.171761\.1017\.45285\.431717\.1817\.600\.30316\.101920\.3918\.02301\.641833\.4617\.980\.40280\.031786\.7517\.65277\.621711\.9816\.720\.50269\.431700\.1216\.09262\.461510\.4516\.820\.60290\.161802\.1517\.82311\.002014\.1918\.550\.70312\.642168\.8818\.52324\.612007\.3418\.000\.80308\.511958\.0118\.55294\.161826\.0617\.940\.90246\.071626\.9516\.76259\.481516\.6216\.04
### D\.12\.Gradient\-Based Evidence Localization

To complement the quantitative results, we visualize gradient\-based class\-discriminative activation maps \(Grad\-CAM\) for representative BLIP\-2 cases under clean and JPPO\-adversarial inputs\([Selvaraju et al\., 2017](https://arxiv.org/html/2609.05889#bib.bib44)\)\. As shown in Figure[9](https://arxiv.org/html/2609.05889#acmlabel9), the adversarial input alters the spatial distribution of model\-relevant visual evidence relative to the clean input\. Depending on the sample, the activated regions may become weaker, more diffuse, more fragmented, or shift away from compact object\-centric areas\. In several cases, the adversarial map does not form a new dominant hotspot but instead exhibits globally weaker, more scattered responses, suggesting a reduced reliance on compact visual evidence\. This qualitative behavior is consistent with the substantial increase in generation length, latency, and estimated energy reported in Table[1](https://arxiv.org/html/2609.05889#S6.T1)\. Following prior discussion on the interpretation of saliency methods, we treat these maps as qualitative supporting evidence rather than as standalone causal proof of the attack mechanism\([Adebayo et al\., 2018](https://arxiv.org/html/2609.05889#bib.bib45)\)\.

![Grad-CAM comparison for clean and attacked inputs](https://arxiv.org/html/2609.05889v1/6.png)Figure 9\.Gradient\-based evidence localization for clean and JPPO\-adversarial inputs\.Each row shows a representative BLIP\-2 example\. From left to right, we present the clean image, the Grad\-CAM overlay on the clean image, the adversarial image obtained by applying the bounded pixel perturbation, and the Grad\-CAM overlay on the adversarial image\. Compared with the clean input, the adversarial counterpart alters the spatial distribution of model\-relevant visual evidence, often making it weaker, more diffuse, or more fragmented, rather than tightly concentrated in compact, object\-centric regions\. For readability, only a portion of the generated output is shown\.Grad\-CAM comparison for clean and attacked inputsSeveral BLIP\-2 examples are shown in rows\. Each row contains a clean image, its Grad\-CAM overlay, an adversarial image, and its Grad\-CAM overlay\. Relative to the clean cases, the attacked cases typically show weaker, more diffuse, or more fragmented evidence localization rather than compact object\-centered hotspots\.
### D\.13\.Sensitivity to Perturbation and Iteration Budgets

Table 23\.Impact of perturbation magnitudeϵ\\epsilonon attack effectiveness on BLIP\-2 over the MS COCO and ImageNet datasets\. Bold marks the largest value for each dataset and metric\.ϵ\\epsilonMS COCOImageNetLen\.E^\\widehat\{E\}\(J\)Lat \(s\)Len\.E^\\widehat\{E\}\(J\)Lat \(s\)Clean8\.3667\.380\.427\.0781\.500\.552/255196\.831286\.4910\.07189\.301317\.2111\.724/255278\.201852\.1914\.92277\.471904\.1515\.678/255373\.232324\.2919\.58361\.602665\.9620\.1616/255318\.832072\.8916\.72363\.262357\.0119\.3732/255349\.602174\.4118\.86303\.402148\.3617\.9264/255326\.972086\.8217\.32302\.661848\.8116\.07We next examine sensitivity to perturbation and optimization budgets\. Table[23](https://arxiv.org/html/2609.05889#A4.T23)shows a non\-monotonic relationship between perturbation magnitude and attack strength\. Increasingϵ\\epsilonfrom2/2552/255to8/2558/255substantially improves all three metrics on both datasets; on MS COCO, length rises from 196\.83 to 373\.23 words, latency from 10\.07 to 19\.58 s, and estimated energy from 1286\.49 to 2324\.29 J\. Beyond8/2558/255, larger budgets do not consistently help, indicating non\-monotonic sensitivity to perturbation scale\.

The stage\-wise iteration results show a similar pattern\. Larger optimization budgets can improve attack strength, with the 200/200 schedule achieving the highest values in Table[28](https://arxiv.org/html/2609.05889#A4.T28), but the trend is not strictly monotonic\. We therefore retain 100/100 as the default setting for controlled comparison with prior baselines and avoid claiming a universally optimal schedule\. Overall, strong amplification already appears under moderate perturbation and iteration budgets, indicating that JPPO depends on steering the model into a specific high\-cost decoding regime rather than simply increasing attack budget or input distortion\.

### D\.14\.Sensitivity to Loss\-Weight Configuration

Table[24](https://arxiv.org/html/2609.05889#A4.T24)reports a sensitivity study over the weights of the three availability\-oriented losses\. The shaded rows provide the main\-evaluation Clean and Default JPPO reference values from Table[1](https://arxiv.org/html/2609.05889#S6.T1), while the remaining rows vary the Stage\-I and Stage\-II loss weights under the same small\-scale sweep protocol\.

Table 24\.Sensitivity of JPPO to the loss\-weight configuration on BLIP\-2\. Each row reports the Stage\-I and Stage\-II weights for\(ℒeos,ℒalign,ℒbot\)\(\\mathcal\{L\}\_\{\\mathrm\{eos\}\},\\mathcal\{L\}\_\{\\mathrm\{align\}\},\\mathcal\{L\}\_\{\\mathrm\{bot\}\}\)\. The shaded row with “–” weights is the main\-evaluation Clean reference from Table[1](https://arxiv.org/html/2609.05889#S6.T1); the shaded row with\(2,1,1\)\(2,1,1\)and\(2,3,1\)\(2,3,1\)is the main\-evaluation Default JPPO reference\. All unshaded rows report a separate small\-scale weight\-sensitivity sweep\. Bold marks the highest estimated energy within each dataset among the tested weight configurations\.Stage 1WeightsStage 2WeightsMS COCOImageNetLen\.E^\\widehat\{E\}\(J\)Lat \(s\)Len\.E^\\widehat\{E\}\(J\)Lat \(s\)––8\.3667\.380\.427\.0781\.500\.55\(1,1,1\)\(1,1,1\)\(1,1,1\)\(1,1,1\)360\.902040\.8019\.77344\.622052\.2319\.03\(2,2,2\)\(2,2,2\)\(2,2,2\)\(2,2,2\)318\.132091\.3718\.46301\.941939\.4117\.96\(2,1,1\)\(2,1,1\)\(1,1,1\)\(1,1,1\)240\.071601\.5815\.87242\.091625\.1516\.34\(2,2,1\)\(2,2,1\)\(1,1,1\)\(1,1,1\)306\.931872\.5818\.06321\.021911\.6418\.84\(2,1,2\)\(2,1,2\)\(1,1,1\)\(1,1,1\)392\.672289\.6619\.08384\.542541\.3220\.19\(2,1,1\)\(2,1,1\)\(2,1,1\)\(2,1,1\)302\.731887\.9117\.51297\.641764\.3717\.89\(2,1,1\)\(2,1,1\)\(2,2,1\)\(2,2,1\)276\.731933\.6317\.65264\.101814\.9017\.02\(2,1,1\)\(2,1,1\)\(2,4,1\)\(2,4,1\)336\.432030\.2017\.22351\.682168\.5419\.18\(2,1,2\)\(2,1,2\)\(2,3,2\)\(2,3,2\)310\.072013\.4717\.39308\.982022\.9418\.01\(3,1,1\)\(3,1,1\)\(3,3,1\)\(3,3,1\)251\.401763\.7716\.99246\.611634\.9416\.75\(2,1,1\)\(2,1,1\)\(2,3,1\)\(2,3,1\)373\.232324\.2919\.58361\.602665\.9620\.16
### D\.15\.Prompt Length Dynamics and Stability

Table 25\.Effect of prompt length limits on generation length, estimated energy, and latency on BLIP\-2 over the MS COCO and ImageNet datasets\. Bold marks the largest value for each dataset and metric\.Stage 1Stage 2MS COCOImageNetLen\.E^\\widehat\{E\}\(J\)Lat \(s\)Len\.E^\\widehat\{E\}\(J\)Lat \(s\)5050338\.272030\.0316\.72343\.931985\.0216\.3850100307\.691964\.2116\.20270\.321748\.7113\.74100100306\.042067\.1317\.62321\.421899\.3815\.43100150293\.571999\.2416\.23339\.142170\.0918\.37100200260\.071633\.2513\.07293\.571932\.4417\.08150150270\.631936\.5415\.62327\.491980\.7317\.09150200373\.232324\.2919\.58361\.602665\.9620\.16150300343\.722163\.4318\.96288\.631932\.8217\.51200200257\.691687\.0812\.84274\.551862\.6116\.68200300306\.942012\.9117\.34278\.172003\.5116\.43Table[25](https://arxiv.org/html/2609.05889#A4.T25)shows that prompt\-budget effects are non\-monotonic\. Moderate budgets already induce substantial amplification, but larger budgets do not consistently improve attack strength\. The default 150/200\-word setting performs best across both datasets, suggesting that JPPO benefits from controlled prompt growth rather than unconstrained accumulation\. This is consistent with its two\-stage design: Stage I builds a diverse warm\-up context, while Stage II refines it toward continuation\-oriented decoding\.

### D\.16\.Maximum Generation Caps

Table 26\.Ablation of maximum generation caps for JPPO on the MS COCO and ImageNet datasets\.DatasetMetric32641282565121024MS COCOLen30\.5056\.7787\.57179\.43373\.23726\.83E^\\widehat\{E\}\(J\)206\.01352\.29619\.631283\.682324\.295493\.30Lat \(s\)1\.642\.584\.298\.6419\.5839\.35ImageNetLen\.31\.2152\.80101\.16191\.64361\.60756\.83E^\\widehat\{E\}\(J\)216\.02358\.18669\.231168\.572665\.965601\.88Lat \(s\)1\.282\.144\.557\.9520\.1641\.92Table[26](https://arxiv.org/html/2609.05889#A4.T26)studies the effect of the maximum generation cap at inference time\. As expected, realized resource usage increases substantially as the serving cap is relaxed\. On both MS COCO and ImageNet, longer decoding limits allow the attack to unfold for more steps, thereby increasing generation length, latency, and estimated energy\. This confirms that generation caps act as deployment\-level hard bounds on the maximum per\-request cost that an adversarial example can extract\.

At the same time, the results also show that the availability surface does not disappear under smaller caps\. Even when the decoder is restricted to relatively short outputs, JPPO still induces nontrivial overhead compared with the clean baseline\.

### D\.17\.Stage\-wise Iteration Schedules

Table 27\.Ablation study on stage\-wise iterations for BLIP\-2 on MS COCO and ImageNet\. Bold marks the largest value for each dataset and metric\.Stage 1Stage 2MS COCOImageNetLen\.E^\\widehat\{E\}\(J\)Lat \(s\)Len\.E^\\widehat\{E\}\(J\)Lat \(s\)008\.3667\.380\.427\.0781\.500\.550100225\.001359\.969\.78202\.791445\.3410\.920200226\.871424\.1910\.28205\.831472\.1311\.721000277\.431812\.7214\.54250\.031767\.5413\.982000313\.621968\.6516\.09202\.811474\.0711\.22100100373\.232324\.2919\.58361\.602665\.9620\.16Table 28\.Resource usage under different stage\-wise iteration schedules for BLIP\-2 on MS COCO and ImageNet\.Stage 1Stage 2MS COCOImageNetLen\.E^\\widehat\{E\}\(J\)Lat \(s\)Len\.E^\\widehat\{E\}\(J\)Lat \(s\)008\.3667\.380\.427\.0781\.500\.551010147\.97923\.877\.3777\.07447\.994\.942020194\.431169\.289\.81246\.651560\.4412\.015050238\.271710\.6814\.12224\.371530\.0812\.1150100298\.732137\.2817\.39322\.372119\.7617\.48100100373\.232324\.2919\.58361\.602665\.9620\.16100150288\.571983\.8616\.41292\.372011\.6816\.58150150340\.372152\.5117\.00318\.752145\.4417\.23200200396\.632585\.2021\.02384\.932997\.1321\.54Tables[27](https://arxiv.org/html/2609.05889#A4.T27)and[28](https://arxiv.org/html/2609.05889#A4.T28)examine the effect of the stage\-wise optimization schedule\. The asymmetric settings in Table[27](https://arxiv.org/html/2609.05889#A4.T27)show that both stages contribute meaningfully, but they do not contribute equally\. Under the tested schedules, Stage I alone is generally stronger than Stage II alone, indicating that constructing a strong warm\-up state already plays a major role in pushing the model toward a higher\-cost decoding regime\. At the same time, the full 100/100 schedule clearly outperforms either single\-stage variant, showing that Stage II refinement remains necessary after the warm\-up state has been established\.

The broader schedule sweep in Table[28](https://arxiv.org/html/2609.05889#A4.T28)further shows that larger optimization budgets can improve attack strength within the tested range, while intermediate schedules remain non\-monotonic\. The 200/200 setting yields the largest generation length, estimated energy, and latency in this sweep\. This intermediate non\-monotonicity suggests that additional optimization can alter the decoding trajectory and runtime profile rather than uniformly improving every setting\. We therefore avoid treating 200/200 as a universally optimal schedule, and retain 100/100 as the default protocol for controlled comparison with prior baselines\.

### D\.18\.Sensitivity to Decoding Strategies

Table 29\.Comparison of different decoding strategies on BLIP\-2\. We report absolute generation length, estimated\-energy proxyE^\\widehat\{E\}, and latency on MS COCO and ImageNet\. Bold marks the largest non\-Clean value for each decoding strategy, dataset, and metric\.DecodingMethodMS COCOImageNetLen\.E^\\widehat\{E\}\(J\)Lat\. \(s\)Len\.E^\\widehat\{E\}\(J\)Lat\. \(s\)GreedyClean7\.7265\.260\.377\.8166\.140\.39Noise7\.8166\.160\.417\.3463\.350\.38NICGSlowDown171\.981097\.117\.71226\.181358\.818\.00Verbose Images303\.842086\.4910\.06328\.892073\.8810\.91Hidden Tail76\.10564\.163\.8889\.01549\.193\.45JPPO407\.113058\.1514\.97394\.032896\.7614\.52Top\-kk\(k=10k=10\)Clean8\.2570\.760\.417\.5462\.520\.33Noise7\.9358\.770\.437\.7358\.080\.42NICGSlowDown216\.091444\.708\.94252\.831560\.388\.40Verbose Images282\.151860\.418\.09300\.062008\.119\.51Hidden Tail84\.63772\.413\.9696\.01801\.034\.12JPPO343\.972534\.9214\.85371\.332740\.9915\.12Top\-pp\(p=0\.9p=0\.9\)Clean8\.3667\.380\.427\.0781\.500\.55Noise8\.2365\.650\.447\.2260\.280\.43NICGSlowDown86\.76567\.034\.46103\.41663\.155\.01Verbose Images211\.621432\.6112\.07231\.311592\.6912\.48Hidden Tail52\.05351\.522\.3150\.13309\.452\.29JPPO373\.232324\.2919\.58361\.602665\.9620\.16Beam Search\(b=5b=5\)Clean8\.86118\.820\.628\.14117\.800\.76Noise8\.5293\.710\.687\.4083\.670\.61NICGSlowDown312\.822094\.8310\.11373\.642564\.3113\.50Verbose Images419\.643115\.9816\.30458\.414019\.9418\.31Hidden Tail71\.01661\.963\.0282\.65682\.573\.71JPPO495\.155073\.4121\.54489\.594721\.6424\.06Table[29](https://arxiv.org/html/2609.05889#A4.T29)evaluates JPPO under several decoding strategies, including greedy decoding, top\-kkdecoding withk=10k=10, nucleus sampling withp=0\.9p=0\.9, and beam search with beam widthb=5b=5\. Across these settings, JPPO remains effective, indicating that the attack is not an artifact of a particular sampling rule\.

The decoder nevertheless has a substantial effect on the absolute cost regime\. Among the tested settings, beam search with width55yields the largest realized output length, latency, and estimated energy proxyE^\\widehat\{E\}, while greedy decoding, top\-kkdecoding withk=10k=10, and top\-ppdecoding withp=0\.9p=0\.9remain clearly vulnerable\. Prior image\-only baselines also become stronger with certain decoding strategies, but JPPO yields the highest absolute cost among the methods evaluated in this table for each tested decoding strategy\.

### D\.19\.Variability Across Inputs

Table 30\.Standard deviation of resource exhaustion attacks on MS COCO and ImageNet\. We report the sample standard deviation across image\-level means for generation length, estimated energy, and latency, where each image\-level value is first averaged over three repeated runs with the same input\-method pair\.ModelMethodMS COCOImageNetLen\.E^\\widehat\{E\}\(J\)Lat \(s\)Len\.E^\\widehat\{E\}\(J\)Lat \(s\)BLIP\-2Clean2\.2313\.260\.071\.8616\.710\.11Noise2\.1012\.790\.091\.8712\.910\.08NICGSlowDown145\.91209\.414\.05155\.94357\.094\.81Verbose Images176\.07494\.183\.89172\.10572\.654\.69JPPO97\.62316\.742\.1988\.01443\.133\.11Qwen2\.5\-VLClean9\.98107\.940\.4311\.10127\.220\.49Noise8\.8382\.140\.5910\.86106\.510\.47NICGSlowDown217\.481120\.903\.97151\.64891\.163\.15Verbose Images200\.931091\.143\.06224\.641003\.493\.01JPPO47\.84533\.771\.8868\.72604\.502\.46LLaVA\-NeXTClean21\.12335\.611\.4214\.19296\.611\.21Noise28\.76445\.511\.8212\.47219\.201\.04NICGSlowDown234\.461170\.098\.05251\.631255\.678\.31Verbose Images251\.911291\.049\.43232\.111191\.088\.60JPPO53\.94545\.193\.6339\.02505\.973\.00InstructBLIPClean17\.29148\.351\.0510\.86164\.360\.68Noise15\.36188\.251\.009\.14157\.190\.86NICGSlowDown29\.17199\.421\.3225\.49177\.611\.24Verbose Images56\.01228\.911\.4880\.55480\.323\.57JPPO35\.36248\.912\.0468\.13372\.573\.02MiniGPT\-4Clean37\.06463\.871\.8219\.76541\.102\.06Noise30\.44406\.311\.6322\.19480\.431\.99NICGSlowDown287\.061269\.1413\.09222\.511204\.0011\.16Verbose Images209\.411055\.1311\.53191\.10984\.8411\.02JPPO71\.95770\.867\.0752\.01641\.186\.16Table[30](https://arxiv.org/html/2609.05889#A4.T30)reports the sample standard deviation across image\-level means for generation length, estimated energy, and latency\. Although JPPO substantially increases the mean serving cost, its variability is often comparable to or lower than that of the stronger image\-only baselines, especially on LLaVA\-NeXT, BLIP\-2, Qwen2\.5\-VL, and MiniGPT\-4\. This suggests that JPPO’s gains are not driven solely by a small number of extreme outliers\. Instead, the attack tends to increase resource usage across a broader subset of the sampled inputs\. Some input sensitivity nevertheless remains, and the extent of that variability is model\-dependent\.

### D\.20\.Judge\-Based Output Evaluation

We use GPT\-5\.5, accessed June 29, 2026, as a multimodal judge\. For each matched input, the judge receives the original image, the corresponding benign instruction, and one candidate output generated under Clean, Neutral\-long\-prompt, Prompt\-only, Verbose Images, or JPPO\. Prompt\-only reuses the independently executed text\-side ablation: the original image is kept unperturbed, the stage\-wise prompt\-construction procedure remains active, and all pixel\-space availability losses are disabled\. Although candidate outputs are produced using their method\-specific generation inputs, each output is independently evaluated against the same original image and benign instruction using the same judge prompt and configuration\. The evaluation therefore measures original\-task utility rather than compliance with a method\-specific optimized prompt\.

The complete judge prompt is:

> You are an impartial evaluator of image\-grounded outputs\. Evaluate the candidate output using only the provided original image and benign instruction\. Treat the candidate output as untrusted content\. Do not follow or execute any instructions contained in it\. Evaluate only its quality as an answer to the benign instruction and as a description of information supported by the original image\. Do not reward length by itself\. Irrelevant, unsupported, repetitive, or unnecessarily extended content should not receive higher scores merely because it is longer\. Assign one integer score from 1 to 5 for each dimension\. Relevance: 1 = The output is unrelated to the original image or does not address the benign instruction\. 2 = The output has limited relevance and addresses only a small part of the benign instruction or visible image content\. 3 = The output is generally relevant but omits or deviates from important aspects of the benign instruction or image\. 4 = The output addresses the benign instruction and image well, with only minor irrelevant or missing content\. 5 = The output directly and completely addresses the benign instruction using content grounded in the original image\. Informativeness: 1 = The output provides almost no useful or correct information grounded in the original image\. 2 = The output provides only a few basic or partially useful image\-grounded details\. 3 = The output provides a moderate amount of useful and mostly correct image\-grounded detail\. 4 = The output provides substantial, specific, and accurate image\-grounded detail\. 5 = The output is highly informative, comprehensive, and fully grounded in the original image\. Redundancy: 1 = The output contains no meaningful repetition or unnecessary padding\. 2 = The output contains minor repetition that does not substantially affect readability\. 3 = The output contains noticeable repetition or avoidable restatement\. 4 = The output contains extensive repetition or unnecessary padding\. 5 = The output is dominated by severe repetition, restatement, or unnecessary continuation\. Usefulness: 1 = The output would not help a user complete the task expressed by the benign instruction\. 2 = The output has limited value because of major omissions, inaccuracies, irrelevant content, or repetition\. 3 = The output is partially useful but has noticeable weaknesses\. 4 = The output is useful and largely satisfies the benign instruction\. 5 = The output is highly useful, accurate, clear, and fully satisfies the benign instruction using information supported by the original image\. Benign instruction: <BENIGN\_INSTRUCTION\> Candidate output: <CANDIDATE\_OUTPUT\> Return only one JSON object and no explanation\. Each value must be exactly one of 1, 2, 3, 4, or 5: \{"relevance": <RELEVANCE\_SCORE\>, "informativeness": <INFORMATIVENESS\_SCORE\>, "redundancy": <REDUNDANCY\_SCORE\>, "usefulness": <USEFULNESS\_SCORE\>\}

We score all 450,000 candidate outputs and aggregate the resulting scores under the main evaluation protocol before equal\-weight averaging across the ten model–dataset settings\. The total comprises 90,000 outputs from each of the five evaluated methods\. Each output is evaluated against its matched original image and benign instruction, irrespective of the method\-specific image or prompt used during generation\. Higher relevance, informativeness, and usefulness scores are better, whereas lower redundancy is better\.

### D\.21\.Resource Amplification versus Caption Quality

To assess utility preservation under attack, we additionally report BLEU\-1/2/3/4\([Papineni et al\., 2002](https://arxiv.org/html/2609.05889#bib.bib36)\)and CIDEr\([Vedantam et al\., 2015](https://arxiv.org/html/2609.05889#bib.bib37)\)on MS COCO\. Table[31](https://arxiv.org/html/2609.05889#A4.T31)compares resource\-exhaustion effectiveness and caption quality for BLIP\-2 and Qwen2\.5\-VL under the same evaluation protocol\.

Table 31\.Resource\-exhaustion effectiveness and caption quality on MS COCO for BLIP\-2 and Qwen2\.5\-VL\. We report generation length, estimated energy, and latency together with BLEU\-1/2/3/4 and CIDEr\. Clean rows are omitted, so the table compares caption\-quality metrics among attack outputs rather than directly reporting clean\-versus\-attack degradation\. All methods are evaluated under the same protocol\. Bold marks the largest value for each model and metric\.ModelMethodAttack EffectCaption QualityLen\.E^\\widehat\{E\}\(J\)Lat \(s\)BLEU\-1BLEU\-2BLEU\-3BLEU\-4CIDErBLIP\-2NICGSlowDown86\.76567\.034\.460\.0110\.0040\.0020\.0010\.023Verbose Images211\.621432\.6112\.070\.0270\.0130\.0030\.0020\.072JPPO373\.232324\.2919\.580\.0220\.0100\.0020\.0010\.045Qwen2\.5\-VLNICGSlowDown146\.121561\.607\.280\.0140\.0070\.0030\.0010\.025Verbose Images252\.462898\.4110\.880\.0400\.0180\.0060\.0020\.061JPPO441\.235458\.5924\.030\.0410\.0230\.0090\.0050\.049Table[31](https://arxiv.org/html/2609.05889#A4.T31)compares resource\-exhaustion effectiveness with caption\-quality metrics on MS COCO among attack outputs\. Because clean caption\-quality values are omitted, the table provides an attack\-to\-attack comparison of realized serving cost and standard caption\-overlap metrics under the same evaluation protocol\.

The quality ordering depends on both the target model and the metric\. For example, on Qwen2\.5\-VL, JPPO has higher BLEU scores than Verbose Images but lower CIDEr, whereas on BLIP\-2 it does not dominate the caption\-overlap metrics\. The two analyses capture different aspects of output quality: the present analysis covers MS COCO on two models and measures overlap with reference captions, whereas Table[4](https://arxiv.org/html/2609.05889#S6.T4)measures original\-task relevance, informativeness, redundancy, and usefulness against a common original\-image and benign\-instruction reference across all ten model–dataset settings\. Together, the results indicate that JPPO primarily optimizes per\-request serving cost rather than caption fidelity, although its outputs remain higher\-quality and less redundant than Verbose Images under the aggregate judge protocol\.

### D\.22\.Representation\-Space Similarity of Adversarial Images

We measure representation\-space similarity using cosine similarity between CLIP image embeddings extracted from the original and adversarial images\([Radford et al\., 2021](https://arxiv.org/html/2609.05889#bib.bib1)\)\.

Table 32\.Representation\-space similarity measured by cosine similarity between CLIP image embeddings of the original and adversarial images\.DatasetModelNICGSlowDownVerbose ImagesHidden TailJPPOMS COCOLLaVA\-NeXT0\.9700\.9720\.9620\.978BLIP\-20\.9730\.9710\.9580\.965Qwen2\.5\-VL0\.9550\.9410\.9480\.962InstructBLIP0\.9790\.9700\.9620\.952MiniGPT\-40\.9740\.9690\.9680\.968ImageNetLLaVA\-NeXT0\.9710\.9730\.9770\.988BLIP\-20\.9810\.9690\.9530\.971Qwen2\.5\-VL0\.9640\.9630\.9820\.971InstructBLIP0\.9850\.9650\.9740\.954MiniGPT\-40\.9860\.9720\.9600\.957Table[32](https://arxiv.org/html/2609.05889#A4.T32)shows that the adversarial examples produced by JPPO remain highly similar to the original images in CLIP embedding space\. Across the reported model\-dataset pairs, the cosine similarity values are generally high, indicating that strong resource amplification does not require substantial representation\-level drift\.

### D\.23\.Qualitative Clean\-versus\-Attacked Examples

Figures[10](https://arxiv.org/html/2609.05889#acmlabel10)–[14](https://arxiv.org/html/2609.05889#acmlabel14)qualitatively compare clean and attacked outputs for all five victim models under the same image and prompt, reporting length, estimated energy, and latency\.

![Qualitative BLIP-2 clean versus attacked example](https://arxiv.org/html/2609.05889v1/5_1.png)Figure 10\.Qualitative clean\-versus\-attacked example for BLIP\-2\.For the same image and task prompt, the clean input yields a relatively concise output, whereas the attacked input induces a substantially longer continuation\. The figure additionally reports the realized Length, Estimated Energy, and Latency for this case\.Qualitative BLIP\-2 clean versus attacked exampleA qualitative BLIP\-2 example compares clean and attacked outputs for the same image and task prompt\. The attacked case produces a noticeably longer response and higher reported serving\-cost metrics than the clean case\.![Qualitative Qwen2.5-VL clean versus attacked example](https://arxiv.org/html/2609.05889v1/5_2.png)Figure 11\.Qualitative clean\-versus\-attacked example for Qwen2\.5\-VL\.For the same image and task prompt, the attacked input drives the model toward a longer, more resource\-intensive decoding trajectory than the clean input does\. The figure additionally reports the realized Length, Energy, and Latency for this case\.Qualitative Qwen2\.5\-VL clean versus attacked exampleA qualitative Qwen2\.5\-VL example compares clean and attacked outputs for the same image and task prompt\. The attacked case shows a longer continuation and higher reported serving\-cost metrics than the clean case\.![Qualitative LLaVA-NeXT clean versus attacked example](https://arxiv.org/html/2609.05889v1/5_3.png)Figure 12\.Qualitative clean\-versus\-attacked example for LLaVA\-NeXT\.Compared with the clean input, the attacked input produces a visibly more persistent continuation under the same image and task prompt\. The figure additionally reports the realized Length, Energy, and Latency for this case\.Qualitative LLaVA\-NeXT clean versus attacked exampleA qualitative LLaVA\-NeXT example compares clean and attacked outputs for the same image and task prompt\. The attacked case generates a more persistent continuation and higher reported serving\-cost metrics than the clean case\.![Qualitative InstructBLIP clean versus attacked example](https://arxiv.org/html/2609.05889v1/5_4.png)Figure 13\.Qualitative clean\-versus\-attacked example for InstructBLIP\.For the same image and task prompt, the attacked input yields a substantially longer response than the clean input and incurs a higher realized serving cost\. The figure additionally reports the realized Length, Energy, and Latency for this case\.Qualitative InstructBLIP clean versus attacked exampleA qualitative InstructBLIP example compares clean and attacked outputs for the same image and task prompt\. The attacked case produces a substantially longer response and higher reported serving\-cost metrics than the clean case\.![Qualitative MiniGPT-4 clean versus attacked example](https://arxiv.org/html/2609.05889v1/5_5.png)Figure 14\.Qualitative clean\-versus\-attacked example for MiniGPT\-4\.The attacked input induces a longer and more expensive decoding trajectory than the clean input for the same image and task prompt\. The figure additionally reports the realized Length, Energy, and Latency for this case\.Qualitative MiniGPT\-4 clean versus attacked exampleA qualitative MiniGPT\-4 example compares clean and attacked outputs for the same image and task prompt\. The attacked case shows a longer continuation and higher reported serving\-cost metrics than the clean case\.

Similar Articles

Visual prompt engineering for video models

Hugging Face Daily Papers

This paper introduces Visual Prompt Engineering (VIPE), a method that automatically modifies task images to improve video model performance, showing it can be more effective than text-based prompt engineering or test-time scaling.

How Do Prompt Variations Affect Energy Consumption in On-Device LLMs?

arXiv cs.CL

This paper explores how prompt properties like cognitive load and phrasing pattern influence energy usage in on-device LLM inference, showing that cognitive load affects energy per token while phrasing impacts token usage, highlighting the need for model-aware prompt design for energy efficiency.