用于视觉Token通信的全预算反事实推理选择性摊销

arXiv cs.AI 论文

摘要

ACV-Gate是一个自适应候选评估框架,通过选择性地对信息丰富的视觉Token应用反事实推理,在生成式图像通信中提高重建质量,从而降低计算成本。

arXiv:2609.30756v1 Announce Type: new Abstract: Generative image communication transmits compact semantic tokens under a limited packet budget, where token selection directly affects the final reconstruction quality after the complete packet is decoded. However, accurately estimating the terminal value of every candidate token requires repeated receiver-side reconstruction, resulting in substantial encoder-side computation. To address this problem, we propose ACV-Gate, an adaptive candidate evaluation framework that learns to approximate full-budget counterfactual evaluation and selectively assigns exact evaluations to the most informative candidates. Specifically, a set-aware student is trained using terminal advantages and regrets to predict candidate rankings directly, while a selective refinement mechanism evaluates only a bounded candidate set containing both Local-MDL and direct actions; cost-based thresholds further enable explicit control of the average evaluation workload. Experiments on CIFAR-10 show that ACV-Gate consistently improves reconstruction quality while substantially reducing candidate evaluations; at 0.20 bpp, the primary adaptive configuration improves PSNR over LocalMDL by 0.636 dB with only 2.13 candidate evaluations per image, corresponding to 27.60% of the calls required by the Exact-Full expert. Matched-candidate comparisons, synchronized GPU measurements, and evaluations on STL-10 and 384 *384 scale transfer further demonstrate consistent quality computation trade-offs, with particularly pronounced gains at low bit rates. These results show that combining terminal-value learning with selective candidate evaluation provides an effective and controllable mechanism for allocating encoder computation in packet-constrained generative image communication.
查看原文
查看缓存全文

缓存时间: 2026/09/28 09:45

# Selective Amortization of Full-Budget Counterfactual Reasoning for Visual Token Communication
Source: [https://arxiv.org/html/2609.30756](https://arxiv.org/html/2609.30756)
## Selective Amortization of Full\-Budget Counterfactual Reasoning for Visual Token CommunicationThanks:This work was supported by the National Natural Science Foundation of China under Grant No\. 62002263, the Tianjin Science and Technology Program Projects under Grant No\. 24YDTPJC00630, and the Tianjin Research Innovation Project for Postgraduate Students under Grant No\. 2026KYCX001F\.Thanks:Q\.Qi, Z\.Liang are with the School of Artificial Intelligence, Nanyang Normal University, Nanyang, China\. \( e\-mail: qiqinglei@nynu\.edu\.cn, liangzhihe329@gmail\.com\)Thanks:F\.Jing, S\.Zhu, L\.Zhang Y\.Zhang and J\. Guo are with the School of Computer and Information Engineering, Tianjin Normal University, Tianjin, China \(e\-mail: 2611090041@stu\.tjnu\.edu\.cn ,zhushenao34@gmail\.com,2611090026@stu\.tjnu\.edu\.cn, 2411090038@stu\.tjnu\.edu\.cn, c04s316@bupt\.cnThanks:S\. He is with the School of Information Science and Engineering, Linyi University, Linyi 276000, China \(email:heshuqing@lyu\.edu\.cn\)Thanks:\*Corresponding Author: Jia GuoThanks:The source code for this work is publicly available at: https://github\.com/c04s316/\-ACV\-Gate\-

Zhihe LiangFengzhan JingShenao ZhuLei ZhangAffiliation:Chenyang Zhang, Shuqing He, Jia Guo\*,

###### Abstract

Generative image communication transmits compact semantic tokens under a limited packet budget, where token selection directly affects the final reconstruction quality after the complete packet is decoded\. However, accurately estimating the terminal value of every candidate token requires repeated receiver\-side reconstruction, resulting in substantial encoder\-side computation\. To address this problem, we propose ACV\-Gate, an adaptive candidate evaluation framework that learns to approximate full\-budget counterfactual evaluation and selectively assigns exact evaluations to the most informative candidates\. Specifically, a set\-aware student is trained using terminal advantages and regrets to predict candidate rankings directly, while a selective refinement mechanism evaluates only a bounded candidate set containing both Local\-MDL and direct actions; cost\-based thresholds further enable explicit control of the average evaluation workload\. Experiments on CIFAR\-10 show that ACV\-Gate consistently improves reconstruction quality while substantially reducing candidate evaluations; at 0\.20 bpp, the primary adaptive configuration improves PSNR over Local\-MDL by0\.636±0\.1190\.636\\pm 0\.119dB with only2\.13±0\.532\.13\\pm 0\.53candidate evaluations per image, corresponding to 27\.60% of the calls required by the Exact\-Full expert\. Matched\-candidate comparisons, synchronized GPU measurements, and evaluations on STL\-10 and384×384384\\times 384scale transfer further demonstrate consistent quality–computation trade\-offs, with particularly pronounced gains at low bit rates\. These results show that combining terminal\-value learning with selective candidate evaluation provides an effective and controllable mechanism for allocating encoder computation in packet\-constrained generative image communication\.

###### Index Terms:

Visual token communication, counterfactual evaluation, selective computation, knowledge distillation, resource allocation\.

## IIntroduction

Bandwidth\-constrained visual communication is important for applications in which images must be delivered under limited or time\-varying transmission resources, such as wireless links, edge visual systems, and progressive image delivery\. Generative receivers provide a promising mechanism for such settings by reconstructing missing visual content from transmitted information together with learned image priors\. Instead of describing every image component explicitly, the transmitter can allocate its packet budget to a subset of informative visual tokens and allow the receiver to complete the remaining content\. This communication paradigm makes the selection of transmitted tokens a central design problem: under a fixed packet budget, the transmitted subset should provide the greatest improvement in the final reconstructed image\.

The value of an individual token, however, is determined by more than its local image content\. It depends on the tokens that have already been transmitted, the tokens that will be selected subsequently, and the way in which the generative receiver completes the untransmitted content\. Consequently, effective token selection requires an estimate of the reconstruction quality obtained after the available transmission budget has been fully used\. This terminal perspective is particularly important at low bit rates, where each transmitted token occupies a substantial fraction of the packet budget and its interaction with the receiver prior can strongly influence the final reconstruction\.

Learned image compression has established a direct connection between rate allocation and reconstruction quality\[[1](https://arxiv.org/html/2609.30756#bib.bib1),[2](https://arxiv.org/html/2609.30756#bib.bib2),[3](https://arxiv.org/html/2609.30756#bib.bib3)\], while discrete representations and masked generation provide a natural token\-based interface for image representation and completion\[[4](https://arxiv.org/html/2609.30756#bib.bib4),[5](https://arxiv.org/html/2609.30756#bib.bib5),[6](https://arxiv.org/html/2609.30756#bib.bib6)\]\. Deep joint source–channel coding and semantic transmission further exploit learned representations and receiver\-side priors to improve visual communication under constrained channels\[[7](https://arxiv.org/html/2609.30756#bib.bib7),[8](https://arxiv.org/html/2609.30756#bib.bib8),[9](https://arxiv.org/html/2609.30756#bib.bib9)\]\. Progressive, content\-adaptive, masked, and sparse transmission methods extend these ideas to variable communication conditions by allocating transmission resources according to image content and communication requirements\[[10](https://arxiv.org/html/2609.30756#bib.bib10),[11](https://arxiv.org/html/2609.30756#bib.bib11),[12](https://arxiv.org/html/2609.30756#bib.bib12),[13](https://arxiv.org/html/2609.30756#bib.bib13)\]\. Together, these developments provide increasingly flexible mechanisms for representing and transmitting visual information\. For a fixed tokenizer and generative receiver, an important remaining question is how to identify the token positions whose transmission contributes most to the final reconstruction under the available packet budget\.

Existing token\-pruning and selection methods commonly estimate token importance from attention, saliency, diversity, or coverage\[[14](https://arxiv.org/html/2609.30756#bib.bib14),[15](https://arxiv.org/html/2609.30756#bib.bib15),[16](https://arxiv.org/html/2609.30756#bib.bib16),[17](https://arxiv.org/html/2609.30756#bib.bib17),[18](https://arxiv.org/html/2609.30756#bib.bib18),[24](https://arxiv.org/html/2609.30756#bib.bib24),[25](https://arxiv.org/html/2609.30756#bib.bib25),[26](https://arxiv.org/html/2609.30756#bib.bib26)\]\. Learned selectors and adaptive compression methods further reduce visual computation by predicting which tokens or features deserve additional processing\[[27](https://arxiv.org/html/2609.30756#bib.bib27),[28](https://arxiv.org/html/2609.30756#bib.bib28),[29](https://arxiv.org/html/2609.30756#bib.bib29),[30](https://arxiv.org/html/2609.30756#bib.bib30),[31](https://arxiv.org/html/2609.30756#bib.bib31)\]\. These approaches provide effective mechanisms for identifying informative visual elements, but communication introduces a different evaluation target: the usefulness of a transmitted token is ultimately determined by the quality of the image reconstructed after the complete packet has been received\. A locally important token may contribute differently when combined with other transmitted tokens or when the receiver can infer its content from contextual information\. Token selection for generative communication therefore benefits from a terminal\-value criterion that captures both sequential token interactions and receiver\-side completion\.

Full\-budget counterfactual evaluation provides such a criterion by evaluating each candidate action through the reconstruction obtained after completing the remaining transmission budget\. GCR\-C\[[33](https://arxiv.org/html/2609.30756#bib.bib33)\]implements this principle by comparing candidate tokens under the same Local\-MDL continuation, packet budget, and receiver\. Each resulting value measures the contribution of a candidate within a complete transmission sequence and therefore directly reflects terminal reconstruction quality\. This evaluation also introduces a substantial encoder\-side computational workload: every candidate query requires repeated prior inference, continuation of the remaining transmission decisions, and final image reconstruction\. As a result, the transmitter must allocate two coupled resources—the communication bits used to represent the image and the computation used to determine how those bits should be spent\.

This computation\-allocation problem is related to selective prediction and learning\-to\-defer, where an expensive expert is invoked only for inputs that benefit most from additional evaluation\[[19](https://arxiv.org/html/2609.30756#bib.bib19),[20](https://arxiv.org/html/2609.30756#bib.bib20),[22](https://arxiv.org/html/2609.30756#bib.bib22),[23](https://arxiv.org/html/2609.30756#bib.bib23)\]\. In generative image communication, however, selective evaluation operates at both the state and candidate levels\. The transmitter must determine which communication states merit additional refinement and, within each selected state, which candidate tokens should receive exact terminal evaluation\. The resulting controller must therefore combine an estimate of terminal reconstruction value with the computational cost associated with candidate queries\.

To address this problem, we propose ACV\-Gate, a selective amortization framework that learns terminal candidate rankings from full\-budget counterfactual evaluations and allocates exact evaluations to the candidates and communication states where they are most informative\. A set\-aware student learns candidate rankings from terminal advantages and regrets produced by the full\-budget evaluator, allowing it to select a direct action without exact candidate evaluation and to rank alternative actions for subsequent refinement\. A bounded candidate screen then retains the Local\-MDL action, the direct student action, and highly ranked alternatives for exact evaluation\. Cost\-based state allocation further controls when this refinement is applied, enabling the encoder to trade candidate\-query workload for reconstruction quality under explicit computational budgets\.

Relation to the terminal evaluator:GCR\-C\[[33](https://arxiv.org/html/2609.30756#bib.bib33)\]serves as the fixed full\-budget evaluator from which ACV\-Gate learns terminal candidate preferences\. ACV\-Gate preserves this terminal\-quality objective while amortizing its repeated use during token selection\. The learned student provides an immediate candidate ranking, and selective refinement concentrates exact GCR\-C evaluations on a bounded set of promising alternatives\. This formulation turns terminal evaluation from a computation applied uniformly across candidates into a controllable encoder resource that can be allocated according to its expected reconstruction benefit\.

The main contributions are as follows:

1. 1\.We formulate token selection in generative image communication under three coupled resources: serialized packet bits, a per\-image cap on exact candidate evaluations, and an average candidate\-query budget\. This formulation makes encoder computation an explicit resource alongside communication rate\.
2. 2\.We develop a set\-aware student that learns terminal candidate rankings from full\-budget advantages and regrets\. The learned ranking supports direct token selection without exact candidate evaluation and provides an ordered candidate set for bounded refinement\.
3. 3\.We develop a reference\-preserving selective controller that combines bounded candidate screening with cost\-based state allocation\. Matched\-budget comparisons separate the effects of terminal\-value ranking, candidate screening, and state allocation, while synchronized runtime measurements and evaluations across image systems characterize the resulting computation–quality trade\-off\.

## IIRelated Work

### II\-ASemantic communication and learned image transmission

Learned image compression optimizes analysis and synthesis transforms together with entropy models\[[1](https://arxiv.org/html/2609.30756#bib.bib1),[2](https://arxiv.org/html/2609.30756#bib.bib2),[3](https://arxiv.org/html/2609.30756#bib.bib3)\]\. Vector quantization and masked generation represent images as discrete tokens whose missing entries can be predicted from context\[[4](https://arxiv.org/html/2609.30756#bib.bib4),[5](https://arxiv.org/html/2609.30756#bib.bib5),[6](https://arxiv.org/html/2609.30756#bib.bib6)\]\. Deep joint source–channel coding learns image transmission over a noisy link, while semantic and generative communication exploit task structure and receiver priors\[[7](https://arxiv.org/html/2609.30756#bib.bib7),[8](https://arxiv.org/html/2609.30756#bib.bib8),[9](https://arxiv.org/html/2609.30756#bib.bib9)\]\.

Recent communication methods address progressive transmission, content adaptation, semantic masking and sparse visual representations\[[10](https://arxiv.org/html/2609.30756#bib.bib10),[11](https://arxiv.org/html/2609.30756#bib.bib11),[12](https://arxiv.org/html/2609.30756#bib.bib12),[13](https://arxiv.org/html/2609.30756#bib.bib13)\]\. Their design choices concern the representation, channel mapping and rate allocation\. For a fixed representation and receiver, token selection still requires the encoder to estimate the effect of each transmitted position\. ACV\-Gate allocates the computation used to make that decision\.

### II\-BVisual token selection and adaptive computation

DynamicViT, EViT, TokenLearner, A\-ViT and Token Merging reduce visual computation through learned importance, token reorganization, adaptive selection or merging\[[14](https://arxiv.org/html/2609.30756#bib.bib14),[15](https://arxiv.org/html/2609.30756#bib.bib15),[16](https://arxiv.org/html/2609.30756#bib.bib16),[17](https://arxiv.org/html/2609.30756#bib.bib17),[18](https://arxiv.org/html/2609.30756#bib.bib18)\]\. More recent methods use diversity, visual cues, saliency–coverage, graph structure and attention centrality\[[24](https://arxiv.org/html/2609.30756#bib.bib24),[25](https://arxiv.org/html/2609.30756#bib.bib25),[26](https://arxiv.org/html/2609.30756#bib.bib26),[30](https://arxiv.org/html/2609.30756#bib.bib30),[31](https://arxiv.org/html/2609.30756#bib.bib31)\]\. SparseVLM, LLaVA\-PruMerge and FitPrune extend efficient token reduction to multimodal inference\[[27](https://arxiv.org/html/2609.30756#bib.bib27),[28](https://arxiv.org/html/2609.30756#bib.bib28),[29](https://arxiv.org/html/2609.30756#bib.bib29)\]\. Analysis of encoder\-layer information further shows why a pruning rule’s effectiveness depends on where and how token information is represented\[[32](https://arxiv.org/html/2609.30756#bib.bib32)\]\.

Attention, diversity and coverage provide inexpensive estimates of token importance\. For packetized reconstruction, these estimates must reflect how a selected token affects subsequent transmission and receiver completion\. We evaluate adapted DivPrune, VisPruner and SCOPE rules with a common proposal and receiver\. ACV\-Gate learns its ranking from completed\-image quality under that same communication process\.

### II\-CSelective prediction and decision distillation

Selective prediction equips a predictor with a reject option, and learning\-to\-defer extends this idea to costly expert access\[[19](https://arxiv.org/html/2609.30756#bib.bib19),[20](https://arxiv.org/html/2609.30756#bib.bib20),[21](https://arxiv.org/html/2609.30756#bib.bib21),[22](https://arxiv.org/html/2609.30756#bib.bib22),[23](https://arxiv.org/html/2609.30756#bib.bib23)\]\. In visual\-token communication, the expert is itself a computation: GCR\-C completes the packet under a fixed continuation and reconstructs the image to obtain terminal candidate value\[[33](https://arxiv.org/html/2609.30756#bib.bib33)\]\.

These strands motivate selective computation based on terminal reconstruction value\. ACV\-Gate adds an amortization and budgeting layer to the GCR\-C evaluator\. The student learns terminal candidate rankings, and the controller allocates exact evaluations to bounded candidate sets\. This links token selection and encoder computation under a common packet budget and receiver\.

## IIISystem Model and Problem Formulation

Letxxbe an image andz=\(z1,…,zN\)z=\(z\_\{1\},\\ldots,z\_\{N\}\)its discrete visual\-token sequence\. A transmitted setSSis represented by a packet containing a mode field, a position description, token payload, a cyclic redundancy check \(CRC\) and optional forward error correction \(FEC\)\. LetB⁡\(x\)B\(x\)denote the bit budget assigned toxxat a given operating rate\. With packet function𝖻𝗂𝗍𝗌⁡\(S\)\\mathsf\{bits\}\(S\), the feasible candidate set is

ℱ⁡\(S,B\)=\{a∉S:𝖻𝗂𝗍𝗌⁡\(S∪\{a\}\)≤B\}\.\\mathcal\{F\}\(S,B\)=\\\{a\\notin S:\\mathsf\{bits\}\(S\\cup\\\{a\\\}\)\\leq B\\\}\.\(1\)The receiver uses a masked priorDθD\_\{\\theta\}to complete unknown positions and producesx^​\(S\)\\hat\{x\}\(S\)\. A communication state isst=\(St,Brem,t,𝐟t,ct\)s\_\{t\}=\(S\_\{t\},B\_\{\\mathrm\{rem\},t\},\\mathbf\{f\}\_\{t\},c\_\{t\}\), where𝐟t\\mathbf\{f\}\_\{t\}contains recoverability and local\-score features andctc\_\{t\}encodes the operating rate, unselected\-token fraction and stage\. The experiments use a fixed channel configuration\. The Local\-MDL action is

aL​\(st\)=arg⁡maxa∈ℱ⁡\(St,B\)−log⁡pθ​\(za∣zSt\)\.a\_\{\\mathrm\{L\}\}\(s\_\{t\}\)=\\arg\\max\_\{a\\in\\mathcal\{F\}\(S\_\{t\},B\)\}\-\\log p\_\{\\theta\}\(z\_\{a\}\\mid z\_\{S\_\{t\}\}\)\.\(2\)To evaluate candidateaa, the transmitter appends it toStS\_\{t\}and completes the packet using Local\-MDL\. The resulting receiver PSNR definesQB​\(a∣st\)Q\_\{B\}\(a\\mid s\_\{t\}\)\. This exact teacher uses the same continuation, packet syntax and reconstruction as the communication system\. Relative to Local, the teacher advantage and within\-state regret are

AB​\(a∣st\)\\displaystyle A\_\{B\}\(a\\mid s\_\{t\}\)=QB​\(a∣st\)−QB​\(aL∣st\),\\displaystyle=Q\_\{B\}\(a\\mid s\_\{t\}\)\-Q\_\{B\}\(a\_\{\\mathrm\{L\}\}\\mid s\_\{t\}\),\(3\)R⁡\(a∣st\)\\displaystyle R\(a\\mid s\_\{t\}\)=maxb∈𝒫⁡\(st\)⁡AB​\(b∣st\)−AB​\(a∣st\)\.\\displaystyle=\\max\_\{b\\in\\mathcal\{P\}\(s\_\{t\}\)\}A\_\{B\}\(b\\mid s\_\{t\}\)\-A\_\{B\}\(a\\mid s\_\{t\}\)\.ThusR⁡\(a∣st\)≥0R\(a\\mid s\_\{t\}\)\\geq 0within the compact proposal, anda∗a^\{\*\}denotes its minimum\-regret candidate\. Advantage and regret supervise the student during training\.

The compact proposal is the deduplicated union of Top\-3 Local, Top\-3 selected\-set importance and Top\-2 coverage candidates\. We denote it by𝒫⁡\(st\)\\mathcal\{P\}\(s\_\{t\}\)and writeKs=\|𝒫⁡\(st\)\|≤8K\_\{s\}=\|\\mathcal\{P\}\(s\_\{t\}\)\|\\leq 8\. The four\-candidate screen introduced below is a subset of this proposal\. We refer to the full\-proposal expert as Exact\-Full throughout the paper\.

For a policyπ\\pi, letCmaxC\_\{\\max\}be the per\-image hard cap on exact candidate evaluations and letqqbe the average candidate\-query budget\. The constrained objective is

maxπ\\displaystyle\\max\_\{\\pi\}1\|𝒳\|​∑x∈𝒳Q⁡\(π,x\)\\displaystyle\\frac\{1\}\{\|\\mathcal\{X\}\|\}\\sum\_\{x\\in\\mathcal\{X\}\}Q\(\\pi;x\)\(4\)subject to\\displaystyle\\text\{subject to \}𝖻𝗂𝗍𝗌\(Sπ\(x\)\)≤B\(x\),∀x∈𝒳,\\displaystyle\\mathsf\{bits\}\(S\_\{\\pi\}\(x\)\)\\leq B\(x\),\\quad\\forall x\\in\\mathcal\{X\},Ncf\(x\)≤Cmax,∀x∈𝒳,\\displaystyle N\_\{\\mathrm\{cf\}\}\(x\)\\leq C\_\{\\max\},\\quad\\forall x\\in\\mathcal\{X\},1\|𝒳\|​∑x∈𝒳Ncf​\(x\)≤q\.\\displaystyle\\frac\{1\}\{\|\\mathcal\{X\}\|\}\\sum\_\{x\\in\\mathcal\{X\}\}N\_\{\\mathrm\{cf\}\}\(x\)\\leq q\.whereQ⁡\(π,x\)Q\(\\pi;x\)is terminal receiver quality for imagexx,Ncf​\(x\)N\_\{\\mathrm\{cf\}\}\(x\)counts candidate\-level exactQB​\(a∣s\)Q\_\{B\}\(a\\mid s\)evaluations,Sπ​\(x\)S\_\{\\pi\}\(x\)is the terminal transmitted set, andCmaxC\_\{\\max\}is the hard per\-image cap\. HereB⁡\(x\)B\(x\)counts serialized communication bits,CmaxC\_\{\\max\}bounds exact candidate workload for one image, andqqcontrols average workload across an image set\. The controller maintainsCrem,0=CmaxC\_\{\\mathrm\{rem\},0\}=C\_\{\\max\}and decrements it by the number of candidates evaluated after each expert call\. Threshold calibration sets the average operating point, and each experiment reports its realized meanNcfN\_\{\\mathrm\{cf\}\}/image together with encoder runtime\.

Online evaluation uses one eligible refinement decision per image, at the initial state, where the available candidate budget equalsCmaxC\_\{\\max\}\. After this decision, both the direct and refined branches complete the packet with Local\-MDL\.

For a fixed patch sizeppand image dimensions\(H,W\)\(H,W\), we define the token grid and the three scale\-aware accounting quantities

Ntok​\(H,W\)\\displaystyle N\_\{\\mathrm\{tok\}\}\(H,W\)=⌊Hp⌋​⌊Wp⌋,\\displaystyle=\\left\\lfloor\\frac\{H\}\{p\}\\right\\rfloor\\left\\lfloor\\frac\{W\}\{p\}\\right\\rfloor,\(5\)cMP\\displaystyle c\_\{\\mathrm\{MP\}\}=NcfH​W/106,ctok=NcfNtok​\(H,W\)\.\\displaystyle=\\frac\{N\_\{\\mathrm\{cf\}\}\}\{HW/10^\{6\}\},\\qquad c\_\{\\mathrm\{tok\}\}=\\frac\{N\_\{\\mathrm\{cf\}\}\}\{N\_\{\\mathrm\{tok\}\}\(H,W\)\}\.The absoluteNcfN\_\{\\mathrm\{cf\}\}/image measures candidate workload;cMPc\_\{\\mathrm\{MP\}\}andctokc\_\{\\mathrm\{tok\}\}normalize this workload by image area and token count\. Measured runtime captures the cost of each evaluation on a given token grid\.

## IVACV\-Gate

Fig\. 1:ACV\-Gate architecture\. \(a\) Offline full\-budget rollouts provide candidate advantages and regrets for student and allocation\-score learning\. \(b\) At the initial communication state, the student ranks the compact proposal\. The allocation score and remaining candidate budget select either the direct action or a screened exact evaluation\. \(c\) The selected action is followed by Local\-MDL continuation, packet transmission and receiver reconstruction\. Dashed arrows indicate transfer of learned parameters and the calibrated threshold\.### IV\-AFixed full\-budget evaluator and compact proposal

At training time, the full\-budget evaluator labels the compact proposal formed by the union of Top\-3 Local, Top\-3 selected\-set importance and Top\-2 coverage candidates\. Duplicates are removed and Local is retained\. Each label is obtained by appending a candidate, completing the packet with Local\-MDL, and measuring the reconstructed image’s PSNR\. The Exact\-FullQBQ\_\{B\}reference evaluates the entire proposal with this same procedure\[[33](https://arxiv.org/html/2609.30756#bib.bib33)\]\. It provides the candidate optimum for the fixed proposal and continuation\.

### IV\-BState and candidate representation

Each candidate representation combines the masked\-prior hidden state, codebook vector, position, normalized remaining budget, Local score, entropy, packet feasibility and proposal source\. Relative features describe its relationship to the proposal and selected tokens\. These include standardized scores, rank offsets, codebook and hidden\-state distances, coverage and redundancy\. Source–rate interactions and normalized position distances complete the set context\. All input statistics are computed from the current communication state; terminal advantage and regret provide the training targets\.

### IV\-CSet\-aware student and learning objective

The candidate embeddings are processed jointly by a two\-layer Transformer without positional embeddings\. Hence permuting the proposal rows permutes the output scores while preserving the set\-level context\. For a padded candidate group, the student produces logitslal\_\{a\}and two auxiliary cost estimates\(r^t,g^t\)\(\\hat\{r\}\_\{t\},\\hat\{g\}\_\{t\}\):

\(\{la\}a∈𝒫⁡\(st\),r^t,g^t\)=fϕ​\(\{ϕ⁡\(st,a\)\}a∈𝒫⁡\(st\),ct\)\.\(\\\{l\_\{a\}\\\}\_\{a\\in\\mathcal\{P\}\(s\_\{t\}\)\},\\hat\{r\}\_\{t\},\\hat\{g\}\_\{t\}\)=f\_\{\\phi\}\(\\\{\\phi\(s\_\{t\},a\)\\\}\_\{a\\in\\mathcal\{P\}\(s\_\{t\}\)\},c\_\{t\}\)\.\(6\)The direct action isa^=arg⁡maxa⁡la\\hat\{a\}=\\arg\\max\_\{a\}l\_\{a\}\. The auxiliary outputs estimate its regret and lost positive gain, giving the allocation scoreutACV=r^t\+0\.5​g^tu\_\{t\}^\{\\mathrm\{ACV\}\}=\\hat\{r\}\_\{t\}\+0\.5\\hat\{g\}\_\{t\}\. We denote the controller’s state score byut=u⁡\(st\)u\_\{t\}=u\(s\_\{t\}\)and compare this learned score with student margin, Local uncertainty and random allocation under a common candidate budget\. The conditionctc\_\{t\}contains a seven\-way rate ID, normalized rater/max⁡rr/\\max r, unselected\-token fraction1−\|St\|/N1\-\|S\_\{t\}\|/N, and initial/early/mid stage ID\. A small feature\-wise linear modulation \(FiLM\) network conditions each candidate embedding on these variables\.

#### Regret and gain supervision\.

Letqt​a∝exp\[−R\(a∣st\)/Tt\]q\_\{ta\}\\propto\\exp\[\-R\(a\\mid s\_\{t\}\)/T\_\{t\}\]be the regret\-soft teacher target withTt=0\.10T\_\{t\}=0\.10, and letpt​a=softmax⁡\(lt​a/Ts\)p\_\{ta\}=\\operatorname\{softmax\}\(l\_\{ta\}/T\_\{s\}\)withTs=0\.35T\_\{s\}=0\.35\. The hard\-target weight isαt=σ⁡\[\(gt−0\.20\)/0\.06\]\\alpha\_\{t\}=\\sigma\[\(g\_\{t\}\-0\.20\)/0\.06\], wheregtg\_\{t\}is the gap between the teacher’s two highest candidate values in dB\. The complete training objective is

ℒcls\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{cls\}\}=αt​CE⁡\(a∗,p\)\+\(1−αt\)​ℒsoft,\\displaystyle=\\alpha\_\{t\}\\operatorname\{CE\}\(a^\{\*\},p\)\+\(1\-\\alpha\_\{t\}\)\\mathcal\{L\}\_\{\\mathrm\{soft\}\},\(7\)ℒH2\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{H2\}\}=ℒcls\+0\.33​ℒreg\+0\.40​ℒpair,\\displaystyle=\\mathcal\{L\}\_\{\\mathrm\{cls\}\}\+0\.33\\mathcal\{L\}\_\{\\mathrm\{reg\}\}\+0\.40\\mathcal\{L\}\_\{\\mathrm\{pair\}\},ℒACV\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{ACV\}\}=ℒH2\+0\.10​ℒgain\+0\.05​ℒsafe\.\\displaystyle=\\mathcal\{L\}\_\{\\mathrm\{H2\}\}\+0\.10\\mathcal\{L\}\_\{\\mathrm\{gain\}\}\+0\.05\\mathcal\{L\}\_\{\\mathrm\{safe\}\}\.Expected regret penalizes the terminal quality lost by the student, while pairwise ranking gives larger penalties to costly action errors\. Withrt​a=\[maxb⁡AB​\(b∣st\)−AB​\(a∣st\)\]\+r\_\{ta\}=\[\\max\_\{b\}A\_\{B\}\(b\\mid s\_\{t\}\)\-A\_\{B\}\(a\\mid s\_\{t\}\)\]\_\{\+\},qt​a=exp\(−rt​a/0\.10\)/∑bexp\(−rt​b/0\.10\)q\_\{ta\}=\\exp\(\-r\_\{ta\}/0\.10\)/\\sum\_\{b\}\\exp\(\-r\_\{tb\}/0\.10\), andpt​a=softmax⁡\(lt​a/0\.35\)p\_\{ta\}=\\operatorname\{softmax\}\(l\_\{ta\}/0\.35\),

ℒsoft\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{soft\}\}=−∑aqt​alogpt​a,\\displaystyle=\-\\sum\_\{a\}q\_\{ta\}\\log p\_\{ta\},\(8\)ℒreg\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{reg\}\}=∑apt​a​rt​a,\\displaystyle=\\sum\_\{a\}p\_\{ta\}r\_\{ta\},ℒpair\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{pair\}\}=meant,a∈𝒫t∖\{at∗\}⁡wt​a​softplus⁡\(mt​a−\(lt​a∗−lt​a\)\),\\displaystyle=\\operatorname\{mean\}\_\{t,\\;a\\in\\mathcal\{P\}\_\{t\}\\setminus\\\{a\_\{t\}^\{\*\}\\\}\}w\_\{ta\}\\,\\operatorname\{softplus\}\\\!\\left\(m\_\{ta\}\-\(l\_\{ta^\{\*\}\}\-l\_\{ta\}\)\\right\),wt​a\\displaystyle w\_\{ta\}=0\.25\+0\.75​min⁡\(rt​a/0\.50,1\),\\displaystyle=0\.25\+0\.75\\min\(r\_\{ta\}/0\.50,1\),mt​a\\displaystyle m\_\{ta\}=0\.05\+0\.45​min⁡\(rt​a/0\.50,1\)\.\\displaystyle=0\.05\+0\.45\\min\(r\_\{ta\}/0\.50,1\)\.Letht=\[maxa⁡AB​\(a∣st\)\]\+h\_\{t\}=\[\\max\_\{a\}A\_\{B\}\(a\\mid s\_\{t\}\)\]\_\{\+\},ut​a=\[AB​\(a∣st\)\]\+/\(ht\+10−6\)u\_\{ta\}=\[A\_\{B\}\(a\\mid s\_\{t\}\)\]\_\{\+\}/\(h\_\{t\}\+10^\{\-6\}\), and letQ50\+Q^\{\+\}\_\{50\}be the training median of positive headrooms\. The gain and safety terms are

ℒgain\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{gain\}\}=meanht\>0\.05\[clip\(ht/Q\+50,0\.5,2\.0\)\\displaystyle=\\operatorname\{mean\}\_\{h\_\{t\}\>0\.05\}\[\\operatorname\{clip\}\(h\_\{t\}/Q^\{\+\}\_\{50\},0\.5,2\.0\)\(9\)×\(1−∑apt​aut​a\)\],\\displaystyle\\times\(1\-\\sum\_\{a\}p\_\{ta\}u\_\{ta\}\)\],ℒsafe\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{safe\}\}=mean⁡∑at⁡pt​a​\[−AB​\(a∣st\)\]\+0\.50\.\\displaystyle=\\operatorname\{mean\}\_\{t\}\\sum\_\{a\}p\_\{ta\}\\frac\{\[\-A\_\{B\}\(a\\mid s\_\{t\}\)\]\_\{\+\}\}\{0\.50\}\.The gain term rewards recovery of available positive headroom, and the safety term penalizes actions whose terminal quality falls below Local\. Anchored adaptation combines these terms with H2 and a0\.350\.35KL penalty relative to the frozen H2 action distribution\. All padded rows are masked before the softmax and loss reduction\. The pair term is zero for a group with fewer than two active candidates, and the gain term is zero when no active state hasht\>0\.05h\_\{t\}\>0\.05\. The training stages, optimizers, masking and parameter\-freezing rules appear in Supplementary Table S13\.

Training uses a 10\-epoch hard\-CE/pair warm\-up, 30 regret\-decision epochs and a 10\-epoch anchored adaptation phase that updates the conditional student blocks and score head\. After the direct checkpoint is selected, the selector is frozen, and a 15\-epoch allocation\-score phase fits the nonnegative regret and lost\-gain outputs\. We denote this complete configuration by Full\.

### IV\-DBudgeted selective expert allocation

The direct branch selects the highest\-scoring candidate and completes the packet with Local\-MDL\. For selective refinement,utu\_\{t\}ranks communication states and a screen bounds the candidates evaluated at each selected state\. Rate\-specific thresholds set the frequency of refinement\. The learned ACV configuration usesutACVu\_\{t\}^\{\\mathrm\{ACV\}\}; equal\-budget experiments compare it with student margin, Local uncertainty and random allocation\.

LetCrem,tC\_\{\\mathrm\{rem\},t\}denote the remaining candidate\-level budget for the current image and leta^t=arg⁡maxa⁡la\\hat\{a\}\_\{t\}=\\arg\\max\_\{a\}l\_\{a\}be the direct action\. At a selected state, the remaining budget determines the screen size\. We first form the priority listℒt=\[aL​\(st\),a^t,Rankl⁡\(𝒫⁡\(st\)∖\{aL​\(st\),a^t\}\)\]\\mathcal\{L\}\_\{t\}=\[a\_\{\\mathrm\{L\}\}\(s\_\{t\}\),\\hat\{a\}\_\{t\},\\operatorname\{Rank\}^\{l\}\(\\mathcal\{P\}\(s\_\{t\}\)\\setminus\\\{a\_\{\\mathrm\{L\}\}\(s\_\{t\}\),\\hat\{a\}\_\{t\}\\\}\)\]\. The screened proposal contains the firstktk\_\{t\}unique entries of this list:

kt\\displaystyle k\_\{t\}=min⁡\{Crem,t,\|𝒫⁡\(st\)\|\},\\displaystyle=\\min\\\{C\_\{\\mathrm\{rem\},t\},\|\\mathcal\{P\}\(s\_\{t\}\)\|\\\},\(10\)𝒫kt​\(st\)\\displaystyle\\mathcal\{P\}\_\{k\_\{t\}\}\(s\_\{t\}\)=FirstUniquekt⁡\(ℒt\)\.\\displaystyle=\\operatorname\{FirstUnique\}\_\{k\_\{t\}\}\(\\mathcal\{L\}\_\{t\}\)\.HereFirstUniquek\\operatorname\{FirstUnique\}\_\{k\}removes repeated positions in priority order and stops afterkkunique entries\. The screen is filled even when Local and the direct action coincide\. Refinement requires at least two available candidate slots\. Definea~t\(k\)=arg⁡maxa∈𝒫kt​AB​\(a∣st\)\\tilde\{a\}\_\{t\}^\{\(k\)\}=\\arg\\max\_\{a\\in\\mathcal\{P\}\_\{k\_\{t\}\}\}A\_\{B\}\(a\\mid s\_\{t\}\)\. After an expert evaluation, the budget state is updated byCrem,t\+1=Crem,t−\|𝒫kt​\(st\)\|C\_\{\\mathrm\{rem\},t\+1\}=C\_\{\\mathrm\{rem\},t\}\-\|\\mathcal\{P\}\_\{k\_\{t\}\}\(s\_\{t\}\)\|\. The online policy is therefore

πadapt​\(st\)=\{a^t,ut<τu,a~t\(k\),ut≥τu∧kt≥2,a^t,otherwise\.\\pi\_\{\\mathrm\{adapt\}\}\(s\_\{t\}\)=\\begin\{cases\}\\hat\{a\}\_\{t\},&u\_\{t\}<\\tau\_\{u\},\\\\ \\tilde\{a\}\_\{t\}^\{\(k\)\},&u\_\{t\}\\geq\\tau\_\{u\}\\ \\wedge\\ k\_\{t\}\\geq 2,\\\\ \\hat\{a\}\_\{t\},&\\text\{otherwise\.\}\\end\{cases\}\(11\)To relate these decisions to reconstruction quality, define the terminal values of the Local action, direct action, full proposal and screened proposal:

QL​\(st\)\\displaystyle Q\_\{\\mathrm\{L\}\}\(s\_\{t\}\)=QB​\(aL​\(st\)∣st\),\\displaystyle=Q\_\{B\}\(a\_\{\\mathrm\{L\}\}\(s\_\{t\}\)\\mid s\_\{t\}\),\(12\)QD​\(st\)\\displaystyle Q\_\{\\mathrm\{D\}\}\(s\_\{t\}\)=QB​\(a^t∣st\),\\displaystyle=Q\_\{B\}\(\\hat\{a\}\_\{t\}\\mid s\_\{t\}\),QE​\(st\)\\displaystyle Q\_\{\\mathrm\{E\}\}\(s\_\{t\}\)=maxa∈𝒫⁡\(st\)⁡QB​\(a∣st\),\\displaystyle=\\max\_\{a\\in\\mathcal\{P\}\(s\_\{t\}\)\}Q\_\{B\}\(a\\mid s\_\{t\}\),Qk​\(st\)\\displaystyle Q\_\{k\}\(s\_\{t\}\)=maxa∈𝒫kt​\(st\)⁡QB​\(a∣st\),\\displaystyle=\\max\_\{a\\in\\mathcal\{P\}\_\{k\_\{t\}\}\(s\_\{t\}\)\}Q\_\{B\}\(a\\mid s\_\{t\}\),H⁡\(st\)\\displaystyle H\(s\_\{t\}\)=QE​\(st\)−QL​\(st\),\\displaystyle=Q\_\{\\mathrm\{E\}\}\(s\_\{t\}\)\-Q\_\{\\mathrm\{L\}\}\(s\_\{t\}\),RD​\(st\)\\displaystyle R\_\{\\mathrm\{D\}\}\(s\_\{t\}\)=QE​\(st\)−QD​\(st\),\\displaystyle=Q\_\{\\mathrm\{E\}\}\(s\_\{t\}\)\-Q\_\{\\mathrm\{D\}\}\(s\_\{t\}\),Vk​\(st\)\\displaystyle V\_\{k\}\(s\_\{t\}\)=Qk​\(st\)−QD​\(st\)\.\\displaystyle=Q\_\{k\}\(s\_\{t\}\)\-Q\_\{\\mathrm\{D\}\}\(s\_\{t\}\)\.HereHHis the expert’s headroom over Local,RDR\_\{\\mathrm\{D\}\}is the direct action’s regret, andVkV\_\{k\}is the gain recoverable by the screen\. Forkt<2k\_\{t\}<2, the direct branch givesQk=QDQ\_\{k\}=Q\_\{\\mathrm\{D\}\}andVk=0V\_\{k\}=0\. Ifδt\\delta\_\{t\}indicates that the expert branch is evaluated, the adaptive value decomposes as

QAdaptive​\(st\)\\displaystyle Q\_\{\\mathrm\{Adaptive\}\}\(s\_\{t\}\)=QD​\(st\)\+δt​Vk​\(st\),\\displaystyle=Q\_\{\\mathrm\{D\}\}\(s\_\{t\}\)\+\\delta\_\{t\}V\_\{k\}\(s\_\{t\}\),\(13\)QAdaptive​\(st\)−QL​\(st\)\\displaystyle Q\_\{\\mathrm\{Adaptive\}\}\(s\_\{t\}\)\-Q\_\{\\mathrm\{L\}\}\(s\_\{t\}\)=H⁡\(st\)−RD​\(st\)\+δt​Vk​\(st\)\.\\displaystyle=H\(s\_\{t\}\)\-R\_\{\\mathrm\{D\}\}\(s\_\{t\}\)\+\\delta\_\{t\}V\_\{k\}\(s\_\{t\}\)\.Each called screen retains Local and the direct action under the evaluator’s deterministic continuation\. It therefore satisfiesQk​\(st\)≥max⁡\{QD​\(st\),QL​\(st\)\}Q\_\{k\}\(s\_\{t\}\)\\geq\\max\\\{Q\_\{\\mathrm\{D\}\}\(s\_\{t\}\),Q\_\{\\mathrm\{L\}\}\(s\_\{t\}\)\\\}andVk​\(st\)≥0V\_\{k\}\(s\_\{t\}\)\\geq 0\. Equation \([13](https://arxiv.org/html/2609.30756#S4.E13)\) separates the student’s recovered gain,H⁡\(st\)−RD​\(st\)H\(s\_\{t\}\)\-R\_\{\\mathrm\{D\}\}\(s\_\{t\}\), from the additional recoverable valueδt​Vk​\(st\)\\delta\_\{t\}V\_\{k\}\(s\_\{t\}\)supplied by exact screening\.

Each screened evaluation adds\|𝒫kt\|\|\\mathcal\{P\}\_\{k\_\{t\}\}\|toNcfN\_\{\\mathrm\{cf\}\}, so the invariantNcf​\(x\)≤CmaxN\_\{\\mathrm\{cf\}\}\(x\)\\leq C\_\{\\max\}holds by construction\. When the cap is large enough, the endpoint evaluates the full compact proposal\. The fixed\-screen experiments usekt=4k\_\{t\}=4for every eligible call\.

### IV\-ETraining and threshold calibration

Training uses image\-grouped fitting and monitor splits\. The monitor selects the direct checkpoint, which is then fixed during allocation\-head fitting\. For the common\-screen comparison, all four student families use student margin for state allocation\.

Cost\-based calibration chooses a threshold using realized candidate expenditure onMMmonitor images\. Let𝒯=\{\+∞\}∪\{ui:i∈𝒟cal\}\\mathcal\{T\}=\\\{\+\\infty\\\}\\cup\\\{u\_\{i\}:i\\in\\mathcal\{D\}\_\{\\mathrm\{cal\}\}\\\}, where𝒟cal\\mathcal\{D\}\_\{\\mathrm\{cal\}\}is the monitor image set\. Then

bcal​\(τ\)\\displaystyle b\_\{\\mathrm\{cal\}\}\(\\tau\)=1M​∑i=1Mki​𝟏​\{ui≥τ,ki≥2\},\\displaystyle=\\frac\{1\}\{M\}\\sum\_\{i=1\}^\{M\}k\_\{i\}\\mathbf\{1\}\\\{u\_\{i\}\\geq\\tau,\\ k\_\{i\}\\geq 2\\\},\(14\)τ∗​\(b0\)\\displaystyle\\tau^\{\*\}\(b\_\{0\}\)∈arg⁡maxτ∈𝒯:bcal​\(τ\)≤b0bcal\(τ\),\\displaystyle\\in\\underset\{\\tau\\in\\mathcal\{T\}:\\,b\_\{\\mathrm\{cal\}\}\(\\tau\)\\leq b\_\{0\}\}\{\\arg\\max\}\\;b\_\{\\mathrm\{cal\}\}\(\\tau\),wherekik\_\{i\}is the realized screen size\. Ties are resolved by choosing the largest threshold\. The\+∞\+\\inftythreshold provides a zero\-call option, and the selected point maximizes realized expenditure withinb0b\_\{0\}\. Applying the fixed threshold to validation yields the transferred average workloadqtestq\_\{\\mathrm\{test\}\}, while the per\-image cap remains enforced during execution\.

## VExperimental Protocol

### V\-ADatasets and comparison protocol

The main CIFAR\-10 study uses the frozen discrete tokenizer, masked\-prior receiver, codebook, compact proposal, packet syntax and reconstruction code from GCR\-C\. It contains 200 development and 200 image\-disjoint validation images, seven rates\{0\.16,0\.20,0\.28,0\.32,0\.40,0\.44,0\.52\}\\\{0\.16,0\.20,0\.28,0\.32,0\.40,0\.44,0\.52\\\}and2×200×7×2=5,6002\\times 200\\times 7\\times 2=5,600state groups across data roles, images, rates and initial/early states\. Development groups are used for parameter fitting, checkpoint selection, allocation\-head fitting and threshold calibration\. Validation groups provide the reported performance evaluation\. The primary end\-to\-end policies use three independently initialized selector seeds \(2026081720260817,2026081820260818, and2026081920260819\) on the same 200 validation images\. The independent threshold\-selection evaluation uses 100 Calibration and 400 disjoint Test images; measured\-cost transfer uses a 40\-image development monitor and the 200\-image validation split\. The STL\-10 and high\-resolution evaluations are described below\. The supplement specifies image assignments, training settings and the matched component comparisons\.

Within each image system, all methods share the tokenizer, packet syntax, CRC/FEC accounting, receiver and Local\-MDL continuation\. The primary end\-to\-end comparison evaluates the full proposal when the adaptive branch is selected \(kt=Ksk\_\{t\}=K\_\{s\},Cmax=8C\_\{\\max\}=8\)\. The common\-screen and allocation comparisons useCmax=4C\_\{\\max\}=4, retaining Local and the direct action before adding student\-ranked alternatives\. A common screening rule therefore gives each student its own four\-candidate set\. In the quota comparisons, the top⌊M​q/4⌋\\lfloor Mq/4\\rfloorstates receive a call; measured\-cost transfer instead applies a fixed calibration threshold\. Larger student margins receive higher priority in the margin control\.

The unified timing experiment measures Local, Exact\-Full, Local\-ranked\-4 and Full/A1 direct and adaptive paths in one warmed CUDA process\. A second experiment repeats the Local and expert branches five times\. Supplementary Table S14 defines their measurement boundaries and aggregation; Table S15 summarizes the data roles\. Image identity is the splitting and bootstrap unit\.

The scale\-transfer study uses a frozen LlamaGen VQ\-16 tokenizer and decoder at 384×\\times384, a 24×\\times24 token grid and a locally trained masked prior\. For the expert comparison, deterministic DIV2K crops train the prior, and eight DIV2K validation crops calibrate the packet budgets\. All 24 Kodak center crops are evaluated at two operating points\. Local\-MDL and Exact\-Full share an initial\-state proposal of Top\-3 Local and Top\-5 coverage candidates\. Each token uses 14 code bits; the packet has a 32\-bit header, one adaptive\-min position description, a 16\-bit CRC and a1\.25×1\.25\\timesFEC factor\. The calibrated budgets are 2,101 and 3,804 bits for target fractions 0\.15 and 0\.30, respectively\. The receiver completes the token sequence using the frozen masked prior and LlamaGen decoder under error\-free packet delivery\. Supplementary Table S12 reports a separately trained 256\-source student on Internal32 and Tecnick40, with training and allocation settings in the accompanying text\.

### V\-BMetrics

We report peak signal\-to\-noise ratio \(PSNR\), structural similarity \(SSIM\), actual bits per pixel \(bpp\), mean teacher regret, retained positive gain, catastrophic regret \(regret\>0\.5\>0\.5dB\), candidate\-level exact evaluationsNcfN\_\{\\mathrm\{cf\}\}/image,NcfN\_\{\\mathrm\{cf\}\}/MP,NcfN\_\{\\mathrm\{cf\}\}/token, expert invocations/image, internalNpriorN\_\{\\mathrm\{prior\}\}andNrollN\_\{\\mathrm\{roll\}\}counters, and synchronized CUDA runtime\. Gain recovery and candidate\-level query fraction are defined relative to the fixed Exact\-FullQBQ\_\{B\}reference as

GR⁡\(m\)=Qm−QLocalQExact−QLocal,QF⁡\(m\)=Ncf​\(m\)Ncf​\(Exact\)\.\\mathrm\{GR\}\(m\)=\\frac\{Q\_\{m\}\-Q\_\{\\mathrm\{Local\}\}\}\{Q\_\{\\mathrm\{Exact\}\}\-Q\_\{\\mathrm\{Local\}\}\},\\qquad\\mathrm\{QF\}\(m\)=\\frac\{N\_\{\\mathrm\{cf\}\}\(m\)\}\{N\_\{\\mathrm\{cf\}\}\(\\mathrm\{Exact\}\)\}\.\(15\)An expert invocation sends one communication state to the exact branch and can incur several candidate evaluations\. Encoder runtime is measured with synchronized CUDA boundaries and complements the candidate countNcfN\_\{\\mathrm\{cf\}\}\. Across\-seed results are reported as means±\\pmstandard deviations\. Paired image\-bootstrap intervals resample image IDs while retaining all associated states and rates\. The primary quality results use three selector seeds; the unified runtime comparison uses a fixed seed and paired image\-bootstrap intervals\.

## VIResults

We first examine how the terminal objective changes candidate ranking and end\-to\-end reconstruction quality\. We then evaluate candidate screening, budget allocation, measured runtime and transfer across image systems\.

### VI\-ATerminal horizon changes candidate ranking

We use the same compact proposal, state and receiver to compare an immediate one\-step targetQimmQ\_\{\\mathrm\{imm\}\}with the full\-budget terminal targetQBQ\_\{B\}\. The immediate branch stops after adding the candidate, whereas the terminal branch completes the remaining packet with Local\-MDL\. We report the terminal quality of the immediate choice, terminal regret and action agreement\.

At 0\.20 bpp, immediate\-value selection achieves a terminal gain of\+0\.198\+0\.198dB over Local, compared with\+1\.353\+1\.353dB for terminal\-value selection\. The immediate choice incurs 1\.155 dB terminal regret and agrees with the terminal choice on 23\.0% of images\. Accounting for the remaining transmission therefore changes both the preferred action and its reconstruction benefit\.

Fig\. 2:Effect of evaluation horizon on candidate selection\. Immediate and full\-budget selection use the same proposal, state and receiver\. Both selected actions are evaluated using the full\-budget terminal objective\.
### VI\-BEnd\-to\-end reconstruction quality

Table[I](https://arxiv.org/html/2609.30756#S6.T1)evaluates Local, Direct, Adaptive and Exact\-FullQBQ\_\{B\}under the same packet and receiver settings\. Three independently selected Full checkpoints are evaluated on the same 200 validation images; Local and Exact\-Full are fixed references\.

At 0\.20 bpp, Direct gains0\.313±0\.0880\.313\\pm 0\.088dB with zero exact candidate evaluations\. In the primary full\-proposal adaptive configuration, the gain reaches0\.636±0\.1190\.636\\pm 0\.119dB with2\.13±0\.532\.13\\pm 0\.53candidate evaluations per image, recovering 46\.99% of the Exact\-Full expert gain with 27\.60% of its candidate calls\. The synchronized runtime comparison in Fig\.[7](https://arxiv.org/html/2609.30756#S6.F7)evaluates Direct and the four\-candidate Adaptive\-4 configuration with fixed Full and A1 checkpoints\.

Reconstruction gains decrease as the packet budget increases\. At 0\.32 and 0\.44 bpp, Adaptive recovers 18\.22% and 12\.18% of the expert gain using 16\.05% and 8\.14% of its calls\. The expert’s available gain itself falls from 1\.353 dB at 0\.20 bpp to 0\.060 dB at 0\.44 bpp\.

TABLE I:End\-to\-end quality and candidate workload under identical packet and receiver settings\. Direct and Adaptive ACV\-Gate entries are means±\\pmstandard deviations over three selector seeds on the same 200 image\-disjoint validation images; Local and Exact\-FullQBQ\_\{B\}are fixed 200\-image protocol references\. Adaptive uses the full proposal on called images \(Cmax=8C\_\{\\max\}=8\)\.NcfN\_\{\\mathrm\{cf\}\}counts candidate\-level exact calls\. Gain recovery \(GR\) and exact\-query fraction \(QF\) use the corresponding fixed Exact\-FullQBQ\_\{B\}row as denominator\. Actual bpp is 0\.1986, 0\.3156 and 0\.4375 at the three rates\.The paired image intervals for the adaptive gain are positive at 0\.20 bpp and cross zero at 0\.32 and 0\.44 bpp\. Low\-rate communication therefore offers the clearest benefit from allocating terminal reasoning, while higher rates provide less reconstruction headroom\.

Selective refinement also reduces the frequency of reconstruction losses relative to Local\. The fraction of validation images withΔ\\DeltaPSNR below−0\.1\-0\.1dB is 31\.5% for Direct and 23\.0% for Adaptive at 0\.20 bpp; the corresponding fractions are 19\.0%/18\.0% at 0\.32 bpp and 4\.0%/3\.5% at 0\.44 bpp\. The largest reduction occurs at 0\.20 bpp, consistent with the larger mean benefit of refinement at low rate\.

### VI\-CCommon\-screen terminal\-value comparison

Four student families use the same four\-candidate screening rule and student\-margin allocation\. Their candidate rankings determine the direct action and the alternatives in the screen\. We evaluate these choices with cached terminalQBQ\_\{B\}values\. Each point averages 200 validation images and three selector seeds\.

At 0\.20 bpp, Full, A1 Set\-context, Row\-valueMSE and DeepSets\-regret\-soft achieve direct gains of 0\.313, 0\.430, 0\.402 and 0\.414 dB, respectively\. These direct decisions require zero exact candidate evaluations\. At a target of one candidate evaluation per image, the gains rise to 0\.545, 0\.595, 0\.593 and 0\.603 dB\. At the four\-call endpoint, they reach 1\.119, 1\.142, 1\.077 and 1\.166 dB\. At 0\.32 bpp, the one\-call gains are 0\.093, 0\.128, 0\.124 and 0\.032 dB; at 0\.44 bpp, all methods remain within 0\.006 dB of Local at this budget\.

The four\-call endpoint isolates candidate screening because every image receives the same evaluation budget\. At 0\.20 bpp, learned screens achieve 1\.077–1\.166 dB across the four students, compared with 0\.271 dB for Local\-ranked\-4\. Ranking candidates by learned terminal value thus retains substantially more reconstruction gain within the same four exact evaluations\.

Fig\. 3:Common\-screen reconstruction gain versus candidate workload\. The four student families share a four\-candidate screening rule and student\-margin allocation\. Values are meanΔ\\DeltaPSNR relative to Local over three seeds; bars show seed standard deviations\. Terminal outcomes are evaluated from cachedQBQ\_\{B\}values\.Figure[4](https://arxiv.org/html/2609.30756#S6.F4)compares Full and DeepSets\-regret\-soft on paired images\. DeepSets\-regret\-soft reaches 0\.874 dB at two calls per image and 0\.20 bpp, with the relative ranking changing across rates\. The subsequent allocation experiments use Full, which combines the set\-aware student and learned allocation head\.

Fig\. 4:Paired direct gains for Full and DeepSets\-regret\-soft on the same validation images\. The diagonal denotes equal gain, and the zero axes mark the Local baseline\. Annotations report the mean gains\.
### VI\-DMeasured candidate expenditure and threshold transfer

We calibrate a threshold for a target of one candidate evaluation per image using 40 development images per rate, reserving every fifth development image\. Each selected state contributes its realized screen size to the calibration cost\. We then apply the fixed threshold to the 200 validation images\. For Full, the realized workloads are 0\.71, 0\.60 and 0\.99 evaluations per image at 0\.20, 0\.32 and 0\.44 bpp, respectively\. The corresponding paired gains are 0\.482, 0\.072 and 0\.014 dB\. The low\-rate interval excludes zero, and every validation image satisfies the four\-call cap\.

Fig\. 5:Actual\-cost threshold transfer for Full ACV\-Gate\. Thresholds are calibrated on a reserved development monitor at a one\-call target and frozen before validation\. Error bars are paired image\-bootstrap intervals\.
### VI\-EFixed\-screen allocation\-score controls

To isolate state allocation, we fix the Full direct student and its four\-candidate screen, then vary the score used before exact evaluation\. We compare random, student\-margin, Local\-uncertainty and ACV\-risk allocation with an oracle that ranks states by the realized screen valueV4V\_\{4\}\. The oracle uses the same four candidates and measures the available state\-ranking headroom\. At 0\.20 bpp and one candidate call per image, Random, Student margin, Local uncertainty, and ACV risk achieve 0\.526, 0\.544, 0\.491, and 0\.528 dB, respectively, while the same\-screenV4V\_\{4\}oracle reaches 1\.048 dB\. These matched\-budget results place the practical allocation scores in a similar operating range and identify state ranking as the main remaining source of allocation headroom\. At 0\.32 bpp the corresponding practical scores give 0\.111, 0\.093, 0\.074 and 0\.118 dB, while the oracle gives 0\.437 dB\. At 0\.44 bpp all practical scores remain at or below 0\.015 dB at this budget, reflecting the small terminal headroom\. At the low\-rate one\-call budget, random allocation adds 0\.213 dB to the 0\.313 dB direct gain, and learned\-risk allocation adds a further 0\.002 dB\. Thus, screened refinement accounts for most of the observed gain at this operating point, while the oracle gap quantifies the potential of improved state ranking\. The complete allocation curves are shown in Fig\.[6](https://arxiv.org/html/2609.30756#S6.F6); the reported one\-call values use the same validation images and seed aggregation as the common\-screen replay\.

Fig\. 6:Fixed\-screen allocation\-score controls\. The four practical scores share the same Full direct student, four\-candidate screen and target budgets; the post\-hoc oracle ranks by the realizedV4V\_\{4\}screen value and gives the allocation upper bound\. Error bars denote seed standard deviations\.
### VI\-FMeasured runtime–quality trade\-off

We measure Local\-MDL, Exact\-FullQBQ\_\{B\}, student\-free Local\-ranked\-4, Full Direct, Full Adaptive\-4, A1 Direct and A1 Adaptive\-4 in one warmed CUDA process\. The same 200 validation images and three operating rates are used throughout; the frozen Full and A1 checkpoints are the development\-selected seed\-2026081720260817instances\. The student paths use one timed execution per image/rate\. A separate branch\-timing experiment uses five repeated measurements and the per\-image median\. All paths use batch size one and synchronized boundaries\. CPU tokenization is measured once per image and added to each path; the student paths share one measured Local/proposal prefix, whose duration is included in every student row\. Supplementary Table S14 specifies the timing protocols\.

Figure[7](https://arxiv.org/html/2609.30756#S6.F7)reports the resulting matched measurements\. The upper panels plot paired PSNR gain against synchronized total runtime; the lower panels plot the same gain against realizedNcfN\_\{\\mathrm\{cf\}\}/image\. The dashed curves mix Local with Exact\-Full \(B1\) or Local\-ranked\-4 \(B2\) using measured branch times and five fixed routing seeds\. Each operating point therefore specifies both the reconstruction gain and its measured computational cost\. The timing outputs also record the 95th\-percentile per\-image runtime\.

In the unified run at 0\.20 bpp, Full Direct achieves\+0\.386\+0\.386dB at 337\.0 ms with zero exact candidate evaluations\. Full Adaptive\-4 achieves\+0\.459\+0\.459dB at 365\.1 ms with 0\.66 evaluations per image\. The corresponding A1 gains are\+0\.433\+0\.433dB at 322\.0 ms and\+0\.638\+0\.638dB at 406\.4 ms; A1 Adaptive\-4 uses 1\.32 evaluations per image\. At 0\.32 bpp, Full Direct/Adaptive\-4 gains are\+0\.080\+0\.080/\+0\.130\+0\.130dB at 464\.9/577\.5 ms with 0/0\.82 calls\. A1 Direct/Adaptive\-4 gains are\+0\.087\+0\.087/\+0\.180\+0\.180dB at 452\.5/585\.9 ms with 0/0\.94 calls\. At 0\.44 bpp, all student gains are below\+0\.013\+0\.013dB while the Exact\-Full reference remains\+0\.060\+0\.060dB\.

The repeated branch\-timing experiment gives Local runtimes of 123\.5, 264\.0 and 387\.2 ms at 0\.20, 0\.32 and 0\.44 bpp, respectively\. Exact\-Full gives\+1\.353/\+0\.521/\+0\.060\+1\.353/\+0\.521/\+0\.060dB and 830\.3/1587\.2/2194\.3 ms, while Local\-ranked\-4 gives\+0\.271/\+0\.061/\+0\.006\+0\.271/\+0\.061/\+0\.006dB and 416\.1/900\.8/1319\.2 ms with four calls per image\. At a routing probability ofp=0\.5p=0\.5and 0\.20 bpp, B1 obtains\+0\.680\+0\.680dB at 482\.8 ms with 3\.94 calls per image, and B2 obtains\+0\.138\+0\.138dB at 273\.8 ms with 2\.05 calls\. These operating points quantify the extra computation required to improve on Local reconstruction\.

Fig\. 7:Measured runtime–quality and workload–quality frontiers\. Upper panels use synchronized total time; lower panels use realized candidate calls per image\. Error bars are paired image\-bootstrap intervals\. Solid markers are measured branches; dashed curves are B1 Local/Exact\-Full and B2 Local/Local\-ranked\-4 route mixtures over the fixed route seeds\. Full and A1 student paths use their frozen seed\-2026081720260817instances; the repeated five\-pass branch measurements are given in Supplementary Table S4\.
### VI\-GComparison with adapted token\-selection methods

We adapt the diversity, visual\-cue and saliency–coverage rules of DivPrune, VisPruner and SCOPE from large vision–language models to packetized reconstruction\. Each adapter selects from the ACV proposal using current\-state features, then completes the packet with Local\-MDL\. All methods use the same 200 validation images, tokenizer, receiver and packet accounting\. Random\-proposal provides a proposal\-level control, and the Exact\-FullQBQ\_\{B\}expert provides the terminal\-value reference\.

Figure[8](https://arxiv.org/html/2609.30756#S6.F8)compares the reconstruction gains at each candidate workload\. At 0\.20 bpp, VisPruner\-adapted improves over Local by 0\.128 dB, while DivPrune\-adapted, SCOPE\-adapted and Random\-proposal change PSNR by \+0\.015,−0\.062\-0\.062and \+0\.061 dB, respectively\. At 0\.32 bpp, all three structured adapters are at or below Local and DivPrune\-adapted is−0\.148\-0\.148dB\. At 0\.44 bpp, all adapters using zero exact candidate evaluations are within 0\.006 dB of Local, consistent with the small terminal headroom\. The Exact\-FullQBQ\_\{B\}expert gains 1\.353, 0\.521 and 0\.060 dB at the three rates with 7\.705 candidate evaluations per image on average\. At 0\.20 bpp, ACV\-Gate Direct achieves \+0\.313±\\pm0\.088 dB with zero exact candidate evaluations, exceeding the gains of the three adapted rules\. Adaptive reaches \+0\.636±\\pm0\.119 dB with2\.13±0\.532\.13\\pm 0\.53calls\. Terminal supervision thus improves the low\-rate direct decision, while selective evaluation recovers additional quality at an intermediate candidate cost\.

Fig\. 8:Reconstruction gain of adapted token\-selection methods and ACV\-Gate under the same packet and receiver\. Direct and Adaptive are three\-seed Full means, with error bars denoting seed standard deviations\. TheNcfN\_\{\\mathrm\{cf\}\}column in each panel reports mean candidate calls per image\.
### VI\-HEvaluation on STL\-10

We train the selector and allocation mechanism for a separate 96×\\times96 STL\-10 tokenizer/receiver and evaluate them on 200 official test images\. Another 200 images form the development set, which is divided into fitting and monitoring subsets\. All three selector seeds are included in the end\-to\-end evaluation\. The budget isB=round⁡\(r×1024\)B=\\operatorname\{round\}\(r\\times 1024\), whererris the normalized budget parameter shown on the horizontal axis\. Actual bpp equals the serialized packet length divided by96296^\{2\}\.

Figure[9](https://arxiv.org/html/2609.30756#S6.F9)reports paired gains relative to the STL\-10 Local baseline\. Full Adaptive gains\+0\.442±0\.041\+0\.442\\pm 0\.041,\+0\.140±0\.046\+0\.140\\pm 0\.046and\+0\.110±0\.020\+0\.110\\pm 0\.020dB atr=0\.20r=0\.20,0\.320\.32and0\.440\.44, respectively\. The paired image intervals exclude zero at all three operating points\. Exact\-Full achieves larger gains with 7\.78 candidate evaluations per image\. This separately trained system reproduces the gain from selective computation and its dependence on the transmission budget\.

Fig\. 9:Full adaptive ACV\-Gate and the fixed Exact\-FullQBQ\_\{B\}reference on the independent STL\-10 evaluation\. Adaptive points are three\-seed means with seed standard deviations; the dashed Exact\-Full curve is annotated with its 7\.78 candidate\-level calls per image\. The reported paired image intervals for the adaptive branch are positive at all three rates\.
### VI\-IHigh\-resolution terminal\-value learning and selective computation

The Kodak\-24 evaluation measures the reconstruction benefit and cost of terminal evaluation on a 24×\\times24 token grid at 384×\\times384\. The initial\-state Exact\-Full expert improves PSNR over Local by 1\.033 dB at 0\.0142 bpp and 0\.847 dB at approximately 0\.0257 bpp, using eight candidate calls per image\. Its measured selection runtime is 8\.9 and 8\.4 times that of Local, respectively \(Supplementary Fig\. S3\)\.

At this resolution, a separately trained 256\-source student achieves low\-rate adaptive gains of 0\.553 dB on Internal32 and 0\.520 dB on Tecnick40\. The corresponding workloads are 0\.90 and 0\.98 candidate evaluations per image \(Supplementary Table S12\)\. Both low\-rate adaptive image\-bootstrap intervals exclude zero; higher\-rate gains decrease and vary by dataset\. Internal32 tests generalization to sources held out from student training, and Tecnick40 provides an external source distribution\. The supplement reports the five\-seed results, allocation settings and source assignments for prior and student training\.

## VIIDiscussion

Terminal reconstruction value links token selection to encoder computation\. The student ranks candidates with zero exact evaluations, and selective refinement evaluates a compact set of alternatives\. Equation \([13](https://arxiv.org/html/2609.30756#S4.E13)\) separates the student’s recovered gain from the additional value obtained through screening\.

The return on computation depends on the transmission budget\. At low rates, the receiver generates more content, so a selected token can substantially affect completion\. Larger budgets leave less Exact\-Full headroom over Local, reducing the benefit of additional evaluation\. The separately trained STL\-10 and384×384384\\times 384systems likewise concentrate their gains at low rates\.

Compact students retain useful terminal rankings: A1 Set\-context and DeepSets\-regret\-soft match or exceed Full at several common\-screen operating points\. State allocation provides further headroom\. At 0\.20 bpp and one candidate evaluation per image, practical scores achieve 0\.491–0\.544 dB, compared with 1\.048 dB for the same\-screen oracle\. This gap identifies the potential of learning when exact refinement is most valuable\.

Deployment requires both workload and runtime measurements\. The cap bounds exact evaluations; runtime also includes feature extraction, student inference, continuation and reconstruction\. Their joint frontiers guide the choice of computational budget\. Extensions include budget reservation across transmission states, perceptual terminal objectives and channel\-conditioned value estimates\.

## VIIIConclusion

ACV\-Gate combines terminal\-value learning, bounded screening and budgeted state allocation for visual token communication\. On CIFAR\-10 at 0\.20 bpp, Direct improves PSNR over Local\-MDL by0\.313±0\.0880\.313\\pm 0\.088dB with zero exact candidate evaluations\. The primary full\-proposal Adaptive configuration achieves0\.636±0\.1190\.636\\pm 0\.119dB using2\.13±0\.532\.13\\pm 0\.53evaluations per image, or 27\.60% of the Exact\-Full workload\. Matched four\-candidate experiments identify terminal ranking and screening as the main sources of gain; allocation controls quantify further headroom in state ranking\. Separately trained STL\-10 and384×384384\\times 384systems also achieve their largest gains at low rates\. Measured runtime and candidate workload provide explicit operating points for allocating encoder computation under a fixed packet budget\.

## Data and code availability

The experiments use CIFAR\-10, STL\-10, DIV2K, Kodak and Tecnick images\. The accompanying source repository contains the selector, packet accounting, training and evaluation scripts, with instructions for supplying datasets and model artifacts\.

## References

- \[1\]J\. Ballé, V\. Laparra, and E\. P\. Simoncelli, “End\-to\-end optimized image compression,” in*Proc\. ICLR*, 2017\.
- \[2\]J\. Ballé, D\. Minnen, S\. Singh, S\. J\. Hwang, and N\. Johnston, “Variational image compression with a scale hyperprior,” in*Proc\. ICLR*, 2018\.
- \[3\]D\. Minnen, J\. Ballé, and G\. D\. Toderici, “Joint autoregressive and hierarchical priors for learned image compression,” in*Proc\. NeurIPS*, vol\. 31, 2018\.
- \[4\]A\. van den Oord, O\. Vinyals, and A\. Kavukcuoglu, “Neural discrete representation learning,” in*Proc\. NeurIPS*, vol\. 30, 2017\.
- \[5\]P\. Esser, R\. Rombach, and O\. Ommer, “Taming transformers for high\-resolution image synthesis,” in*Proc\. CVPR*, 2021, pp\. 12873–12883\.
- \[6\]H\. Chang, H\. Zhang, L\. Jiang, C\. Liu, and W\. T\. Freeman, “MaskGIT: Masked generative image transformer,” in*Proc\. CVPR*, 2022, pp\. 11315–11325\.
- \[7\]E\. Bourtsoulatze, D\. B\. Kurka, and D\. Gunduz, “Deep joint source\-channel coding for wireless image transmission,”*IEEE Trans\. Cogn\. Commun\. Netw\.*, vol\. 5, no\. 3, pp\. 567–579, 2019\.
- \[8\]H\. Xie, Z\. Qin, G\. Y\. Li, and B\.\-H\. Juang, “Deep learning enabled semantic communication systems,”*IEEE Trans\. Signal Process\.*, vol\. 69, pp\. 2663–2675, 2021\.
- \[9\]T\. Han, J\. Tang, Q\. Yang, Y\. Duan, Z\. Zhang, and Z\. Shi, “Generative model based highly efficient semantic communication approach for image transmission,” arXiv:2211\.10287, 2022\.
- \[10\]G\. Zhang, H\. Li, Y\. Cai, Q\. Hu, G\. Yu, and Z\. Qin, “Progressive learned image transmission for semantic communication using hierarchical VAE,”*IEEE Trans\. Cogn\. Commun\. Netw\.*, vol\. 11, no\. 6, pp\. 3640–3654, 2025, doi: 10\.1109/TCCN\.2025\.3546935\.
- \[11\]Y\. Li, X\. Chen, X\. Deng, and J\. Gui, “Content adaptive distributed joint source\-channel coding for image transmission with hyperprior,”*IEEE Trans\. Cogn\. Commun\. Netw\.*, vol\. 11, no\. 1, pp\. 105–117, 2025, doi: 10\.1109/TCCN\.2024\.3438371\.
- \[12\]F\. Pezone, S\. Barbarossa, and G\. Caire, “SQ\-GAN: Semantic image communications using masked vector quantization,”*IEEE Trans\. Cogn\. Commun\. Netw\.*, early access, 2025, doi: 10\.1109/TCCN\.2025\.3620819\.
- \[13\]S\. Tong, X\. Yu, R\. Li, K\. Lu, Z\. Zhao, and H\. Zhang, “Alternate learning\-based SNR\-adaptive sparse semantic visual transmission,”*IEEE Trans\. Wireless Commun\.*, vol\. 24, no\. 2, pp\. 1737–1752, Feb\. 2025, doi: 10\.1109/TWC\.2024\.3512652\.
- \[14\]Y\. Rao, W\. Zhao, B\. Liu, J\. Lu, J\. Zhou, and C\.\-J\. Hsieh, “DynamicViT: Efficient vision transformers with dynamic token sparsification,” in*Proc\. NeurIPS*, vol\. 34, 2021\.
- \[15\]Y\. Liang, C\. Ge, Z\. Tong, Z\. Song, J\. Wang, and P\. Xie, “Not all patches are what you need: Expediting vision transformers via token reorganizations,” in*Proc\. ICLR*, 2022\.
- \[16\]M\. S\. Ryoo, A\. Piergiovanni, A\. Arnab, M\. Dehghani, and A\. Angelova, “TokenLearner: What can 8 learned tokens do for images and videos?” in*Proc\. NeurIPS*, vol\. 34, 2021\.
- \[17\]H\. Yin, A\. Vahdat, J\. M\. Alvarez, A\. Mallya, A\. Kautz, and P\. Molchanov, “A\-ViT: Adaptive tokens for efficient vision transformer,” in*Proc\. CVPR*, 2022, pp\. 10809–10818\.
- \[18\]D\. Bolya, C\.\-Y\. Fu, X\. Dai, P\. Zhang, C\. Feichtenhofer, and J\. Hoffman, “Token merging: Your ViT but faster,” in*Proc\. ICLR*, 2023\.
- \[19\]Y\. Geifman and R\. El\-Yaniv, “Selective classification for deep neural networks,” in*Proc\. NeurIPS*, vol\. 30, 2017\.
- \[20\]B\. Lakshminarayanan, A\. Pritzel, and C\. Blundell, “Simple and scalable predictive uncertainty estimation using deep ensembles,” in*Proc\. NeurIPS*, vol\. 30, 2017\.
- \[21\]A\. N\. Angelopoulos, S\. Bates, A\. Fisch, L\. Lei, and T\. Schuster, “Conformal risk control,” arXiv:2208\.02814, 2022\.
- \[22\]H\. Mozannar and H\. Sontag, “Consistent estimators for learning to defer to an expert,” in*Proc\. ICML*, 2020, pp\. 7076–7087\.
- \[23\]P\. Hemmer, L\. Thede, M\. Vössing, J\. Jakubik, and N\. Kühl, “Learning to defer with limited expert predictions,” arXiv:2304\.07306, 2023\.
- \[24\]S\. R\. Alvar, G\. Singh, M\. Akbari, and Y\. Zhang, “DivPrune: Diversity\-based visual token pruning for large multimodal models,” in*Proc\. IEEE/CVF Conf\. Comput\. Vis\. Pattern Recognit\. \(CVPR\)*, 2025, pp\. 9392–9401\.
- \[25\]Q\. Zhang, A\. Cheng, M\. Lu, R\. Zhang, Z\. Zhuo, J\. Cao, S\. Guo, Q\. She, and S\. Zhang, “Beyond text\-visual attention: Exploiting visual cues for effective token pruning in VLMs,” in*Proc\. IEEE/CVF Int\. Conf\. Comput\. Vis\. \(ICCV\)*, 2025, pp\. 20857–20867\.
- \[26\]J\. Deng, W\. Li, J\. T\. Zhou, and Y\. He, “SCOPE: Saliency\-coverage oriented token pruning for efficient multimodel LLMs,” in*Advances in Neural Information Processing Systems*, vol\. 38, 2025\.
- \[27\]Y\. Zhang, C\.\-K\. Fan, J\. Ma, W\. Zheng, T\. Huang, K\. Cheng, D\. A\. Gudovskiy, T\. Okuno, Y\. Nakata, K\. Keutzer, and S\. Zhang, “SparseVLM: Visual token sparsification for efficient vision\-language model inference,” in*Proc\. Int\. Conf\. Mach\. Learn\. \(ICML\)*, vol\. 267, 2025, pp\. 74840–74857\.
- \[28\]Y\. Shang, M\. Cai, B\. Xu, Y\. J\. Lee, and Y\. Yan, “LLaVA\-PruMerge: Adaptive token reduction for efficient large multimodal models,” in*Proc\. IEEE/CVF Int\. Conf\. Comput\. Vis\. \(ICCV\)*, 2025, pp\. 22857–22867\.
- \[29\]W\. Ye, Q\. Wu, W\. Lin, and Y\. Zhou, “Fit and Prune: Fast and training\-free visual token pruning for multi\-modal large language models,” in*Proc\. AAAI Conf\. Artif\. Intell\.*, vol\. 39, no\. 21, 2025, pp\. 22128–22136\.
- \[30\]Y\. Jiang, Q\. Wu, W\. Lin, W\. Yu, and Y\. Zhou, “What kind of visual tokens do we need? Training\-free visual token pruning for multi\-modal large language models from the perspective of graph,” in*Proc\. AAAI Conf\. Artif\. Intell\.*, vol\. 39, no\. 4, 2025, pp\. 4075–4083\.
- \[31\]M\. Marchetti, D\. Traini, D\. Ursino, and L\. Virgili, “Efficient token pruning in vision transformers using an attention\-based multilayer network,”*Expert Syst\. Appl\.*, vol\. 279, Art\. no\. 127449, 2025, doi: 10\.1016/j\.eswa\.2025\.127449\.
- \[32\]Y\. Wang, J\. Wu, Z\. Ni, L\. Yang, Y\. Liu, C\. Yang, Y\. Wen, L\. He, X\. Tang, H\. Liu, and Y\. Zhou, “When token pruning is worse than random: Understanding visual token information in VLLMs,” in*Proc\. IEEE/CVF Conf\. Comput\. Vis\. Pattern Recognit\. \(CVPR\)*, 2026, pp\. 31910–31919\.
- \[33\]J\. Guo*et al\.*, “Baseline\-relative counterfactual refinement for bit\-aware visual token communication,” arXiv:2608\.16192, 2026, doi: 10\.48550/arXiv\.2608\.16192\.

相似文章

学习自适应推理路径以实现高效视觉推理

Hugging Face Daily Papers

AVR是一种自适应视觉推理框架,能够动态选择最优推理格式,在视觉推理任务中减少50-90%的token使用量同时保持准确性。该方法通过将视觉推理分解为三种认知功能并使用FS-GRPO训练来鼓励高效格式选择,从而解决推理路径冗余问题。

视觉语言模型推理中的视觉访问边界

arXiv cs.AI

本文介绍了Visual Access Sweep,一种因果干预方法,用于衡量视觉语言模型推理所需的最小图像标记访问量,并发现思维链(Chain-of-Thought)提示并非主要通过延长直接图像访问来提升性能,而是通过在视觉信息上扩展语言侧计算来实现。

通过闭环验证推理解锁复杂视觉生成

Hugging Face Daily Papers

介绍CLVR(闭环视觉推理),一种将文本到图像生成从单步过程重构为闭环多步视觉推理方法的框架,使用VLM控制器和扩散模型,在组合提示上实现了改进的性能。