Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models

Hugging Face Daily Papers Papers

Summary

This paper investigates when state adaptation during inference matters for masked diffusion language models (MDMs), organizing inference into five axes and showing that selective adaptation (using lightweight detectors to identify high-opportunity states) captures a large share of oracle gains, e.g., 56.9% of opportunity while adapting only the top 10% of states on LLaDA-8B constrained JSON filling.

Masked diffusion language models (MDMs) admit flexible generation orders, making the unmasking strategy an inference decision. Existing methods vary in how they prioritize positions, control parallelism, restrict selection regions, revise predictions, or plan future denoising, yet it remains unclear when these choices should change during generation. We study this question through strategy reversals, where an alternative action becomes preferable to a fixed choice. We organize MDM inference into five axes--score, cardinality, region, commitment, and planning--and define adaptation opportunity as the one-step utility advantage of the best candidate action over a validation-selected fixed action. This view shows that adaptation value depends on both the frequency and magnitude of such reversals. Across three MDMs and ten tasks, adaptation opportunities are highly heterogeneous, with some regimes exhibiting concentrated and predictable one-step gains. This motivates selective adaptation: lightweight detectors calibrated on validation prompts identify high-opportunity states, capturing, for example, 56.9 percent of the candidate-set oracle opportunity by adapting only the top 10 percent of states on LLaDA-8B constrained JSON filling. Our transition-level results suggest that state adaptation is most useful when applied selectively rather than uniformly.
Original Article
View Cached Full Text

Cached at: 10/01/26, 08:23 AM

Paper page - Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models

Source: https://huggingface.co/papers/2609.33355

Abstract

Maskeddiffusionlanguagemodels(MDMs)admitflexiblegenerationorders,makingtheunmaskingstrategyaninferencedecision.Existingmethodsvaryinhowtheyprioritizepositions,controlparallelism,restrictselectionregions,revisepredictions,orplanfuturedenoising,yetitremainsunclearwhenthesechoicesshouldchangeduringgeneration.Westudythisquestionthroughstrategyreversals,whereanalternativeactionbecomespreferabletoafixedchoice.WeorganizeMDMinferenceintofiveaxes--score,cardinality,region,commitment,andplanning--anddefineadaptationopportunityastheone-steputilityadvantageofthebestcandidateactionoveravalidation-selectedfixedaction.Thisviewshowsthatadaptationvaluedependsonboththefrequencyandmagnitudeofsuchreversals.AcrossthreeMDMsandtentasks,adaptationopportunitiesarehighlyheterogeneous,withsomeregimesexhibitingconcentratedandpredictableone-stepgains.Thismotivatesselectiveadaptation:lightweightdetectorscalibratedonvalidationpromptsidentifyhigh-opportunitystates,capturing,forexample,56.9percentofthecandidate-setoracleopportunitybyadaptingonlythetop10percentofstatesonLLaDA-8BconstrainedJSONfilling.Ourtransition-levelresultssuggestthatstateadaptationismostusefulwhenappliedselectivelyratherthanuniformly.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.33355

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.33355 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.33355 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.33355 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Masked Diffusion Decoding as $x$-Prediction Flow

arXiv cs.CL

This paper reinterprets masked diffusion language model decoding as continuous clean-state prediction, introducing a flow-based framework where tokens are updated continuously and asynchronously based on confidence, achieving 97% of LLaDA's performance with 25% of the decoding budget.

Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models

arXiv cs.CL

Proposes AdaLook, an adaptive multi-step lookahead decoding framework for masked diffusion language models that dynamically determines rollout depth and branch expansion based on candidate-score variance, achieving better accuracy-decoding steps trade-off compared to existing one-step lookahead decoding methods.