Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models
Summary
This paper investigates when state adaptation during inference matters for masked diffusion language models (MDMs), organizing inference into five axes and showing that selective adaptation (using lightweight detectors to identify high-opportunity states) captures a large share of oracle gains, e.g., 56.9% of opportunity while adapting only the top 10% of states on LLaDA-8B constrained JSON filling.
View Cached Full Text
Cached at: 10/01/26, 08:23 AM
Paper page - Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models
Source: https://huggingface.co/papers/2609.33355
Abstract
Maskeddiffusionlanguagemodels(MDMs)admitflexiblegenerationorders,makingtheunmaskingstrategyaninferencedecision.Existingmethodsvaryinhowtheyprioritizepositions,controlparallelism,restrictselectionregions,revisepredictions,orplanfuturedenoising,yetitremainsunclearwhenthesechoicesshouldchangeduringgeneration.Westudythisquestionthroughstrategyreversals,whereanalternativeactionbecomespreferabletoafixedchoice.WeorganizeMDMinferenceintofiveaxes--score,cardinality,region,commitment,andplanning--anddefineadaptationopportunityastheone-steputilityadvantageofthebestcandidateactionoveravalidation-selectedfixedaction.Thisviewshowsthatadaptationvaluedependsonboththefrequencyandmagnitudeofsuchreversals.AcrossthreeMDMsandtentasks,adaptationopportunitiesarehighlyheterogeneous,withsomeregimesexhibitingconcentratedandpredictableone-stepgains.Thismotivatesselectiveadaptation:lightweightdetectorscalibratedonvalidationpromptsidentifyhigh-opportunitystates,capturing,forexample,56.9percentofthecandidate-setoracleopportunitybyadaptingonlythetop10percentofstatesonLLaDA-8BconstrainedJSONfilling.Ourtransition-levelresultssuggestthatstateadaptationismostusefulwhenappliedselectivelyratherthanuniformly.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2609\.33355
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.33355 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.33355 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.33355 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Masked Diffusion Decoding as $x$-Prediction Flow
This paper reinterprets masked diffusion language model decoding as continuous clean-state prediction, introducing a flow-based framework where tokens are updated continuously and asynchronously based on confidence, achieving 97% of LLaDA's performance with 25% of the decoding budget.
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL [R]
This paper proposes using Masked Diffusion Language Models (MDLMs) as text-based world models for agentic reinforcement learning, showing that their any-order denoising objective avoids prefix mode collapse and leads to stronger performance than autoregressive baselines.
Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
This paper introduces ADAS, a training-free reranking rule for parallel masked diffusion decoding that uses attention to discount tokens that strongly attend to uncertain positions, improving low-NFE performance on reasoning and code tasks with minimal runtime overhead.
Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware
This paper characterizes the serving behavior of masked diffusion language models (dLLMs) using real hardware measurements, identifying key differences from autoregressive models and deriving design principles for efficient inference systems.
Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models
Proposes AdaLook, an adaptive multi-step lookahead decoding framework for masked diffusion language models that dynamically determines rollout depth and branch expansion based on candidate-score variance, achieving better accuracy-decoding steps trade-off compared to existing one-step lookahead decoding methods.