Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
Summary
This paper introduces UMM-Reflection, a reinforcement learning method for unified multimodal models that enables self-repair of generated images, improving performance on benchmarks like GenEval, WISE, and T2I-CompBench++ without external verifiers.
View Cached Full Text
Cached at: 09/29/26, 08:13 AM
Paper page - Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
Source: https://huggingface.co/papers/2609.35767
Abstract
Unifiedmultimodalmodelscanbothlookatandrenderimages,soinprincipletheycanrepairtheirowngenerations:diagnosewhatanimagegetswrong,reviseit,observetheresult,anddiagnoseagain.Whetherarevisionhelpsisknownonlyafteritisrendered,sothereflectiontextandtheimagegenerationmustbelearnedjointly,overthewholeloop.Supervisedfine-tuning(SFT)onreflectiontrajectoriesgivesacoldstartbutdoesnotfindthehigh-successrepairpaths,andnaiveRLthatoptimizesonlytherendereroronlyoneheadleavesmostofthegainuntapped.WeintroduceUMM-Reflection,whichappliesreinforcementlearning(RL)tocompletereflectiontrajectoriesinsideoneunifiedmodel:siblingtrajectoriesshareoneinitialimage,sothegroup-relativeadvantagecomparesreflectionstrategies,andonetrajectory-leveladvantageupdatesboththereflectiontokensandtheflow-basedrevisions,avoidingthecombinatorialblow-upofper-roundcreditassignment.Unlikesingle-roundeditingorpipelineswithanexternalcritic,creditflowsacrossroundsandtobothrolesofthesamemodel,andnoverifierisneededatinference.OnBAGEL,UMM-ReflectionimprovesGenEvalby12.05pointsoverSFT,andthegainstransfertoWISE(+10.97),OneIG-Bench(+3.48),andT2I-CompBench++(+4.63),noneofwhichisusedintraining.
View arXiv pageView PDFProject pageGitHub4Add to collection
Get this paper in your agent:
hf papers read 2609\.35767
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.35767 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.35767 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.35767 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
MIRAGE: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning Guidance
The paper introduces MIRAGE, an inference-time framework for enhancing LLM reasoning by dynamically switching perspectives using reinforcement learning guidance, outperforming existing prompting methods on various benchmarks.
ReflectMT: Internalizing Reflection for Efficient and High-Quality Machine Translation
ReflectMT introduces a two-stage RL method that trains LRMs to internalize reflection, enabling single-pass high-quality translation with 94% fewer tokens than multi-step reasoning models like DeepSeek-R1.
iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning
Introduces iVGR, a reinforcement learning framework that internalizes visual localization into textual reasoning for multimodal language models, eliminating the need for explicit visual grounding during inference while improving fine-grained perception performance.
Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System
This paper explores the synergy between visual understanding and generation in unified multimodal models, showing that task-decoupled architectures and end-to-end optimization can enhance performance by turning coexistence into synergy.
AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
AlphaGRPO is a new framework that applies Group Relative Policy Optimization to Unified Multimodal Models, enhancing generation through self-reflective refinement and decompositional verifiable rewards.