Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Hugging Face Daily Papers Papers

Summary

This paper introduces UMM-Reflection, a reinforcement learning method for unified multimodal models that enables self-repair of generated images, improving performance on benchmarks like GenEval, WISE, and T2I-CompBench++ without external verifiers.

Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.
Original Article
View Cached Full Text

Cached at: 09/29/26, 08:13 AM

Paper page - Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Source: https://huggingface.co/papers/2609.35767

Abstract

Unifiedmultimodalmodelscanbothlookatandrenderimages,soinprincipletheycanrepairtheirowngenerations:diagnosewhatanimagegetswrong,reviseit,observetheresult,anddiagnoseagain.Whetherarevisionhelpsisknownonlyafteritisrendered,sothereflectiontextandtheimagegenerationmustbelearnedjointly,overthewholeloop.Supervisedfine-tuning(SFT)onreflectiontrajectoriesgivesacoldstartbutdoesnotfindthehigh-successrepairpaths,andnaiveRLthatoptimizesonlytherendereroronlyoneheadleavesmostofthegainuntapped.WeintroduceUMM-Reflection,whichappliesreinforcementlearning(RL)tocompletereflectiontrajectoriesinsideoneunifiedmodel:siblingtrajectoriesshareoneinitialimage,sothegroup-relativeadvantagecomparesreflectionstrategies,andonetrajectory-leveladvantageupdatesboththereflectiontokensandtheflow-basedrevisions,avoidingthecombinatorialblow-upofper-roundcreditassignment.Unlikesingle-roundeditingorpipelineswithanexternalcritic,creditflowsacrossroundsandtobothrolesofthesamemodel,andnoverifierisneededatinference.OnBAGEL,UMM-ReflectionimprovesGenEvalby12.05pointsoverSFT,andthegainstransfertoWISE(+10.97),OneIG-Bench(+3.48),andT2I-CompBench++(+4.63),noneofwhichisusedintraining.

View arXiv pageView PDFProject pageGitHub4Add to collection

Get this paper in your agent:

hf papers read 2609\.35767

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.35767 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.35767 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.35767 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles