Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On
Summary
Oxygen-TryOn is a unified foundation model for any-item virtual try-on, achieving state-of-the-art consistency and realism across single and multi-item try-on tasks through a dedicated data engine and three-stage training pipeline.
View Cached Full Text
Cached at: 07/28/26, 06:33 AM
Paper page - Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On
Source: https://huggingface.co/papers/2607.21694 Authors:
,
,
,
,
,
,
,
,
,
,
,
Abstract
WepresentOxygen-TryOn,aunifiedfoundationmodelforany-itemvirtualtry-on.Ratherthanrepurposingageneral-purposeimageeditor,Oxygen-TryOnisfashion-native,builtfortry-onthroughadedicateddataengineandtry-on-specifictraining.Givenoneormorereferenceitems(cleanproductshotsorin-the-wildworn-onphotos)andasingletargetsubjectimage,itsynthesizesaphotorealisticimageofthesubjectwearingtheitemsacrossvirtuallyanyfashioncategory.Priorsystemshandleasinglegarmentcategoryinastudiosetting,andrecentmulti-referencemethodsremaingarment-centric;incontrast,Oxygen-TryOnsupportsdiverseitemsandscenarios,includingfull-andhalf-bodyviews,avariablenumberofreferences,andfreemulti-itemcomposition,whilefaithfullypreservingbothsubjectidentityanditemappearance.Insteadofmask-basedinpainting,wereformulatetry-onasamulti-reference,understanding-drivengenerationtask.Webuildadataenginethatcollects,manufactures,annotates,andfiltershigh-qualitytry-ondataatscale,anddesignathree-stagerecipeofcontinuedpre-training(CPT),supervisedfine-tuning(SFT),andreinforcementlearning(RL).TheRLstageusesahybridrewardcombininganin-housetry-onrewardmodelwithaproprietary,rubric-guidedgeneral-purposemodel,jointlysupervisingfine-grainedconsistencyandinstruction-levelquality.Italsofollowsgeneraleditinginstructions(e.g.,posechanges)inthesamepass.Acrosspublicbenchmarksandourin-houseOxygen-TryOnBench,itachievesstate-of-the-artconsistencyandrealismonsingle-itemtry-onandleadsonmulti-itemtry-on,matchingorsurpassingbothleadingproprietarysystems(NanoBananaPro,GPT-Image-2,Seedream5Lite)andopen-sourcemodels(FLUX.2).
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2607\.21694
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.21694 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.21694 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.21694 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items
Tstars-Tryon 1.0 is a commercial-scale virtual try-on system delivering photorealistic, real-time garment visualization across diverse fashion categories, now deployed on Taobao serving millions of users.
TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy
This paper presents TryOnCrafter, a novel framework for camera-controllable video virtual try-on that uses a renderable 4D try-on proxy and DiT-based video generation to achieve omnidirectional viewpoint exploration, overcoming the limitations of existing methods that depend on fixed source camera trajectories.
CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation
This paper introduces VIP-SAM for instance-level garment segmentation and CtrlVTON, a controllable virtual try-on framework that treats try-on as an image editing problem, allowing precise control over garment layout, style, and placement. Both methods achieve state-of-the-art results on their respective tasks.
Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
Introduces Target Viewpoint Reproduction (TVR) task and TVRBench benchmark for evaluating foundation models' ability to actively adjust 3D viewpoints to match target images. Experiments reveal significant limitations in current open and closed-source models, with a unified post-training framework boosting success rates from ~12% to ~51%.
OpenMHC: Accelerating the Science of Wearable Foundation Models
OpenMHC introduces the largest open-access wearable health dataset with over 60 million hours of data and open-source implementations of wearable foundation models, including a unified benchmark for prediction, imputation, and forecasting.