QQWorld: Quantile-Quantile Matching for World Model Regularization

Hugging Face Daily Papers Papers

Summary

This paper proposes QQWorld, a quantile-quantile matching objective that replaces the Epps-Pulley objective in LeWorldModel for better regularization of latent distributions, improving planning success in control environments.

Latent world models enable efficient planning by predicting future states in a compact representation space, but their performance depends critically on the quality of the learned latent distribution. LeWorldModel (LeWM) regularizes its latents toward an isotropic Gaussian using the Epps-Pulley (EP) objective. We show that the corrective gradients of EP rapidly vanish for isolated tail samples, leaving heavy-tailed deviations insufficiently controlled. To address this limitation, we propose QQWorld, which replaces EP with a quantile-quantile matching objective that directly aligns projected latent samples with rank-matched Gaussian quantiles, thereby maintaining effective corrective gradients in the tails. We further develop cross-batch QQ, which enlarges the effective ranking pool using detached samples from previous batches, and characterize its bias-variance trade-off. Across four control environments, QQWorld effectively improves the average planning success rate of LeWM, while consistently yielding better Gaussian alignment and thinner latent tails.
Original Article
View Cached Full Text

Cached at: 08/03/26, 05:30 AM

Paper page - QQWorld: Quantile-Quantile Matching for World Model Regularization

Source: https://huggingface.co/papers/2607.28415

Abstract

Latentworldmodelsenableefficientplanningbypredictingfuturestatesinacompactrepresentationspace,buttheirperformancedependscriticallyonthequalityofthelearnedlatentdistribution.LeWorldModel(LeWM)regularizesitslatentstowardanisotropicGaussianusingtheEpps-Pulley(EP)objective.WeshowthatthecorrectivegradientsofEPrapidlyvanishforisolatedtailsamples,leavingheavy-taileddeviationsinsufficientlycontrolled.Toaddressthislimitation,weproposeQQWorld,whichreplacesEPwithaquantile-quantilematchingobjectivethatdirectlyalignsprojectedlatentsampleswithrank-matchedGaussianquantiles,therebymaintainingeffectivecorrectivegradientsinthetails.Wefurtherdevelopcross-batchQQ,whichenlargestheeffectiverankingpoolusingdetachedsamplesfrompreviousbatches,andcharacterizeitsbias-variancetrade-off.Acrossfourcontrolenvironments,QQWorldeffectivelyimprovestheaverageplanningsuccessrateofLeWM,whileconsistentlyyieldingbetterGaussianalignmentandthinnerlatenttails.

View arXiv pageView PDFAdd to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2607.28415 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2607.28415 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2607.28415 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Q-Learning With World Models

arXiv cs.LG

This paper introduces QWM, a framework that integrates world models with Q-learning to enhance sample efficiency in reinforcement learning by using imagined trajectories for action selection without compromising training on real data. It demonstrates significant improvements over state-of-the-art methods on manipulation benchmarks.

PROWL: Prioritized Regret-Driven Optimization for World Model Learning

arXiv cs.LG

Introduces PROWL, a prioritized regret-driven optimization framework that uses an adversarial curriculum to improve diffusion-based world model robustness by focusing on high-error trajectories, achieving better performance on out-of-distribution scenarios in MineRL.