FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
Summary
Introduces DENSEWORLD, a 1,000-hour dataset of crowded Global South urban scenes, and FactorJEPA, a JEPA variant that factorizes future prediction into layout, agents, and interactions, improving accuracy and robustness under occlusion and heterogeneity.
View Cached Full Text
Cached at: 08/07/26, 09:59 PM
Paper page - FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
Source: https://huggingface.co/papers/2608.01049
Abstract
Worldmodelshaveattractedsignificantattentionfortheirabilitytocaptureandpredictthestructureanddynamicsofthephysicalworld.Inthisemerginglandscape,JointEmbeddingPredictiveArchitectures(JEPA)offeraparticularlycompellingdirection.Westudyalargelyunexploredregime:populous,crowded,andchaoticGlobalSouthurbanenvironments,whichwecallDENSEWORLD.Unlikethelower-density,lane-structuredsettingsthatdominateexistingevaluations,thesescenesexhibitsoftspatialboundaries,extremeagentheterogeneity,persistentocclusion,andrapidsocialnegotiationundermixedtraffic.Weintroducethefirstlarge-scaledatasetforthisregime:1,000hoursofdrive-through,walk-through,andaerialvideoacross22cities.ExistingJEPAformulationsstruggletopreservedenseinteractiondynamicsunderheterogeneityandpartialobservability.WeintroduceFactorJEPA,whichmakesworldstructureafirst-classpredictiveprimitive.Ratherthanencodingthefutureinamonolithiclatent,itcomposeslayout,entities,andinteractions,usingavisibilitygateandseparatedsubspacestopreservepartiallyobservedagentsanddiscouragecross-factorshortcuts.FactorJEPAimproves(i)future-latentaccuracy(Future-frameL1),(ii)intervention-sensitiveprediction(CausalL1),and(iii)robustnesstoreducedvisualevidence(Mask-ratioslope),whileexposing(iv)areproduciblemotion-informationtrade-off(Motioncosine).Methodrankingsreplicateacross2Band1BV-JEPA2.1backbones,withrho=0.895to0.978.WepubliclyreleasetheDENSEWORLD-115kdataset(https://huggingface.co/datasets/anonymousML123/denseworld-115k)andthesurgery-trainedFactorJEPAcheckpoints(https://huggingface.co/datasets/anonymousML123/factorjepa-outputs/tree/main/outputs/full/vjepa_2_1_vitg_1B/train/m09c_surgery_3stage_DI_diheavy_encoder).
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2608\.01049
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.01049 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.01049 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.01049 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
DVD-JEPA: an open-source, fully-reproducible JEPA world model [P]
DVD-JEPA is an open-source, minimal JEPA world model that learns representations from video by predicting future embeddings rather than pixels. It uses a bouncing DVD logo to demonstrate position recovery, dreaming, and anomaly detection, all running in a browser.
I built Micro-JEPA: A lightweight JEPA (Joint Embedding Predictive Architecture) in Python
Micro-JEPA is a lightweight Python implementation of the Joint Embedding Predictive Architecture (JEPA), enabling an agent to learn environment representations, predict future states in latent space, and plan actions to avoid obstacles.
One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA
This paper introduces Collective-State JEPA (CS-JEPA), a recurrent joint-embedding predictive architecture that lets every robot in a swarm predict the same future collective state from local observations and limited messages. The method shows label-efficient improvements in prediction error and inter-robot agreement, plus planning-relevant value estimation.
Is the future of coding agents JEPA? [D]
The author discusses applying Yann LeCun's JEPA (Joint Embedding Predictive Architecture) to coding agents, proposing that instead of treating code as text, agents should learn compact state representations and predict future states, potentially achieving orders of magnitude efficiency improvements over current LLM-based approaches.
@AbdelStark: It’s time to JEPA pill the world! awesome-jepa: A curated list of papers, models, code, datasets, and learning resource…
A curated list of papers, models, code, datasets, and learning resources for Joint Embedding Predictive Architectures (JEPA), the self-supervised approach to world models proposed by Yann LeCun.