Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation
Summary
This paper introduces GEAR, a Geometry-Enabled Attention Routing framework that uses geometry as explicit token-level addresses for visual memory to enhance long-horizon camera-controlled video generation, achieving state-of-the-art results.
View Cached Full Text
Cached at: 09/29/26, 04:11 PM
Paper page - Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation
Source: https://huggingface.co/papers/2609.34722 Authors:
,
,
,
,
,
,
,
,
,
,
,
Abstract
Long-horizoncamera-controlledvideogenerationrequiresrecoveringpreviouslyobservedcontentfromanever-growingvisualhistory.Existingapproacheseithersearchhistoricalcontextimplicitlyorreconstructitintopersistent3Dmemory,facinginefficientmemoryaccessoraccumulatedgeometricerrors.Ourkeyinsightisthatgeometryneednotexplainthescene--itonlyneedstodeterminewherevisualmemoryshouldbereadfrom,whileattentiondecideswhatshouldberecovered.Basedonthisinsight,weintroduceGEAR,aGeometry-EnabledAttentionRoutingframeworkthatusesgeometryasanexplicittoken-leveladdressforvisualmemory.Ratherthanfusinghistoricalobservationsintoapersistentglobal3Drepresentation,GEARretainsthemasframelatentsandusesper-framegeometryonlytoestablishtoken-levelcorrespondenceswithtargetviews,therebyavoidingpersistenterroraccumulationfromglobalfusion.Guidedbythesecorrespondences,GeometricCorrespondenceAttention(GCA)selectivelyinjectsgeometricallymatchedhistoricalfeaturesintonoisytargetpatchesduringdenoising.WefurtherintroduceanInvisibleOctreetoaccumulatevisibilityevidenceandrejectgeometricallyplausiblebutoccludedcorrespondences.ExtensiveexperimentsdemonstratethatGEARachievesstate-of-the-artvisualquality,precisecameracontrol,andrevisitconsistency,enablingminute-longvideogenerationalongchallengingtrajectories.
View arXiv pageView PDFProject pageGitHubAdd to collection
Get this paper in your agent:
hf papers read 2609\.34722
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.34722 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.34722 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.34722 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Geo-Align: Video Generation Alignment via Metric Geometry Reward
Geo-Align presents a reinforcement learning framework for camera-controlled video re-rendering that improves generalization through scale-aware perceptual rewards and metric 3D estimation for camera trajectory extraction.
Memory Has Geometry: Non-Uniform Geometric Memory for Long-Horizon Personalized AI
The paper proposes modeling memory as a user-specific dynamical state space with non-uniform geometry to enhance long-horizon personalization in AI, enabling trajectory-conditioned reconstruction for better memory access.
Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry
The paper proposes credit-addressable reasoning with executable code traces and localized reinforcement learning to enhance multimodal geometry reasoning, achieving significant accuracy improvements over baseline models like Qwen3-VL-8B.
Stable Geometry with Divergent Task Evidence for Efficient Long-Horizon Agent Compression
This paper introduces Geometry Guided Evidence Preserving Memory (GEM), a training-free compressor for long-horizon agents that optimizes for preserved task evidence over geometric coverage, reducing token usage while maintaining performance.
GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation
The paper introduces GAE, a geometry-native autoencoder that creates a compact latent space for generating 3D-consistent scenes, enhancing visual quality and coherence over existing methods.