Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

Hugging Face Daily Papers Papers

Summary

This paper introduces Grounded Entity Biographies (GEB), a framework for augmenting long-video memory to track entities across events, demonstrating improved performance on question-answering benchmarks.

Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:19 AM

Paper page - Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

Source: https://huggingface.co/papers/2609.38155

Abstract

Answeringquestionsaboutlongvideosoftenrequiresconnectingeventsinvolvingthesameobjectsacrosshoursordays.Chronologicaldescriptionsandtext-derivedentitiescanleavephysicalidentityunresolved:differentobjectsmayshareadescription,whileobservationsofthesameobjectremaindisconnectedacrossevents.Retrievingrelevanteventsthereforedoesnotnecessarilyrecoverthe“biography“oftheparticularentityaquestionconcerns.Toaddressthis,weintroduceGroundedEntityBiographies(GEB),along-videomemoryframeworkthatgroupsvisuallygroundedobservationsofthesamephysicalinstanceacrossclipsintoretrievablebiographieswhilepreservingthecontextofeachmoment.Duringquestionanswering,thebiographyisretrievedalongsideepisodicevidence,allowingthemodeltofollowanentitythrougheventsusingidentitylinksestablishedduringmemoryconstruction.Evaluationsacrossfourbenchmarks,includingday-longandweek-longrecordings,demonstrateimprovementsoverpriormemoryframeworksinbothmultiple-choiceandopen-endedquestionanswering.OnEgoLifeQA,GEBachieves72.0%accuracy,4.4percentagepointsabovethebestpublishedresult.Ablationsshowthatgroundedidentityassociationandbiographyreadingbothcontributetothegains,whichadditionaldescriptionsalonedonotfullyrecover.

View arXiv pageView PDFProject pageGitHub20Add to collection

Get this paper in your agent:

hf papers read 2609\.38155

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.38155 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.38155 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.38155 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles