Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
Summary
This paper introduces Grounded Entity Biographies (GEB), a framework for augmenting long-video memory to track entities across events, demonstrating improved performance on question-answering benchmarks.
View Cached Full Text
Cached at: 09/30/26, 04:19 AM
Paper page - Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
Source: https://huggingface.co/papers/2609.38155
Abstract
Answeringquestionsaboutlongvideosoftenrequiresconnectingeventsinvolvingthesameobjectsacrosshoursordays.Chronologicaldescriptionsandtext-derivedentitiescanleavephysicalidentityunresolved:differentobjectsmayshareadescription,whileobservationsofthesameobjectremaindisconnectedacrossevents.Retrievingrelevanteventsthereforedoesnotnecessarilyrecoverthe“biography“oftheparticularentityaquestionconcerns.Toaddressthis,weintroduceGroundedEntityBiographies(GEB),along-videomemoryframeworkthatgroupsvisuallygroundedobservationsofthesamephysicalinstanceacrossclipsintoretrievablebiographieswhilepreservingthecontextofeachmoment.Duringquestionanswering,thebiographyisretrievedalongsideepisodicevidence,allowingthemodeltofollowanentitythrougheventsusingidentitylinksestablishedduringmemoryconstruction.Evaluationsacrossfourbenchmarks,includingday-longandweek-longrecordings,demonstrateimprovementsoverpriormemoryframeworksinbothmultiple-choiceandopen-endedquestionanswering.OnEgoLifeQA,GEBachieves72.0%accuracy,4.4percentagepointsabovethebestpublishedresult.Ablationsshowthatgroundedidentityassociationandbiographyreadingbothcontributetothegains,whichadditionaldescriptionsalonedonotfullyrecover.
View arXiv pageView PDFProject pageGitHub20Add to collection
Get this paper in your agent:
hf papers read 2609\.38155
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.38155 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.38155 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.38155 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
EM^2Mem: Event-Centric Multimodal Memory for Large Language Models
EM^2Mem proposes an event-centric multimodal memory framework that binds heterogeneous evidence to event anchors for compact, generation-ready memory in long-video question answering, improving accuracy and reducing latency.
ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
ViSAGE is a multimodal agentic memory framework for long-form video understanding that builds self-correcting, entity-centric memories via cross-modal binding, bidirectional memory refinement, and multi-agent cross-verification, achieving 5.9% higher accuracy than baselines.
MBench: A Comprehensive Benchmark on Memory Capability for Video World Models
This paper introduces MBench, a benchmark for evaluating the memory capabilities of video world models across entity, environment, and causal consistency over long temporal horizons.
Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model
A 2B vision-language model distilled from a tool-using agent achieves first-place performance on long egocentric video question answering by pruning its multilingual embedding table to meet parameter limits, reaching 89% accuracy of the larger pipeline with only 1.1% of parameters.
GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience
GROVE is a training-free framework that grows a temporally stratified memory from continuous video streams, supporting both reactive QA and proactive assistance. It achieves state-of-the-art results on benchmarks like MM-lifelong and EgoServe.