🤖 AI Summary
This study addresses the challenges of entity identity confusion and disrupted cross-event associations caused by temporal descriptions in long video question answering. To this end, we propose GEB, a visually grounded Entity Biography memory framework. By leveraging visual grounding techniques to establish identity links across video segments, GEB aggregates multi-temporal observations of the same entity into retrievable biographies, thereby overcoming the limitations of conventional temporal memory. Furthermore, it enhances large language model reasoning by integrating biography retrieval with contextual evidence augmentation. Extensive experiments demonstrate that GEB yields significant performance improvements across four benchmarks, achieving an accuracy of 72.0% on EgoLifeQA and surpassing the state-of-the-art method by 4.4 percentage points.
📝 Abstract
Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.