iSEE: Object Permanence Through Self-Supervision

πŸ“… 2026-10-01
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the problem of identity loss and positional discontinuity caused by occlusion during object tracking in unlabeled videos. To this end, we propose iSEE, a purely self-supervised framework for permanent tracking. Methodologically, iSEE introduces an evidential modeling mechanism and designs a dual-stream slot attention architecture that decouples appearance from spatial information. Furthermore, it incorporates a position extrapolator relying solely on recurrence signals to predict trajectories during occluded periods. Evaluated on the LA-CATER dataset, the framework achieves an 86% re-identification rate after occlusion with minimal localization error during hidden phases. It significantly outperforms existing self-supervised methods and effectively supports downstream planning tasks, thereby realizing, for the first time, complete permanent reasoning without reliance on manual annotations.
πŸ“ Abstract
Object permanence, keeping track of an object's identity and position while it is occluded, is central to video representations that track, predict and plan. Trackers that achieve it learn from boxes, track identities and visibility labels. On the other hand, self-supervised object-centric methods discover objects without labels: through slot attention, it represents a video as slots that bind to objects and follow them across frames. However, these slots are lost under occlusion, making the desired permanence impossible. Reasoning permanence is a hard problem because it requires to detect when an object becomes occluded, re-identify when object reappears, and keep the object's hidden position continuous, using reapperance as the only learning cue. To address this, we propose iSEE, a novel framework that offers all three aforementioned requirements, without any labels whatsoever. We built iSEE using the following three proposed components: (i) Object evidence modelling: a slot's attention, compared with its own past, reveals when its object is hidden. (ii) Appearance-position separation: two slot streams let the appearance be held for re-identification while the position keeps changing. (iii) Permanence from reappearance: a walker follows the hidden object's position, trained only on where the object reappears. On LA-CATER static, iSEE returns a reappearing object to its own slot after 86% of occlusions, against 32% for SlotContrast, and localises it while hidden within 4.1 mAP of the label-trained SoTA RAM. The two streams also allow downstream planning, with the position stream as the action of a world model. Project page: https://insait-institute.github.io/iSEE/
Problem

Research questions and friction points this paper is trying to address.

Object Permanence
Self-supervised Learning
Occlusion
Video Representation
Slot Attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Object Permanence
Self-Supervised Learning
Slot Attention
Occlusion Reasoning
Appearance-Position Separation