🤖 AI Summary
This study addresses the absence of promptable instance segmentation methods leveraging feed-forward 4D geometry for dynamic scenes and the inherent difficulty in preserving spatiotemporal consistency. To this end, we propose a Segment Anything Model built upon feed-forward 4D visual geometry. By incorporating a spatiotemporal query decoder and a persistent object query mechanism, the model generates class-agnostic 4D masks that support both independent and conditioned segmentation. Furthermore, we construct an agent framework integrating active tree search, a dual-stream localizer, and vision-language model (VLM) priors to achieve natural language-guided grounding in 4D scenes. The proposed approach attains state-of-the-art performance in 4D instance segmentation across both dynamic and static environments while demonstrating exceptional language-guided grounding capabilities.
📝 Abstract
Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Anything models operate primarily on 2D image or video masks and preserve identity through sequential memory, while promptable 4D instance segmentation built upon feed-forward visual geometry remains underexplored. We introduce S4VY, a Segment Anything model built on feed-forward 4D visual geometry. From a set of RGB observations, S4VY transforms shared visual-geometric features into an exhaustive set of class-agnostic 4D instance masks through a space-time query decoder, with each persistent object query binding one entity across all observations. This representation supports prompt-independent segmentation as well as point- and box- conditioned selection, without requiring a seed mask or temporal ordering. We further develop an agentic harness for natural-language grounding in the large observation space of a 4D scene. Active tree search identifies relevant frames without scanning every fixed window; a dual-stream grounder combines fine-grained VLM visual priors with geometry-consistent instance features through complementary bounding-box prediction and object-query matching; and an independent critic selects the final 4D instance mask from their predictions. Extensive experiments demonstrate state-of-the-art 4D instance segmentation and strong language-guided grounding performance under a unified evaluation spanning static and dynamic scenes.