spatio-temporal grounding

Designs and evaluates systems that detect and link semantic concepts to spatial locations and temporal intervals in spatio‑temporal data (for example, video), producing grounded annotations such as per‑frame regions and labeled temporal segments. Builds pipelines to extract atomic spatio‑temporal evidence units and organize them into structured evidence chains used for supervised fine‑tuning, interpretability, and downstream reasoning.

spatio-temporalgrounding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence

Oct 23, 2025
JM
Jiahao Meng
🏛️ Peking University | ByteDance | CASIA | WHU | NUS

Existing video reasoning models generate only textual reasoning chains, lacking explicit localization of spatiotemporal evidence (i.e., timestamps and spatial bounding boxes), resulting in poor interpretability and a disconnect between explanations and visual grounding. To address this, we propose an explicit spatiotemporal evidence–based video reasoning framework featuring a novel joint temporal localization and spatial tracking mechanism that synchronously annotates key frames, target objects, and their bounding boxes during inference. We introduce STGR-CoT-30k/STGR-RL-36k—the first high-quality dataset supporting joint spatiotemporal supervision—and design a cold-start reinforcement learning strategy with multi-stage end-to-end training and a spatiotemporally aware reward function to jointly optimize answer accuracy, temporal alignment, and spatial localization precision. Our method achieves +14.4% mAM and +24.2% mLGM on V-STAR, and sets new state-of-the-art results across VideoMME and WorldSense. It is the first to enable verifiable, confidence-aware, fine-grained evidence tracing.

Existing datasets lack unified spatio-temporal supervision and reasoning tracesModels require joint temporal tracking and spatial localization in videosVideo reasoning lacks explicit spatio-temporal evidence localization

This work addresses the lack of explicit modeling of object motion trajectories in existing video reasoning methods, which hinders the verification of motion patterns in temporal observations. It formalizes spatio-temporal-trajectory (STT) reasoning for the first time, introducing an explicit trajectory representation and a Motion Chain of Thought (MCoT) reasoning pathway. To provide strong supervision, the authors construct a trajectory-annotated dataset and propose a motion-aware training mechanism that requires no architectural modifications, along with trajectory-level bounding box tracking and a vision-evidence-based reward function. Experiments demonstrate that the approach significantly improves performance in spatio-temporal localization and trajectory prediction, while remaining fully compatible with existing video understanding frameworks, thereby validating the critical role of motion reasoning in evidence-driven video comprehension.

motion patternsobject trajectoryspatio-temporal grounding

1 + 1 > 2: Detector-Empowered Video Large Language Model for Spatio-Temporal Grounding and Reasoning

Dec 07, 2025
SG
Shida Gao
🏛️ Beijing University of Posts and Telecommunications | University of Trento | Institute of Automation, Chinese Academy of Sciences | Hong Kong University of Science and Technology | ZTE Corporation

Existing video multimodal large language models (MLLMs) model bounding boxes as autoregressive text sequences, leading to verbose outputs, accumulating spatial errors over time, and localization drift. This work proposes a collaborative framework integrating a video LLM with an open-vocabulary detector. Its core innovations are: (1) a Reference-Semantic Token (RST) mechanism, which leverages the user query’s semantics both as a control signal and as a substitute for textual embeddings, enabling end-to-end referring understanding and grounding; and (2) Tubular Temporal Regularization (TTReg), which enforces temporal consistency of object trajectories across frames. By circumventing error-prone autoregressive coordinate generation, the method significantly improves spatiotemporal localization accuracy and enhances complex semantic reasoning—such as causal and sequential inference—on fine-grained video understanding benchmarks including STVG and GroundedVQA. Results validate the efficacy of co-modeling detection priors with large language models.

Autoregressive spatial decoding causes error accumulation in video localizationCurrent models inefficiently handle spatio-temporal grounding and reasoning tasksExisting methods struggle with temporal consistency in object tracking

Current small-scale large language models exhibit limited performance in fine-grained spatial relations, metric distance reasoning, and temporal sequencing, hindering the capabilities of embodied agents in dynamic 3D environments. To address this, this work proposes the FESTS framework, which introduces SpRE—a novel spatiotemporal regular expression formalism that integrates regular expressions with S4u spatial logic and supports both universal and existential quantifiers. The framework compiles natural language queries into formal spatiotemporal specifications and automatically generates large-scale, aligned training data by matching structured video logs, eliminating the need for manual annotation. Fine-tuning a 3-billion-parameter model on 27k synthesized samples boosts frame-level F1 score from 48.5% to 87.5%, achieving performance comparable to GPT-4.1 while being two orders of magnitude smaller in model size.

Embodied AILarge Language ModelsSpatial Relations

LASER: A Neuro-Symbolic Framework for Learning Spatial-Temporal Scene Graphs with Weak Supervision

Apr 15, 2023
JH
Jiani Huang
🏛️ University of Pennsylvania | University of Central Florida

This work addresses the high annotation cost of video spatio-temporal scene graphs (STSGs) by proposing a weakly supervised learning framework that relies solely on video-caption pairs. Methodologically, it introduces the first differentiable symbolic reasoning module jointly optimized with contrastive, temporal, and semantic losses to generate logic-guided STSGs; additionally, it leverages large language models (LLMs) to automatically distill spatio-temporal logical rules, forming a neuro-symbolic architecture. Contributions include: (1) the first end-to-end weakly supervised paradigm for STSG generation without manual STSG annotations; (2) an LLM-driven mechanism for automatic spatio-temporal logical rule induction; and (3) state-of-the-art performance on Something-Something V2, MUGEN, and OpenPVSG, demonstrating substantial improvements in fine-grained video semantic representation.

Aligning predicted graphs with logical specifications from captionsLearning spatio-temporal scene graphs without annotated videosUsing video captions as weak supervision for training

Latest Papers

What's happening recently
View more

This work addresses the limitation of existing video multimodal large language models, which often rely on irrelevant frames or objects for fine-grained spatiotemporal reasoning due to a lack of reliable semantic alignment evidence. To this end, the authors reformulate spatiotemporal evidence localization as a constrained verification task and introduce a Semantic Evidence Reward (SER) mechanism. This mechanism employs a referee vision-language model (VLM) to assess the relevance and localization quality of generated evidence, augmented with a temporal consistency penalty. Notably, the approach is trained solely on standard video question-answering data without requiring dense bounding box annotations. By replacing pixel-level overlap with semantic alignment as the evaluation criterion, the method significantly enhances both interpretability and accuracy, achieving 49.6% mLGM on the V-STAR benchmark—3.0 percentage points higher than the strong baseline Open-o3-Video.

evidence groundingreward designsemantic alignment

This work addresses the limited capacity of existing vision-language models (VLMs) and tool-augmented agents to perform effective spatial reasoning in continuous, dynamic 3D environments, as they are largely confined to static, single-frame understanding. We propose S-Agent, a novel paradigm that reframes the VLM as a semantic planner driven by hierarchical tool invocation. By orchestrating a spatial toolchain that integrates 2D perception, 3D geometric reconstruction, and spatiotemporal memory, S-Agent enables cross-frame evidence accumulation and high-level spatial knowledge construction without requiring any additional training. The framework also generates high-quality spatial reasoning trajectories suitable for supervised fine-tuning. Experiments demonstrate that S-Agent substantially improves both open- and closed-source VLMs on multiview and video-based spatial reasoning benchmarks. Fine-tuning on the S-300K dataset generated by our method yields the S-Agent-8B model, which outperforms same-scale baselines and rivals advanced closed-source systems such as GPT-5.4 and Gemini 3.

continuous 3D reasoningmulti-view perceptionspatial intelligence

This work addresses the inefficiency of manual inspection in urban surveillance video analysis by proposing a visual analytics system that segments long videos into short clips and leverages vision-language models to generate semantic descriptions for indexing. The system integrates retrieval-augmented generation (RAG), domain-specific knowledge graphs, and video-entity alignment to enable semantic-driven event retrieval and visual verification. It innovatively combines taxonomy-aware entity extraction with video grounding mechanisms to enhance consistency between textual reasoning and visual evidence. Evaluations on the StreetAware dataset for hazardous scene detection and pedestrian crossing analysis demonstrate that the system substantially reduces analysts’ cognitive load while improving both analytical efficiency and result reliability.

event retrievallong-duration videoscene interpretation

Existing spatial semantic representations struggle to effectively reason about structured temporal dynamics—such as the periodic movement of household objects—in semi-static environments. This work proposes PredictiveGraphs, a predictive 3D scene graph that integrates spatiotemporal and semantic information by embedding Perpetua* Bayesian filters directly into inter-node relationships, enabling temporal modeling and future prediction of object states. By jointly modeling spatiotemporal-semantic relations and performing recurrent state inference, the approach maintains robustness under distributional shifts. Evaluated over three-week navigation tasks in both simulation and real-world settings—with environmental changes occurring every two hours—the method significantly outperforms current baselines in accurately forecasting the dynamic evolution of the environment.

environment state predictionpredictive scene graphssemi-static scenes

Hot Scholars

JX

Junbin Xiao

National University of Singapore
Video and LanguageEmbodied InteractionTrustworthy Multimodality
AY

Angela Yao

National University of Singapore
computer visiondeep learningmachine learning
YY

Yifan Yang

Senior Research SDE, Microsoft Research Asia
Multi-modalityComputer VisionMachine LearningArtificial Intelligence
YZ

Yue Zhou

Associate Professor, East China Normal University
Remote Sensing Vision-Language ModelOriented Object Detection
KL

Kan Li

Huazhong University of Science and Technology
3D AssemblyStretchable ElectronicsMetamaterials