Score
Designs and evaluates systems that detect and link semantic concepts to spatial locations and temporal intervals in spatio‑temporal data (for example, video), producing grounded annotations such as per‑frame regions and labeled temporal segments. Builds pipelines to extract atomic spatio‑temporal evidence units and organize them into structured evidence chains used for supervised fine‑tuning, interpretability, and downstream reasoning.
Existing video reasoning models generate only textual reasoning chains, lacking explicit localization of spatiotemporal evidence (i.e., timestamps and spatial bounding boxes), resulting in poor interpretability and a disconnect between explanations and visual grounding. To address this, we propose an explicit spatiotemporal evidence–based video reasoning framework featuring a novel joint temporal localization and spatial tracking mechanism that synchronously annotates key frames, target objects, and their bounding boxes during inference. We introduce STGR-CoT-30k/STGR-RL-36k—the first high-quality dataset supporting joint spatiotemporal supervision—and design a cold-start reinforcement learning strategy with multi-stage end-to-end training and a spatiotemporally aware reward function to jointly optimize answer accuracy, temporal alignment, and spatial localization precision. Our method achieves +14.4% mAM and +24.2% mLGM on V-STAR, and sets new state-of-the-art results across VideoMME and WorldSense. It is the first to enable verifiable, confidence-aware, fine-grained evidence tracing.
This work addresses the lack of explicit modeling of object motion trajectories in existing video reasoning methods, which hinders the verification of motion patterns in temporal observations. It formalizes spatio-temporal-trajectory (STT) reasoning for the first time, introducing an explicit trajectory representation and a Motion Chain of Thought (MCoT) reasoning pathway. To provide strong supervision, the authors construct a trajectory-annotated dataset and propose a motion-aware training mechanism that requires no architectural modifications, along with trajectory-level bounding box tracking and a vision-evidence-based reward function. Experiments demonstrate that the approach significantly improves performance in spatio-temporal localization and trajectory prediction, while remaining fully compatible with existing video understanding frameworks, thereby validating the critical role of motion reasoning in evidence-driven video comprehension.
Existing video multimodal large language models (MLLMs) model bounding boxes as autoregressive text sequences, leading to verbose outputs, accumulating spatial errors over time, and localization drift. This work proposes a collaborative framework integrating a video LLM with an open-vocabulary detector. Its core innovations are: (1) a Reference-Semantic Token (RST) mechanism, which leverages the user query’s semantics both as a control signal and as a substitute for textual embeddings, enabling end-to-end referring understanding and grounding; and (2) Tubular Temporal Regularization (TTReg), which enforces temporal consistency of object trajectories across frames. By circumventing error-prone autoregressive coordinate generation, the method significantly improves spatiotemporal localization accuracy and enhances complex semantic reasoning—such as causal and sequential inference—on fine-grained video understanding benchmarks including STVG and GroundedVQA. Results validate the efficacy of co-modeling detection priors with large language models.
Current small-scale large language models exhibit limited performance in fine-grained spatial relations, metric distance reasoning, and temporal sequencing, hindering the capabilities of embodied agents in dynamic 3D environments. To address this, this work proposes the FESTS framework, which introduces SpRE—a novel spatiotemporal regular expression formalism that integrates regular expressions with S4u spatial logic and supports both universal and existential quantifiers. The framework compiles natural language queries into formal spatiotemporal specifications and automatically generates large-scale, aligned training data by matching structured video logs, eliminating the need for manual annotation. Fine-tuning a 3-billion-parameter model on 27k synthesized samples boosts frame-level F1 score from 48.5% to 87.5%, achieving performance comparable to GPT-4.1 while being two orders of magnitude smaller in model size.
This work addresses the high annotation cost of video spatio-temporal scene graphs (STSGs) by proposing a weakly supervised learning framework that relies solely on video-caption pairs. Methodologically, it introduces the first differentiable symbolic reasoning module jointly optimized with contrastive, temporal, and semantic losses to generate logic-guided STSGs; additionally, it leverages large language models (LLMs) to automatically distill spatio-temporal logical rules, forming a neuro-symbolic architecture. Contributions include: (1) the first end-to-end weakly supervised paradigm for STSG generation without manual STSG annotations; (2) an LLM-driven mechanism for automatic spatio-temporal logical rule induction; and (3) state-of-the-art performance on Something-Something V2, MUGEN, and OpenPVSG, demonstrating substantial improvements in fine-grained video semantic representation.
This work addresses the limitation of existing video multimodal large language models, which often rely on irrelevant frames or objects for fine-grained spatiotemporal reasoning due to a lack of reliable semantic alignment evidence. To this end, the authors reformulate spatiotemporal evidence localization as a constrained verification task and introduce a Semantic Evidence Reward (SER) mechanism. This mechanism employs a referee vision-language model (VLM) to assess the relevance and localization quality of generated evidence, augmented with a temporal consistency penalty. Notably, the approach is trained solely on standard video question-answering data without requiring dense bounding box annotations. By replacing pixel-level overlap with semantic alignment as the evaluation criterion, the method significantly enhances both interpretability and accuracy, achieving 49.6% mLGM on the V-STAR benchmark—3.0 percentage points higher than the strong baseline Open-o3-Video.
This work addresses the limited capacity of existing vision-language models (VLMs) and tool-augmented agents to perform effective spatial reasoning in continuous, dynamic 3D environments, as they are largely confined to static, single-frame understanding. We propose S-Agent, a novel paradigm that reframes the VLM as a semantic planner driven by hierarchical tool invocation. By orchestrating a spatial toolchain that integrates 2D perception, 3D geometric reconstruction, and spatiotemporal memory, S-Agent enables cross-frame evidence accumulation and high-level spatial knowledge construction without requiring any additional training. The framework also generates high-quality spatial reasoning trajectories suitable for supervised fine-tuning. Experiments demonstrate that S-Agent substantially improves both open- and closed-source VLMs on multiview and video-based spatial reasoning benchmarks. Fine-tuning on the S-300K dataset generated by our method yields the S-Agent-8B model, which outperforms same-scale baselines and rivals advanced closed-source systems such as GPT-5.4 and Gemini 3.
This work addresses the inefficiency of manual inspection in urban surveillance video analysis by proposing a visual analytics system that segments long videos into short clips and leverages vision-language models to generate semantic descriptions for indexing. The system integrates retrieval-augmented generation (RAG), domain-specific knowledge graphs, and video-entity alignment to enable semantic-driven event retrieval and visual verification. It innovatively combines taxonomy-aware entity extraction with video grounding mechanisms to enhance consistency between textual reasoning and visual evidence. Evaluations on the StreetAware dataset for hazardous scene detection and pedestrian crossing analysis demonstrate that the system substantially reduces analysts’ cognitive load while improving both analytical efficiency and result reliability.
Existing spatial semantic representations struggle to effectively reason about structured temporal dynamics—such as the periodic movement of household objects—in semi-static environments. This work proposes PredictiveGraphs, a predictive 3D scene graph that integrates spatiotemporal and semantic information by embedding Perpetua* Bayesian filters directly into inter-node relationships, enabling temporal modeling and future prediction of object states. By jointly modeling spatiotemporal-semantic relations and performing recurrent state inference, the approach maintains robustness under distributional shifts. Evaluated over three-week navigation tasks in both simulation and real-world settings—with environmental changes occurring every two hours—the method significantly outperforms current baselines in accurately forecasting the dynamic evolution of the environment.