AgentSTAR: Agentic Shape Tracking and Reconstruction from Monocular Videos

📅 2026-09-21
📈 Citations: 0
✹ Influential: 0
📄 PDF
🀖 AI Summary
本文提出了䞀种通过分析-合成方法从单目视频䞭进行圢状跟螪和重建的方法䜿甚视觉-语蚀暡型代理迭代䌘化物䜓的3D暡型和姿态䌰计。
📝 Abstract
In this work, we present a method for shape reconstruction and tracking from video via agentic analysis-by-synthesis. Unlike prior methods which first estimate dense pixel correspondences and then recover object motion from them, our method infers a structured 3D object model, including its geometry and kinematic structure, and uses this model to optimise object track estimates over time. In our optimisation loop, a Vision-Language Model (VLM) agent iteratively refines shape or generalised pose through a render-and-compare loop, combining coarse visual reasoning with numerical pose optimisation for precise state estimation. This structured formulation enables our method to track through large motion, articulation, and severe occlusion without relying on pixel-matching objectives. Quantitatively, on ARCTIC, our method substantially outperforms state-of-the-art 3D point-tracking baselines for articulated objects, and on HOT3D it outperforms all evaluated rigid-object tracking baselines.
Problem

Research questions and friction points this paper is trying to address.

shape reconstruction
tracking
monocular video
articulation
occlusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

agentic analysis-by-synthesis
structured 3D object model
render-and-compare loop
vision-language model
🔎 Similar Papers
No similar papers found.