GAUGE: Group-Wise View-Inconsistency Rectification for Feed-Forward 4D Tracking

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant residual errors along the viewing direction in feedforward 4D tracking, which have previously lacked a structural explanation. We reveal for the first time that these errors originate from group-level view inconsistency and propose GAUGE, a training-free post-processing module designed to correct them. Specifically, GAUGE performs motion grouping based on directional consistency and integrates sparse metric anchors with adaptive scale estimation to recover the underlying motion group structure while rectifying radial scale and translational biases. Extensive experiments demonstrate that GAUGE reduces endpoint errors by 15.1% to 62.6% across eight mainstream trackers, outperforming conventional fine-tuning approaches. By requiring no gradient updates, our method achieves efficient, plug-and-play correction for 4D tracking pipelines.
📝 Abstract
Feed-forward models regress dense 3D point trajectories directly from monocular video, yet the residual after global alignment is substantial and lacks a structural explanation. Measured on dynamic query points across models and datasets, the error concentrates along the view direction, while the scale correction each motion group requires differs. The predicted displacement direction nevertheless supports reliable grouping, with a median angle far below the 90{\deg} random baseline. The systematic part of the residual is therefore a family of radial degrees of freedom per motion group, along directions 2D observations cannot constrain. We call it group-wise view inconsistency. We present GAUGE (Group-wise Adaptive Unsupervised Gauge Estimation), a training-free and model-agnostic post-hoc module. It recovers motion groups from direction consistency and spatial connectivity, then estimates a per-frame radial scale and group-level translation from 1% to 5% metric anchors, four degrees of freedom per group and frame. On dynamic query points of eight trackers, including D4RT, 4RC and SM4RT, our correction lowers endpoint error by 15.1% to 62.6% over the uncorrected predictions, while spending the same anchors on gradient fine-tuning improves the same models by only -1.1% to 15.2%. Code is publicly available at https://github.com/HCPLab-SYSU/GAUGE.
Problem

Research questions and friction points this paper is trying to address.

4D tracking
feed-forward model
view inconsistency
monocular video
trajectory estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

4D Tracking
View-Inconsistency Rectification
Training-free Post-hoc Module
Motion Grouping
Feed-Forward Models
🔎 Similar Papers
No similar papers found.