Score
Design and implement representations, estimators, and preprocessing that decompose and parametrize camera poses and rig geometry into distinct components (for example separating ego‑motion from fixed rig transforms, normalizing extrinsics, or factorizing a trajectory into per‑rig and per‑motion parts). This work produces concrete parameterizations and algorithms to estimate, compose, and normalize those pose components so that static rig priors can be applied independently and trajectory‑specific temporal models can be built.
Existing visual geometry models in autonomous driving struggle to effectively leverage static multi-camera geometric priors from the vehicle due to their coupled modeling of time-varying ego-motion and fixed camera rig geometry. This work proposes TRIG, a novel framework that explicitly decouples camera pose into two components: the ego-vehicle trajectory and the camera rig configuration, thereby separately modeling dynamic self-motion and static multi-camera topology. TRIG introduces a decoupled pose representation, independent supervision mechanisms, and sparse spatio-temporal attention, which together preserve geometric reasoning capabilities while substantially reducing computational overhead. Evaluated across five autonomous driving benchmarks, the method achieves state-of-the-art performance in pose estimation, metric depth prediction, and 3D reconstruction.
Existing methods for joint 3D reconstruction, camera pose estimation, and rigid-body (rig) structure discovery under multi-camera rigid mounts neglect rig geometry by treating images as unordered sets. Method: We propose rig-aware latent-space modeling—enabling both explicit geometric conditioning and implicit rig-structure inference—and design a bi-ray graph decoder that jointly predicts global camera poses and the rig center. Our end-to-end multi-view framework fuses camera IDs, timestamps, and rig pose metadata, and jointly decodes point clouds and bi-ray graphs. Results: Evaluated on real-world rig datasets, our method achieves state-of-the-art performance across all three tasks: 3D reconstruction, pose estimation, and rig discovery. It improves mean Average Accuracy (mAA) by 17–45% over prior work, attains optimal results in a single forward pass, and requires no post-processing.
State-of-the-art rigging methods assume a canonical rest pose--an assumption that fails for sequential data (e.g., animal motion capture or AIGC/video-derived mesh sequences) that lack the T-pose. Applied frame-by-frame, these methods are not pose-invariant and produce topological inconsistencies across frames. Thus We propose SPRig, a general fine-tuning framework that enforces cross-frame consistency losses to learn pose-invariant rigs on top of existing models. We validate our approach on rigging using a new permutation-invariant stability protocol. Experiments demonstrate SOTA temporal stability: our method produces coherent rigs from challenging sequences and dramatically reduces the artifacts that plague baseline methods. The code will be released publicly upon acceptance.
Existing methods for automatic rigging of 3D meshes often fail to model plausible joint pose distributions effectively, leading to anatomically implausible or geometrically self-intersecting poses. This work proposes ViPS, a framework that, for the first time, distills motion priors from pretrained 2D video diffusion models into a general-purpose 3D pose distribution, enabling zero-shot generalization to unseen species and skeletal topologies without relying on scarce 4D data. By integrating a differentiable geometric validator with latent-space pose modeling, ViPS supports effective pose sampling, inverse kinematics projection, and temporally coherent keyframe generation. Experiments demonstrate that ViPS, trained solely on video priors, matches state-of-the-art methods based on synthetic 4D data in both pose plausibility and diversity, while exhibiting superior cross-domain generalization capabilities.
This paper addresses the relative pose estimation problem for three calibrated cameras given only four correspondences across all views. To overcome limitations of conventional methods—namely, their reliance on more correspondences or insufficient robustness—we propose a novel strategy that approximates a fifth correspondence using the centroid of the four observed points. We further introduce the first joint three-view pose estimation framework integrating a 4-point affine fundamental matrix solver, a standard 5-point relative pose solver, and a P3P solver. Geometric modeling enhances robustness against noise and outliers, while local optimization refines accuracy. Evaluated on real-world datasets, our method achieves state-of-the-art performance: the centroid-based strategy significantly outperforms pure affine approaches, striking a superior balance among accuracy, robustness, and computational efficiency, with straightforward implementation.
This work addresses the challenge of flexibly supporting multimodal camera motion control in video generation. To this end, the authors propose a modality-agnostic framework that maps video, pose, and text inputs into a unified motion embedding space, enabling consistent and precise viewpoint manipulation. Key contributions include the construction of a Motion Triplet Dataset, the introduction of a geometry-driven motion representation based on camera extrinsics, and the design of a motion consistency objective in the latent space. The proposed method not only unifies multimodal inputs under a single processing pipeline but also enables novel capabilities such as motion sequence composition and cross-modal interpolation. Experiments demonstrate that the approach generates high-quality videos across all three modalities, accurately adhering to target camera trajectories and validating its effectiveness and generalization.
This work addresses the high computational complexity and substantial point correspondence requirements that often hinder practical deployment of relative pose estimation in multi-camera systems. The authors propose two efficient minimal solvers that, for the first time, incorporate IMU-derived priors—either the vertical direction or a known rotation axis—into a minimal solver framework. Requiring only four point correspondences, the approach leverages a novel parametrization and algebraic geometry techniques to reduce the problem to solving a univariate sextic polynomial, a significant simplification over existing octic formulations. Integrated within a RANSAC-based robust estimation pipeline, extensive experiments on synthetic data and the KITTI benchmark demonstrate that the proposed method achieves comparable accuracy while substantially lowering computational cost and data requirements.
Existing relative pose estimation algorithms incur high computational costs and rely heavily on numerous feature matches, making them ill-suited for the real-time and robustness demands of autonomous driving. This work proposes a unified and efficient framework for relative pose estimation that introduces a novel translation parameterization and a first-order rotation approximation to derive three minimal solvers tailored for ground vehicles. By integrating multi-source priors—such as IMU-provided gravity direction, rotational axis constraints during steering, and the planar motion assumption—the method substantially reduces both the required number of point correspondences and algebraic complexity. Experiments on synthetic data and the KITTI benchmark demonstrate that the proposed approach achieves a superior trade-off between accuracy and speed compared to state-of-the-art methods.