sequence alignment

Designs and implements methods to compute and refine spatial and temporal correspondences between sequences of observations, including estimating per-sequence transforms (e.g., from structure‑from‑motion) and computing per‑sequence or per‑frame alignments. Builds algorithms that aggregate multi‑view geometry and noisy camera poses to align sequence geometry to a canonical frame or to other sequences.

sequencealignment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of existing panoramic stitching methods, which rely on pairwise feature matching and often fail to maintain multi-view geometric consistency in complex scenes characterized by weak textures, large disparities, or repetitive patterns, leading to misalignments and distortions. To overcome these challenges, the authors propose a photogrammetry-driven global alignment framework that leverages estimated camera poses to align images in 3D space. They introduce a novel 3D-aware Transformer architecture that explicitly models multi-view geometric consistency through joint feature optimization and cross-view information aggregation. Key contributions include the first formulation of multi-view consistency in 3D space, a Transformer-based 3D-aware stitching network, and the creation of the first large-scale real-world panoramic stitching dataset. Experiments demonstrate that the proposed method significantly outperforms state-of-the-art approaches in both alignment accuracy and visual quality, particularly exhibiting superior robustness and consistency in challenging scenarios.

3D photogrammetrygeometric consistencyimage distortion

Geometry-aware Feature Matching for Large-Scale Structure from Motion

Sep 03, 2024
GC
Gonglin Chen
🏛️ University of Southern California | The Ohio State University

In large-scale Structure-from-Motion (SfM), sparse inter-view overlap and drastic viewpoint changes—especially in aerial-to-ground scenarios—lead to low cross-image feature matching density and weak geometric consistency. To address this, we propose a geometry-guided hybrid matching paradigm: (1) geometric verification is formulated as an optimization problem based on Sampson distance; (2) detector-agnostic dense matching is fused with detector-driven sparse anchor guidance, where sparse anchors constrain and enhance the geometric consistency of dense matches; and (3) multi-view geometric consistency is explicitly modeled. Our method significantly improves both matching density and accuracy, outperforming state-of-the-art approaches in extreme large-scale settings. Consequently, camera pose estimation becomes more accurate, and the reconstructed 3D point cloud achieves higher completeness and fidelity.

Combining detector-free and detector-based methods for geometric consistencyEnhancing feature matching with geometry cues for large-scale SfMImproving correspondence density and accuracy in sparse view overlap

A Linear N-Point Solver for Structure and Motion from Asynchronous Tracks

Jul 30, 2025
HS
Hang Su
🏛️ ShanghaiTech University | University of Pennsylvania | Amap | Alibaba Group

In asynchronous imaging systems—such as rolling-shutter cameras and event cameras—2D point correspondences possess non-synchronized timestamps, posing fundamental challenges for structure-from-motion (SfM). Method: This paper proposes the first unified linear N-point framework for simultaneous structure and motion recovery, grounded in a first-order constant-velocity dynamic model. It derives linear geometric constraints from asynchronous point trajectories, enabling flexible handling of arbitrary numbers of views, multi-sensor fusion, and cross-modal modeling across global-shutter, rolling-shutter, and event cameras. Contribution/Results: The method is fully linear and admits closed-form solutions, accompanied by rigorous degeneracy analysis and criteria for solution multiplicity. Extensive evaluations on synthetic and real-world datasets demonstrate significant improvements in accuracy and robustness over existing nonlinear or single-frame approaches. The implementation is publicly available.

Developing a linear solver for efficient velocity and 3D point recoveryEstimating structure and motion from asynchronous point correspondencesHandling arbitrary timestamps and multiple sensor modalities

Practical solutions to the relative pose of three calibrated cameras

Mar 28, 2023
CT
C. Tzamos
🏛️ Czech Technical University in Prague | ETH Zürich

This paper addresses the relative pose estimation problem for three calibrated cameras given only four correspondences across all views. To overcome limitations of conventional methods—namely, their reliance on more correspondences or insufficient robustness—we propose a novel strategy that approximates a fifth correspondence using the centroid of the four observed points. We further introduce the first joint three-view pose estimation framework integrating a 4-point affine fundamental matrix solver, a standard 5-point relative pose solver, and a P3P solver. Geometric modeling enhances robustness against noise and outliers, while local optimization refines accuracy. Evaluated on real-world datasets, our method achieves state-of-the-art performance: the centroid-based strategy significantly outperforms pure affine approaches, striking a superior balance among accuracy, robustness, and computational efficiency, with straightforward implementation.

Estimating relative pose of three calibrated camerasImproving robustness with approximate mean-point correspondencesUsing four point correspondences for efficient solutions

This study addresses the insufficient robustness and accuracy of local feature matching in overlapping regions of satellite imagery. To this end, the authors construct a manually curated satellite image dataset annotated with GPS coordinates and conduct a systematic evaluation of SIFT and ORB algorithms across the entire matching pipeline—including keypoint detection, descriptor extraction, feature matching, and RANSAC-based geometric verification. Using the inlier ratio as the primary metric for matching quality, the work quantitatively analyzes the impact of keypoint quantity on matching performance. The results reveal a nonlinear relationship between the number of detected keypoints and the inlier ratio, offering empirical evidence and theoretical guidance for algorithm selection and parameter tuning in remote sensing image matching tasks.

Image MatchingInlier RatioORB

Latest Papers

What's happening recently
View more

While existing multi-frame models achieve cross-frame consistency, their single-frame accuracy often lags behind that of single-frame methods. Through systematic ablation studies, this work demonstrates that data diversity and quality are critical for 3D geometry estimation and reveals that commonly used loss functions may inadvertently suppress performance. To address these issues, the authors propose CARVE, a novel approach integrating a high-resolution network architecture, joint sequence- and frame-level supervision, a consistency loss, and alignment between depth maps and camera parameters. CARVE achieves state-of-the-art and robust performance across multiple benchmarks in tasks including point cloud reconstruction, video depth estimation, and estimation of camera pose and intrinsics.

3D reconstructiondepth estimationmulti-frame consistency

This study addresses the susceptibility of alternating minimization to local optima and its computational inefficiency in correspondence-free point set alignment. We propose a global optimization method based on support vectors derived from the convex hull vertices of permuted polygons. By proving a tight bound of $n(n-1)$ vertices, we resolve an open problem posed by Rote. Integrating the Procrustes-Wasserstein framework with a branch-and-bound algorithm, our approach achieves exact solutions in 2D and extends naturally to 3D. Evaluated on the MPEG-7 benchmark, the method requires only 12ms on average, achieving a 50-fold speedup over grid search while delivering superior accuracy. These improvements substantially enhance shape retrieval performance, demonstrating both theoretical rigor and practical efficiency for robust point set registration.

correspondence-freeglobal optimizationpoint set alignment

This study addresses the challenge of unreliable manual feature matching in traditional RPC bundle adjustment for multi-temporal satellite imagery, which suffers from seasonal variations, illumination changes, and surface cover dynamics that degrade uncontrolled geometric positioning accuracy. To overcome this limitation, the authors propose an appearance-aware RPC refinement method that, for the first time, jointly leverages learned local features and global image descriptors to robustly extract season-invariant correspondences. Furthermore, the approach employs visual compatibility metrics to select optimal image pairs, thereby enhancing the match graph structure. Evaluated on a multi-season WorldView-3 dataset, the method significantly outperforms open-source baselines, achieving notably reduced geometric consistency errors and substantially improved matching efficiency across image blocks of 39–42 scenes.

feature matchinggeolocation accuracymulti-date satellite imagery

Existing methods struggle to simultaneously achieve geometric accuracy and temporal consistency in long-duration monocular video reconstruction. This work proposes an end-to-end neural network that generates affine-invariant 3D point maps through a single forward pass, enabling scale-consistent reconstruction within a unified reference frame. The approach introduces three core innovations: viewpoint-invariant geometric alignment, appearance-invariant learning, and frequency-modulated localization, which collectively facilitate consistent modeling across exponentially varying time scales and support robust extrapolation over ultra-long sequences. Experiments on datasets such as ScanNet demonstrate a 24.2% reduction in point map error and a 34.9% decrease in temporal alignment error, significantly enhancing reconstruction robustness under complex camera trajectories and varying illumination conditions.

3D geometry estimationgeometric accuracymonocular video

Hot Scholars

JZ

Jiawei Zhou

Assistant Professor, Stony Brook University | TTIC, Harvard
Natural Language ProcessingMachine Learning
DS

Dale Schuurmans

Google DeepMind & University of Alberta
Machine Learning
ET

Earl T. Barr

Professor, University College London
software engineeringcomputer securityprogramming languages
KZ

Kexin Zhang

Tsinghua University
Data MiningMachine Learning
JP

Jiahao Pan

Hong Kong University of Science and Technology
Speech ProcessingSpeech EnhancmentMusic Generation