multi-camera synchronization

Coordinating capture timing and sensor calibration across multiple cameras and modalities to produce synchronized multi-view recordings in controlled or in-the-wild settings, including designs for sensor setups that enable precise gaze, manipulation, or dexterous-data collection without specialized infrastructure.

multi-camerasynchronization

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

RocSync: Millisecond-Accurate Temporal Synchronization for Heterogeneous Camera Systems

Nov 18, 2025
JM
Jaro Meyer
🏛️ ETH Zurich | Balgrist University Hospital | University of Zurich

Heterogeneous camera systems—e.g., visible-light/infrared, professional/consumer-grade, or audio-equipped/audio-less setups—lack hardware synchronization in real-world scenarios, leading to significant spatiotemporal misalignment across multi-view videos. Method: This paper proposes a vision-based time-encoding method leveraging a custom-designed LED Clock. By embedding temporal exposure timestamps within frames using red and infrared LEDs, the approach achieves cross-modal, audio-free, and external-timecode-free millisecond-level synchronization. It further integrates RMSE-optimized temporal alignment with joint multi-device calibration. Contribution/Results: The method reduces synchronization residuals to 1.34 ms—substantially outperforming existing optical signal, audio-based, and timecode synchronization schemes. Validated in large-scale surgical recordings involving over 25 heterogeneous cameras, it significantly improves downstream tasks including multi-view 3D reconstruction and pose estimation.

Achieving millisecond temporal alignment for RGB and IR camerasEnabling accurate multi-view applications in unconstrained real-world environmentsSynchronizing heterogeneous camera systems lacking hardware sync

Existing systems struggle to support strict audio-visual synchronization, limiting the analysis of fine-grained temporal features in dialogue such as turn-taking, overlapping speech, and prosody. To address this challenge, this work proposes an end-to-end multimodal acquisition and calibration framework that treats synchronized audio and video as equally central modalities for the first time. By integrating a multi-camera array with multi-channel microphones under a unified temporal architecture, the system enables scalable, reproducible, high-quality recording. Standardized calibration and quality control procedures ensure high temporal consistency across modalities, yielding data that effectively supports fine-grained analysis of conversational behavior and data-driven modeling.

audio-visual synchronizationconversational interactionhuman motion recording

Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion

Dec 01, 2025
SL
Shaowei Liu
🏛️ University of Illinois Urbana-Champaign

This paper addresses the problem of high-precision automatic synchronization of multi-camera video streams from consumer-grade cameras under uncontrolled environments, without dedicated hardware or manual intervention. We propose VisualSync, the first framework to jointly model generic 3D reconstruction, cross-view feature matching, and dense motion trajectory tracking—leveraging epipolar geometry constraints from co-visible dynamic objects to directly estimate millisecond-level inter-camera time offsets via end-to-end optimization. The method relies entirely on off-the-shelf algorithms, requiring no camera calibration, external synchronization signals, or scene instrumentation. Evaluated on four diverse real-world datasets, VisualSync achieves a median synchronization error below 50 ms, significantly outperforming existing baselines. It establishes a scalable, calibration-free synchronization paradigm for low-cost multi-view motion analysis and collaborative perception.

Aligning cross-view recordings using moving object motion and epipolar constraintsEstimating time offsets for unposed cameras with millisecond accuracySynchronizing unsynchronized multi-camera videos without manual input

Spatiotemporal Multi-Camera Calibration using Freely Moving People

Feb 18, 2025
SL
Sang-Eun Lee
🏛️ Kyoto University | Kyoto Institute of Technology

This paper addresses the spatiotemporal calibration challenge for multi-view video in dynamic, multi-person scenes. We propose an end-to-end, markerless calibration method leveraging freely moving pedestrians. By modeling cross-view human motion as probabilistic point-set registration on the unit sphere, our approach jointly estimates camera rotation, translation, inter-camera time offsets, and person-level cross-view correspondences. Our key contribution is a novel temporal-geometric joint optimization framework that integrates monocular 3D pose estimation, unit-sphere projection, soft assignment matching, coplanarity constraints, and multi-view consistency regularization. Evaluated on both synthetic and real-world datasets, the method achieves sub-degree rotational accuracy (<1°) and sub-hundred-millisecond temporal synchronization precision—significantly outperforming existing calibration-free approaches. The framework is robust to unconstrained human motion and requires no specialized calibration objects, making it suitable for flexible deployment in practical surveillance and human motion analysis applications.

Marker-free calibration using human motionSpatiotemporal multi-camera calibration challengeUnified framework for camera calibration

This work addresses the challenge of time synchronization among heterogeneous sensors in roadside and vehicle-mounted multi-LiDAR–multi-camera systems by proposing an open-source, modular, and scalable hardware synchronization solution. Using the LiDAR synchronization pulse as a reference, the system employs programmable delay circuits to generate independent trigger signals for each camera, enabling flexible and precise spatiotemporal alignment. The architecture supports arbitrary combinations of sensor counts and has been validated on both a three-camera roadside platform and a seven-camera vehicular setup. Experimental results demonstrate significantly improved spatial consistency between point clouds and images, while the design ensures robustness, reproducibility, and ease of deployment.

heterogeneous sensorsmulti-cameramulti-lidar

Latest Papers

What's happening recently
View more

This work addresses the challenge of extrinsic calibration for non-overlapping multi-camera systems by proposing a novel method that requires only pure rotational motion and a single static calibration target. By introducing an implicit turntable coordinate frame and formulating a 3D reprojection error on the SE(3) manifold, the approach integrates observations of the same calibration board captured by different cameras at distinct time instances into a unified global nonlinear optimization framework. The method eliminates the need for large calibration patterns or complex motion estimation, thereby avoiding scale ambiguity and drift issues. High accuracy, strong robustness, and ease of deployment are demonstrated on both controlled rigs and real-world vehicle platforms, marking the first successful realization of high-precision extrinsic calibration for non-overlapping multi-camera setups using only pure rotation and a single static target.

calibrationextrinsic calibrationmulti-camera systems

This work addresses the challenges in extrinsic calibration between motion capture systems and external cameras—particularly fisheye cameras—in large-scale data collection, where errors arising from calibration board misalignment, ambiguous initialization, and temporal drift are difficult to detect promptly. To overcome these issues, the authors propose a robust joint calibration and independent validation framework that simultaneously estimates camera extrinsics and the transformation from the calibration board to motion capture markers. The method employs a staged nonlinear optimization strategy insensitive to initialization and introduces a fully independent validation pipeline—the “lollypop” module—that effectively handles non-uniform fisheye distortion. Experiments on the Meta Quest 3 demonstrate superior calibration accuracy over existing approaches, with the lollypop module reliably detecting calibration degradation during long-term operation. The system has been successfully deployed in real-world data acquisition pipelines.

calibration verificationcamera-to-mocap calibrationextrinsic calibration

This work addresses geometric inconsistencies, error accumulation, and poor scalability in real-time multi-view 3D reconstruction caused by the tight coupling of extrinsic calibration, point cloud fusion, and global optimization. To overcome these limitations, we propose FUSE-Flow, a novel framework that decouples calibration and fusion into two synergistic modules for the first time. The GMAC module leverages geometric constraints and a multi-view reconstruction Transformer to estimate sparse extrinsics without requiring calibration targets, while the FUSE module enables stateless, real-time point cloud fusion through confidence-weighted integration and adaptive spatial hashing. These modules mutually refine each other via a confidence feedback mechanism. Extensive experiments on public datasets and real-world systems demonstrate that our approach significantly outperforms existing methods in accuracy, dynamic stability, and scalability, enabling large-scale real-time 3D reconstruction.

extrinsic calibrationmulti-view fusionpoint cloud

This work addresses the limitations of traditional marker-based pose estimation in multi-camera dynamic augmented reality systems, which rely on continuously visible fiducials and struggle with collaborative localization across non-overlapping camera views. The authors propose a novel markerless approach for dynamic multi-camera pose estimation that leverages spatiotemporal overlaps of known objects in the scene to construct and incrementally update a spatiotemporal scene graph, enabling cross-camera pose co-optimization. By fusing multi-view object observations within this graph structure, the method significantly enhances pose accuracy. Extensive experiments on YCB-V, T-LESS, and a newly introduced multi-camera, multi-object overlapping dataset demonstrate consistent superiority over existing methods, validating the approach’s effectiveness and robustness for markerless AR applications.

augmented realitymarker-less trackingmulti-camera pose estimation

Indoor visual localization is hindered by detection noise, occlusions, and limited camera coverage, leading to multi-stage uncertainties that existing fusion methods fail to explicitly model. This work proposes a component-level error quantification and calibration mechanism that explicitly characterizes the uncertainty in homography calibration, human detection, and motion tracking, and leverages these estimates to optimize multi-camera fusion weights. By transforming the fusion process from a black-box into an interpretable framework, the method significantly enhances trajectory stability and motion smoothness. Experimental results demonstrate that, while yielding only marginal gains in absolute localization accuracy over single-camera baselines, the proposed strategy effectively reduces trajectory variance and substantially improves the continuity and robustness of motion estimation.

detection noiseerror characterizationindoor localization

Hot Scholars

BG

Banglei Guan

National University of Defense Technology
PhotomechanicsVideometrics
MP

Marc Pollefeys

Professor of Computer Science, ETH Zurich, and Director Spatial AI Lab, Microsoft
Computer VisionComputer GraphicsRoboticsMachine Learning
DC

Daniel Cremers

Technical University of Munich
Computer VisionMachine LearningOptimizationRobotics
ZL

Zibin Liu

National University of Defense Technology
Neuromorphic vision sensorsEvent cameraCamera calibrationPose estimation
PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics