perform real-time motion tracking

Designs, builds, or evaluates systems that continuously estimate and report 6-DoF poses, translations, rotations, and applied forces of rigid, articulated, or deformable objects in real time using motion-capture and electromagnetic (EMT) tracking sensors. This work includes sensor fusion and hardware–software integration with simulators or guidance systems, empirical validation under realistic conditions, and alignment of tools or tracked objects to planned trajectories.

performreal-timemotiontracking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$194K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the limitations of traditional human dynamics analysis, which relies on contact-based force/torque sensors and controlled environments, rendering it unsuitable for non-contact scenarios. To overcome this, the authors propose an optics-mechanics integrated framework for simultaneous kinematic and dynamic estimation. By formulating a constrained multibody dynamics model, the method leverages vision-based kinematic measurements as non-contact inputs and employs a genetic algorithm to optimize joint torque identification. This approach represents the first demonstration of non-contact dynamic parameter estimation for multibody systems using only visual data and a mechanical model, eliminating dependence on force sensors. Experimental validation on an air-bearing platform shows a mean absolute error of 0.46 Nm in wrist joint torque estimation and a forward-predicted angular velocity error as low as 0.006 rad/s.

contactless measurementdynamic estimationjoint torque estimation

6D Object Pose Tracking in Internet Videos for Robotic Manipulation

Mar 13, 2025
GP
Georgy Ponimatkin
🏛️ Czech Technical University in Prague | H Company

This work addresses the challenging task of extracting 6D pose trajectories of manipulated objects from unstructured internet-sourced instructional videos—characterized by unconstrained camera motion, unknown object CAD models, and subtle object dynamics causing temporal inconsistency. We propose the first RGB-only, end-to-end framework requiring no object priors: it jointly performs cross-modal CAD retrieval, image-level 6D pose alignment, and scene-scale anchoring for initial pose estimation; then refines trajectories via video-based smoothing and robot configuration-space optimization for action retargeting. Our method achieves significant improvements over state-of-the-art on YCB-V, HOPE-Video, and a newly curated instructional video dataset. It successfully drives a 7-DOF robotic arm to replicate demonstrated manipulations in both simulation and real-world settings. Furthermore, we demonstrate strong generalization to first-person videos from EPIC-KITCHENS, validating its potential for embodied AI applications.

Estimate 6D pose without prior object knowledge.Extract 6D pose trajectories from Internet videos.Transfer 6D object motion to robotic manipulators.

This study addresses the limitations of existing sonomyography (SMG) interfaces, which rely on extensive user- and position-specific training data and support only single-degree-of-freedom or task-specific control, thereby failing to meet the needs of individuals with tetraplegia for high-dimensional, continuous, and rapidly adaptable systems. The authors propose a real-time, sensor-position-invariant SMG control framework that leverages sparse optical flow to track muscle deformation, enabling continuous one-degree-of-freedom control with just three calibration poses and extending to two degrees of freedom via lightweight computer-assisted calibration. This approach achieves, for the first time, generalized multi-degree-of-freedom continuous SMG control across multiple anatomical sites—including the arm, neck, and upper torso—significantly reducing calibration burden. In experiments with nine participants, including three with cervical spinal cord injury, all achieved high-accuracy 1-DOF control (error typically <4%) across six sensor locations and successfully performed 2D cursor navigation and drawing tasks, demonstrating the system’s efficacy and generalizability.

continuous controlhigh-dimensional controlsensor-placement-agnostic

This study addresses the challenge of performing accurate, lab-free 3D human kinematic assessment during Activities of Daily Living (ADL) for telemedicine, sports science, and rehabilitation. We systematically benchmark monocular video-based and IMU-based approaches using state-of-the-art models—including MotionAGFormer, MotionBERT, MMPose (2D-to-3D), and NVIDIA BodyTrack—and unify evaluation via OpenSim inverse dynamics and Human3.6M joint-angle metrics. Results demonstrate that MotionAGFormer achieves the highest accuracy (RMSE = 9.27° ± 4.80°, MAE = 7.86° ± 4.18°, correlation r = 0.86 ± 0.15, R² = 0.67 ± 0.28), confirming the clinical feasibility of monocular video in real-world settings. We introduce a new benchmark for in-the-wild human motion capture, explicitly characterizing the trade-offs among accuracy, cost, and deployment practicality between video and IMU modalities. This work provides empirical validation and methodological guidance for scalable, low-cost remote motion monitoring.

Benchmarking monocular video 3D pose estimation against IMUs for kinematic assessmentComparing video and sensor trade-offs for telehealth movement analysisEvaluating joint angle accuracy in daily living activities using deep learning

UPETrack: Unidirectional Position Estimation for Tracking Occluded Deformable Linear Objects

Dec 09, 2025
FW
Fan Wu
🏛️ Huazhong University of Science and Technology | Foshan Institute of Intelligent Equipment Technology | Hunan University

Real-time tracking of deformable linear objects (DLOs) is fundamentally challenging due to their high-dimensional configuration space, strongly nonlinear dynamics, and frequent partial occlusions. This paper proposes a model-free, markerless real-time tracking method: first, a Gaussian Mixture Model (GMM) is constructed from the point cloud of the visible segment to capture local geometry; second, a novel Unidirectional Position Estimation (UPE) algorithm is introduced, which leverages geometric continuity and historical curvature modeling to derive a closed-form deformation extrapolation—without requiring physical modeling, simulation, or iterative optimization; finally, robust occluded-segment state prediction is achieved by jointly enforcing proximal linearity constraints and local displacement composition. Evaluated under diverse occlusion scenarios, the method achieves superior localization accuracy and computational efficiency compared to TrackDLO and CDCPD2, enabling stable, millisecond-level tracking.

Estimates occluded positions using geometric continuity without physical models.Improves accuracy and efficiency over existing state-of-the-art methods.Tracks deformable linear objects in real-time despite occlusions.

Latest Papers

What's happening recently
View more

Achieving high-precision, camera-free full-body 3D motion capture in unconstrained outdoor environments remains challenging. This work proposes the first shape-agnostic, cross-species generalizable method based solely on sparse pairwise distance (PWD) measurements from wearable ultra-wideband (UWB) sensors, eliminating the need for subject-specific body parameters or environmental calibration. At its core is Wild-Poser (WiP), a lightweight, real-time Transformer architecture that directly predicts 3D joint positions from noisy PWD data and jointly learns to recover joint rotations. Experiments demonstrate that the approach achieves low joint position errors and high-fidelity 3D motion reconstruction on both human and non-human subjects, validating its practical potential as a low-cost, scalable solution for complex real-world scenarios.

3D reconstructionin-the-wildmotion capture

该研究通过PRISMA-ScR综述方法,评估了基于视频的无标记运动捕捉技术在临床和康复生物力学中的应用现状及验证情况,指出了现有技术在特定人群和运动分析方面的不足。

Clinical and rehabilitation biomechanicsJoint kinematicsPathological populations

This study addresses the challenge of configuration estimation for continuum robots in the absence of external visual feedback by proposing a distributed sensing scheme that integrates inertial measurement units (IMUs) with active magnetic fields. Through a multi-source data fusion algorithm, the proposed method achieves embedded self-perception, enabling precise localization without external cameras while supporting high-frequency state updates at 16.7 Hz. Experimental results demonstrate that the system exhibits excellent real-time performance, effectively maintaining the desired pose of the end-effector and facilitating stable closed-loop control. This work establishes a novel paradigm for vision-free autonomous perception of continuum robots operating within confined spaces.

Configuration estimationContinuum robotPose estimation

This work addresses the challenge that trajectory representations often fail to maintain consistent identification and generalization performance across different coordinate systems, as existing coordinate-invariant formulations are susceptible to measurement noise and suffer from singularities. To overcome these limitations, the paper proposes a Dual Upper-Triangular Invariant Representation (DUTIR), which leverages differential geometry and invariant theory to construct computable local coordinate-invariant features. This approach enables unified modeling of rigid-body motion and interaction-force trajectories under arbitrary coordinate frames. The method demonstrates significantly enhanced robustness against both singularities and sensor noise, achieving stable and consistent performance in trajectory segmentation, recognition, and prediction across multiple coordinate systems. Its effectiveness and broad applicability are validated through experiments in robotics and biomechanics.

coordinate-invariant representationforce trajectoriesmotion trajectories

Current assessment of early motor impairments in infants relies heavily on expert visual observation, lacking objective and automated methods. This study presents the first integration of markerless 3D pose estimation with biomechanical modeling to systematically evaluate the performance of three algorithms—MeTRAbs-ACAE, SAM 3D Body, and Sapiens—on multi-view infant videos, followed by full-body motion reconstruction via an inverse kinematics framework. Comprehensive analysis based on multi-view geometric consistency, reprojection error, and Procrustes alignment error demonstrates that SAM 3D Body achieves the best performance (Procrustes error: 19–28 mm) and effectively discriminates between typical and atypical movement patterns as annotated by clinical experts. These findings validate the feasibility of a video-driven, scalable approach for automated early developmental assessment in infants, establishing a critical technical foundation for clinical screening applications.

3D pose estimationinfant biomechanicskinematic estimation

Hot Scholars

NN

Nassir Navab

Professor of Computer Science, Technische Universität München
JY

Jie Ying Wu

Assistant Professor in CS, Vanderbilt University
Medical RoboticsModelling and SimulationMachine LearningTelerobotics
TH

Tobias Höllerer

Professor, Computer Science, UC Santa Barbara
human-computer interactionaugmented realityvirtual realityinformation visualization
SK

Seungryong Kim

Associate Professor, KAIST
Computer VisionMachine Learning