extract keypoints

Designs and implements methods that detect and localize 2D and 3D keypoints or landmarks by predicting or regressing their coordinates from input signals; these systems handle both predefined and category-agnostic/open-domain targets, support prompt-conditioning, normalize and denoise measurements, estimate uncertainty, and refine predictions (for example using depth cues) to produce landmark outputs for downstream computations such as pose estimation. Work includes building models and training/evaluation pipelines across sensor modalities (e.g., RGB, RGB‑D, wearable sensors), and implementing filtering, representation normalization, and task-specific metrics (e.g., PCK) for keypoint quality assessment.

extractkeypoints

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

In visual localization, devices struggle to accurately match landmarks and estimate their positions when wireless signals are unavailable or unreliable and nearby landmarks exhibit high visual similarity. This paper proposes a purely vision-aided localization method leveraging multi-source distance measurements and geometric constraints. First, landmarks are modeled as a marked Poisson point process (PPP); we theoretically prove that only three noise-free distance measurements suffice to uniquely identify the landmark set in 2D. Second, we construct a joint distribution model for key random variables and derive a closed-form analytical expression for the landmark-set identification probability under noisy measurements. Experiments demonstrate that, even when landmarks are visually indistinguishable, the method achieves robust identification and precise localization with high confidence.

Similar Visual Landmarks IdentificationVisual LocalizationWireless Signal Unavailability

This work addresses the unreliability of existing keypoint-based pose estimation methods, which often neglect geometric constraints inherent to object shape, leading to undetectable failures. To overcome this limitation, the authors propose a failure detection approach that does not rely on keypoint uncertainty estimates. Instead, it leverages handcrafted geometric consistency features—specifically pairwise keypoint distances, reprojection errors, and consistency between rendered and observed masks—to characterize spatial relationships among 2D keypoints. These features are fed into a logistic regression classifier to determine whether a given pose estimate has failed. Experimental results demonstrate that the proposed method significantly outperforms existing confidence-based failure detection schemes, such as conformal keypoint prediction, offering superior reliability and practical utility in real-world applications.

failure detectiongeometric consistencykeypoint prediction

Good Keypoints for the Two-View Geometry Estimation Problem

Mar 24, 2025
KP
Konstantin Pakulev
🏛️ Skolkovo Institute of Science and Technology | Slamcore

This work addresses the critical impact of local feature quality on downstream geometric estimation performance in two-view geometry. For the first time, it theoretically models keypoint quality, identifying repeatability and low measurement error as the two fundamental determinants of geometric estimation accuracy. To this end, we propose BoNeSS-ST—a novel keypoint detector featuring: (i) a sub-pixel refinement mechanism for high-precision localization; (ii) a self-supervised keypoint scoring function; and (iii) enhanced robustness to low-salience regions. BoNeSS-ST is jointly trained on both homography and fundamental matrix estimation tasks to promote geometric consistency. Evaluated on planar homography and epipolar geometry estimation benchmarks, BoNeSS-ST significantly outperforms existing self-supervised detectors, achieving state-of-the-art performance in both matching quality and geometric estimation accuracy.

Design a keypoint detector improving homography and epipolar geometry accuracyDevelop a model for scoring keypoints in homography estimationIdentify keypoint properties for better two-view geometry estimation

This work addresses the lack of spatial uncertainty modeling in YOLO-Pose for keypoint localization. The authors propose a lightweight, post-hoc probabilistic extension that introduces an additional probability head to predict input-dependent 2×2 covariance matrices, enabling calibrated bivariate Gaussian or Student-t distributions over original keypoints. This is the first approach to equip YOLO-Pose with keypoint-level predictive distributions. A novel evaluation protocol is introduced, combining distribution calibration diagnostics with Average Keypoint Precision (AKP). Experiments on COCO demonstrate that the method effectively supports reliability-based keypoint ranking, with the Student-t formulation yielding more accurate residual distribution fitting. Furthermore, in an aircraft visual landing task, the calibrated covariance enables uncertainty-aware pose estimation and sensor fusion.

keypoint uncertaintypredictive distributionsspatial uncertainty

Latest Papers

What's happening recently
View more

This study addresses the challenge of accurately identifying the target object indicated by a human pointing gesture using only a single RGB image. To this end, the authors propose a modular pipeline that integrates object detection, human pose estimation, monocular depth estimation, and a vision-language model. By reconstructing 3D spatial relationships and leveraging image captioning to correct classification errors, the approach effectively resolves ambiguities in pointing direction. This work presents the first systematic evaluation of the synergistic role between 3D spatial information reconstructed from a single image and vision-language models for pointing target recognition. Experiments on a newly curated dataset demonstrate that incorporating depth cues significantly improves accuracy in complex occlusion scenarios, without requiring specialized depth sensors, thereby offering strong deployment flexibility.

human-robot interactionobject recognitionpointing gesture

Visual localization is a core technology for augmented reality and autonomous navigation. Recent methods combine the efficient rendering of 3D Gaussian Splatting (3DGS) with feature-based localization. These methods rely on direct matching between 2D query features and the 3D Gaussian feature field, but this often results in mismatches due to an inherent bias in the learned Gaussian feature. We theoretically analyze the feature learning process in 3DGS, revealing that the widely adopted $α$-blending optimization inherently introduces bias into 3D point features. This bias stems from the entanglement between individual Gaussians and their neighboring Gaussians, making the learned features unsuitable for precise matching tasks. Motivated by these findings, we propose ULF-Loc, an unbiased landmark feature framework that replaces biased feature optimization with geometry-weighted feature fusion. We further introduce keypoint-consensus landmark sampling to select reliable Gaussians and local geometric consistency verification to reject mismatches caused by rendering artifacts. On the Cambridge Landmarks dataset, ULF-Loc reduces the mean median translation error by 17\% compared to the state-of-the-art, while achieving superior efficiency with only 1/10 the training time and 1/6 the GPU memory of STDLoc.

3D Gaussian Splattingaugmented realityfeature bias

This work addresses the susceptibility of existing vision-language models (VLMs) to landmark bias in image geolocation, which leads to spurious correlations and localization inaccuracies. To mitigate this, the authors propose HoloGeo, an evidence-driven reasoning framework that guides models to equitably leverage diverse visual cues through structured multi-evidence reasoning chains for unbiased geolocation. The study introduces Bias Intensity and Bias Harmfulness—novel quantitative metrics for assessing landmark bias—and presents LandmarkBias-3K, a new benchmark for evaluation. Furthermore, a multidimensional reinforcement learning reward mechanism is designed, trained on the high-quality BF-30k dataset to foster joint reasoning. Experiments demonstrate that HoloGeo achieves state-of-the-art performance on IM2GPS3K and YFCC4k, and significantly outperforms existing open-source VLMs on LandmarkBias-3K, confirming its robustness and effectiveness.

geo-localizationgeographical cueslandmark bias

This work addresses the challenge of pose estimation in space-constrained environments, where conventional multi-point PnP methods are difficult to deploy and fail to exploit available ego-motion priors such as known height and tilt angle. The authors propose a minimal pose solver requiring only two active LED markers, uniquely incorporating height and tilt constraints into a two-point geometric model. They derive both a closed-form solution and a linear least-squares formulation, and provide a systematic analysis of degenerate configurations. By fusing event camera data with IMU and altimeter measurements within the proposed geometric framework, the method significantly outperforms existing P2P approaches on both synthetic and real-world datasets, achieving accuracy comparable to P3P while demonstrating superior efficiency, accuracy, and robustness.

active LED markersevent camerasheight constraint

Hot Scholars

BD

Bo Du

Department of Management, Griffith Business School
Sustainable TransportTravel BehaviourUrban Data AnalyticsLogistics and Supply Chain
JW

Jiajun Wu

Stanford University
Computer VisionRoboticsArtificial IntelligenceMachine Learning
BG

Banglei Guan

National University of Defense Technology
PhotomechanicsVideometrics
SL

Stefan Leutenegger

Associate Professor at ETH Zurich
Mobile RoboticsSpatial AIUnmanned Aerial Systems
DS

Danail Stoyanov

Professor of Robot Vision, University College London
Surgical VisionSurgical AISurgical RoboticsComputer Assisted Interventions