Know-Your-Scene (KYS)-SLAM: Hierarchical Semantic-Motion Priors for Feature Matching in Stereo Visual SLAM

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决立体视觉SLAM中的语义模糊、实例混淆及独立运动物体问题,KYS-SLAM通过引入分层语义-运动先验来连续调整特征匹配,而非简单拒绝特征。
📝 Abstract
Stereo visual SLAM systems built on local descriptors suffer from semantic ambiguity, instance-level confusion, and independently moving objects, each corrupting data association and accumulating as trajectory drift. Prevailing semantic and dynamic SLAM methods address this through binary feature rejection, sacrificing correspondence density for outlier suppression. We contend that contextual implausibility is better expressed as a graded quantity than an exclusion criterion. We present Know-Your-Scene (KYS)-SLAM, a modular extension of ORB-SLAM3 that supplants feature rejection with continuous correspondence modulation. The contribution is the reframing of contextual evidence as correspondence cost, applied within feature matching and leaving the geometric backend unmodified. Each keypoint is augmented with semantic, panoptic, and motion priors fused through a hierarchical compatibility formulation, in which semantic class and instance identity enforce structural plausibility while a zero-shot motion score down-weights features on independently moving objects. That score comes from a training-free module fitting a depth-aware ego-motion model to background optical flow and classifying panoptic segments via self-calibrating, coverage-aware thresholds, so only segments with sufficient motion evidence are penalized and static structure is left unpenalized. Penalizing correspondences rather than discarding them preserves the geometric support bundle adjustment depends on. Under one fixed configuration, no coefficient retuned per sequence or dataset, KYS-SLAM reduces per-sequence ATE RMSE by 17.4% on outdoor KITTI and 27.7% on indoor EuRoC across 21 stereo sequences with no regressions, and by 6.6% on dynamic subsets of KITTI Tracking and 17.8%, up to 31.2%, on Virtual KITTI 2 -- cross-domain transfer across outdoor driving, indoor flight, and synthetic imagery under one set of constants.
Problem

Research questions and friction points this paper is trying to address.

stereo visual SLAM
semantic ambiguity
instance-level confusion
independently moving objects
trajectory drift
Innovation

Methods, ideas, or system contributions that make the work stand out.

continuous correspondence modulation
hierarchical compatibility formulation
zero-shot motion score
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Preeti Chatterjee
School of Computing, University of Georgia, Athens, GA 30602, USA
Jin Lu
Jin Lu
University of Georgia
Machine LearningData MiningHealth informatics
Jin Sun
Jin Sun
Assistant Professor, University of Georgia
Computer Vision
S
Suchendra M. Bhandarkar
School of Computing, University of Georgia, Athens, GA 30602, USA