PGL-3D: Towards Progressive Geometric Learning for 3D Visual Query Localization

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the underutilization of geometric information for feature refinement in 3D visual query localization by proposing a progressive geometric learning framework. The framework introduces a "predict-select-refine-repredict" pipeline and a QTM mechanism that leverages intermediate predicted bounding boxes to guide evidence aggregation and representation updates, alongside ST-D9O supervision for geometric learning. Training is further optimized by integrating RGB-point cloud sequences, soft pooling, and multi-task boundary and distance loss functions. Experimental results demonstrate that the proposed method achieves a stAP of 0.270 on the 3DVQL benchmark and improves mAO to 25.78% on GSOT3D, significantly enhancing 3D localization performance.
📝 Abstract
3D Visual Query Localization (3DVQL) retrieves the latest contiguous occurrence of a queried object in an RGB--point-cloud sequence and predicts a 9-DoF cuboid for every response frame. The query is captured independently of the search sequence, so its annotated pose may differ from how the object appears in the search frames. The benchmark baseline predicts cuboids after feature modeling, leaving their geometry unused for subsequent feature refinement. We investigate whether complete intermediate cuboids can improve query and proposal representations before final decoding. We introduce Progressive Geometric Learning for 3DVQL (PGL-3D), a predict--select--refine--re-predict framework that uses intermediate cuboids to guide the aggregation of search evidence and update query and proposal representations. A shared head first predicts a complete cuboid for every proposal. Query--Tube--Memory (QTM) then selects reference observations by combining proposal association, cuboid quality, frame response, and target absence, since association confidence alone establishes neither target presence nor geometric accuracy. The center, size, and orientation of each selected cuboid define soft pooling weights over query-conditioned proposal features. The pooled memory updates the query and proposal representations, and the head re-predicts from the updated features. A training-only objective, ST-D9O, supervises cuboid geometry at every stage by adding boundary, signed-distance, and soft-overlap terms to parameter regression. PGL-3D achieves a mean stAP of $0.270 \pm 0.004$ on 3DVQL, compared with $0.044$ reported for LaF. Ablations support the benefits of geometry-guided feature updates, while stage-wise analyses show improved cuboid accuracy. Replacing the geometry objective in our PROT3D reproduction with ST-D9O improves mAO on GSOT3D from $21.63\%$ to $25.78\%$. Our code and models will be released.
Problem

Research questions and friction points this paper is trying to address.

3D Visual Query Localization
Progressive Geometric Learning
Cuboid Prediction
Feature Refinement
9-DoF Cuboid
Innovation

Methods, ideas, or system contributions that make the work stand out.

Progressive Geometric Learning
3D Visual Query Localization
Query-Tube-Memory
ST-D9O Objective
Predict-Select-Refine
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Liang Peng
School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University
S
Shizhuo Mu
School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University
B
Bohan Tan
School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University
W
Wenyuan Wang
School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University
C
Chen Zhao
School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University
X
Xingping Dong
School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University
Heng Fan
Heng Fan
Assistant Professor, University of North Texas
Computer VisionMachine LearningArtificial Intelligence
Libo Zhang
Libo Zhang
Institute of Software, Chinese Academy of Sciences
Computer VisionMachine LearningDeep Learning
Bo Du
Bo Du
Department of Management, Griffith Business School
Sustainable TransportTravel BehaviourUrban Data AnalyticsLogistics and Supply Chain