Point-Focused Attention Meets Context-Scan State Space: Robust Biological Visual Perception for Point Cloud Representation

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of jointly modeling local geometry and global context in point cloud representation learning. To this end, we propose PointLearner, a novel bio-inspired "focus-context" architecture that mimics the foveal vision and saccadic mechanisms of biological visual systems. Methodologically, PointLearner achieves fine-grained local feature extraction through point-focused attention and learnable induced point pooling, while efficiently capturing long-range dependencies by integrating competitive normalization attention, Hilbert curve-guided scanning, and a bidirectional S6 state space model. Extensive experiments demonstrate that the proposed method exhibits exceptional robustness across multiple point cloud tasks, achieving state-of-the-art performance.
📝 Abstract
Synergistically capturing intricate local structures and global contextual dependencies has become a critical challenge in point cloud representation learning. To address this, we introduce PointLearner, a point cloud representation learning network that closely aligns with biological vision which employs an active, foveation-inspired processing strategy, thus enabling local geometric modeling and long-range dependency interactions simultaneously. Specifically, we first design a point-focused attention, which simulates foveal vision at the visual focus through a competitive normalized attention mechanism between local neighbors and spatially downsampled features. The spatially downsampled features are extracted by a pooling method based on learnable inducing points, which can flexibly adapt to the non-uniform distribution of point clouds as the number of inducing points is controlled and they interact directly with point clouds. Second, we propose a context-scan state space that mimics eye's saccade inference, which infers the overall semantic structure and spatial content in the scene through a scan path guided by the Hilbert curve for the bidirectional S6. With this focus-then-context biomimetic design, PointLearner demonstrates remarkable robustness and achieves state-of-the-art performance across multiple point cloud tasks.
Problem

Research questions and friction points this paper is trying to address.

Point Cloud Representation
Local Structure
Global Contextual Dependencies
Representation Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Point Cloud Representation
Foveation-inspired Attention
State Space Model
Biomimetic Vision
Hilbert Curve
K
Kanglin Qu
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics; Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education
Pan Gao
Pan Gao
Professor, Nanjing University of Aeronautics and Astronautics;
Image/Video/Point cloudsdeep learningMultimedia
Q
Qun Dai
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics; Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education
Y
Yuanhao Sun
School of Mathematical Sciences, Beijing University of Posts and Telecommunications