Score
Designs and implements masked autoencoder models, masking strategies, encoders/decoders, and training pipelines for 3D point-cloud data that learn to reconstruct masked portions or predict geometric evolution. Builds loss functions, data augmentations, and evaluation metrics to produce self-supervised spatiotemporal 3D feature representations that internalize shape and dynamic structure.
This work addresses the limitation of conventional random masking in point cloud self-supervised learning, which disregards intrinsic geometric structure and thus hinders representation learning. To this end, we propose GeoMask3D—a geometry-aware masking strategy. Its core innovations are twofold: (1) the first learnable geometric complexity metric, which guides the selection of mask regions based on structural intricacy; and (2) a feature-level full-partial knowledge distillation mechanism within a teacher-student framework, enabling context-guided geometric complexity prediction. GeoMask3D is seamlessly integrated into the masked autoencoder (MAE) paradigm, substantially enhancing the model’s capacity to capture fine-grained local geometry. Extensive experiments on benchmarks including ModelNet40 demonstrate state-of-the-art performance in both classification and few-shot recognition tasks, validating that geometry-aware masking delivers substantial and consistent gains for downstream performance.
Existing 3D multimodal masked autoencoders (MAEs) require joint input of 2D images and 3D point clouds; however, point clouds inherently encode multi-view geometric information, making explicit 2D supervision not only inefficient but also detrimental to pure 3D geometric representation learning. To address this, we propose 3D-to-Multi-View MAE: a self-supervised framework that takes only masked point clouds as input, encodes them via a point cloud Transformer, and employs a cross-modal decoder to jointly reconstruct the original point cloud and multi-pose depth maps—achieving, for the first time, end-to-end 3D-only encoding for multi-view depth rendering. Our method deeply integrates geometric structure understanding with cross-view semantic alignment, eliminating reliance on 2D imagery. Extensive experiments demonstrate state-of-the-art performance across downstream tasks—including classification, few-shot learning, part segmentation, and object detection—with consistent gains exceeding 5% on multiple metrics.
Existing research on masked image modeling (MIM) for self-supervised visual representation learning lacks a unified formalization of pretraining paradigms and standardized, comparable evaluation. Method: We formally categorize MIM into two principal paradigms—reconstruction-based and contrastive-based—and construct an interpretable, hierarchical taxonomy via expert curation and agglomerative clustering. We conduct systematic, unified benchmarking of over 20 state-of-the-art models on ImageNet and other standard datasets. Contribution/Results: Our analysis identifies critical open challenges—including cross-paradigm integration, long-tailed masking strategies, and compute-accuracy trade-offs. To foster reproducibility and standardization, we publicly release a structured literature repository and an extensible evaluation framework on GitHub. This work establishes foundational infrastructure for rigorous, comparable advancement in MIM research.
Existing MAE-based pretraining methods, designed for ViT architectures, struggle to capture the critical geometric structures and spatial relationships inherent in medical images, thereby limiting 3D segmentation performance. To address this, we propose a topology- and spatially-aware self-supervised pretraining framework: (1) a novel topological signature-based loss function that explicitly preserves anatomical structural integrity; (2) two new auxiliary tasks—3D cropping center localization and octant-point regression—to enhance spatial position understanding; and (3) joint pretraining of a ViT backbone with state-of-the-art segmentation networks in a hybrid architecture. Evaluated on five public 3D medical segmentation benchmarks, our method achieves average Dice score improvements of 2.1–4.7 percentage points over prior MAE approaches, with significantly enhanced generalization and robustness.
To address poor robustness in regional representation caused by strong noise and sparse labels in urban spatiotemporal graph data, this paper introduces the Spatiotemporal Heterogeneous Graph Neural Encoder (ST-HGAE), the first application of masked autoencoding to spatiotemporal graph learning. ST-HGAE jointly masks node features and graph structure, enabling generative self-supervised learning to automatically distill dynamic spatiotemporal dependencies. Its core innovations include a structure-aware masking strategy tailored for heterogeneous spatiotemporal graphs, and a dual reconstruction objective integrating node-level feature recovery with topology reconstruction. Evaluated on traffic flow, pedestrian flow, and crime prediction tasks, ST-HGAE consistently outperforms state-of-the-art methods—particularly under high noise levels and low label rates—while significantly enhancing modeling of dynamic spatial correlations among regions.
This work addresses the limitation of existing 3D Masked Autoencoders (MAEs), which overly rely on positional information during spatial coordinate reconstruction, thereby compromising semantic representation learning. To mitigate this “positional leakage” issue, the authors propose the MPL-MAE framework, featuring two key innovations: a novel positional embedding module that suppresses metric-dominant signals while preserving geometric topology, and a gated positional interface that dynamically modulates the injection of positional cues during reconstruction. By effectively balancing the interplay between spatial priors and semantic features, the proposed method significantly enhances the robustness and informativeness of learned representations across diverse downstream tasks, demonstrating its superiority over prior approaches.
Existing rotation-invariant point cloud masked autoencoders employ random masking, neglecting geometric structure and semantic coherence, thereby failing to model spatial relationships consistent across orientations or rotation-robust semantic parts. To address this, we propose a dual-stream masking strategy: (1) a 3D spatial grid masking based on coordinate sorting, explicitly preserving local geometric structure; and (2) an attention-driven clustering-based semantic masking that focuses on identity-stable semantic regions under arbitrary rotations. These two masking schemes are dynamically weighted via curriculum learning, enabling progressive, geometry-to-semantics cooperative training. Our method is plug-and-play—requiring no backbone modification. Extensive experiments on ModelNet40, ScanObjectNN, and OmniObject3D demonstrate significant improvements over baselines, achieving state-of-the-art performance under diverse rotation settings.
This work addresses the scarcity of annotated parametric CAD data and the lack of effective self-supervised learning methods by proposing Masked Topological Modeling (MTM), a self-supervised pretraining approach tailored for native Boundary Representations (B-Reps). MTM leverages the face adjacency graph structure inherent in B-Reps for the first time, jointly optimizing masked edge geometry reconstruction with contrastive learning through a region-based masking objective. It introduces a BFS-connected-region masking strategy and CAD-aware data augmentation, integrated within a graph neural network framework enhanced by momentum contrastive learning. Extensive experiments demonstrate that MTM significantly outperforms existing methods across multiple CAD understanding benchmarks, confirming its effectiveness and strong generalization capability in low-label regimes.
为解决图像与点云配准中的错误对应问题,提出基于相似性和强化学习的双掩码自编码器框架ID-MAE,增强跨模态特征表示一致性。
This work addresses the limited representational capacity of point cloud learning by proposing a cross-modal representation learning framework that incorporates 2D visual priors. Methodologically, it innovatively leverages pre-trained 2D vision models to guide 3D point cloud network training via feature alignment—enabling effective 2D→3D knowledge transfer without naïve modality conversion. The framework integrates a point cloud encoder-decoder architecture, self-supervised pre-training, and primitive-level supervised segmentation into a unified learning paradigm. Extensive experiments on ScanNet and S3DIS benchmarks demonstrate substantial improvements in point cloud segmentation and scene understanding performance, validating the efficacy of 2D semantic priors in enhancing 3D representation learning. The proposed approach establishes a scalable, multimodal representation learning pathway for efficient 3D perception in autonomous driving and robotics applications.