point-cloud masked autoencoding

Designs and implements masked autoencoder models, masking strategies, encoders/decoders, and training pipelines for 3D point-cloud data that learn to reconstruct masked portions or predict geometric evolution. Builds loss functions, data augmentations, and evaluation metrics to produce self-supervised spatiotemporal 3D feature representations that internalize shape and dynamic structure.

point-cloudmaskedautoencoding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.43
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

GeoMask3D: Geometrically Informed Mask Selection for Self-Supervised Point Cloud Learning in 3D

May 20, 2024
AB
Ali Bahri
🏛️ École de Technologie Supérieure

This work addresses the limitation of conventional random masking in point cloud self-supervised learning, which disregards intrinsic geometric structure and thus hinders representation learning. To this end, we propose GeoMask3D—a geometry-aware masking strategy. Its core innovations are twofold: (1) the first learnable geometric complexity metric, which guides the selection of mask regions based on structural intricacy; and (2) a feature-level full-partial knowledge distillation mechanism within a teacher-student framework, enabling context-guided geometric complexity prediction. GeoMask3D is seamlessly integrated into the masked autoencoder (MAE) paradigm, substantially enhancing the model’s capacity to capture fine-grained local geometry. Extensive experiments on benchmarks including ModelNet40 demonstrate state-of-the-art performance in both classification and few-shot recognition tasks, validating that geometry-aware masking delivers substantial and consistent gains for downstream performance.

Enhances self-supervised learning for 3D point cloudsFocuses on geometrically complex areas for robust featuresImproves performance in classification and few-shot tasks

Point Cloud Self-supervised Learning via 3D to Multi-view Masked Autoencoder

Nov 17, 2023
ZC
Zhimin Chen
🏛️ Clemson University | Johns Hopkins University | The City University of New York

Existing 3D multimodal masked autoencoders (MAEs) require joint input of 2D images and 3D point clouds; however, point clouds inherently encode multi-view geometric information, making explicit 2D supervision not only inefficient but also detrimental to pure 3D geometric representation learning. To address this, we propose 3D-to-Multi-View MAE: a self-supervised framework that takes only masked point clouds as input, encodes them via a point cloud Transformer, and employs a cross-modal decoder to jointly reconstruct the original point cloud and multi-pose depth maps—achieving, for the first time, end-to-end 3D-only encoding for multi-view depth rendering. Our method deeply integrates geometric structure understanding with cross-view semantic alignment, eliminating reliance on 2D imagery. Extensive experiments demonstrate state-of-the-art performance across downstream tasks—including classification, few-shot learning, part segmentation, and object detection—with consistent gains exceeding 5% on multiple metrics.

Inefficient use of 2D and 3D modalities in self-supervised learningNeed for better 3D geometric representation learningReconstruction learning overly relies on visible 2D information

Masked Image Modeling: A Survey

Aug 13, 2024
VH
Vlad Hondru
🏛️ University of Bucharest | Amazon | University of Trento

Existing research on masked image modeling (MIM) for self-supervised visual representation learning lacks a unified formalization of pretraining paradigms and standardized, comparable evaluation. Method: We formally categorize MIM into two principal paradigms—reconstruction-based and contrastive-based—and construct an interpretable, hierarchical taxonomy via expert curation and agglomerative clustering. We conduct systematic, unified benchmarking of over 20 state-of-the-art models on ImageNet and other standard datasets. Contribution/Results: Our analysis identifies critical open challenges—including cross-paradigm integration, long-tailed masking strategies, and compute-accuracy trade-offs. To foster reproducibility and standardization, we publicly release a structured literature repository and an extensible evaluation framework on GitHub. This work establishes foundational infrastructure for rigorous, comparable advancement in MIM research.

Automatic LearningComputer VisionMasked Image Modeling

Self Pre-training with Topology- and Spatiality-aware Masked Autoencoders for 3D Medical Image Segmentation

Jun 15, 2024
PG
Pengfei Gu
🏛️ University of Texas Rio Grande Valley | University of Notre Dame

Existing MAE-based pretraining methods, designed for ViT architectures, struggle to capture the critical geometric structures and spatial relationships inherent in medical images, thereby limiting 3D segmentation performance. To address this, we propose a topology- and spatially-aware self-supervised pretraining framework: (1) a novel topological signature-based loss function that explicitly preserves anatomical structural integrity; (2) two new auxiliary tasks—3D cropping center localization and octant-point regression—to enhance spatial position understanding; and (3) joint pretraining of a ViT backbone with state-of-the-art segmentation networks in a hybrid architecture. Evaluated on five public 3D medical segmentation benchmarks, our method achieves average Dice score improvements of 2.1–4.7 percentage points over prior MAE approaches, with significantly enhanced generalization and robustness.

Aggregating spatial relationships for medical image segmentationCapturing geometric shape information in 3D medical imagesEnhancing masked autoencoders for hybrid segmentation architectures

Graph Masked Autoencoder for Spatio-Temporal Graph Learning

Oct 14, 2024
QZ
Qianru Zhang
🏛️ University of Hong Kong | University of California, Los Angeles | University of Queensland

To address poor robustness in regional representation caused by strong noise and sparse labels in urban spatiotemporal graph data, this paper introduces the Spatiotemporal Heterogeneous Graph Neural Encoder (ST-HGAE), the first application of masked autoencoding to spatiotemporal graph learning. ST-HGAE jointly masks node features and graph structure, enabling generative self-supervised learning to automatically distill dynamic spatiotemporal dependencies. Its core innovations include a structure-aware masking strategy tailored for heterogeneous spatiotemporal graphs, and a dual reconstruction objective integrating node-level feature recovery with topology reconstruction. Evaluated on traffic flow, pedestrian flow, and crime prediction tasks, ST-HGAE consistently outperforms state-of-the-art methods—particularly under high noise levels and low label rates—while significantly enhancing modeling of dynamic spatial correlations among regions.

Handling real-world urban data noise and sparsity challengesLearning meaningful region representations in spatial-temporal graphsOvercoming noisy sparse spatial-temporal urban data limitations

Latest Papers

What's happening recently
View more

This work addresses the limitation of existing 3D Masked Autoencoders (MAEs), which overly rely on positional information during spatial coordinate reconstruction, thereby compromising semantic representation learning. To mitigate this “positional leakage” issue, the authors propose the MPL-MAE framework, featuring two key innovations: a novel positional embedding module that suppresses metric-dominant signals while preserving geometric topology, and a gated positional interface that dynamically modulates the injection of positional cues during reconstruction. By effectively balancing the interplay between spatial priors and semantic features, the proposed method significantly enhances the robustness and informativeness of learned representations across diverse downstream tasks, demonstrating its superiority over prior approaches.

3D masked autoencoderspoint cloudpositional leakage

Existing rotation-invariant point cloud masked autoencoders employ random masking, neglecting geometric structure and semantic coherence, thereby failing to model spatial relationships consistent across orientations or rotation-robust semantic parts. To address this, we propose a dual-stream masking strategy: (1) a 3D spatial grid masking based on coordinate sorting, explicitly preserving local geometric structure; and (2) an attention-driven clustering-based semantic masking that focuses on identity-stable semantic regions under arbitrary rotations. These two masking schemes are dynamically weighted via curriculum learning, enabling progressive, geometry-to-semantics cooperative training. Our method is plug-and-play—requiring no backbone modification. Extensive experiments on ModelNet40, ScanObjectNN, and OmniObject3D demonstrate significant improvements over baselines, achieving state-of-the-art performance under diverse rotation settings.

Ensures compatibility with existing frameworks without architectural modificationsOvercomes random masking's neglect of geometric structure and semantic coherenceProposes dual-stream approach for rotation-invariant point cloud representation learning

This work addresses the scarcity of annotated parametric CAD data and the lack of effective self-supervised learning methods by proposing Masked Topological Modeling (MTM), a self-supervised pretraining approach tailored for native Boundary Representations (B-Reps). MTM leverages the face adjacency graph structure inherent in B-Reps for the first time, jointly optimizing masked edge geometry reconstruction with contrastive learning through a region-based masking objective. It introduces a BFS-connected-region masking strategy and CAD-aware data augmentation, integrated within a graph neural network framework enhanced by momentum contrastive learning. Extensive experiments demonstrate that MTM significantly outperforms existing methods across multiple CAD understanding benchmarks, confirming its effectiveness and strong generalization capability in low-label regimes.

B-Repboundary representationdata-efficient learning

Representation Learning for Point Cloud Understanding

Dec 05, 2025
SY
Siming Yan
🏛️ The University of Texas at Austin

This work addresses the limited representational capacity of point cloud learning by proposing a cross-modal representation learning framework that incorporates 2D visual priors. Methodologically, it innovatively leverages pre-trained 2D vision models to guide 3D point cloud network training via feature alignment—enabling effective 2D→3D knowledge transfer without naïve modality conversion. The framework integrates a point cloud encoder-decoder architecture, self-supervised pre-training, and primitive-level supervised segmentation into a unified learning paradigm. Extensive experiments on ScanNet and S3DIS benchmarks demonstrate substantial improvements in point cloud segmentation and scene understanding performance, validating the efficacy of 2D semantic priors in enhancing 3D representation learning. The proposed approach establishes a scalable, multimodal representation learning pathway for efficient 3D perception in autonomous driving and robotics applications.

Self-supervised learning methods for point cloud understandingSupervised learning for point cloud primitive segmentationTransfer learning from 2D to 3D using pre-trained models

Hot Scholars

XH

Xiaoshui Huang

Shanghai Jiao Tong University
Artificial intelligenceComputer visionAI for healthcare
GM

Guofeng Mei

Fondazione Bruno Kessler, University of Technology Sydney (Ph.D.), Wuhan University
Artificial Intelligence(NLPRecommendationComputer Vision)Complex network
ZY

Zhifei Yang

Peking University
3D GenerationGenerative Models