Point Ladder Tuning: Parameter-Efficient Hierarchical Adaptation for 3D Point Cloud Understanding

📅 2026-07-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing parameter-efficient fine-tuning methods struggle to recover the multi-scale local geometric information lost during downsampling in 3D point cloud backbone networks, thereby limiting dense prediction performance. This work proposes a local-aware parameter-efficient fine-tuning framework that, while keeping the backbone frozen, introduces for the first time a hierarchical local feature reconstruction mechanism. This mechanism leverages a multi-resolution local feature pyramid, a local-global semantic fusion module, and a dynamic multi-scale prompt generator to restore fine-grained geometric structures, coupled with a lightweight upsampling segmentation head for efficient adaptation. The approach achieves state-of-the-art performance with only 2.71% (for classification) and 7.69% (for dense prediction) trainable parameters, and scales effectively on PointGPT-L with merely 0.36% additional parameters.
📝 Abstract
Fine-tuning pre-trained point-cloud backbones typically updates all parameters, resulting in substantial computation and memory overhead. More importantly, modern point backbones rely on aggressive tokenization and downsampling, which yields compact global tokens but irreversibly discards fine-grained local geometry, an inherent bottleneck for parameter-efficient adaptation. Consequently, existing PEFT methods that operate only on these coarsened tokens can modulate global semantics but struggle to recover the missing multi-scale locality. We present Point Ladder Tuning (PLT), a locality-aware PEFT framework that performs hierarchical, instance-conditioned adaptation while keeping the backbone frozen. PLT forms a lightweight closed loop: (i) a Hierarchical Ladder Network (HLN) constructs a multi-resolution local feature pyramid directly from raw points; (ii) a Local-Global Fusion (LGF) aligns and fuses local pyramids with intermediate backbone semantics; and (iii) a Dynamic Prompt Generator produces instance-aware multi-scale prompts to modulate the frozen backbone effectively. For dense prediction, we further introduce a lightweight segmentation head that progressively upsamples fused features and leverages backbone priors to refine fine structures. Extensive experiments on classification and dense prediction show that PLT consistently surpasses prior PEFT baselines with minimal tunable parameters. PLT achieves state-of-the-art performance using only 2.71% trainable parameters for classification and 7.69% for dense prediction, and scales favorably to larger backbones, requiring merely 0.36% parameters on PointGPT-L. The code is released at https://github.com/JunLinChang/ECCV2026-PLT.
Problem

Research questions and friction points this paper is trying to address.

point cloud understanding
parameter-efficient fine-tuning
local geometry recovery
multi-scale locality
3D representation learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parameter-Efficient Fine-Tuning
3D Point Cloud Understanding
Hierarchical Adaptation
Local-Global Fusion
Dynamic Prompting
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Junlin Chang
Beihang University, Beijing; Pengcheng Laboratory, Shenzhen
L
Longhao Zou
Pengcheng Laboratory, Shenzhen
R
Rui Li
Pengcheng Laboratory, Shenzhen