🤖 AI Summary
This study addresses the privacy leakage and accuracy bottlenecks arising from individual data association in multimodal 3D human pose estimation by proposing a unified privacy-preserving framework that integrates kinematic sensors. Methodologically, we design a black-box membership inference attack and point-level maximum leakage analysis to quantify privacy risks, alongside an action-temporal hierarchical differential privacy strategy. Furthermore, private training is achieved through multimodal feature alignment, skeletal structure injection, and an adaptive aggregation mechanism. Multi-protocol experiments on the MM-Fi dataset demonstrate that the proposed framework effectively maintains model accuracy while ensuring subject-level privacy.
📝 Abstract
Multimodal 3D Human Pose Estimation (3D HPE) combines complementary information from RGB, LiDAR, and mmWave radar, but models trained on correlated observations from the same individuals, raise privacy risks overlooked by record level analysis. We present a unified framework for multimodal 3D HPE that couples kinematics-induced sensor fusion with subject level privacy auditing and private training. First, our multimodal model aligns modality specific joint representation, injects skeletal structure and adaptively aggregates complementary sensor evidence for accurate pose prediction. Second, we formulate a black-box subject membership inference attack for 3D HPE, complemented by an empirical pointwise maximal leakage analysis, which characterizes how individual attack score outcomes change inference about the membership outcome. Third, we instantiate user-level differential privacy via Action Temporal Stratification, a population weighted within-subject sampling strategy that enforces action and temporal coverage. We evaluate our framework on the MM-Fi dataset across three diverse experimental protocols. Source-code will be released upon acceptance.