🤖 AI Summary
This study addresses the systematic misremoval of long-tail, low-posture workers, such as those in squatting positions, during robot training for construction sites due to quality filtering. We audit the bias inherent in data filtering pipelines toward low-posture samples, revealing their scarcity and associated detection bottlenecks. Specifically, this work proposes a multi-dimensional postural cue auditing method integrating NLF joint estimation, pose clustering, simulated scenario testing, and multimodal feature analysis. Our analysis demonstrates that conventional bounding boxes and pose estimation fail to reliably preserve long-tail data. Results indicate that low-posture samples constitute merely 2% of the dataset, and existing filters significantly exacerbate their attrition. These findings expose critical limitations in current data curation strategies and provide essential evidence for enhancing visual robustness under long-tailed distributions.
📝 Abstract
Robots on construction sites must detect workers who are kneeling or bending, which we call low poses. These workers can be lost from training datasets during automatic labeling. We study a pipeline that detects people, estimates their body joints using NLF, and groups similar poses. Low poses account for only about 2 percent of the retained examples. This low share may partly reflect the pipeline's quality filter, which rejects examples with low detection confidence or uncertain joint estimates. We examine this filtering using four alternative pose clues: bounding-box shape, vertical body span, pose grouping aligned to the scene's vertical direction, and image appearance. All four suggest that low poses are rejected by the filter more often. Separately, controlled simulated scenes show that a person detector fine-tuned on a public construction dataset misses more workers in these poses even when we correct their bounding box height is matched to that of standing workers. These findings suggest that low poses are scarce and hard to find, and we cannot rely on bounding boxes or poses for long-tail human data curation.