Localize Any Object in X-Ray Security Scans without Human Annotation

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of universal object localization in X-ray imaging, which is constrained by the scarcity of annotations and the limited transferability of RGB-based foundation models. To overcome these limitations, this work proposes LAO-X, a self-supervised framework that adapts SAM2 to the X-ray domain without manual annotation. Specifically, it introduces saliency-guided object mining and absorption-based physical domain synthesis to generate training data, coupled with an occlusion-controlled curriculum learning strategy for efficient fine-tuning. Evaluated across six benchmarks, the proposed method achieves 2%–23% mAP improvements in complex cluttered scenes, enabling high-precision, category-agnostic X-ray object localization with zero annotations.
📝 Abstract
Universal object localization in X-ray security inspection is critical for automated threat detection in safety-critical venues. However, unlike everyday RGB images that dominate web-scale visual data, X-ray scans exhibit distinct color patterns, ambiguous boundaries, and compositional structures caused by volumetric superposition. These gaps hinder the direct zero-shot transfer of dense perception foundation models trained on web-scale RGB data. Moreover, annotated X-ray data is scarce and requires expert labeling, limiting both the training of generalizable X-ray native models and the adaptation of RGB foundation models for X-ray data via fine-tuning. Given these challenges, the bright promise of highly generalizable perception models, enabled by data scaling laws in the RGB domain, remains largely out of reach for X-ray inspection. To this end, we introduce LAO-X, a self-supervised adaptation framework that Locates Any Object in X-ray scans using diverse synthesized image--annotation pairs with granularity-aware supervision. LAO-X first designs a saliency-guided X-ray object mining module to separate diverse object instances, which are then used for physics-guided synthesis in the absorbance domain. LAO-X further incorporates an occlusion-controlled curriculum strategy to fine-tune a Segment Anything Model 2 (SAM2) localizer, progressively adapting it to X-ray scans with increasing object counts and overlap levels. Experiments on six X-ray benchmarks show that LAO-X substantially improves category-agnostic localization, achieving 2\% to 23\% mAP gains over SAM2 and X-ray specific baselines in heavily cluttered scenarios, entirely without human-annotated labels.
Problem

Research questions and friction points this paper is trying to address.

X-ray security inspection
object localization
domain adaptation
foundation models
zero-shot transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-supervised adaptation
X-ray object localization
Physics-guided synthesis
Curriculum learning
Segment Anything Model 2