🤖 AI Summary
This study addresses the challenges of scarce annotations, minute defects, and real-time edge deployment in X-ray inspection for metal additive manufacturing. We propose a defect-conditioned LoRA Mixture-of-Experts architecture built upon a frozen SAM backbone, which achieves prompt-free segmentation by fine-tuning only a small number of parameters on synthetic data. Furthermore, we overcome ViT-B memory bottlenecks on NPUs through numerically equivalent attention rewriting, enabling fully on-device inference. Experimental results demonstrate that our method comprehensively outperforms baselines, achieving a 64.2% IoU on real porosity defects. With FP16 quantization, it processes 1024×1024 images on the Qualcomm Hexagon NPU with less than 0.01% accuracy loss, effectively balancing computational efficiency and high precision.
📝 Abstract
Metal additive manufacturing parts are inspected by X-ray computed tomography, where labelled data is scarce, the pores and inclusions that matter span a few pixels, and inspection must happen at the machine. We present DCM-SAM, a defect-conditioned adaptive mixture of LoRA experts: one frozen Segment Anything backbone carries a separate Conv-LoRA expert bank and mask decoder per defect class, each trained in its own pass, without prompts, on synthetic slices alone, updating only 4.4% of the parameters. On benchmarks that XCT-SAM reports, DCM-SAM improves on every baseline for both classes from a ViT-B backbone against their ViT-H, and reaches 64.2% pore IoU on real NIST scans having seen no real images during training. Deployment then exposes what adaptation work rarely measures: on a Qualcomm Hexagon NPU, ViT-H and ViT-L compile yet cannot allocate at 1024x1024 image resolution, since activations rather than weights exceed the device ceiling, and quantizing weights does not help. ViT-B alone runs, but the adapted encoder then fails to allocate where the stock one succeeds, until a numerically identical rewrite of the attention lets the complete DCM-SAM run in FP16 at 1024x1024, with no operator falling back to the CPU, masks within 0.01% of pixels of the FP32 reference. Code: https://github.com/MushfiqShovon/DCM-SAM.