ThyCLIPNet: A BiomedCLIP-Guided Lightweight Attention-Enhanced DeepLabV3+ Framework for Robust Thyroid Nodule Segmentation

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of low contrast, noise interference, high computational overhead, and the lack of global semantic guidance in thyroid ultrasound image segmentation by proposing ThyCLIPNet, a lightweight segmentation framework. This work pioneers the integration of the BiomedCLIP vision encoder into a multi-scale CNN pipeline, providing image-only global semantic guidance without text prompts and enabling efficient fusion of semantic and local features via a gating mechanism. The decoder constructs an enhanced DeepLabV3+ architecture incorporating MobileNetV2, ASPP, CBAM, and hierarchical skip connections. Evaluated on four datasets, ThyCLIPNet achieves Dice coefficients ranging from 80.7% to 96.2% with only 8.55M parameters, significantly improving both segmentation accuracy and computational efficiency.
📝 Abstract
Accurate thyroid ultrasound segmentation is often challenged by low contrast, speckle noise, and unclear boundaries. Although recent methods have improved segmentation accuracy, many rely on resource-intensive architectures or lack explicit integration of multiscale features with global biomedical visual guidance. In this paper, we introduce ThyCLIPNet, a lightweight semantic-guided hybrid encoder-decoder framework that integrates BiomedCLIP-derived biomedical semantic guidance into a lightweight multi-scale CNN segmentation pipeline. The encoder integrates MobileNetV2 with efficient channel attention, while atrous spatial pyramid pooling and a custom convolutional block attention module enrich bottleneck features. The decoder combines hierarchical skip connections and lightweight attention refinement with a BiomedCLIP-guided gated fusion pathway that projects vision-only global biomedical embeddings into decoder feature space and selectively integrates them through semantic-local fusion and spatial gating. To the best of our knowledge, ThyCLIPNet is among the first lightweight thyroid ultrasound segmentation frameworks to use BiomedCLIP's vision encoder alone for image-only global semantic guidance without text prompting. Experiments on TG3K, TN3K, DDTI, and PKTN achieve dice similarity coefficients of 96.22%, 87.58%, 84.73%, and 80.70%; intersection over union scores of 92.72%, 77.91%, 73.51%, and 67.64%; and 95th-percentile hausdorff distances of 3.75, 16.38, 18.23, and 10.86, respectively. ThyCLIPNet uses 8.55M parameters and 22.99G FLOPs. Overall, the results support integrating global biomedical semantic guidance with lightweight multi-scale CNN representations for robust and computationally efficient thyroid ultrasound segmentation. Source code: https://github.com/Tasnim-Jahan/ThyCLIPNet. [Abstract shortened for arXiv. See PDF for full abstract.]
Problem

Research questions and friction points this paper is trying to address.

Thyroid nodule segmentation
Ultrasound image segmentation
Lightweight model
Biomedical semantic guidance
Multiscale feature integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

BiomedCLIP
Lightweight Segmentation
Attention Mechanism
DeepLabV3+
Thyroid Nodule
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Tasnim Jahan
Department of Computer Science and Engineering, United International University, United City, Madani Ave, Dhaka 1212, Bangladesh
M
Md Easin Arafat
Faculty of Informatics, Institute of Industry-Academia Innovation, Department of Data Science and Engineering, Eötvös Loránd University, Pázmány Péter Sétány 1/C, Budapest 1117, Hungary
Swakkhar Shatabda
Swakkhar Shatabda
Professor, School of Data and Sciences, BRAC University
optimizationmachine learningcomputational biologybioinformatics