Integrating Local Detail and Global Context: A Dual-Input Multi-Task Learning Framework for Bone Tumor Diagnosis

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of single-view models in X-ray diagnosis of bone tumors caused by morphological heterogeneity, ambiguous boundaries, and structural overlap. We propose a multi-task learning framework built upon a dual-stream DenseNet121 architecture. The method employs YOLO for lesion detection and introduces a novel cross-modal bidirectional attention mechanism to fuse lesion-cropped regions with whole-image information. Combined with hierarchical multi-scale feature fusion, this approach enables joint optimization of segmentation and classification. Experimental results demonstrate that the proposed model achieves a Dice coefficient of 0.896, a classification F1-score of 0.928, and an AUC of 0.999 for osteosarcoma, significantly outperforming existing baseline methods.
📝 Abstract
Primary bone tumors are rare but clinically aggressive neoplasms whose diagnosis from radiographs is challenged by heterogeneous morphology, subtle lesion margins, and overlapping bone structures. To address the limitations of existing single-view models, we present a dual-input, multi-task learning framework that, to our knowledge, is the first to apply bidirectional cross-modal attention between a lesion crop and the full radiograph for joint segmentation and subtype classification. Using the multi-institutional Bone Tumor X-ray Radiograph Dataset (BTXRD, n=3,746), we employ a YOLO-based detector to generate regions of interest, which are paired with full images as inputs to a dual-stream DenseNet121 architecture. Features are integrated via a novel cross-modal attention fusion strategy, refined by Hierarchical Multi-scale Feature Fusion, effectively balancing fine-grained lesion detail with global anatomical context. Evaluated on a held-out patient-level test split, the model demonstrates superior performance over single-input baselines, achieving an overall Dice Similarity Coefficient of 0.896 and a macro-averaged classification F1-score of 0.928. Notably, the system exhibits exceptional sensitivity for malignant osteosarcoma (AUC 0.999), validating the potential of dual-stream context modeling to support radiologists in accurate, early decision-making.
Problem

Research questions and friction points this paper is trying to address.

Bone Tumor Diagnosis
X-ray Radiograph
Lesion Segmentation
Subtype Classification
Single-view Limitations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dual-input multi-task learning
Cross-modal attention
Hierarchical multi-scale feature fusion
Bone tumor diagnosis
Dual-stream architecture
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
S. M. Nasif Uddin
Department of Electrical and Electronic Engineering, Ahsanullah University of Science and Technology, Bangladesh
Rusab Sarmun
Rusab Sarmun
University of Dhaka
Artificial IntelligenceDeep LearningComputer VisionReinforcement LearningRobotics
M
Muhammad E. H. Chowdhury
Department of Electrical Engineering, Qatar University, Doha 2713, Doha, Qatar
A
Adam Mushtak
Department of Radiology, Hamad Medical Corporation, Doha, Qatar
I
Israa Al-Hashimi
Department of Radiology, Hamad Medical Corporation, Doha, Qatar
S
Sohaib Bassam Zoghoul
Department of Radiology, Hamad Medical Corporation, Doha, Qatar