🤖 AI Summary
Current nuclear segmentation in breast cancer histopathological images is hindered by the absence of standardized benchmark datasets, impeding fair and reproducible method comparison. To address this, we introduce and publicly release the first high-quality, task-specific dataset: 100 high-resolution whole-slide image (WSI) patches from 25 patients, containing approximately 36,000 nuclei meticulously annotated by board-certified pathologists. We propose a patient-level fixed-ratio split (75% training / 25% testing) to prevent data leakage and ensure consistent, reproducible cross-method evaluation. All images adhere to uniform acquisition and annotation protocols, enabling rigorous boundary-precision assessment and robust generalization analysis. This dataset fills a critical gap in the field and establishes a community-standard benchmark for automated nuclear analysis in precision oncology.
📝 Abstract
The NuSeC dataset is created by selecting 4 images with the size of 1024*1024 pixels from the slides of each patient among 25 patients. Therefore, there are a total of 100 images in the NuSeC dataset. To carry out a consistent comparative analysis between the methods that will be developed using the NuSeC dataset by the researchers in the future, we divide the NuSeC dataset 75% as the training set and 25% as the testing set. In detail, an image is randomly selected from 4 images of each patient among 25 patients to build the testing set, and then the remaining images are reserved for the training set. While the training set includes 75 images with around 30000 nuclei structures, the testing set includes 25 images with around 6000 nuclei structures.