🤖 AI Summary
This study investigates whether commonly used evaluation metrics in multilevel image thresholding—namely Structural Similarity Index (SSIM) and Peak Signal-to-Noise Ratio (PSNR)—exhibit implicit preferences toward classical objective functions such as Otsu’s method and Kapur’s entropy. By exhaustively enumerating the full threshold space on the BSDS500 dataset, the work systematically analyzes the correlation between objective functions and evaluation metrics. Empirical results reveal, for the first time, that SSIM and PSNR demonstrate strong and consistent positive correlations with Otsu’s criterion, whereas their correlations with Kapur’s entropy are weak and unstable. Notably, Otsu outperforms Kapur in PSNR correlation across all images and achieves superior SSIM correlation in over 91% of cases. These findings challenge the widely held assumption of metric neutrality, exposing a systematic bias in current evaluation practices.
📝 Abstract
Multilevel image thresholding is widely used for segmentation in applications ranging from medical imaging to remote sensing. Classical objective functions, such as Otsu's between-class variance and Kapur's entropy, are often optimized using metaheuristic algorithms, with performance evaluated via metrics like Structural Similarity Index (SSIM) and Peak Signal-to-Noise Ratio (PSNR). These evaluations implicitly assume that SSIM and PSNR provide unbiased measures of segmentation quality. In this study, we examine this assumption by analyzing the correlation between thresholding objective functions and quality metrics across all possible thresholds for images in the BSDS500 dataset. Results show that Otsu's criterion consistently exhibits high correlation with both SSIM and PSNR, while Kapur's entropy demonstrates weaker and more variable correlation. Otsu outperforms Kapur in correlation with PSNR for all images and with SSIM for over 91%. Our findings reveal an inherent metric-objective-function bias. This work highlights the need for more neutral evaluation frameworks and motivates extending the analysis to additional thresholding criteria and domains. Source code of this paper can be found at https://w3id.org/met-dp/icpr26-95