A Boundary-Metric Evaluation Protocol for Whiteboard Stroke Segmentation Under Extreme Imbalance

📅 2026-02-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the extreme class imbalance in whiteboard stroke binary segmentation, where foreground pixels constitute only 1.79% on average, rendering conventional region-based metrics inadequate for evaluating fine-stroke performance. To this end, the authors propose a comprehensive evaluation protocol that integrates region- and boundary-based metrics, fairness analysis across core and fine-stroke subsets, and robustness statistics from multi-run training. Using a DeepLabV3-MobileNetV3 architecture, five loss functions are systematically compared. Results show that overlap-aware losses (e.g., Dice+Focal) improve F1 by over 20 points (0.663 vs. 0.438) compared to cross-entropy, while Tversky loss significantly outperforms the Sauvola method in worst-case F1 (0.565 vs. 0.452). Increasing training resolution further boosts F1 by up to 12.7 points. This work is the first to incorporate non-parametric significance testing and worst-case evaluation, uncovering hidden trade-offs among loss functions in fine-grained segmentation tasks.

Technology Category

Computer Vision: SegmentationMachine Learning: Evaluation and AnalysisNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP Models

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methodsSystems and Infrastructure for Web, Mobile and WoT: Web performance, measurement, and characterization
📝 Abstract
The binary segmentation of whiteboard strokes is hindered by extreme class imbalance, caused by stroke pixels that constitute only $1.79%$ of the image on average, and in addition, the thin-stroke subset averages $1.14% \pm 0.41%$ in the foreground. Standard region metrics (F1, IoU) can mask thin-stroke failures because the vast majority of the background dominates the score. In contrast, adding boundary-aware metrics and a thin-subset equity analysis changes how loss functions rank and exposes hidden trade-offs. We contribute an evaluation protocol that jointly examines region metrics, boundary metrics (BF1, B-IoU), a core/thin-subset equity analysis, and per-image robustness statistics (median, IQR, worst-case) under seeded, multi-run training with non-parametric significance testing. Five losses -- cross-entropy, focal, Dice, Dice+focal, and Tversky -- are trained three times each on a DeepLabV3-MobileNetV3 model and evaluated on 12 held-out images split into core and thin subsets. Overlap-based losses improve F1 by more than 20 points over cross-entropy ($0.663$ vs $0.438$, $p < 0.001$). In addition, the boundary metrics confirm that the gain extends to the precision of the contour. Adaptive thresholding and Sauvola binarization at native resolution achieve a higher mean F1 ($0.787$ for Sauvola) but with substantially worse worst-case performance (F1 $= 0.452$ vs $0.565$ for Tversky), exposing a consistency-accuracy trade-off: classical baselines lead on mean F1 while the learned model delivers higher worst-case reliability. Doubling training resolution further increases F1 by 12.7 points.
Problem

Research questions and friction points this paper is trying to address.

class imbalance
whiteboard stroke segmentation
thin-stroke segmentation
evaluation metrics
binary segmentation
Innovation

Methods, ideas, or system contributions that make the work stand out.

boundary-aware metrics
class imbalance
thin-stroke segmentation
evaluation protocol
robustness analysis
N
Nicholas Korcynski
Rowan University