🤖 AI Summary
This study addresses the challenge of word segmentation in Bengali handwritten text captured via mobile devices, where performance is often degraded by paper-ink color variations, complex illumination, and shadow interference. To overcome these limitations, this work proposes a robust automatic segmentation method that transcends conventional color-dependent constraints. A dedicated dataset encompassing multi-colored backgrounds and complex lighting conditions is constructed, and an integrated pipeline combining adaptive thresholding, morphological dilation filtering, and bounding box detection is developed to achieve resilient segmentation. Experimental evaluations on 7,374 test words demonstrate that the proposed approach attains a recall of 90.60%, a precision of 91.80%, and an F1-score of 91.20%. These results indicate a substantial improvement in word segmentation performance for mobile optical character recognition systems operating under unconstrained imaging conditions.
📝 Abstract
An optical character recognition(OCR) system can scan paper and extract text, making people's jobs easier. While numerous OCR systems are accessible in the software sector, finding a dependable equivalent solution for Bangla is tough. When it comes to handwritten texts, the case is even more rare. The first fundamental step to any OCR is to segment words from text images. If this stage fails, the total OCR's performance will be poor no matter how promising the later stages perform. This research aims to segment words in a handwritten Bangla text image. This research can be implemented on any smartphone-captured image, irrespective of the color and type of paper and ink. Furthermore, as smartphone-captured images can create shadow interferences, the custom dataset built for this research is created in such a way that every possible obstacle that can be faced is included. For 7374 words, a total of 7278 bounding boxes are generated, which have recall of 90.60 %, precision of 91.80 %, and F1-score of 91.20 %. The system can be further improved with nested operations on bounding boxes containing several words or by adjusting the adaptive thresholding and dilation filter sizes to a more precise level.