Automated document processing system for government agencies using DBNET++ and BART models

📅 2025-10-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address real-world challenges in government document processing—including illumination variation, text distortion, occlusion, and low resolution—this paper proposes an end-to-end automatic document classification method. The approach integrates DBNet++ for robust text detection with the BART model for semantic understanding, supporting both offline images and real-time camera input. A lightweight image preprocessing module enhances system robustness, while extracted textual content enables accurate classification into four categories: invoices, reports, letters, and tables. A cross-platform interactive interface is implemented using PyQt5. Trained on the Total-Text dataset for 10 hours, the text detector achieves 92.88% accuracy. This work constitutes the first integration of DBNet++ and BART for semantic document-type classification, demonstrating strong generalization across heterogeneous input sources and challenging imaging conditions, thereby offering practical value for real-world deployment.

Technology Category

Natural Language Processing: Text Classification & Sentiment AnalysisComputer Vision: Object Detection & CategorizationMachine Learning: Multi-class/Multi-label Learning & Extreme Classification

Application Category

Web Mining and Content Analysis: Normalization, clustering, classification, and summarization of Web textSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
An automatic document classification system is presented that detects textual content in images and classifies documents into four predefined categories (Invoice, Report, Letter, and Form). The system supports both offline images (e.g., files on flash drives, HDDs, microSD) and real-time capture via connected cameras, and is designed to mitigate practical challenges such as variable illumination, arbitrary orientation, curved or partially occluded text, low resolution, and distant text. The pipeline comprises four stages: image capture and preprocessing, text detection [1] using a DBNet++ (Differentiable Binarization Network Plus) detector, and text classification [2] using a BART (Bidirectional and Auto-Regressive Transformers) classifier, all integrated within a user interface implemented in Python with PyQt5. The achieved results by the system for text detection in images were good at about 92.88% through 10 hours on Total-Text dataset that involve high resolution images simulate a various and very difficult challenges. The results indicate the proposed approach is effective for practical, mixed-source document categorization in unconstrained imaging scenarios.
Problem

Research questions and friction points this paper is trying to address.

Automated document classification system for government agencies
Detects text in images under challenging conditions
Classifies documents into four predefined categories
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses DBNet++ for robust text detection
Employs BART model for document classification
Integrates PyQt5 interface for mixed-source processing
💼 Related Jobs
No related jobs found.
Informatics Institute for Postgraduate Studies
A
Aya Kaysan Bahjat
Informatics Institute for Postgraduate Studies, Baghdad, Iraq