ERPA: Efficient RPA Model Integrating OCR and LLMs for Intelligent Document Processing

📅 2024-11-13
🏛️ 2024 International Mobile, Intelligent, and Ubiquitous Computing Conference (MIUCC)
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the low OCR accuracy, poor layout adaptability, and inefficiency in large-scale document processing inherent in traditional RPA systems for immigration document handling, this paper proposes an LLM-augmented intelligent RPA framework. The method pioneers the integration of fine-tuned large language models (LLMs) across the entire RPA pipeline, synergistically combining OCR engines with rule-enhanced workflows and context-aware textual post-processing to enable fuzzy character correction, complex layout parsing, and end-to-end structured information extraction. Experiments demonstrate that ID data extraction time is reduced to an average of 9.94 seconds—up to 94% faster than UiPath and Automation Anywhere—while accuracy and cross-document robustness are significantly improved. This work establishes a scalable technical paradigm for automated understanding of high-noise, multi-layout government documents.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Intelligent Robots: Multimodal Perception & Sensor FusionNatural Language Processing: Information Extraction

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGEconomics, Online Markets and Human Computation: Architectures and workflows that use LLMs for crowd workSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
This paper presents ERPA, an innovative Robotic Process Automation (RPA) model designed to enhance ID data extraction and optimize Optical Character Recognition (OCR) tasks within immigration workflows. Traditional RPA solutions often face performance limitations when processing large volumes of documents, leading to inefficiencies. ERPA addresses these challenges by incorporating Large Language Models (LLMs) to improve the accuracy and clarity of extracted text, effectively handling ambiguous characters and complex structures. Benchmark comparisons with leading platforms like UiPath and Automation Anywhere demonstrate that ERPA significantly reduces processing times by up to 94%, completing ID data extraction in just 9.94 seconds. These findings highlight ERPA's potential to revolutionize document automation, offering a faster and more reliable alternative to current RPA solutions.
Problem

Research questions and friction points this paper is trying to address.

Document Processing Efficiency
Immigration File Information Extraction
Automation Tool Effectiveness
Innovation

Methods, ideas, or system contributions that make the work stand out.

ERPA
Advanced Language Models
Document Processing Efficiency
💼 Related Jobs
No related jobs found.
MSA University
O
Osama Hosam Abdellaif
Computer Science, MSA University, Cairo, Egypt
A
Abdelrahman Nader Hassan
Computer Science, MSA University, Cairo, Egypt
Ali Hamdi
Ali Hamdi
Computer Science, MSA University
Computer VisionDeep LearningText Mining