Augmented Reinforcement Learning Framework For Enhancing Decision-Making In Machine Learning Models Using External Agents

📅 2025-08-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the “garbage-in, garbage-out” problem in reinforcement learning (RL) that degrades decision-making quality, this paper proposes an enhanced RL framework integrated with external supervision. Methodologically, it introduces a dual-external-agent architecture: one agent delivers real-time human feedback and action correction, while the other performs dynamic data filtering and constructs high-quality trajectory loops. The framework unifies RL, reinforcement learning from human feedback (RLHF), online evaluation, and data cleansing into a human–machine collaborative hybrid intelligence training paradigm. Its key innovation lies in the first integration of dual-agent supervision into the RL decision pipeline, enabling simultaneous dynamic policy correction and training data quality enhancement. Experiments on banking document recognition and information extraction demonstrate a 12.7% accuracy improvement over baselines, alongside significantly enhanced robustness and cross-scenario adaptability.

Technology Category

Humans and AI: Human-in-the-loop Machine LearningMachine Learning: Reinforcement LearningNatural Language Processing: Safety and Robustness

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingResponsible Web: Machine-in-the-loop, human agency and autonomySemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
This work proposes a novel technique Augmented Reinforcement Learning framework for the improvement of decision-making capabilities of machine learning models. The introduction of agents as external overseers checks on model decisions. The external agent can be anyone, like humans or automated scripts, that helps in decision path correction. It seeks to ascertain the priority of the "Garbage-In, Garbage-Out" problem that caused poor data inputs or incorrect actions in reinforcement learning. The ARL framework incorporates two external agents that aid in course correction and the guarantee of quality data at all points of the training cycle. The External Agent 1 is a real-time evaluator, which will provide feedback light of decisions taken by the model, identify suboptimal actions forming the Rejected Data Pipeline. The External Agent 2 helps in selective curation of the provided feedback with relevance and accuracy in business scenarios creates an approved dataset for future training cycles. The validation of the framework is also applied to a real-world scenario, which is "Document Identification and Information Extraction". This problem originates mainly from banking systems, but can be extended anywhere. The method of classification and extraction of information has to be done correctly here. Experimental results show that including human feedback significantly enhances the ability of the model in order to increase robustness and accuracy in making decisions. The augmented approach, with a combination of machine efficiency and human insight, attains a higher learning standard-mainly in complex or ambiguous environments. The findings of this study show that human-in-the-loop reinforcement learning frameworks such as ARL can provide a scalable approach to improving model performance in data-driven applications.
Problem

Research questions and friction points this paper is trying to address.

Enhancing decision-making in ML models using external agents
Addressing Garbage-In Garbage-Out problem in reinforcement learning
Improving document identification and information extraction accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Augmented Reinforcement Learning with external agents
Real-time evaluator for feedback and rejected data
Selective feedback curation for approved datasets
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sandesh Kumar Singh
MASTERS OF COMPUTER SCIENCE