A Cascaded Unsupervised-Supervised NLP Pipeline for Detecting Accusatory Language in Public Procurement

📅 2026-08-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the underutilization of stakeholder comments and government open data in public procurement, which hinders timely detection of procedural violations. To overcome this challenge, the authors propose a lightweight, domain-adapted cascaded unsupervised–supervised NLP framework. The approach first employs domain-finetuned Word2Vec embeddings combined with Gaussian Mixture Model (GMM) clustering to uncover latent anomalous patterns; it then applies a Random Forest classifier to identify comments with accusatory intent. Despite severe class imbalance, the method achieves high precision and recall without requiring extensive computational resources, enabling effective identification of procurement risks and significantly enhancing regulatory transparency.
📝 Abstract
Public procurement involves the allocation of substantial financial resources; therefore, continuous oversight through audits, controls, and monitoring mechanisms is essential. However, stakeholder comments and publicly available government data are often underutilized, despite their potential to reveal procedural irregularities. To address this gap, this paper analyzes metadata from Ecuador's Sistema Oficial de Contratación Pública (SOCE, Official Public Procurement System), with particular emphasis on participant comments generated during the pre-contractual phase. We propose a hybrid modeling framework that integrates unsupervised clustering and supervised classification within a natural language processing (NLP) pipeline to uncover latent patterns and detect potentially irregular procurement processes. Semantic embeddings are generated using Word2Vec, LLaMA, and RoBERTa, followed by Gaussian Mixture Models (GMMs) for unsupervised clustering. A supervised classification stage is then applied to identify accusatory or whistleblowing-style comments. Experimental results show that the combination of domain-trained Word2Vec embeddings, GMM-based clustering, and a Random Forest classifier achieves high precision and recall, even under severe class imbalance. These findings demonstrate that lightweight, domain-adapted NLP architectures can effectively support risk identification and enhance transparency in public procurement systems without requiring large-scale computational infrastructure.
Problem

Research questions and friction points this paper is trying to address.

public procurement
accusatory language
whistleblowing
procedural irregularities
stakeholder comments
Innovation

Methods, ideas, or system contributions that make the work stand out.

hybrid NLP pipeline
unsupervised-supervised learning
accusatory language detection
domain-adapted embeddings
public procurement transparency
🔎 Similar Papers
No similar papers found.
B
Bryan Torres
Colegio de Ciencias e Ingenierías "El Politécnico", Universidad San Francisco de Quito (USFQ), Diego de Robles S/N, Quito 170157
Daniel Riofrío
Daniel Riofrío
Colegio de Ciencias e Ingenierías "El Politécnico", Universidad San Francisco de Quito (USFQ), Diego de Robles S/N, Quito 170157
J
José Vega-Sánchez
Colegio de Ciencias e Ingenierías "El Politécnico", Universidad San Francisco de Quito (USFQ), Diego de Robles S/N, Quito 170157
N
Nathaly Orozco
Faculty of Engineering and Applied Sciences, Telecommunications Engineering, ETEL Research Group, Universidad de Las Américas (UDLA), 170503 Quito, Ecuador
C
Carla Parra
Departamento de Estudios Organizacionales y Desarrollo Humano, Escuela Politécnica Nacional, Quito 170525, Ecuador
Karen Rosero
Karen Rosero
Carnegie Mellon University
AI in HealthcareMultimodal Speech Processing
Felipe Grijalva
Felipe Grijalva
Associate Professor at USFQ
Signal ProcessingSpatial AudioMachine LearningComputer VisionAssistive Technologies