Score
Designs, builds, and evaluates systems that assign discrete category labels to text, documents, and tasks—covering document-type and structural detection (e.g., text_only|scanned|special_char), intent/task categorization, and multiclass label prediction for domain-specific labels. Implements end-to-end pipelines to filter, route, or select documents (e.g., for OCR or downstream extractors), detects content attributes such as political messaging or medical diagnoses, and optimizes and evaluates multiclass performance using metrics like macro‑F1 and top‑k while comparing modeling strategies.
Pre-trained language models (PLMs) often underperform in domain-specific text classification due to domain-specific terminology, syntactic idiosyncrasies, and class imbalance. Method: We conduct a systematic literature review (2018–early 2024) of 41 studies, adhering to PRISMA guidelines and augmented by AI-assisted tools for rigorous screening; we propose the first taxonomy of PLM adaptation techniques for domain text classification and establish a cross-domain, multi-dimensional performance evaluation framework. Contribution/Results: Empirical analysis of Transformer-based models—including BERT, SciBERT, and BioBERT—across biomedical and other domains identifies domain adaptation and data bias as critical bottlenecks. We validate the efficacy of fine-tuning strategies, domain-aware self-supervised pre-training, and balanced sampling techniques. This work provides both theoretical foundations and practical guidelines for designing domain-adaptive PLMs, advancing reproducible and robust domain-specific NLP.
Existing natural language processing resources often lack task-specific information for niche or emerging entities, hindering accurate classification in domains such as business or healthcare provider categorization. To address this limitation, this work proposes a dynamic classification framework that requires no additional labeled text: given only entity names and their corresponding labels, the method retrieves web-based information and leverages large language models (LLMs) to generate task-relevant descriptions, which are then used to train a text classifier. This end-to-end approach achieves strong performance on low-resource entity classification, attaining macro-averaged F1 scores of 82.3% on Standard Industrial Classification (SIC) coding and 72.9% on healthcare provider categorization, thereby demonstrating its effectiveness and practical utility.
This study addresses the challenge of balancing accuracy and computational efficiency in automated document classification under class-imbalanced conditions by systematically evaluating three representative models: Naive Bayes, bidirectional LSTM (BiLSTM), and fine-tuned BERT. Experimental results demonstrate that while BERT achieves over 99% accuracy, it incurs substantial computational overhead; Naive Bayes offers the fastest training speed but yields only approximately 94.5% accuracy; BiLSTM strikes the best trade-off between performance and efficiency, attaining 98.56% accuracy. Leveraging these insights, the authors implement a lightweight, robust demonstration system for automated technical request routing, confirming BiLSTM’s practical suitability for real-world deployment and offering an effective solution for text classification in resource-constrained environments.
Existing document classification benchmarks are largely confined to single-domain settings and flat label structures, failing to capture the hierarchical, multimodal, and cross-domain characteristics of real-world business documents. This work proposes MMM-Bench—the first industrial-scale benchmark for multi-level, multi-domain, and multimodal document classification—comprising 5,990 authentic documents across 12 commercial domains, annotated with a five-level hierarchical label taxonomy and complete human-verified classification paths. We systematically identify four core challenges inherent to this task, establish comprehensive baselines leveraging both open-source models and commercial APIs, and validate their efficacy through expert evaluation and empirical experiments. The MMM-Bench dataset and accompanying evaluation toolkit are publicly released to advance research in document intelligence.
To address real-world challenges in government document processing—including illumination variation, text distortion, occlusion, and low resolution—this paper proposes an end-to-end automatic document classification method. The approach integrates DBNet++ for robust text detection with the BART model for semantic understanding, supporting both offline images and real-time camera input. A lightweight image preprocessing module enhances system robustness, while extracted textual content enables accurate classification into four categories: invoices, reports, letters, and tables. A cross-platform interactive interface is implemented using PyQt5. Trained on the Total-Text dataset for 10 hours, the text detector achieves 92.88% accuracy. This work constitutes the first integration of DBNet++ and BART for semantic document-type classification, demonstrating strong generalization across heterogeneous input sources and challenging imaging conditions, thereby offering practical value for real-world deployment.
Hierarchical Extreme Multi-Label Classification (HEMLC) faces significant challenges due to the complexity and scale of label taxonomies. To address this, we propose Hierarchical Multi-label Generation (HMG), a novel paradigm that reformulates HEMLC as end-to-end generation of cross-level relevant labels within a given taxonomy. We introduce the first Probabilistic Level Constraint (PLC) mechanism, explicitly controlling the number of generated labels, path length, and hierarchical depth—enabling strong controllability without relying on clustering or other preprocessing steps. Our method jointly leverages taxonomy structural priors and a PLC-guided probabilistic loss, augmented by a taxonomy-aware decoding strategy. Evaluated on standard HEMLC benchmarks, HMG achieves new state-of-the-art performance, improving hierarchical compliance rate by 23.6% over prior methods while demonstrating superior controllability and generation quality.
This study addresses the challenges of multimodal fusion in visually rich documents and the lack of standardized evaluation in existing approaches by conducting controlled comparative and ablation experiments on representative models—including LayoutLMv3, Donut, Qwen3-VL-32B-Instruct, and Qwen3-32B—within a unified experimental framework. For the first time, it systematically compares OCR-dependent and OCR-free multimodal methods, revealing that specialized multimodal Transformers significantly outperform large language models. The findings further demonstrate that visual information plays a dominant role in layout-intensive document classification, whereas OCR-derived text provides only auxiliary value. These results offer empirical guidance for model selection and feature composition in multimodal document understanding.
This work addresses the gap between research and production deployment in large-scale multi-page document processing by proposing a microservice architecture tailored for high-throughput scenarios, integrating a multi-stage pipeline of document classification, optical character recognition (OCR), and large language model (LLM) inference. The system employs a hybrid classification strategy, decouples GPU-based inference from CPU-driven orchestration, leverages asynchronous I/O, and supports independent horizontal scaling, enabling stable processing of thousands of documents per hour. Empirical analysis reveals that OCR constitutes the primary bottleneck in end-to-end latency and that system concurrency is constrained by the inference capacity of shared GPUs rather than the number of nodes. This study offers a reusable, efficient deployment paradigm for industrial-scale document understanding systems.
This study addresses the inconsistency in human annotation caused by ambiguous category definitions in traditional content moderation. To resolve this, the authors propose an AI-driven constitutional annotation framework: large language models first assist humans in formulating structured, interpretable category “constitutions,” which then guide automated dual-axis labeling of intent and content safety. This approach shifts human effort from case-by-case judgments to high-level semantic definition. Evaluated on harassment, hate speech, and non-violent criminal conduct tasks, the method reduces cross-model annotation inconsistency by up to 57-fold compared to conventional paragraph-based rules and effectively exposes latent gaps in existing policy formulations.