content categorization

Designs, builds, and evaluates systems that assign discrete category labels to text, documents, and tasks—covering document-type and structural detection (e.g., text_only|scanned|special_char), intent/task categorization, and multiclass label prediction for domain-specific labels. Implements end-to-end pipelines to filter, route, or select documents (e.g., for OCR or downstream extractors), detects content attributes such as political messaging or medical diagnoses, and optimizes and evaluates multiclass performance using metrics like macro‑F1 and top‑k while comparing modeling strategies.

contentcategorization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.77
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing natural language processing resources often lack task-specific information for niche or emerging entities, hindering accurate classification in domains such as business or healthcare provider categorization. To address this limitation, this work proposes a dynamic classification framework that requires no additional labeled text: given only entity names and their corresponding labels, the method retrieves web-based information and leverages large language models (LLMs) to generate task-relevant descriptions, which are then used to train a text classifier. This end-to-end approach achieves strong performance on low-resource entity classification, attaining macro-averaged F1 scores of 82.3% on Standard Industrial Classification (SIC) coding and 72.9% on healthcare provider categorization, thereby demonstrating its effectiveness and practical utility.

entity classificationentity coveragelesser-known entities

This study addresses the challenge of balancing accuracy and computational efficiency in automated document classification under class-imbalanced conditions by systematically evaluating three representative models: Naive Bayes, bidirectional LSTM (BiLSTM), and fine-tuned BERT. Experimental results demonstrate that while BERT achieves over 99% accuracy, it incurs substantial computational overhead; Naive Bayes offers the fastest training speed but yields only approximately 94.5% accuracy; BiLSTM strikes the best trade-off between performance and efficiency, attaining 98.56% accuracy. Leveraging these insights, the authors implement a lightweight, robust demonstration system for automated technical request routing, confirming BiLSTM’s practical suitability for real-world deployment and offering an effective solution for text classification in resource-constrained environments.

class imbalanceclassification accuracycomputational efficiency

Existing document classification benchmarks are largely confined to single-domain settings and flat label structures, failing to capture the hierarchical, multimodal, and cross-domain characteristics of real-world business documents. This work proposes MMM-Bench—the first industrial-scale benchmark for multi-level, multi-domain, and multimodal document classification—comprising 5,990 authentic documents across 12 commercial domains, annotated with a five-level hierarchical label taxonomy and complete human-verified classification paths. We systematically identify four core challenges inherent to this task, establish comprehensive baselines leveraging both open-source models and commercial APIs, and validate their efficacy through expert evaluation and empirical experiments. The MMM-Bench dataset and accompanying evaluation toolkit are publicly released to advance research in document intelligence.

document classificationenterprise content managementhierarchical taxonomy

Automated document processing system for government agencies using DBNET++ and BART models

Oct 15, 2025
AK
Aya Kaysan Bahjat
🏛️ Informatics Institute for Postgraduate Studies

To address real-world challenges in government document processing—including illumination variation, text distortion, occlusion, and low resolution—this paper proposes an end-to-end automatic document classification method. The approach integrates DBNet++ for robust text detection with the BART model for semantic understanding, supporting both offline images and real-time camera input. A lightweight image preprocessing module enhances system robustness, while extracted textual content enables accurate classification into four categories: invoices, reports, letters, and tables. A cross-platform interactive interface is implemented using PyQt5. Trained on the Total-Text dataset for 10 hours, the text detector achieves 92.88% accuracy. This work constitutes the first integration of DBNet++ and BART for semantic document-type classification, demonstrating strong generalization across heterogeneous input sources and challenging imaging conditions, thereby offering practical value for real-world deployment.

Automated document classification system for government agenciesClassifies documents into four predefined categoriesDetects text in images under challenging conditions

Hierarchical Multi-Label Generation with Probabilistic Level-Constraint

Apr 30, 2025
LC
Linqing Chen
🏛️ PatSnap Co., LTD.

Hierarchical Extreme Multi-Label Classification (HEMLC) faces significant challenges due to the complexity and scale of label taxonomies. To address this, we propose Hierarchical Multi-label Generation (HMG), a novel paradigm that reformulates HEMLC as end-to-end generation of cross-level relevant labels within a given taxonomy. We introduce the first Probabilistic Level Constraint (PLC) mechanism, explicitly controlling the number of generated labels, path length, and hierarchical depth—enabling strong controllability without relying on clustering or other preprocessing steps. Our method jointly leverages taxonomy structural priors and a PLC-guided probabilistic loss, augmented by a taxonomy-aware decoding strategy. Evaluated on standard HEMLC benchmarks, HMG achieves new state-of-the-art performance, improving hierarchical compliance rate by 23.6% over prior methods while demonstrating superior controllability and generation quality.

Generating multi-label outputs without preliminary clustering stepsHandling complex hierarchical label relationships in classificationPrecisely controlling model output count, length, and level

Latest Papers

What's happening recently
View more

This study addresses the challenges of multimodal fusion in visually rich documents and the lack of standardized evaluation in existing approaches by conducting controlled comparative and ablation experiments on representative models—including LayoutLMv3, Donut, Qwen3-VL-32B-Instruct, and Qwen3-32B—within a unified experimental framework. For the first time, it systematically compares OCR-dependent and OCR-free multimodal methods, revealing that specialized multimodal Transformers significantly outperform large language models. The findings further demonstrate that visual information plays a dominant role in layout-intensive document classification, whereas OCR-derived text provides only auxiliary value. These results offer empirical guidance for model selection and feature composition in multimodal document understanding.

document type classificationlayout structuremultimodal modeling

This work addresses the gap between research and production deployment in large-scale multi-page document processing by proposing a microservice architecture tailored for high-throughput scenarios, integrating a multi-stage pipeline of document classification, optical character recognition (OCR), and large language model (LLM) inference. The system employs a hybrid classification strategy, decouples GPU-based inference from CPU-driven orchestration, leverages asynchronous I/O, and supports independent horizontal scaling, enabling stable processing of thousands of documents per hour. Empirical analysis reveals that OCR constitutes the primary bottleneck in end-to-end latency and that system concurrency is constrained by the inference capacity of shared GPUs rather than the number of nodes. This study offers a reusable, efficient deployment paradigm for industrial-scale document understanding systems.

Document AILLM pipelinesmicroservice architecture

This study addresses the inconsistency in human annotation caused by ambiguous category definitions in traditional content moderation. To resolve this, the authors propose an AI-driven constitutional annotation framework: large language models first assist humans in formulating structured, interpretable category “constitutions,” which then guide automated dual-axis labeling of intent and content safety. This approach shifts human effort from case-by-case judgments to high-level semantic definition. Evaluated on harassment, hate speech, and non-violent criminal conduct tasks, the method reduces cross-model annotation inconsistency by up to 57-fold compared to conventional paragraph-based rules and effectively exposes latent gaps in existing policy formulations.

annotation driftcategory definitionscontent moderation

Hot Scholars

JZ

Jifan Zhang

University of Wisconsin-Madison
Label-Efficient LearningActive LearningDeep Learning
JC

Jiachi Chen

Associate Professor, Sun Yat-Sen University
Smart ContractsBlockchainLarge Language ModelsSoftware Security
HX

Hui Xiong

Senior Scientist, Candela Corporation
Ultrafast dynamicsatomic molecular physicsfree electron laser
ZZ

Zibin Zheng

IEEE Fellow, Highly Cited Researcher, Sun Yat-sen University, China
BlockchainSmart ContractServices ComputingSoftware Reliability
RD

Ronnie de Souza Santos

Assistant Professor, University of Calgary
Human Aspects of Software EngineeringSoftware TestingSoftware FairnessSoftware Development