Score
Designs and implements classification systems that predict labels arranged in a taxonomy, ensuring outputs respect parent–child relationships and other hierarchical constraints and optimizing for hierarchical evaluation metrics. This includes building loss functions, inference/decoding procedures, or post-processing rules that enforce top-level consistency and coherent multi-level label assignments.
In multi-level hierarchical classification (MLHC), conventional methods suffer from inconsistent predictions and poor fairness compliance due to independent output layers that ignore inter-class hierarchical relationships. This work proposes a model-agnostic masked output layer, enabling hierarchical consistency and sensitive-attribute fairness enforcement during inference—without modifying the backbone architecture or requiring retraining. Our approach unifies hierarchical constraints and group fairness objectives within a single optimization framework, leveraging a plug-and-play masking mechanism and multi-objective optimization. Evaluated on multiple benchmark datasets, the method significantly improves prediction consistency over baselines, reduces the demographic parity gap (ΔSPD) by 37%, and achieves higher accuracy than both unadjusted models and LLM-based debiasing approaches. The solution is particularly suited for high-reliability applications such as e-commerce and healthcare.
Hierarchical Extreme Multi-Label Classification (HEMLC) faces significant challenges due to the complexity and scale of label taxonomies. To address this, we propose Hierarchical Multi-label Generation (HMG), a novel paradigm that reformulates HEMLC as end-to-end generation of cross-level relevant labels within a given taxonomy. We introduce the first Probabilistic Level Constraint (PLC) mechanism, explicitly controlling the number of generated labels, path length, and hierarchical depth—enabling strong controllability without relying on clustering or other preprocessing steps. Our method jointly leverages taxonomy structural priors and a PLC-guided probabilistic loss, augmented by a taxonomy-aware decoding strategy. Evaluated on standard HEMLC benchmarks, HMG achieves new state-of-the-art performance, improving hierarchical compliance rate by 23.6% over prior methods while demonstrating superior controllability and generation quality.
Addressing the challenge of multi-label automatic annotation for large-scale, hierarchical classification systems in software requirements engineering, this study proposes a sentence-level zero-shot classification paradigm to circumvent the high annotation costs associated with supervised training. We introduce the first industrial-scale requirements annotation benchmark comprising 769 taxonomy labels and systematically demonstrate a strong negative correlation between the number of taxonomy leaf nodes and classification recall. We further propose a zero-shot multi-label classification method leveraging SBERT sentence embeddings, achieving significant improvements in recall. Empirical evaluation reveals that hierarchical strategies yield no consistent performance gain across settings. Our work validates the effectiveness and feasibility of zero-shot learning for large-scale requirements classification, offering a scalable, low-human-effort automation solution for requirements tracing. (138 words)
To address the limitation of prior-dependent constraints undermining generalizability in hierarchical multi-label classification (HMC), this paper proposes EDR, a prior-free error-driven constraint discovery framework. EDR automatically identifies model misprediction patterns to induce interpretable, structured logical constraints; it further integrates a constraint-driven post-processing mechanism with a neuro-symbolic joint modeling architecture to jointly perform error detection, constraint recovery, and multi-level consistency verification. For the first time, EDR achieves fully automated, interpretable knowledge discovery and robust cross-domain constraint transfer without predefined constraints—enabling effective constraint learning even under label noise. Evaluated on multiple public benchmarks and a newly constructed military vehicle recognition dataset, EDR achieves an error detection F1-score exceeding 0.89 and constraint recovery accuracy above 92%, significantly improving both hierarchical consistency and overall classification performance.
To address cross-granularity prediction errors in hierarchical image classification caused by visual inconsistency at test time, this paper proposes the first hierarchical classification paradigm grounded in *intra-image visual consistency*. Our method requires no external semantic supervision or pixel-level annotations; instead, it employs a self-supervised segmentation alignment mechanism to visually align fine-grained predictions with coarse-grained regions within the same image. By integrating multi-scale feature modeling with CLIP’s zero-shot transfer capability, the framework enforces both semantic and visual consistency. Evaluated on multiple hierarchical classification benchmarks, our approach significantly outperforms zero-shot CLIP and existing state-of-the-art methods, achieving higher classification accuracy and improved prediction coherence. Notably, it simultaneously enhances unsupervised image segmentation quality—thereby strengthening model interpretability and robustness—without additional supervision.
This work addresses the challenge of effectively leveraging large-scale unlabeled data containing unknown classes in semi-supervised hierarchical open-set classification. To this end, we propose a novel pseudo-labeling approach based on a teacher–student framework that introduces subtree pseudo-labels to provide structure-aware, reliable supervision signals. Additionally, we design an age-gated mechanism to mitigate overconfident pseudo-label predictions. As the first study to successfully apply semi-supervised learning to hierarchical open-set classification, our method achieves strong performance on the iNaturalist19 benchmark using only 20 labeled samples per class—surpassing self-supervised pretraining followed by fine-tuning and approaching fully supervised performance.
This work addresses fine-grained hierarchical classification in regulatory-intensive domains such as customs tariff codes and export controls, where predictions must strictly adhere to hierarchical structures and rule-based boundaries. Existing approaches struggle to simultaneously ensure hierarchical validity, rule consistency, and boundary-aware reasoning. To bridge this gap, the paper formalizes, for the first time, a regulation-driven hierarchical classification task and introduces a constraint-aware hierarchical search framework. The method parses regulatory documents into a searchable tree and, at each step, retrieves only legally permissible candidate nodes, guiding path decisions through structured fields and evidential text snippets. Evaluated on four expert-validated datasets, the approach significantly outperforms baselines in average accuracy—particularly excelling in distinguishing adjacent categories and handling boundary cases—while producing auditable and traceable decision paths.
This study addresses the limitation that classification trees generated by generative AI, despite appearing plausible, often suffer from leaf-node redundancy and cross-branch leakage, thereby lacking practical operational utility. To overcome the constraints of conventional local naming checks, this work proposes a holistic tree-level evaluation framework incorporating a structural discriminator and a team partitionability metric. Furthermore, an agent-based framework is designed to process proprietary customer feedback corpora, combining statistical analysis with semantic deduplication for multidimensional quantitative assessment. Experimental results reveal that 97.7% of generated leaf nodes duplicate ancestor names, accompanied by severe cross-branch leakage, demonstrating that surface plausibility alone cannot guarantee taxonomy quality. These findings establish a new paradigm for evaluating the practical utility of classification systems.
Traditional clustering methods struggle to model the infinitely fine recursive structures inherent in real-world geometric hierarchies. This work proposes Classification Fields, a novel framework that recursively generates hierarchical clusterings of unbounded depth through local parent–child refinement rules, and for the first time enables learning a generator from finite observations that extrapolates to arbitrary depths. The approach integrates a residual tuple recursion mechanism, Voronoi cell construction, and metric DAG encoding, and is implemented via ReLU networks with provable theoretical convergence. Experiments demonstrate that the method effectively preserves child ordering, geometric structure, and path metrics across synthetic context-free grammar (CFG) hierarchies, iterated function system (IFS) fractals, and image-induced clustering tasks, exhibiting strong generalization capabilities.
This work addresses the limitation of existing vision-language models in fine-grained classification, where predictions at leaf nodes are often correct but inconsistent with their parent categories due to a lack of hierarchical reasoning. To resolve this, the authors propose VL-Taxon, a novel framework that explicitly enforces hierarchical consistency during both training and inference. The approach operates in two stages: first, a top-down strategy enhances leaf-node classification accuracy; second, supervised fine-tuning combined with reinforcement learning ensures logical coherence across the entire taxonomic hierarchy. Evaluated on a small-scale subset of iNaturalist-2021 using Qwen2.5-VL-7B, VL-Taxon achieves an average improvement of over 10% in both leaf-node accuracy and hierarchical consistency—outperforming the original 72B model—without relying on externally generated data.