🤖 AI Summary
Existing multi-level hierarchical classification (MLHC) methods often neglect parent-child category relationships, leading to predictions that violate hierarchical constraints. To address this, we propose a taxonomy-structured, transition-based cross-modal product classification framework. Our approach introduces a novel taxonomy-embedded transition mechanism that explicitly models hierarchical dependencies to ensure inter-layer prediction consistency. It integrates category-tree encoding, hierarchical prompt engineering, and multimodal feature alignment, while leveraging frozen large language models (LLMs) with lightweight adapters for end-to-end training—enabling LLM-agnostic deployment. Evaluated on the MEP-3M dataset, our method achieves a 23.6% improvement in hierarchical consistency and a 9.4% gain in overall accuracy over conventional LLM-based baselines. The framework effectively balances structural awareness with generalization capability, advancing robust, constraint-aware hierarchical classification.
📝 Abstract
Multi-level Hierarchical Classification (MLHC) tackles the challenge of categorizing items within a complex, multi-layered class structure. However, traditional MLHC classifiers often rely on a backbone model with independent output layers, which tend to ignore the hierarchical relationships between classes. This oversight can lead to inconsistent predictions that violate the underlying taxonomy. Leveraging Large Language Models (LLMs), we propose a novel taxonomy-embedded transitional LLM-agnostic framework for multimodality classification. The cornerstone of this advancement is the ability of models to enforce consistency across hierarchical levels. Our evaluations on the MEP-3M dataset - a multi-modal e-commerce product dataset with various hierarchical levels - demonstrated a significant performance improvement compared to conventional LLM structures.