Score
Design and build hierarchical labeling schemes and supporting substrates that produce and manage multi-granularity labels and derive coherent hierarchical taxonomies or prefix-code style label structures. Implement annotation and encoding procedures (coarse-to-fine assignment, staged residual/refinement codes, and single-pass reusable annotations) that enable efficient labeling, lookup, and reuse across granularities.
Existing data mixing approaches are constrained by fixed semantic labels and a single granularity level, limiting their ability to flexibly explore combinatorial effects across multiple granularities. This work proposes a hierarchical, data-driven multi-granularity annotation framework that leverages learnable semantic transformations and a three-stage residual vector quantization scheme to generate up to 130,000 reusable hierarchical document codes, enabling dynamic navigation from coarse to fine granularities. Evaluated in a pretraining setting with 1B parameters and 25B tokens, the method—combined with an equal-subbucket coverage strategy—achieves an average performance gain of +0.0253 across 16 tasks at specific granularities, demonstrating a significant interaction effect between granularity selection and mixing strategy.
Existing fine-grained corpus analysis relies either on labor-intensive manual annotation or opaque statistical methods, compromising scalability and interpretability. This paper proposes a two-stage inductive coding framework powered by large language models (LLMs): (1) bottom-up prompt engineering for automated, fine-grained label generation; and (2) semantic embedding–guided hierarchical clustering to construct interpretable, multi-level topic structures. Grounded in qualitative research logic, the approach ensures analytical transparency and human controllability while substantially enhancing scalability for large-scale qualitative text analysis. Experiments across three heterogeneous datasets demonstrate high alignment with expert annotations (average F1 = 0.82). Applied to opioid litigation texts, the method successfully uncovered systemic, aggressive marketing strategies employed by pharmaceutical companies—validating its theoretical insightfulness and practical utility.
Software developers face inefficient and time-consuming challenges in comprehending large, multifunctional codebases; existing README-based, coarse-grained project-level categorization fails to support fine-grained functional understanding. To address this, we propose AutoFL—the first automated, cross-granularity functional domain labeling method supporting file-, package-, and project-level annotations without relying on non-code documentation (e.g., READMEs). AutoFL directly models source code semantics via a weakly supervised learning framework that integrates code text parsing, multi-granularity semantic embedding, and hierarchical aggregation for end-to-end label generation. Evaluated across multilingual open-source projects, AutoFL significantly improves the accuracy, consistency, and interpretability of functional labels compared to baselines. It effectively alleviates key bottlenecks in software comprehension by enabling precise, scalable, and documentation-agnostic functional awareness.
Existing code search methods primarily focus on function-level alignment, overlooking the reusability of finer-grained code fragments—such as blocks and statements—thereby limiting cross-granularity retrieval performance. To address this, we introduce MGCodeSearchNet, the first multi-granularity code search dataset covering functions, blocks, and statements. We further propose MGS3, a novel framework featuring a Hierarchical Multi-Granularity Representation (HMGR) module. HMGR integrates syntax-aware hierarchical encoding, granularity-consistent positive sample construction, and intra-function negative sampling to enable self-supervised contrastive learning for cross-granularity alignment. MGS3 is compatible with mainstream pre-trained code models and achieves state-of-the-art performance across all granularities on multi-granularity benchmarks—outperforming prior approaches by +12.7% average Mean Reciprocal Rank (MRR) at both block- and statement-level retrieval.
Identifying appropriate parent classes for novel concepts in small-scale existing taxonomies (<100 nodes) remains challenging due to sparse structural signals and lack of labeled training data. Method: This paper proposes a label-free, large language model (LLM)-driven taxonomy expansion method. Its core innovation is *code-style prompting*: explicitly encoding hierarchical semantic structure via indentation, nesting, and functional abstraction—programming conventions that enable zero- or few-shot LLM comprehension and relational reasoning over taxonomy hierarchies. The approach integrates taxonomy-aware input representations with structured code prompts, eliminating reliance on large-scale annotated or self-supervised data construction. Results: Evaluated on five cross-domain real-world benchmarks, the method achieves an average 12.6% absolute improvement in parent-class prediction accuracy over state-of-the-art methods. Gains are especially pronounced in extremely small-scale settings (30–80 nodes), demonstrating robustness where conventional supervised or embedding-based approaches falter.
This work addresses the challenge that existing retrieval-augmented generation methods struggle to model high-level architecture and cross-file dependencies in theory-driven codebases, resulting in a semantic gap between theoretical specifications and implementations. To bridge this gap, we propose HCAG, a novel framework that integrates hierarchical abstraction with architecture-guided generation. HCAG constructs a multi-granularity semantic knowledge base offline, performs hierarchical retrieval online, and prioritizes architectural coherence during modular code generation. It further incorporates adaptive node compression to optimize computational cost and employs a multi-agent collaborative discussion mechanism to enhance planning capabilities. We also introduce the first large-scale dataset explicitly aligned between theoretical constructs and their implementations. Evaluated on algorithmic game theory system generation, HCAG substantially outperforms current approaches, achieving significant improvements in code quality, architectural consistency, and requirement fulfillment, while effectively boosting the performance of domain-specific large language models.
This work addresses the high cost of re-annotation in document layout analysis caused by evolving label categories by proposing a plug-and-play pseudo-labeling framework tailored for object detection. It introduces label propagation to document layout analysis for the first time, constructing multimodal object representations through the fusion of visual, textual, and positional embeddings. This enables efficient semi-supervised category propagation using only a small set of annotated samples. Experimental results on the D4LA dataset demonstrate that with merely 10% of the labeled data, the method achieves a mean average precision (mAP) of 54.0%, equivalent to 81.6% of the fully supervised performance, thereby substantially reducing manual annotation effort.
This work addresses the label granularity skew arising in federated hierarchical image classification due to heterogeneous annotation granularities across clients. It formally characterizes this previously unexplored form of statistical heterogeneity and proposes Federated Branch-Decoupled Fine-Tuning (FedBDFT). FedBDFT constructs client-specific local label hierarchies and integrates WordNet-guided hierarchy coarsening with conditional Softmax classifiers, enabling decoupled fine-tuning and aggregation of branch classifiers within a federated learning framework. Experimental results demonstrate that FedBDFT substantially improves generalization performance on CIFAR-100, TinyImageNet, and ImageNet, achieving average accuracy gains of 27.9% and 56.4% under granularity skew levels of 0.6 and 0.9, respectively, while effectively preserving hierarchical semantic structure even in zero-shot settings.
This work addresses fine-grained hierarchical classification in regulatory-intensive domains such as customs tariff codes and export controls, where predictions must strictly adhere to hierarchical structures and rule-based boundaries. Existing approaches struggle to simultaneously ensure hierarchical validity, rule consistency, and boundary-aware reasoning. To bridge this gap, the paper formalizes, for the first time, a regulation-driven hierarchical classification task and introduces a constraint-aware hierarchical search framework. The method parses regulatory documents into a searchable tree and, at each step, retrieves only legally permissible candidate nodes, guiding path decisions through structured fields and evidential text snippets. Evaluated on four expert-validated datasets, the approach significantly outperforms baselines in average accuracy—particularly excelling in distinguishing adjacent categories and handling boundary cases—while producing auditable and traceable decision paths.