Score
Design and build hierarchical taxonomies that categorize diagnostic tests or benchmark tasks into multi-level structures and orthogonal splits (e.g., semantic/analytical/robustness), enumerate specific failure modes, and define the axes and granularity needed to generate, select, and score diagnostic benchmarks so practitioners can systematically compare and analyze system behavior across targeted capabilities.
This study addresses the high cost and expert dependency of manual taxonomy construction in software engineering (SE) by conducting the first systematic, multi-dimensional empirical evaluation of large language model (LLM)-driven automatic classification in this domain. Leveraging two representative approaches—TnT-LLM and CLIMB—and five state-of-the-art LLMs across seven human-annotated SE paper datasets, the work analyzes performance along key dimensions including classification quality, alignment with expert judgments, reliability, and efficiency. Results reveal that TnT-LLM achieves near-human classification quality but incurs high computational cost and structural complexity, whereas CLIMB offers 15–40× faster inference and 8–49× lower cost at the expense of reduced accuracy in tasks requiring deep technical reasoning. The findings elucidate critical trade-offs among quality, cost, and complexity, providing actionable guidance for method selection in practice.
This work addresses the limitations of the Category-Partition (CP) testing method, which is often hindered by tedious manual execution and error-prone processes due to a lack of automation and visualization support. To overcome these challenges, the authors design and implement a CP testing tool featuring an integrated graphical user interface that fully automates the entire workflow—from defining parameters, environment variables, categories, and options (including constraints) to constructing test frames and generating test cases. The tool introduces type-aware option specifications (supporting Boolean, integer, real, and string types), a robust constraint-handling mechanism, and multiple combinatorial generation strategies, significantly enhancing both expressiveness and usability. Empirical validation through nine case studies demonstrates that the tool efficiently produces valid CP-compliant test cases, effectively supporting systematic test design.
In complex equipment fault diagnosis, domain expertise is difficult to formalize, and manual fault tree construction is inefficient. Method: This paper proposes a knowledge graph–based approach for automated fault tree synthesis. It introduces a lightweight, semantically rich knowledge graph representation that enables semi-automatic extraction of failure logic relationships from unstructured documents (e.g., maintenance manuals) and structured/functional models. Leveraging hierarchical modeling and semantic reasoning, the method generates fault trees in a fully structured manner—without requiring historical fault data, relying solely on engineering knowledge. Contribution/Results: The synthesized fault trees are inherently interpretable and accurately capture system-level failure propagation paths. Experimental validation on the Lycoming O-320 aircraft engine demonstrates substantial improvements in diagnostic modeling efficiency and engineering applicability.
Existing benchmarks for knowledge work evaluation largely adhere to traditional NLP task paradigms, failing to capture systems’ capabilities in real-world knowledge-intensive settings. This work proposes a three-step framework—explicitly defining work activities, establishing realistic test environments, and focusing evaluation on deliverable outputs—and derives 18 core knowledge work activities from the O*NET database. Innovatively integrating role responsibilities, local tool usage, and downstream usability into benchmark design, the approach establishes a coherent “work activity–test setup–scoring artifact” alignment. Validation through three case studies (GDPval, OfficeQA Pro, and APEX-SWE) exposes critical misalignments in current benchmarks between tasks, environments, and actual work objectives, offering a new paradigm for evaluating knowledge work systems in practical, application-oriented contexts.
This work addresses the lack of a systematic hierarchical taxonomy for GitHub repositories, where existing tag-based mechanisms are flat, inconsistent, and sparsely populated. The authors propose the first end-to-end framework for automatically generating a hierarchical classification of repositories by integrating knowledge from large language models with empirical repository distributions. Their approach employs a multi-agent architecture—comprising designer agents that construct taxonomic dimensions and classifier agents that assign projects—augmented with an iterative self-correction mechanism and a novel hierarchical path evaluation strategy. Evaluated on a benchmark of 2,001 repositories, the method achieves a Taxonomy Quality Factor (TQF) of 83.13%, outperforming the best baseline by 15 percentage points. In downstream tasks, it attains 85.71% precision at rank 1 for alternative discovery, surpassing human-curated lists and substantially improving retrieval efficiency, while also uncovering evolutionary trends in domains such as AI and machine learning.
This work addresses the challenge of efficiently constructing a comprehensive and well-structured taxonomy of artificial intelligence skills and tasks from massive hiring data. To this end, the authors propose TaxonomyBuilder, a framework that integrates systematic data filtering, clustering algorithms, and large language model–enhanced hierarchical label generation to automatically derive domain-specific taxonomies from curated, high-quality data subsets. Experimental results demonstrate that taxonomies built from filtered data exhibit significantly broader coverage and superior structural coherence compared to those generated from raw, unfiltered data using existing methods. The study thus establishes a novel paradigm for data-driven, automated taxonomy construction in specialized domains.
Current evaluation methods for large language models (LLMs) primarily identify failing samples or categories but struggle to uncover underlying capability deficiencies, thereby limiting targeted model improvement. This work proposes CRAFT, a novel framework that diagnoses model weaknesses at the scoring-criterion level. CRAFT constructs a hierarchical capability tree by extracting capability descriptions and applying hierarchical clustering, then dynamically identifies low-performance nodes across multiple granularities to generate targeted fine-tuning data. Evaluated on financial and legal domains as well as 13 standard benchmarks, CRAFT significantly outperforms prompt-clustering and random data generation baselines. Fine-tuning four open-source LLMs with CRAFT-generated data consistently enhances their performance, demonstrating more precise localization of capability gaps and enabling efficient, targeted model refinement.
Scientific process descriptions are often embedded in unstructured text, hindering reproducibility, comparison, and automation. To address this challenge, this work presents the first cross-disciplinary, expert-driven repository of structured scientific process schemas, encompassing 16 expert-annotated patterns across five domains. Through a human-in-the-loop workflow, candidate schemas generated by large language models were iteratively refined via domain expert feedback, yielding reusable fields such as inputs, outputs, steps, and parameters. The resulting schemas are formalized in both JSON Schema and SHACL formats and accompanied by an integrated toolchain. The project also releases a comprehensive dataset—including schemas, intermediate artifacts, review records, and analysis scripts—to support knowledge graph construction, semantic publishing, and cross-study comparison.
This work addresses the challenge of effectively reusing failure feedback from existing agent execution trajectories, which are often lengthy, instance-specific, and lack standardized failure descriptions. To overcome this, the authors propose an unsupervised method that automatically distills raw trajectories into a structured, evidence-backed failure taxonomy. This taxonomy forms an adaptive failure glossary organized along three axes—system-level, role-level, and domain-level—and serves as a unified feedback interface integrated into trajectory selection, runtime monitoring, and system search processes. Requiring no manual annotation, the glossary achieves a 10× compression ratio while exhibiting semantics closely aligned with expert annotations. Empirical results demonstrate significant performance gains across multiple benchmarks: SWE-agent’s resolution rate improves from 60% to 70%, Claude Code reaches 70.7%, and Terminal-Bench 2.0 accuracy increases by 8–15 points.