Automatic Multi-level Feature Tree Construction for Domain-Specific Reusable Artifacts Management

📅 2025-06-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the reliance on domain expertise and low efficiency in feature tree construction for reusable components in multi-domain software, this paper proposes FTBUILDER, a fully automated, multi-level feature tree construction framework. FTBUILDER integrates repository metadata crawling, hierarchical clustering, and large language model (GPT-4)-driven semantic induction via prompt engineering to enable bottom-up, end-to-end, semantics-aware feature tree generation—eliminating human domain expert involvement for the first time. The framework is transferable across open-source ecosystems (e.g., Linux) and industrial domains (e.g., aerospace). Experimental evaluation in the Linux ecosystem demonstrates significant improvements: silhouette coefficient increases by 9%, GValue by 11%, component selection time decreases by 26%, and GPT-4–based recommendation accuracy improves by 235%.

Technology Category

Machine Learning: Feature Construction/ReformulationConstraint Satisfaction and Optimization: Satisfiability Modulo TheoriesData Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Vertical and domain-specific searchSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
📝 Abstract
With the rapid growth of open-source ecosystems (e.g., Linux) and domain-specific software projects (e.g., aerospace), efficient management of reusable artifacts is becoming increasingly crucial for software reuse. The multi-level feature tree enables semantic management based on functionality and supports requirements-driven artifact selection. However, constructing such a tree heavily relies on domain expertise, which is time-consuming and labor-intensive. To address this issue, this paper proposes an automatic multi-level feature tree construction framework named FTBUILDER, which consists of three stages. It automatically crawls domain-specific software repositories and merges their metadata to construct a structured artifact library. It employs clustering algorithms to identify a set of artifacts with common features. It constructs a prompt and uses LLMs to summarize their common features. FTBUILDER recursively applies the identification and summarization stages to construct a multi-level feature tree from the bottom up. To validate FTBUILDER, we conduct experiments from multiple aspects (e.g., tree quality and time cost) using the Linux distribution ecosystem. Specifically, we first simultaneously develop and evaluate 24 alternative solutions in the FTBUILDER. We then construct a three-level feature tree using the best solution among them. Compared to the official feature tree, our tree exhibits higher quality, with a 9% improvement in the silhouette coefficient and an 11% increase in GValue. Furthermore, it can save developers more time in selecting artifacts by 26% and improve the accuracy of artifact recommendations with GPT-4 by 235%. FTBUILDER can be extended to other open-source software communities and domain-specific industrial enterprises.
Problem

Research questions and friction points this paper is trying to address.

Automates multi-level feature tree construction for reusable artifacts
Reduces reliance on domain expertise in artifact management
Improves efficiency and accuracy in artifact selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automatically crawls and merges software repositories metadata
Uses clustering to identify artifacts with common features
Employs LLMs to summarize features for tree construction
🔎 Similar Papers
No similar papers found.