Score
Designs and implements integration components that wrap and connect foundation models to applications, such as adapter layers, model adapters, and unified invocation wrappers that standardize model initialization, invocation, and capability/metadata exposure. Builds minimal-dependency APIs and consistency layers for heterogeneous model packages (including input/output adapters for tabular foundation models), plus tooling to manage model lifecycle and invocation.
General-purpose foundation models struggle to adapt to domain-specific patterns and requirements. Method: This study systematically establishes a theoretical and methodological framework for Domain-Specific Foundation Models (DSFMs), proposing the first cross-industry reusable DSFM customization framework that integrates pretraining-finetuning, instruction alignment, domain adaptation, knowledge injection, and efficient parameter updating, accompanied by a multi-dimensional evaluation system. Results: The framework is validated across ten+ domains—including finance, healthcare, and manufacturing—yielding a comprehensive application landscape; it identifies three universal bottlenecks: data scarcity, computational constraints, and regulatory compliance barriers. This work fills a critical gap in DSFM surveys and delivers the first authoritative, academically rigorous methodology guide and practical reference for industry–academia–research collaboration in DSFM development.
This work addresses the lack of a scalable, traceable, and systematic approach to modernizing large-scale legacy systems while preserving both functional and non-functional characteristics. The authors propose a four-phase model-driven method that leverages a semantically rich intermediate model to uniformly abstract a legacy system’s structure, dependencies, and metadata. By designing semantics-preserving transformation rules, the approach enables semi-automated migration to modern platforms such as web-based architectures. The method establishes an end-to-end model-driven pipeline that integrates semantic metadata modeling with automated code synthesis. Evaluated on an industrial-scale .NET system, it successfully migrated core UI components, significantly enhancing maintainability and scalability while reducing modernization risks and manual effort.
Current tool-augmented large language model (LLM) ecosystems suffer from fragmentation—characterized by coexisting heterogeneous protocols (e.g., OpenAI Function Calling, Toolformer), manual schema definition, and complex execution orchestration—leading to low development efficiency and high integration overhead. To address this, we propose a protocol-agnostic unified tool integration framework. Our approach introduces an abstract protocol layer for cross-standard compatibility, an automated schema inference mechanism to eliminate manual specification, and a dual-mode concurrent scheduler enabling seamless synchronous and asynchronous tool execution. Experimental evaluation demonstrates that, compared to baseline approaches, our framework reduces implementation code volume by 60–80%, achieves up to 3.1× improvement in end-to-end execution latency, and maintains full backward compatibility with mainstream LLM tool-calling ecosystems.
Current micro-frontend architectures heavily rely on specific bundlers (e.g., Webpack), leading to inflexible module composition, constrained cross-team collaboration, and bottlenecks in error detection, runtime observability, and loading performance. To address these limitations, we propose Bundler-Independent Module Federation (BIMF)—the first runtime module federation framework decoupled from build-time bundlers. BIMF enables dynamic module loading, type-safe inter-module collaboration, and cross-team dependency sharing. It integrates runtime dependency resolution, distributed tracing, server-side rendering (SSR), and intelligent prefetching to significantly enhance observability and first-contentful-paint (FCP) performance. Experimental evaluation of a prototype implementation demonstrates: (1) full preservation of TypeScript type contracts across modules; (2) 100% dependency deduplication; (3) a 37% reduction in average module loading latency; and (4) a 42% improvement in parallel development efficiency across distributed teams.
This study addresses the lack of standardized practices in Model Context Protocol (MCP) regarding configuration, communication, and human oversight, noting that prior research has predominantly focused on server-side implementations while neglecting application-level usage. To bridge this gap, we introduce MCPAppTax, the first taxonomy for MCP applications, and conduct a large-scale empirical analysis of 1,723 MCP applications on GitHub, leveraging large language model–assisted annotation and static code analysis. Our findings reveal both convergent practices—such as 85.2% adopting file-based configuration and 81.1% using official SDKs—and divergent ones, notably the absence of standardized naming conventions for configuration parameters. Furthermore, while human supervision mechanisms are prevalent (90.8% log interactions and 77.2% offer start/stop controls), only 37.2% implement blocking-style human approval, highlighting significant variability in safety-critical oversight.
This study bridges the gap between foundational model (FM) academic research and industrial practice by systematically investigating bidirectional interactions: FM-for-software-engineering (FM4SE) and software-engineering-for-FM (SE4FM). Method: We conduct a gray literature analysis of 1,152 technical blog posts from industry, introducing the novel “Model Jury” framework—integrating multi-LLM collaborative annotation, prompt-engineering-driven semantic classification, and abstractive summarization for automated large-scale analysis of unstructured text. Results: We find code generation dominates FM4SE applications; SE4FM primarily targets deployment, operations, and system orchestration; and edge-aware lightweight FM deployment is an emerging trend. We propose eight interdisciplinary research directions and open-source our dataset, prompt templates, and analysis code—establishing the first empirical benchmark and methodological foundation for FM4SE/SE4FM research.
论文提出基于RM-ODP的多视角建模框架,结合大语言模型辅助分析,解决数字孪生系统中模型重用时的兼容性问题。
This study addresses the interoperability challenges in automotive domain modeling arising from the coexistence of heterogeneous tools, multiple modeling languages, and a mix of proprietary and open-source environments. To tackle this issue, the work proposes a novel automated approach that leverages large language models (LLMs) to map and merge source model instances into target metamodels based on Ecore and SysML v2. A structural validation mechanism is integrated to ensure semantic consistency and syntactic correctness of the generated models. Experimental evaluation on real-world automotive cases demonstrates that the method substantially reduces manual transformation effort while efficiently producing target models that are both structurally valid and aligned with user requirements, thereby establishing a viable new paradigm for cross-tool modeling interoperability.
This study addresses the limitation of existing research in effectively measuring whether open-source models are genuinely translated into publicly accessible applications, as metrics based solely on release, visibility, or technical reuse inadequately capture real-world impact. To bridge this gap, the paper introduces “public application translation” as a distinct dimension of model influence and constructs a large-scale dataset leveraging Model-Space links from the Hugging Face platform. Through systematic analysis combining metadata readiness assessment with heterogeneous space configuration, the work reveals that only a small fraction of models are linked to Spaces—and these are highly concentrated. Models successfully translated into public applications exhibit higher metadata quality and are deeply embedded within a diverse ecosystem encompassing datasets, SDKs, and task categories, thereby extending the evaluative framework for open-source model impact.
This study addresses a structural misalignment between producers and consumers of pretrained language models (PTLMs) on platforms like Hugging Face, which manifests as mismatches in model discovery, documentation, lineage tracing, and governance. Through surveys and qualitative analysis involving 50 model producers and 95 GitHub-based consumers, this work reveals significant discrepancies in how the two groups perceive the placement of critical metadata, motivations for lineage tracking, and priorities in model governance. These findings provide empirical grounding for improving model documentation standards, lineage-tracking tools, and governance frameworks, offering a novel perspective on optimizing the PTLM reuse ecosystem.
This study addresses the significant yet underexplored impact of framework design on LLM agent performance in software engineering. We present the first systematic quantification of framework effects, empirically analyzing the interaction mechanisms among core components—including tool registration, context compression, and sub-agents—using the SWE-bench benchmark with Qwen and DeepSeek models. To facilitate this analysis, we introduce NanoHarness, a lightweight, modular evaluation framework. Our findings establish framework design as a primary determinant of agent performance and reveal diminishing marginal returns when applying complex frameworks to highly capable models. Notably, NanoHarness successfully replicates the performance gains achieved by most production-grade frameworks, yielding a substantial 7.37% improvement in agent effectiveness.