Score
Design and implement the organization of features that are computed, retrieved, and served at model deployment (inference) time, including schemas, metadata, serving APIs, freshness and latency constraints, fallback and transformation logic, and orchestration for real‑time or near‑real‑time serving. Build and analyze feature registries, serving pipelines, and integrations with model servers and monitoring to ensure consistent, low‑latency, and reliable availability of deployment‑time features.
To address the lack of a unified knowledge framework in MLOps, this paper conducts a multi-source literature review (MLR), systematically synthesizing 150 academic publications and 48 grey literature sources to overcome single-perspective limitations. Through thematic coding and cross-source evidence triangulation, it establishes the first comprehensive MLOps conceptual model and practice map spanning the full ML lifecycle and integrating consensus from both industry and academia. Key contributions include: (1) a widely adopted, rigorous definition of MLOps; (2) distillation of 12 core MLOps practices; and (3) identification of seven recurrent implementation challenges alongside empirically grounded mitigation strategies. The resulting knowledge base is modular, reusable, and rigorously validated—serving as a foundational reference for MLOps standardization, tooling development, and empirical research.
This study addresses the lack of systematic, large-scale analyses of structural properties in software feature models, which has hindered the understanding and evolution of variability models. For the first time, it systematically applies large-scale network analysis to 5,709 variability models drawn from 20 repositories. By constructing graphs capturing transitive dependencies and conflicts among features, and integrating graph modeling with network-theoretic and statistical analyses, the work uncovers cross-domain structural commonalities—such as dependency dominance, high centralization, and characteristic degree distributions—as well as domain-specific deviations. These findings provide novel empirical insights and a foundation for identifying pivotal features, guiding modular decomposition, and assessing structural fragility in variability-intensive systems.
This study addresses the lack of systematic empirical investigation into the deployment of large language model (LLM) inference frameworks in real-world software systems. Through large-scale analysis of open-source projects, combining static code analysis with metadata mining, this work systematically characterizes the adoption patterns and integration strategies of prominent frameworks—including vLLM, SGLang, TensorRT-LLM, LMDeploy, and FlashInfer—and examines their relationships with model type, scale, modality, and deployment environment. The findings reveal that vLLM exhibits the highest adoption rate, while multi-framework co-deployment remains limited yet holds complementary potential. Framework selection is strongly driven by model characteristics and application scenarios, effectively supporting diverse system designs such as reinforcement learning inference, multimodal generation, and microservice architectures.
This work addresses the lack of a scalable, traceable, and systematic approach to modernizing large-scale legacy systems while preserving both functional and non-functional characteristics. The authors propose a four-phase model-driven method that leverages a semantically rich intermediate model to uniformly abstract a legacy system’s structure, dependencies, and metadata. By designing semantics-preserving transformation rules, the approach enables semi-automated migration to modern platforms such as web-based architectures. The method establishes an end-to-end model-driven pipeline that integrates semantic metadata modeling with automated code synthesis. Evaluated on an industrial-scale .NET system, it successfully migrated core UI components, significantly enhancing maintainability and scalability while reducing modernization risks and manual effort.
Inconsistent definitions of “feature” across software engineering domains—particularly requirements engineering (RE) and software product lines (SPL)—impede communication, trigger rework, and reduce cross-team collaboration efficiency. Method: We conducted an empirical study across 27 mainstream open-source projects, integrating repository mining, branch behavior analysis, qualitative coding, and pattern induction to derive a data-driven, cross-disciplinary definition of feature. Contribution/Results: This work introduces the first empirically grounded, unified feature definition framework bridging RE and SPL. It identifies recurring collaboration patterns and critical bottlenecks in feature description, implementation, and management, and proposes a roadmap linking academic theory with industrial practice. The findings yield actionable guidelines for project planning, resource allocation, and inter-team coordination, advancing feature conceptual standardization and engineering practice optimization.
Existing online feature modeling tools suffer from functional limitations, discontinued maintenance, or reliance on local installations, hindering collaborative development of configurable systems. To address this, we propose the first fully web-based, lightweight, and collaborative feature modeling toolbox. Built upon FeatureIDE’s core library using TypeScript and React, it enables online loading, visual editing, format conversion, persistent storage, and real-time multi-user collaboration of feature models. The tool requires no local installation and executes natively in browsers with low-latency synchronization, significantly improving accessibility and team productivity. A preview version has been open-sourced and empirically validated for usability and extensibility, effectively bridging a critical gap in current online feature engineering infrastructure.
While large language models (LLMs) can generate executable multi-service application environments, they often deviate from the architectural and security requirements essential for production deployment. This work proposes a method to automatically generate Dockerfiles and Docker Compose configurations solely from code repositories, evaluating deployment fidelity through end-to-end HTTP testing and structural comparison. It explicitly distinguishes between functional correctness and fidelity to deployment intent, deriving a minimal set of explicit deployment specifications that cannot be inferred automatically from source code alone. Experiments successfully reproduce the topology and dependencies of three heterogeneous multi-service systems, confirming functional feasibility; however, critical production-grade features—such as network isolation and multi-stage builds—are consistently absent, revealing fundamental limitations in current LLMs’ ability to model deployment intent.
Automatically generating YAML configuration files that are both structurally valid and compliant with multiple continuous integration (CI) service specifications remains a significant challenge, and the capabilities of current large language models (LLMs) on this task are not well understood. This work introduces DOC2CI, the first cross-CI benchmark dataset comprising 3,363 document–YAML pairs, and systematically evaluates 14 open-source models alongside GPT-series models. A novel failure taxonomy is proposed to uncover the root causes of model discrepancies, and this study provides the first empirical evidence that document similarity and structural validity constitute distinct optimization objectives. Experiments reveal that even the largest models achieve an Exact Match rate below 3.1%; while 97% of generated outputs are syntactically parseable, only 71% conform to the target service schema. Schema-guided post-hoc repair without additional training boosts structural validity to 94%, whereas fine-tuning improves document similarity at the expense of standalone structural correctness.
This work addresses a critical limitation in existing large language model (LLM) routing strategies, which rely solely on model-level labels to determine substitutions while neglecting the influence of a model’s role within multi-call workflows and its deployment context on actual performance. To remedy this, the authors propose a role-conditioned substitution principle that decouples replacement decisions into “whether to replace” and “effect evaluation” via predicate-action decomposition. Through a controlled solve-merge-validate pipeline, they systematically investigate the conditional dependencies governing effective substitutions. Empirical results from multi-call LLM workflow experiments—including input-matching interventions, allocation ablations, and cross-model comparisons across Qwen and GPT families—demonstrate that substitution efficacy is jointly determined by the model’s procedural role and deployment context. Notably, in mixed Qwen/GPT configurations, sparse, role-aware replacements reduce RMSE from 4.818 to 1.538, significantly outperforming indiscriminate full-model upgrades.