Score
Designs and implements software frameworks and tooling that manage the lifecycle of large language models, including model loading and optimization, prompt and chain abstractions, inference serving, fine‑tuning or adapter workflows, distributed training/inference orchestration, and evaluation/monitoring pipelines. Builds developer-facing APIs, integrations, and runtime primitives to enable reproducible deployment, scalable performance tuning, and safety/guardrails for systems that use LLMs.
There is a lack of systematic, software engineering (SE)-driven analysis of the full lifecycle challenges associated with large language models (LLMs). Method: We conduct a structured systematic literature review grounded in the SE lifecycle model, decomposing LLM development into six phases—requirements, data, development, testing, deployment, and maintenance—and perform stage-wise problem identification. Contribution/Results: We uncover phase-specific SE challenges—including prompt engineering maintainability, insufficient test coverage under data drift, and misalignment between model versioning and code evolution—and propose the first SE-oriented research framework for LLMs. This framework delineates stage-specific research directions and integrated technical pathways. Our work bridges a critical theoretical gap at the intersection of LLMs and SE, delivering an actionable research roadmap and practical guidance for building efficient, reliable, and evolvable LLM-based software systems.
This study addresses fragmented task definitions, inconsistent evaluation protocols, and dataset biases hindering systematic progress in applying large language models (LLMs) to source code analysis. We conduct a systematic review of 217 publications from 2019–2024 and propose the first three-dimensional knowledge graph—spanning *tasks*, *models*, and *datasets*—specifically for code analysis, covering 12 analytical task categories, 8 mainstream LLM variants, and 19 core benchmarks. Methodologically, we introduce a reproducible, standardized evaluation framework grounded in empirical analysis. Through bibliometric analysis and architectural feature modeling, we identify key evolutionary trends and persistent bottlenecks—including limited generalizability and inadequate contextual modeling. Our contributions provide a structured conceptual foundation and methodological guidance for advancing LLM-driven code intelligence research.
This study addresses the lack of systematic empirical investigation into the deployment of large language model (LLM) inference frameworks in real-world software systems. Through large-scale analysis of open-source projects, combining static code analysis with metadata mining, this work systematically characterizes the adoption patterns and integration strategies of prominent frameworks—including vLLM, SGLang, TensorRT-LLM, LMDeploy, and FlashInfer—and examines their relationships with model type, scale, modality, and deployment environment. The findings reveal that vLLM exhibits the highest adoption rate, while multi-framework co-deployment remains limited yet holds complementary potential. Framework selection is strongly driven by model characteristics and application scenarios, effectively supporting diverse system designs such as reinforcement learning inference, multimodal generation, and microservice architectures.
To address the limited accuracy of general-purpose large language models (LLMs) in domain-specific modeling, this paper proposes a lightweight, fine-tuning-free optimization framework tailored for Llama 3.1 to enhance its capability in automatically generating domain models—particularly medical data models—from natural language descriptions. Methodologically, the framework integrates search-based hyperparameter tuning with structured prompt engineering: it systematically optimizes inference parameters (e.g., temperature, top-p, max_tokens) and employs stepwise, domain-aware prompt templates to improve output controllability and semantic consistency. By avoiding parameter updates, the approach eliminates the computational overhead and catastrophic forgetting associated with fine-tuning. Experiments across ten heterogeneous domains demonstrate substantial improvements, especially in healthcare (+23.6% F1 over baselines), alongside robust cross-domain generalization—validating both effectiveness and practical applicability.
To address high manual effort, delayed response, and incomplete coverage in industrial software test maintenance, this paper proposes and empirically validates two novel multi-agent architectures leveraging large language models (LLMs) to accurately identify test cases requiring maintenance—and to assist in their automated addition, deletion, or modification—following code changes. The study systematically distills critical triggering conditions and practical deployment constraints for LLMs in industrial settings, and conducts empirical evaluation using real-world data from Ericsson AB. Results demonstrate significant reductions in manual test maintenance effort, improved timeliness of issue response, and enhanced test coverage completeness. This work represents the first application of multi-agent collaboration to industrial-scale test maintenance, establishing a reusable methodological framework and empirical foundation for deploying LLMs in high-reliability software engineering contexts.
Large language models (LLMs) suffer from hallucination, unreliability, and uncontrolled behavior, hindering their trustworthy deployment in safety-critical workflows; existing reliability-enhancement tools are fragmented and lack a systematic framework. This paper introduces LSL (LLM Scripting Language), a domain-specific scripting language that embeds formal specifications, verifiable constraints, and explainability mechanisms directly into the LLM interaction process—enabling structured output constraints, programmable behavioral control, and decoupled execution governance. LSL unifies domain-specific language (DSL) design, formal verification, and runtime checking, significantly improving output reliability, consistency, and traceability. Experiments demonstrate that LSL effectively mitigates hallucination across diverse tasks, supports safe and controllable LLM integration, and establishes a novel interaction paradigm for trustworthy AI systems.
Empirical software engineering faces significant challenges, including large-scale data, methodological complexity, and poor reproducibility, while the application of large language models (LLMs) in this domain lacks systematic integration. This study conducts a systematic literature review of 50 studies published between 2020 and 2025 across 12 leading conferences and journals, offering the first comprehensive taxonomy of 69 LLM-supported auxiliary tasks in empirical software engineering. The analysis reveals that LLMs are predominantly employed in data processing and analysis phases to enhance automation, efficiency, and scalability. However, their use is frequently hindered by issues such as hallucination, inconsistent outputs, sensitivity to prompting, insufficient reporting of reproducibility, and a notable absence of human-centered collaboration and transparency. Building on these findings, this work proposes a research agenda oriented toward the responsible application of LLMs in empirical software engineering.
To address the challenge of adapting large language models (LLMs) to proprietary industrial programming languages—such as ABB RAPID—in automation domains, this paper proposes a fine-tuning-free, few-shot prompting method enabling locally deployed LLMs to directly comprehend and modify RAPID programs. By eliminating reliance on large-scale annotated datasets or custom model training, the approach preserves data privacy and enhances deployment flexibility. Experimental evaluation demonstrates its effectiveness on elementary tasks including code repair and logical adaptation, substantially lowering the adoption barrier for LLMs in non-general-purpose industrial language settings. Key contributions include: (i) the first systematic investigation into LLM support for closed industrial languages like RAPID; (ii) a lightweight, secure, and plug-and-play prompting framework; and (iii) a low-overhead, highly controllable paradigm for AI-assisted programming tailored to high-sensitivity industrial environments.
Despite growing adoption of large language models (LLMs) in software engineering, their real-world impact, benefits, and associated risks remain poorly understood. Method: We conducted a mixed-methods study with 46 industry practitioners from diverse technical backgrounds, combining structured surveys, open-ended interviews, and thematic analysis. Contribution/Results: This is the first systematic empirical investigation revealing LLM usage patterns—particularly in coding assistance, documentation generation, and system maintenance—alongside quantified benefits (e.g., +42% improvement in technical Q&A efficiency and enhanced documentation quality) and critical risks: credential leakage, overreliance, and knowledge atrophy. Notably, 68% of participants expressed concern about diminished programming autonomy. We propose the “supervised adoption” framework—a human-in-the-loop, incremental integration strategy—to guide responsible, secure, and sustainable LLM deployment in practice, grounded in empirical evidence and actionable insights.
This study addresses the unclear human-AI collaboration mechanisms in specification-driven software development with large language models (LLMs). We propose CURRANTE, a structured three-stage collaborative paradigm that guides developers through sequential refinement of requirements specifications, test cases, and function implementations. Implemented as a Visual Studio Code extension, CURRANTE integrates LLM assistance, fine-grained interaction logging, and automated test-based evaluation. By collecting interaction data and multidimensional performance metrics—including pass rates and completion time—on medium-difficulty tasks from LiveCodeBench, our work provides the first systematic empirical analysis of how iterative specification and testing dynamically influence LLM-generated code quality. These findings offer evidence-based insights for designing effective AI-augmented programming environments.