Score
Designs, trains, fine-tunes, evaluates, and deploys large-scale language models and the systems that use them, including model architectures and scaling, pretraining/fine-tuning pipelines, data curation, inference/serving frameworks, and integrations that generate, understand, or manipulate natural language. Builds and analyzes LLM-based applications and capabilities—such as prompt engineering, adapters, alignment and safety mitigations, and performance/cost optimizations—to produce reliable, efficient, and controllable language-powered systems.
This study systematically investigates the capability boundaries and emergence mechanisms of large language models (LLMs), focusing on the origins of reasoning, code generation, and arithmetic abilities. Methodologically, it first uncovers statistically significant correlations between pretraining data distribution and the emergence of Chain-of-Thought (CoT) and Plan-of-Thought (PoT) capabilities; proposes the LLM-modulo collaborative paradigm—integrating modular external tools to overcome inherent limitations of end-to-end model reasoning; and combines Transformer interpretability analysis, CoT/PoT prompting strategies, and empirical evaluation across high-stakes domains (healthcare, finance, law). Key contributions include: (i) clarifying diminishing returns and fundamental bottlenecks in scaling laws; (ii) empirically demonstrating that external tool invocation substantially improves accuracy on complex, multi-step tasks; and (iii) providing theoretically grounded principles and actionable frameworks for controllable LLM evolution and responsible deployment.
There is a lack of systematic, software engineering (SE)-driven analysis of the full lifecycle challenges associated with large language models (LLMs). Method: We conduct a structured systematic literature review grounded in the SE lifecycle model, decomposing LLM development into six phases—requirements, data, development, testing, deployment, and maintenance—and perform stage-wise problem identification. Contribution/Results: We uncover phase-specific SE challenges—including prompt engineering maintainability, insufficient test coverage under data drift, and misalignment between model versioning and code evolution—and propose the first SE-oriented research framework for LLMs. This framework delineates stage-specific research directions and integrated technical pathways. Our work bridges a critical theoretical gap at the intersection of LLMs and SE, delivering an actionable research roadmap and practical guidance for building efficient, reliable, and evolvable LLM-based software systems.
The absence of systematic guidance for selecting open-source large language models (LLMs) hinders efficient and compliant deployment decisions. Method: This paper introduces and maintains an open, dynamically updated checklist assessing the deployment feasibility of LLMs. It systematically catalogs key attributes—publication year, license type (e.g., Apache 2.0, MIT, non-commercial), minimum GPU memory/compute requirements, architectural features (e.g., MoE, quantization support), fine-tuning methods, and ecosystem compatibility—for mainstream foundation and domain-specific models released between 2022 and 2024. Contribution/Results: The work pioneers a unified evaluation framework integrating legal compliance (license constraints) with engineering feasibility (hardware limitations), publishing the structured, machine-readable dataset publicly on GitLab. This resource significantly accelerates model selection for researchers and practitioners, addressing a critical gap in open-source LLM deployment decision support tools.
Current LLM application development lacks systematic, practice-informed guidelines, leading to a growing gap between academic research and industrial engineering. Method: Drawing on transcribed texts from 189 real-world developer practice videos (2022–2024), we integrate BERTopic-based automated topic modeling with iterative human refinement to construct the first empirically grounded, production-oriented thematic map of LLM application development. Contribution/Results: The map identifies eight core themes—including design & architecture, model enhancement, infrastructure, and ethical risk—spanning 20 key issues. Design & Architecture emerges as the most densely populated theme, with RAG at its architectural center; prompt engineering, fine-tuning, deployment toolchains, and AI ethics are recurrent high-frequency concerns. Critically, the map exposes significant lags in academic research relative to industrial practice and delivers an actionable, empirically validated priority framework—thereby bridging a critical empirical gap in the LLM engineering knowledge base.
This study systematically investigates core challenges impeding large language model (LLM) industrial deployment, identifying 12 representative bottlenecks across four critical dimensions: data scarcity, inefficient inference, complex deployment, and inaccurate evaluation. Method: We employ a mixed-methods approach—structured interviews with frontline practitioners, a research-question-driven review of 68 industrial practice papers, and qualitative content analysis. Contribution/Results: We propose the first “industry-perspective-driven” taxonomy for LLM deployment challenges; establish a dynamically updated GitHub knowledge repository of industrial LLM literature; and deliver an actionable, lifecycle-spanning optimization roadmap. The framework has been adopted by multiple enterprises and serves as a key reference benchmark for industrial LLM adoption.
To address high manual effort, delayed response, and incomplete coverage in industrial software test maintenance, this paper proposes and empirically validates two novel multi-agent architectures leveraging large language models (LLMs) to accurately identify test cases requiring maintenance—and to assist in their automated addition, deletion, or modification—following code changes. The study systematically distills critical triggering conditions and practical deployment constraints for LLMs in industrial settings, and conducts empirical evaluation using real-world data from Ericsson AB. Results demonstrate significant reductions in manual test maintenance effort, improved timeliness of issue response, and enhanced test coverage completeness. This work represents the first application of multi-agent collaboration to industrial-scale test maintenance, establishing a reusable methodological framework and empirical foundation for deploying LLMs in high-reliability software engineering contexts.
The high computational cost and resource demands of large language models (LLMs) hinder their practical deployment in generative AI applications. Method: This work investigates the feasibility of small language models (SLMs) for tool-augmented agent tasks, proposing a lightweight adaptation of the 350M-parameter OPT model via single-round supervised fine-tuning (SFT) on the ToolBench benchmark—implemented using the Hugging Face TRL framework, without reinforcement learning or complex chain-of-thought reasoning. Results: The fine-tuned SLM achieves a 77.55% tool-call success rate on ToolBench, substantially outperforming strong LLM baselines including ChatGPT-CoT (26.00%) and ToolLLaMA-DFS (30.18%). It demonstrates robust performance across enterprise-grade tasks such as document summarization, question answering, and structured data parsing, enabling low-cost, high-efficiency production deployment. This study provides the first empirical evidence that domain-specialized SLMs can surpass conventional performance expectations in tool utilization, establishing a novel paradigm for efficient, scalable AI agents.
Large language models (LLMs) suffer from hallucination, unreliability, and uncontrolled behavior, hindering their trustworthy deployment in safety-critical workflows; existing reliability-enhancement tools are fragmented and lack a systematic framework. This paper introduces LSL (LLM Scripting Language), a domain-specific scripting language that embeds formal specifications, verifiable constraints, and explainability mechanisms directly into the LLM interaction process—enabling structured output constraints, programmable behavioral control, and decoupled execution governance. LSL unifies domain-specific language (DSL) design, formal verification, and runtime checking, significantly improving output reliability, consistency, and traceability. Experiments demonstrate that LSL effectively mitigates hallucination across diverse tasks, supports safe and controllable LLM integration, and establishes a novel interaction paradigm for trustworthy AI systems.
This work addresses the challenges of applying large language models (LLMs) in modeling and simulation (M&S), where suboptimal prompt design, improper hyperparameter configuration, or inadequate data handling often lead to performance degradation, information loss, and non-deterministic behavior. For the first time, this study systematically identifies latent pitfalls specific to LLM deployment in M&S and proposes a principled framework centered on rigorous design and empirical evaluation. The framework encompasses key techniques including prompt engineering, retrieval-augmented generation (RAG), low-rank adaptation (LoRA), temperature control, and context management. By offering a structured set of practical guidelines, this research enables practitioners to critically assess the suitability and implementation strategies of LLMs in M&S contexts, thereby substantially enhancing their effectiveness and reliability.
This work addresses the challenge of systematically debugging large language models (LLMs), which is hindered by their black-box nature and stochastic behavior, particularly in the absence of standardized benchmarks. The paper introduces the first general-purpose debugging framework for LLMs, treating them as observable systems and establishing a structured pipeline that spans from issue detection to model optimization. The approach integrates model-agnostic evaluation metrics, interpretability techniques, and error analysis tools to enable coordinated iteration over prompts, parameters, and training data. Experimental results demonstrate that the framework substantially improves debugging efficiency, effectively identifying and rectifying model deficiencies across diverse tasks and non-standardized evaluation settings. This enables reproducible, transparent, and scalable model refinement.
This work addresses the challenge of deploying large language models in production settings, where they often fail to meet low-latency requirements, while smaller models typically suffer from limited reasoning capabilities, hallucinations, and insufficient long-context memory. To overcome these limitations, the authors propose supervised fine-tuning small models such as Mistral on domain-specific natural language–code paired data, thereby internalizing domain knowledge directly into model weights and substantially reducing reliance on runtime context. Experimental results demonstrate that the fine-tuned small models outperform larger counterparts in code generation quality while maintaining lower latency. Load testing and real-world deployment confirm their efficiency and stability. Furthermore, the approach supports additional customer-specific fine-tuning without compromising general-purpose capabilities, offering a practical pathway toward efficient and accurate domain-specific code generation.