Score
Designs and builds systems that let users or models interact with software via natural language, including language-driven tool interfaces, parsers that map utterances to structured tool calls, and runtime connectors that enable and control model-initiated tool invocation. Also develops evaluation and optimization methods for improving tool-calling accuracy, reducing invocation errors, and minimizing token usage for prompt-based tool interactions.
Existing large language models (LLMs) suffer from task interference and rigid format constraints in JSON-based tool calling, leading to degraded performance and poor robustness. To address this, we propose Natural Language Tools (NLT), a novel framework that introduces a pure natural-language paradigm for tool invocation—decoupling tool selection from response generation and eliminating reliance on structured output formats. NLT is fully compatible with any open-weight LLM, requiring no native tool-calling support or architectural modifications. In cross-domain evaluations spanning customer service and mental health assistance, NLT achieves an average accuracy improvement of 18.4 percentage points across 10 models and 6,400 test instances, reduces output variance by 70%, and demonstrates strong robustness against prompt perturbations. Notably, NLT enables open-weight models to outperform state-of-the-art closed-source flagship models in tool-calling performance for the first time.
Large language models (LLMs) suffer from hallucination, unreliability, and uncontrolled behavior, hindering their trustworthy deployment in safety-critical workflows; existing reliability-enhancement tools are fragmented and lack a systematic framework. This paper introduces LSL (LLM Scripting Language), a domain-specific scripting language that embeds formal specifications, verifiable constraints, and explainability mechanisms directly into the LLM interaction process—enabling structured output constraints, programmable behavioral control, and decoupled execution governance. LSL unifies domain-specific language (DSL) design, formal verification, and runtime checking, significantly improving output reliability, consistency, and traceability. Experiments demonstrate that LSL effectively mitigates hallucination across diverse tasks, supports safe and controllable LLM integration, and establishes a novel interaction paradigm for trustworthy AI systems.
To address the “prompting software crisis” arising from the lack of systematic methodologies for prompt development, this paper proposes “Prompt Software Engineering” (PSE)—a novel paradigm that adapts software engineering principles to the full lifecycle management of natural language prompts in non-deterministic large language model (LLM) environments. Methodologically, we introduce, for the first time, a comprehensive framework tailored to LLM characteristics, encompassing prompt requirement analysis, formal design, repeatable testing, interpretable debugging, and iterative evolution—integrating semantic modeling, context-aware testing, and evolutionary mechanisms. Our core contribution lies in extending classical software engineering beyond its traditional boundaries of determinism and precision, thereby establishing the first end-to-end engineering methodology and research roadmap for prompting. This advances LLM-native application development by enabling reusable, verifiable, and maintainable prompt-based systems.
Prior work lacks a systematic understanding of how input parameters—such as prompt design, temperature, number of candidate solutions, and context—affect code generation in language models, hindering their reliable deployment. Method: This paper conducts the first controlled experiments on GitHub Copilot and OpenAI Codex, establishing a reproducible parameter perturbation framework grounded in HumanEval and LeetCode benchmarks. Contribution/Results: We empirically uncover strong, nonlinear couplings among temperature, prompt formulation, and candidate count—demonstrating that optimizing any single parameter in isolation is ineffective and that joint parameter tuning is essential. This challenges conventional manual hyperparameter tuning and provides both theoretical grounding and empirical evidence for automated parameter optimization. Experimental results show substantial correctness improvements under coordinated tuning; however, optimal configurations exhibit high sensitivity to parameter changes, underscoring the necessity of systematic, holistic parameter control.
This study investigates the feasibility of leveraging low-cost, deployable open-source large language models (ranging from 0.5B to 32B parameters) to automatically generate domain-specific language (DSL) representations of UI and data models directly from natural language prompts, without fine-tuning and using only few-shot prompting. It presents the first systematic evaluation of small-scale open-source models on the task of generating multiple, interrelated DSL artifacts. Through a combination of DSL grammar parsing, automated validation, and expert assessment, the work examines model performance in terms of syntactic correctness, semantic completeness, and cross-model referential consistency. Experimental results demonstrate that compact models—such as gemma3:12b and mistral:7b-instruct—achieve generation quality comparable to or even rivaling that of significantly larger models, highlighting their practical viability and cost-effectiveness for model-driven engineering applications.
Current large language model (LLM)-driven programming assistants struggle to accommodate the diverse, ambiguous, and open-ended interaction needs arising from developers’ varying cognitive styles and organizational contexts. This work presents the first framework that systematically integrates developer cognitive traits and organizational context into the personalized design of LLM-based programming assistants. By leveraging user modeling, dialogue systems, and human-computer interaction analysis, the authors develop a prototype conversational programming assistant capable of adaptive personalization. The resulting system demonstrates markedly enhanced inclusivity and practical utility, offering a novel paradigm for intelligent programming tools tailored to heterogeneous developer populations.
The mechanisms by which system prompts influence instruction-tuned models in code generation remain poorly understood, particularly regarding their performance across varying model scales, programming languages, and prompting strategies. This study conducts a large-scale, multi-variable controlled experiment evaluating 360 configurations spanning four instruction-tuned models, five categories of system prompts, three prompting strategies, two programming languages, and multiple temperature settings. The findings reveal that the effectiveness of system prompts is non-monotonic and highly configuration-dependent: notably, few-shot examples can degrade performance in larger models, challenging the conventional wisdom that few-shot prompting consistently outperforms zero-shot. Additionally, Java is found to be more sensitive to prompt design than Python, suggesting the need for language-specific prompting strategies. This work provides empirical foundations and practical guidance for effective prompt engineering in code generation.
This work addresses the limitations of current large language models in tool invocation, which rely heavily on in-context documentation and examples, leading to high inference overhead and susceptibility to hallucination, while conventional fine-tuning struggles to effectively internalize specific tool knowledge. To overcome these challenges, the authors propose ParaTool, a novel framework that shifts tool representation from the context into the model’s parameter space. Through a three-stage process—parameterized pretraining, gated network-driven soft tool selection, and parameterized joint fine-tuning—ParaTool enables dynamic, lightweight tool calling without dependence on contextual cues. Experiments demonstrate that this approach significantly outperforms strong in-context learning baselines on the Stable ToolBench and BFCL benchmarks, achieving higher tool invocation accuracy while reducing computational complexity.
This study addresses the lack of a systematic review on the application of large language models (LLMs) in software engineering documentation and modeling tasks. Through a comprehensive literature survey, it establishes a multi-dimensional taxonomy that categorizes existing research by task type, offering an in-depth analysis of key technical approaches—including prompt engineering, natural language understanding, and structured language processing. The work further synthesizes the distribution of tasks, evaluation metrics, human assessment methodologies, and commonly used datasets across major conferences in the field. By systematically mapping the research landscape and identifying prevailing technical trends, this paper provides a thorough reference and strategic guidance for future investigations at the intersection of LLMs and software engineering.