Score
Designs and implements command-line interfaces and their tooling, specifying option and flag semantics, argument parsing and validation, help and usage messages, and configuration-file integration; ensures CLI behavior, prompts, and output support usability and accessibility requirements.
To address the high learning barrier and inefficient documentation lookup associated with command-line interfaces (CLIs), this paper introduces GUIde—the first AI-powered system that automatically transforms unstructured Unix manual (man) pages into interactive graphical user interfaces (GUIs). Methodologically, GUIde leverages natural language processing and large language models to semantically parse man pages, extract parameter schemas, constraints, and usage examples, and generate executable interface specifications; a lightweight frontend engine then renders these specifications into responsive, visual GUIs in real time. Its core contribution is the first end-to-end, fully automated mapping from non-structured man text to functionally complete GUIs—requiring no manual annotation or tool-specific adaptation. Evaluated on a real-world corpus of 52 common Unix commands, GUIde achieves 96.3% coverage of valid parameter combinations and improves user task completion rates by 41.7%, effectively bridging the interaction paradigm gap between CLI and GUI environments.
Existing CLI modeling datasets lack execution feedback—such as exit codes, stdout/stderr outputs, and environmental side effects—hindering fine-grained behavioral simulation. To address this, we propose ShIOEnv, the first sandboxed environment enabling full interactive CLI modeling, formalizing command generation as a Markov Decision Process (MDP). We introduce a syntax-guided mechanism: automatically constructing context-free grammars from man pages to dynamically mask invalid arguments, and designing a grammar-masked Proximal Policy Optimization (PPO) algorithm to improve data sampling efficiency and quality—especially for small models. Leveraging this framework, we curate a high-quality shell input-output trajectory dataset. Fine-tuning CodeT5 yields +85% BLEU-4 improvement under syntactic constraints and +26% under PPO optimization over baselines, significantly enhancing small models’ faithfulness and generalization in CLI behavioral simulation.
This work addresses the challenge of reliably conveying intent, requirements, and constraints in human–AI–tool collaborative software development by proposing a specification-centric Bosque API (BAPI) ecosystem. The system introduces a highly expressive specification language that, for the first time, enables cross-language interoperability, automated test generation, formal verification, and execution sandboxing across the entire API lifecycle—from requirement definition and implementation to invocation and validation. By providing end-to-end specification guarantees, BAPI significantly enhances system correctness, security, and the efficiency of human–AI collaboration, offering a novel infrastructure for software development in the era of AI agents.
本文解决了自然语言需求转为GUI原型及验证的问题,通过预训练语言模型实现自动化原型生成与需求验证,减少了人工成本。
Automated verification of interactive console I/O programs in Haskell education remains challenging due to the dynamic, history-dependent nature of student implementations. Method: We propose a lightweight, formal behavioral specification language that uniquely integrates global state and execution history, expressed via regex-like syntax; its trace-based semantics enable probabilistic testing and scalable verification through *sampleable validity*. Contribution/Results: Our system automatically validates student submissions against behavioral specifications and supports pedagogical closed-loop applications—including real-time feedback generation, example solution synthesis, and exercise randomization. Empirical evaluation demonstrates substantial improvements in test coverage and pedagogical adaptability while preserving formal rigor. To our knowledge, this is the first framework for verifying interactive behaviors in functional programming education that simultaneously achieves theoretical soundness and practical deployability.
Developing cross-platform graphical user interfaces (GUIs) and plugins for command-line tools in structural bioinformatics is often costly and complex. This work proposes a three-stage automated workflow that leverages a platform-agnostic formal GUI specification, decouples model, view, and presenter components through the Model–View–Presenter (MVP) architectural pattern, and employs a dedicated code generator to automatically produce plugins for target platforms—namely VMD, PyMOL, and the web. To the best of our knowledge, this is the first systematic application of the MVP pattern to the automatic GUI generation for CLI tools, substantially enhancing logic reusability, cross-platform portability, and development efficiency. The framework’s generality, extensibility, and practical utility are demonstrated by successfully generating plugins for multiple tools from the Structural Bioinformatics Library across all three platforms.
Software testing is a fundamental process of software development, and prior work has shown that visualizations of test results support testers' decision-making. However, Human-Computer Interaction research on software testing has yet to explore and understand the shared interface elements and patterns in visualization of testing outputs. To address this, we conducted a visual comparative analysis of the output of 50 software testing tools and harnesses (44 with CLI output, 6 with GUI output) across four popular programming languages. Our analysis reveals the common interface elements in software testing tools, how these tools display and visualize test results, as well as the specific make-up of the output. Our findings provide insight on how visual testing output is formatted and how colour is used across both CLI and GUI environments, identifying trends that can be applied by developers of testing tools.
该文介绍了一种使用大型语言模型辅助开发静态验证软件的工具Eiffel-tools,通过与语言服务器协议结合,并利用形式化验证器提高代码生成和修正的准确性。
This work addresses the security risks inherent in large language model (LLM) agents executing shell commands, where existing safeguards struggle to balance safety, efficiency, and accuracy. The authors propose CARE, a novel “static-first” verification framework that performs static analysis prior to command execution by integrating syntactic, semantic, path-contextual, and risk-pattern information. LLM-based arbitration is invoked only when uncertainty remains high. Through command normalization, deterministic evidence inference, and selective neural judgment, CARE achieves efficient, reproducible, and auditable security control. Experimental results demonstrate that CARE attains an F1 score of 85.64% on the main test set with a false positive rate of merely 0.91% and an average latency of 2.32 milliseconds; in purely static mode, latency drops to 0.34 milliseconds, and RedCode-gen harm is reduced to 37.33%.
Current prompt engineering practices in AI programming assistants lack systematicity and struggle to effectively integrate requirements engineering principles, resulting in a gap between user intent and code implementation. This work introduces a requirements engineering perspective into prompt design, proposing a conceptual “prompt triplet” model that treats prompts as lightweight, evolvable artifacts integrating functional and quality requirements, general solution strategies, and concrete implementation details. Through conceptual modeling, analysis of real-world prompt corpora, dataset construction, and controlled experiments, the study provides preliminary validation of the three proposed components and formulates four testable hypotheses concerning prompt evolution, user variability, requirements validation, and code quality. These contributions lay an empirical foundation for advancing prompt engineering toward a more disciplined, requirements-driven paradigm.