Score
Designs, implements, documents, tests, and maintains application programming interfaces (APIs) that define how software components communicate — specifying endpoints, request/response schemas or types, authentication/authorization, error handling, versioning, and performance or reliability constraints. Builds and analyzes API contracts and specifications (e.g., OpenAPI/GraphQL), client SDKs, integration tests, backward-compatibility strategies, and operational controls such as rate limiting, monitoring, and deployment.
Current RESTful API design quality assessment relies heavily on manual inspection, lacking early, automated validation mechanisms for non-functional requirements—particularly interoperability, modularity, and maintainability. Method: This paper proposes an OpenAPI-based static analysis approach that implements a configurable rule engine. It formalizes 75 design principles derived from scholarly literature and industry standards into structured, machine-checkable constraints, enabling customizable rule activation/deactivation and traceable feedback to align requirements engineering with architectural governance. Contribution/Results: Following the design science research paradigm, we developed and evaluated a prototype tool. Empirical evaluation and expert review demonstrate that the method significantly improves API design compliance and consistency, achieving 82% automation coverage. It effectively supports continuous architectural governance in agile development environments, bridging the gap between design-time assurance and operational API lifecycle management.
This work addresses the challenge of reliably conveying intent, requirements, and constraints in human–AI–tool collaborative software development by proposing a specification-centric Bosque API (BAPI) ecosystem. The system introduces a highly expressive specification language that, for the first time, enables cross-language interoperability, automated test generation, formal verification, and execution sandboxing across the entire API lifecycle—from requirement definition and implementation to invocation and validation. By providing end-to-end specification guarantees, BAPI significantly enhances system correctness, security, and the efficiency of human–AI collaboration, offering a novel infrastructure for software development in the era of AI agents.
This work addresses the challenges of manual Web API integration testing, which is time-consuming, error-prone, and often misaligned with business requirements. The authors propose a novel approach that synergistically combines large language models (LLMs), retrieval-augmented generation (RAG), and prompt engineering to jointly parse natural language business requirements and OpenAPI specifications, thereby automatically generating executable test scripts that are both semantically meaningful and syntactically correct. Evaluated on ten real-world APIs, the method successfully produced valid tests for 89% of the business requirements within three attempts, uncovered multiple previously unknown integration defects, and substantially reduced the manual effort required for test development.
REST API documentation frequently suffers from incompleteness, obsolescence, or inaccessibility, hindering both automated testing efficiency and human comprehension. This paper introduces the first LLM-driven, end-to-end framework for OpenAPI specification inference and black-box API testing—requiring only an API name and an LLM API key. It automatically generates and mutates HTTP requests, then infers specifications and detects defects via response analysis. A novel context-aware prompt masking strategy enables zero-shot discovery of undocumented routes and parameters without model fine-tuning. Evaluated on a standardized benchmark, the framework achieves 85.05% average recall for GET routes and 81.05% for query parameters, successfully uncovering hidden endpoints and diverse server-side errors (e.g., 5xx, logic flaws). The inferred OpenAPI specifications are directly compatible with mainstream API testing tools, enabling seamless integration into existing CI/CD and security validation pipelines.
This work addresses the lack of effective automated testing mechanisms in microservice architectures, where existing API specifications such as OpenAPI suffer from limited semantic expressiveness and thus struggle to support high-coverage automated validation. To overcome this limitation, the authors propose APOSTL—an extension of OpenAPI grounded in restricted first-order logic—that enables formal annotation of semantic properties of APIs. Complementing this specification language, they develop PETIT, a tool that performs fully automated, source-code-free black-box testing using only APOSTL-annotated OpenAPI documents. By embedding formal logic directly into API specifications for the first time, this approach allows interface documentation to drive semantically precise and high-coverage automated tests, significantly enhancing the efficiency and reliability of microservice verification.
This study addresses a critical gap in black-box testing research, which commonly assumes the correctness of OpenAPI specifications while overlooking how their defects impact testing efficacy. The authors propose the first taxonomy of OpenAPI faults, derived from literature, categorizing them into six types, and systematically inject these faults across five severity levels. Using EvoMaster, RESTler, and Schemathesis, they evaluate the effects on two microservice benchmarks through multidimensional metrics—including code and specification coverage, request/response quality, and behavioral diversity. Their findings reveal heterogeneous degradation patterns: method semantics–related faults cause comprehensive performance deterioration, whereas response code modifications have negligible impact. Notably, relaxing schema constraints substantially degrades request/response quality without affecting coverage metrics, demonstrating that reliance solely on coverage can obscure critical quality issues.
This study addresses the limitation that analyzing prompts alone is insufficient for comprehensively evaluating developer interactions with AI programming agents. To overcome this, we propose a novel multidimensional interaction analysis framework termed "Say-Do-Understand," which integrates prompt data, screen activity, and comprehension metrics through a systematic five-stage end-to-end workflow. Employing an observational methodology, the analysis utilizes a prompt codebook, a screen activity coding scheme, and dual scoring rubrics. An empirical study involving ten experienced developers validates the proposed approach. Furthermore, four developer personas synthesizing task performance and comprehension levels are introduced to elucidate behavioral variations. Notably, the findings reveal that excessive reliance on agent self-checking significantly reduces developers' autonomous testing time.
Existing data pipelines often suffer from weak governance, leading to delayed schema validation, inconsistent cross-language execution, and misalignment with business semantics. This work proposes treating data contracts as types, leveraging the “everything-as-code” paradigm to inject schema annotations—encompassing column types, constraints, documentation, and lineage—into input and output tables within a lakehouse architecture via multi-language SDKs. These annotations are parsed across multiple phases of the execution lifecycle, deeply integrating data contracts into the type system. The approach enables both deterministic and non-deterministic reasoning over data flows across languages and execution engines, significantly enhancing the reliability of production data pipelines and ensuring consistent interoperability across systems.
本文介绍了一种名为RESTCov的工具,它通过解析OpenAPI规范和HTTP请求/响应日志来分析REST API的结构覆盖情况,解决了传统方法难以应用于分布式API的问题。
本文提出APIPilot框架,通过执行验证LLM推断的依赖关系并基于响应调整,生成有效的REST API测试序列,提高测试覆盖率和成功率。