Score
Designs, implements, tests, deploys, and maintains interactive applications delivered over the web, encompassing client-side user interfaces, server-side logic and APIs, data persistence, session and state management, and integration with web protocols, authentication, and security controls. Involves building scalable, performant, and accessible frontends and backends (including real-time features and persistence), and establishing deployment, monitoring, and CI/CD practices to operate and evolve the application.
Web application testing (WAT) faces fundamental challenges including dynamic content, asynchronous interactions, and cross-environment compatibility. This paper presents a systematic literature review of WAT research from 2014 to 2024. Methodologically, it introduces the first comprehensive taxonomy—categorizing approaches into model-based, crawler-based, AI-driven, and security-oriented testing—and conducts a longitudinal comparative analysis. Integrating bibliometric analysis, methodological synthesis, and empirical evaluation of widely adopted tools, it constructs a decade-long evolutionary map of WAT. Furthermore, the work proposes a standardized evaluation framework and identifies key open problems, notably asynchronous behavior modeling and cross-platform reliability validation. The contributions provide an authoritative, evidence-based reference for industrial tool selection and academic research direction setting, bridging practice and theory in modern WAT.
This work addresses the fragility, inefficiency, and strong platform coupling commonly found in CI/CD pipelines for legacy COBOL systems, which often result in high maintenance costs and vendor lock-in. To overcome these challenges, the authors propose a portable CI/CD architecture tailored for highly secure and compliance-driven environments. The approach leverages OCI-compliant container images preloaded with COBOL toolchains, introduces a platform abstraction layer, integrates multiple repositories, and employs Groovy script refactoring to achieve platform-agnostic continuous integration and delivery. Empirical evaluation demonstrates that the proposed solution significantly enhances efficiency—reducing pipeline execution time by 82%—while simultaneously improving system portability, security, and maintainability. This architecture offers a reusable paradigm for modernizing legacy COBOL applications within regulated domains.
Current code-generation agents fail to meet functional requirements in over 70% of web application scenarios, primarily due to the absence of automated deployment, browser-level validation, and a feedback loop for iterative repair. This work proposes TDDev, a novel framework that establishes the first fully automated test-driven development (TDD) pipeline by integrating structured acceptance test generation, browser-interaction simulation for validation, and failure-triggered repair signals within a multi-agent collaborative architecture. Empirical evaluation demonstrates that TDDev improves generation quality by 34–48 percentage points, and user studies confirm it eliminates the need for manual intervention entirely. Furthermore, the study reveals that alignment between model generation style and TDD strategy is critical: misalignment not only nullifies performance gains but also inflates token consumption by up to 25-fold.
This study addresses the lack of empirical evidence on the real-world impact of UI testing frameworks in CI/CD pipelines. Using GitHub API data collection, YAML configuration parsing, CI log metric extraction, and controlled time-series analysis across open-source repositories, we systematically quantify the integration patterns and effects of Selenium, Playwright, and Cypress within GitHub Actions workflows. Results show that UI testing significantly improves test pass-rate stability but increases mean build duration by 12% initially. Highly active repositories prefer Playwright—its built-in retry mechanism reduces flaky-test-induced pipeline interruptions by 35%. This work fills a critical gap in understanding UI testing’s practical implications in production CI/CD environments, providing data-driven insights for quality assurance strategy design and framework selection.
Answer Set Programming (ASP) developers traditionally rely on external frontend technologies to construct user interfaces, creating a barrier to rapid prototyping and tight integration of declarative logic with interactive behavior. Method: This paper introduces Clinguin, the first system enabling native interactive UI modeling directly within ASP. It extends ASP with declarative UI predicates—such as `ui/2` for interface structure and `on_event/3` for event handling—unifying UI specification and interaction logic under ASP’s logical rule formalism. Built as an extension of Clingo, Clinguin compiles ASP source code directly into dynamic, browser-executable web interfaces. Contribution/Results: Clinguin significantly lowers the UI development threshold for ASP applications: interactive prototypes can be generated from just a few lines of ASP code. The system is fully integrated into the Clingo toolchain and publicly available as open-source software.
论文提出Agent-Integrated Software模式和Intent-Level Interaction Abstraction方法来解决智能代理与现有应用程序集成时的协调问题,确保任务级交互与应用行为的一致性。
为解决现有评估方法的局限性,提出IWC-Bench,通过代码覆盖率引导交互式探索生成的Web应用,并从视觉美感、可用性和需求一致性三个维度进行评价。
This study addresses the limitations of existing web agent evaluation frameworks, which rely on visual scraping or judge models and consequently suffer from inaccurate verification and unquantifiable interface sensitivity. To overcome these issues, this work proposes an open-source, event-log-based evaluation framework that introduces a novel event-level verification mechanism. By leveraging typed log recording and DOM-controlled variant re-rendering, the method achieves precise, judge-free, and scrape-free task assessment across six simulated websites. Furthermore, it supports single-configuration toggling to generate UI variants for quantifying interface sensitivity. The project releases 152 tasks alongside a public leaderboard, revealing a significant discrepancy of up to 41 points between agents' self-reported completion rates and actual performance measured via event logs.
本文探讨了软件形式的第三次重构,通过引入通用数据库、大模型和代理来解决传统三层架构的问题。
研究评估了三种基于LLM的智能IDE(Copilot、Cursor和Windsurf)在从零开始生成五个全栈Web应用时的表现,发现它们在常见模式生成上成熟度高,但在分布式架构生成中错误较多。