Score
Designs and implements Android mobile applications, including user interfaces, notification/alerting, data storage, and wireless connectivity (e.g., Bluetooth or Wi‑Fi). Builds and integrates instrumentation and automation into Android apps to simulate user interactions, automate UI event sequences, capture network traffic and traces, and trigger feature-specific flows for testing, monitoring, or debugging.
This work addresses the challenges posed by the event-driven nature and development diversity of Android applications, which hinder static analysis from constructing complete control-flow models, while existing dynamic testing approaches suffer from low efficiency and insufficient coverage. To overcome these limitations, the paper proposes DroidGraph, a framework that integrates static code analysis with systematic exploration to build a unified control-flow model bridging low-level method invocations and high-level UI structures, thereby enabling efficient automated test generation. Experimental evaluation on 19 real-world applications demonstrates that DroidGraph achieves 18% higher coverage of application content with 345 fewer interactions on average compared to random exploration, and additionally uncovers 51 previously missed components and 49% more UI callback links.
This study addresses the challenges of unstable end-to-end testing for Android applications in continuous integration (CI) due to fragile emulator configurations. It presents the first large-scale empirical analysis of 4,518 open-source projects, systematically examining how instrumentation tests are configured, how these practices evolve, and their comparative effectiveness in CI environments. Leveraging GitHub Actions metadata, the work evaluates three prevalent approaches: Gradle Managed Devices, community-reusable components, and custom scripts. Findings reveal that only 10.6% of projects adopt such testing; among them, community components demonstrate superior reliability and efficiency, third-party device labs are suitable for regression testing despite higher costs, and custom scripts, while flexible, suffer from high retry rates. The study thus illuminates current practices and critical trade-offs in Android CI testing.
This study addresses the limited effectiveness of Android record-and-replay (R&R) tools in realistic debugging scenarios—including non-crashing functional defects, crashing vulnerabilities, and functional-level user workflows. We conduct the first cross-industry and cross-academic empirical study, systematically evaluating the robustness of mainstream R&R tools using a benchmark comprising 34 representative scenarios, 90 non-crashing defects, and 31 crashing defects, augmented with automated input generation (AIG). Our analysis identifies three fundamental bottlenecks: inaccurate timing precision between UI actions, restricted API compatibility, and constraints imposed by Android’s underlying system architecture. Results show that, on average, 44% of crashing defects, 38% of non-crashing defects, and 17% of user scenarios fail stable replay. These findings provide critical failure root-cause insights and empirically grounded design guidelines for next-generation high-reliability UI testing tools.
Android GUI event sequence generation faces dual challenges: low coverage and difficulty reaching target paths. Traditional model-driven approaches effectively cover GUI states but struggle to generate inputs satisfying complex data dependencies. This paper proposes a synergistic analysis framework integrating concolic execution with GUI modeling—the first to jointly leverage concolic execution for both event sequence construction and GUI traversal guidance. By employing dynamic symbolic execution to identify path constraints, and combining it with a GUI model for event prioritization and data-dependency–aware state transition analysis, the framework systematically explores GUI navigation paths. Experimental evaluation demonstrates significant improvements in code coverage and target-path reachability across multiple benchmark applications, outperforming state-of-the-art model-driven tools (e.g., A3E, DroidBot). The results validate the framework’s advantages in robustness, precision, and systematic GUI exploration.
Mobile test scripts suffer from opaque intent and loose coupling with application logic, leading to high maintenance costs. Method: This paper proposes the first end-to-end joint modeling approach that semantically aligns UI screenshots with source code: it parses interface images via OCR and object detection, identifies responsive code through AST analysis and method localization, and establishes an operation-sequence modeling framework with cross-modal intent alignment to tightly couple widget selectors with functional code; an encoder-decoder model then generates natural-language intent descriptions. Contribution/Results: Evaluated on a real-world app dataset, our method achieves a BLEU-4 score of 58.3%, significantly outperforming all baselines. A user study shows developers’ script comprehension time decreases by 62% on average. This work pioneers coordinated intent inference bridging GUI visual elements and code semantics, establishing a novel paradigm for enhancing test script understandability and maintainability.
研究通过DroidTool框架自动生成工具,结合GUI动作和工具动作提升安卓代理性能,减少人工成本,提高测试覆盖率。
GraphDroid通过历史感知探索和混合意图实现,解决了移动端应用GUI测试中覆盖复杂功能的问题,提高了测试效率与代码覆盖率。
This work addresses the inefficiency and limited coverage of traditional mobile application testing tools, such as Exerciser Monkey, which rely solely on random input generation due to a lack of understanding of application structure. To overcome these limitations, the authors propose Monkey++, a novel approach that integrates a control-flow model of Android applications into the Monkey framework. By modeling interactive UI elements as nodes in a structural graph, Monkey++ replaces random event generation with a model-driven depth-first search strategy. Experimental results demonstrate that this method achieves full coverage of user-interactable elements and improves testing speed by an order of magnitude compared to the original Monkey tool, significantly enhancing both test efficiency and coverage.
This work addresses the challenge of detecting non-crashing functional defects in mobile applications, which often arise due to the absence of explicit test oracles. To tackle this issue, the paper proposes PropGen, a novel approach that integrates large language models with function-guided exploration to establish a closed-loop pipeline from behavioral observation to precise property generation. PropGen leverages behavioral evidence collection, automated property synthesis, and test-feedback-driven refinement to iteratively improve property accuracy. Evaluated on 12 real-world Android applications, PropGen successfully identified 1,210 valid functionalities, generated 912 effective properties, corrected 118 imprecise ones, and uncovered 25 previously unknown functional defects. These results demonstrate a significant advancement in both the effectiveness and practicality of automated property generation for mobile app testing.
This study addresses the high development costs and significant code redundancy associated with traditional institutional mobile applications that rely heavily on native Android development. To overcome these limitations, the authors propose a full-stack solution leveraging a Django backend and an HTMX frontend, integrated via a WebView bridge to deliver a campus management system without writing any Android SDK code. The system supports core functionalities including task scheduling, inventory management, and attendance tracking, and is deployed using a self-hosted Docker Compose setup, eliminating dependence on external cloud services. Evaluated in a real-world institutional setting, this approach demonstrates for the first time that HTMX combined with Django can effectively replace conventional APK-based development, achieving a 54% reduction in development time, a 91% decrease in HTTP payload size, and a user satisfaction score of 4.2 out of 5.0 among 42 participants.