DreamTest: World-Model Surrogates for Search-Based Testing of Deep Reinforcement Learning Agents

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high testing costs of deep reinforcement learning agents and the limitation of traditional surrogate models that overlook execution dynamics. We propose DreamTest, a world model-based search method for agent testing. Departing from black-box prediction paradigms, DreamTest leverages a Recurrent State Space Model (RSSM) to learn environment and agent dynamics from training logs, innovatively modeling the test rollout process. It generates failure scores through imagined episodes to guide efficient search, thereby uncovering diverse failures. Experimental results demonstrate that DreamTest improves AUPRC by up to 97% across multiple benchmark environments. Under equivalent evaluation budgets, it significantly increases the number of novel failures discovered and broadens the coverage of behavioral clusters.
📝 Abstract
Testing deep reinforcement learning (DRL) agents in cyber-physical systems aims to uncover diverse failures before deployment, but each execution can be expensive. Surrogate-assisted testing reduces this cost by learning to predict which test configurations are likely to fail. Prior surrogates treat the system as a black box and predict pass or fail outcomes directly; we instead model how a test unfolds and estimate failure from an imagined episode. We introduce DreamTest, a world-model surrogate for testing DRL agents. DreamTest adapts a recurrent state-space model to learn agent behaviour and environment dynamics from the agent's training log. Given a candidate configuration, imagined rollouts produce a failure score that guides search without executing every candidate in a simulator or real system. We evaluate DreamTest for failure prediction, test generation, and failure diversity on Parking, Humanoid, and DonkeyCar. Mean area under the precision-recall curve (AUPRC) exceeds the strongest baseline by 97%, 12%, and 39%, respectively, and gains on five out-of-distribution test sets reach 145%, 29%, and 44%. Under the same simulator-validation budget, the best"DreamTest + search"combinations find 29%, 22%, and 79% more novel failures on average. Across clusterings with k = 2-40, failures generated with DreamTest cover the most behavioural clusters for almost all k, indicating that DreamTest consistently discovers behaviourally diverse failures.
Problem

Research questions and friction points this paper is trying to address.

Deep Reinforcement Learning
Surrogate-assisted Testing
World Model
Failure Discovery
Cyber-Physical Systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Model Surrogate
Deep Reinforcement Learning Testing
Recurrent State-Space Model
Imagined Rollouts
Search-Based Testing
🔎 Similar Papers
No similar papers found.