DynGraphAgentBench: A Benchmark for Agentic Lifecycle Control in Dynamic Graph Anomaly Detection

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of agent lifecycle control in dynamic graph anomaly detection arising from structural drift and delayed feedback by constructing an executable benchmark framework. Methodologically, it introduces strict leakage-proof validation and temporal causality constraints, enabling a controller to adaptively select detectors based solely on historical context and release results with delay following sandbox training. Leveraging a multi-detector architecture and a deterministic validator, the framework quantifies the utility, cost, and characteristics of ineffective responses across different strategies through complete trajectory analysis. This work establishes a rigorous experimental paradigm for evaluating agents' online adaptive capabilities under non-oracular conditions.
📝 Abstract
Dynamic graph anomaly detection requires repeated decisions as graph structure and class prevalence drift, yet detector benchmarks usually score a fixed pipeline after current labels are known. We introduce DynGraphAgentBench, an executable benchmark for agentic lifecycle control under delayed feedback. It comprises seven temporal graph datasets with node- and edge-level anomaly tasks, eleven selectable detectors, and eight chronological deployment windows per dataset. In each window, a controller sees only time-causal aggregate context, registered model cards, and its own matured history. It must choose a detector before current-window training or candidate scores exist. A sandboxed executor trains the chosen architecture on mature data, scores a hidden deployment window, and releases the outcome after a one-window delay. A deterministic verifier checks decision timing, leakage guards, legal actions, training scope, and persisted artifacts. We measure detection utility with average precision and capture at fixed review depth, and characterize adaptation through model switches and compute. Complete eight-window trajectories from two primary controllers and a no-memory reference on four datasets, together with three additional controllers on three datasets, expose useful, costly, and ineffective reactions to delayed evidence without granting an exhaustive current-window oracle.
Problem

Research questions and friction points this paper is trying to address.

Dynamic graph anomaly detection
Agentic lifecycle control
Delayed feedback
Benchmark
Concept drift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Graph Anomaly Detection
Agentic Lifecycle Control
Delayed Feedback
Benchmark
Sandboxed Executor
🔎 Similar Papers
💼 Related Jobs
No related jobs found.