🤖 AI Summary
This study addresses the challenges of structural sparsity, supervision scarcity, and backbone mismatch in black-box hard-label stealing attacks against graph neural networks (GNNs). We propose a two-stage decoupled attack framework operating under strict constraints. The first stage performs pre-training by decoupling information propagation from manifold-level node mixup augmentation. Subsequently, with the encoder frozen, the classification head is fine-tuned using class-balanced sampling and logit adjustment. This work pioneers a decoupled stealing paradigm that effectively overcomes the obstacles posed by isolated nodes, class imbalance, and architectural discrepancies between surrogate and victim models. Extensive experiments across four benchmarks demonstrate that our method surpasses state-of-the-art approaches, improving fidelity by 18.16% while requiring only 1/12.23 of the query budget consumed by the strongest baseline.
📝 Abstract
As Graph Neural Networks (GNNs) are widely deployed as Machine Learning-as-a-Service (MLaaS) APIs, model stealing attacks have emerged as a critical security threat. By querying a victim model's black-box API, an adversary can construct a functionally equivalent surrogate model, compromising proprietary intellectual property and downstream security. Existing GNN stealing attacks, however, rely on overly permissive assumptions, such as soft-label outputs, large query budgets, full-graph query access, and prior knowledge of victim backbones that rarely hold in real-world deployments. In this work, we formalize a strictly constrained black-box, hard-label and backbone-agnostic threat model for GNN stealing attacks under a tight query budget. Given these realistic restrictions, we identify four fundamental challenges: sparse local structures and isolated nodes that degrade victim label quality, insufficient supervision signals, systematic imbalance with incomplete class coverage, and backbone mismatch. To address these interlocking barriers, we propose Dagger, a novel two-phase decoupling-based attack framework. Specifically, in Phase 1, Dagger pre-trains a surrogate using decoupled information propagation to preserve structural context over sparse local subgraphs while handling isolated nodes, combined with manifold-level node mixup to synthesize continuous supervision signals and smooth decision boundaries. In Phase 2, Dagger freezes the encoder and fine-tunes the classifier head via class-balanced sampling paired with logit adjustment to rectify severe query imbalance without requiring extra victim queries. Extensive experiments across four benchmark graphs and four GNN backbones demonstrate that Dagger consistently outperforms state-of-the-art GNN stealing attacks, achieving up to 18.16\% higher fidelity while only utilizing 12.23$\times$ fewer queries than the strongest baseline.