🤖 AI Summary
This study addresses a critical pitfall in evaluating the robustness of quantum attention models, demonstrating that normalization modules can conflate attack budgets with label preservation, thereby misleading conclusions. To investigate this, the authors conduct an audit based on exact four-qubit simulation, quantum Fisher information regularization, and synthetic power grid trajectory data. The methodology involves matching encoder perturbation upper bounds, verifying the causality of interventions, and distinguishing exploratory from confirmatory evidence. The work reveals the distorting effect of normalization on robustness assessments under fixed physical budgets, showing that single comparisons cannot establish module benefits and that classical baselines outperform quantum counterparts on clean predictions. Ultimately, this paper contributes rigorous auditing protocols that explicitly decouple perturbation budgets from label consistency to ensure valid evaluations of quantum model robustness.
📝 Abstract
Removing an input-scaling module changes both a classifier and the perturbations reaching its encoder. A robustness difference can therefore reflect the comparison rule as well as the module. We demonstrate this problem in a four-qubit quantum-attention detector on generated power-grid trajectories. A learned scaling module appears beneficial at a fixed physical attack budget, but matching an upper bound on perturbations at the encoder reverses the ordering. Neither comparison alone establishes a robustness benefit caused by the module. The initial test also perturbs clean examples into attacked examples while retaining their original labels; tests restricted to already attacked examples do not establish a benefit. Replacing a trained model's input scales disrupts detection. Retraining its linear classification layer restores the detection rate, but changes individual predictions, leaving the comparison descriptive rather than causal. Two further design checks explain why the input quantum Fisher information regularizer cannot train this model's query parameters, and why removing confidence bounds does not establish a larger certified radius. The evidence is limited to ten seeds, exact simulation, synthetic data, and a restricted set of attacks; classical baselines achieve better clean prediction. The practical lesson is to specify which perturbation budget is fixed, check that attacks preserve labels and interventions preserve predictions, and distinguish exploratory controls from confirmatory evidence.