🤖 AI Summary
This study addresses the limitation of existing knowledge distillation-based watermarking methods in jointly preserving feature- and output-level information, which leads to incomplete inheritance by student models. We propose TwinMark, a framework that embeds secret messages into the second-order moments of normalized features and class-averaged logits to achieve dual-channel watermark protection. Furthermore, we introduce a novel "output-matching inheritance" mechanism that integrates second-order moments with class-reuse techniques to ensure synchronized watermark transfer across both feature and output layers. At the detection stage, linear readout, second-order moment decoding, and statistical hypothesis testing are employed. Experiments demonstrate that TwinMark achieves high bit recovery rates on benchmarks such as CIFAR-100 with negligible accuracy degradation, while exhibiting strong robustness across diverse architectures and tasks.
📝 Abstract
We propose TwinMark, a watermarking scheme that reads a single SHAKE128 secret through two complementary linear functionals of model-output summaries: a covariance projector against the carrier-set covariance (cov-Feat) and a class-conditional Fisher-aligned linear carrier decoded from class-mean logits (cc-FALC). The two readouts share one bit vector and cover the two extraction surfaces of a deployed vision model: a classifier API attacked by KL knowledge distillation (KD) (Std. KL-KD), and a representation-only host attacked by feature-matching KD (FM-KD). Each readout admits a teacher-measurable a posteriori certificate that lower-bounds post-distillation detection power, and the two channels combine under a regime-restricted OR rule whose test statistic (calibrated null or bit vote) is selected by the exposed surface. cov-Feat admits a rank-blind operator-norm certificate, cc-FALC admits a centered-logit-gap certificate that decouples bit capacity from class count: at K=1024 in m=100 classes (a 10.24x over-encoding), the bit-vote attains z=23.0 sigma at a teacher-accuracy cost of +0.9+-0.2%p. Across 13 attacks on CIFAR-10, CIFAR-100, and Mini-ImageNet, TwinMark verifies on every cell whose post-attack model retains task utility, survives cross-architecture distillation onto ResNet-18/50, VGG-16, and MobileNet-V3, and ports to GNSS few-shot, VOC detection, ISIC segmentation, and STL-10 SimCLR.