GrayShield: Bit-Level Sanitization for Transformer Model Supply-Chain Security

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the threat to AI supply chain security posed by the exploitation of least significant bits (LSBs) in Transformer model weights as covert channels. We propose a lightweight, zero-data post-training sanitization method that employs Gray code-guided low-transition sequences and keyed tensor phase perturbation to perform fully payload-agnostic bit overwriting on LSBs, fundamentally eliminating hidden space and reducing declared channel capacity to zero. Experimental results demonstrate that our approach decreases malicious payload recovery rates by nearly 50 percentage points to the level of binary random guessing while maintaining accuracy degradation below 1%. The proposed method significantly outperforms existing baselines and substantially reduces weight distribution shift.
📝 Abstract
Transformer models such as BERT and Vision Transformer~(ViT) achieve strong performance via densely parameterized attention backbones. However, the least significant bits~(LSBs) of their 32-bit floating-point weights can be abused as covert channels to conceal malicious payloads, posing a serious threat to the AI model supply chain. We propose \GS (\GSabbr), a lightweight, post-training, zero-data sanitization method that completely replaces the declared mantissa-LSB channel with a Gray-code-guided low-transition sequence. Complete payload-independent overwrite, whether keyed or public, makes the sanitized target bits independent of the embedded payload and gives that declared channel zero capacity. Gray coding supplies overwrite structure, while a keyed per-tensor phase supplies pattern diversity. Benchmarked against seven post-training defenses on four Transformer model presets and two real-world malware payloads, \GSabbr maintains sub-$1\%$ accuracy impact and achieves $49.96\pm0.66$ percentage-point Recovery Reduction (RR) under five implemented attacker variants. Because pre-defense recovery is effectively $100\%$, RR near 50 percentage points corresponds to post-sanitization bit accuracy at binary chance. Its main empirical advantage is stable near-chance sanitization with substantially smaller weight-distribution shift than the evaluated near-chance baselines PatternMask (PM) and Post-Training Quantization (PTQ).
Problem

Research questions and friction points this paper is trying to address.

Transformer models
supply-chain security
covert channels
malicious payloads
least significant bits
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bit-Level Sanitization
Supply-Chain Security
Gray Code
Covert Channel
Zero-Data
🔎 Similar Papers
No similar papers found.
A
Armstrong Foundjem
Mila - Quebec Artificial Intelligence Institute and Polytechnique de Montréal, QC, Canada
T
Tsung-Hsien Chuang
Research Center for Information Technology Innovation, Academia Sinica, Taipei, Taiwan
Foutse Khomh
Foutse Khomh
NSERC Arthur B. McDonald Fellow, CRC Tier 1, Canada CIFAR AI Chair, FRQ-IVADO Chair, Full Professor
Software engineeringMachine learning systems engineeringMining software repositoriesReverse
M
Mohamed Amine Merzouk
Mila - Quebec Artificial Intelligence Institute and McGill University, Montréal, QC, Canada