🤖 AI Summary
This study addresses the limitation of existing large language model (LLM) safety auditing methods, which rely on prompt engineering and struggle to characterize internal malicious representations. We propose an interpretable, mechanism-based detection approach that locates latent safety knowledge via hierarchical linear probing. A three-factor importance scoring scheme combined with a knee-point selection method is employed to identify critical layers and discriminative units for training a lightweight classifier. During inference, pretrained parameters remain frozen, enabling efficient operation with minimal samples. Experimental results demonstrate that our method achieves superior accuracy across three pretrained models, exhibiting strong cross-model generalization and robust few-shot detection capabilities.
📝 Abstract
Detecting malicious smart contracts is essential to safeguarding the Web3.0 ecosystem. However, existing auditing methods based on large language models (LLMs) largely rely on prompt engineering or attack-specific fine-tuning, while the internal representations that support malicious logic detection remain poorly characterized. To address this gap, we propose ContractLens, a framework that localizes and uses latent security knowledge in pretrained LLMs for interpretable malicious smart contract detection. Specifically, we first use layerwise linear probes to identify candidate layers that distinguish malicious from benign contracts, then select critical layers based on their sensitivity to malicious logic removal and stability under behaviour-preserving surface modifications. Next, within these critical layers, we combine activation contrast, malicious-logic specificity, and probe weights into a tri-factor importance score and adaptively select malicious contract-discriminative units (MCDUs) using the knee point of the ranked score curve. Finally, we concatenate the selected units'activations into a compact representation and train a lightweight classifier for detection, keeping the pretrained model frozen throughout. Neuron masking experiments show that masking MCDUs causes larger drops in detection performance than masking randomly selected neurons or neurons matched in activation magnitude. ContractLens achieves higher average detection accuracy than the evaluated methods on all three pretrained models. Further evaluations demonstrate generalization to unseen malicious intents and show that, with the identified MCDUs fixed, a classifier trained on few labelled samples can still achieve strong detection performance.