Breaking BERT: Gradient Attack on Twitter Sentiment Analysis for Targeted Misclassification

📅 2025-04-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work exposes the adversarial vulnerability of BERT in Twitter sentiment analysis and proposes a white-box, gradient-driven targeted adversarial attack. To address the limitations of conventional token-ranking strategies (e.g., based on frequency or heuristic importance), the method introduces a novel gradient-sensitivity-based token-level prioritization mechanism, enabling fine-grained and interpretable perturbation localization. It further integrates synonym-constrained token substitution with semantic consistency optimization to preserve fluency and meaning. Evaluated on SemEval and Sent140, the attack achieves over 92% targeted misclassification rates, modifies only 1.8 tokens per instance on average, and attains 91% human-rated naturalness—substantially outperforming existing baselines. The approach uniquely balances attack efficacy, stealth, and interpretability, establishing a new paradigm for robustness evaluation and defense-aware analysis of transformer-based sentiment models.

Technology Category

Machine Learning: Adversarial Learning & RobustnessComputer Vision: Adversarial Attacks & RobustnessNatural Language Processing: Safety and Robustness

Application Category

User Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Social media platforms like Twitter have increasingly relied on Natural Language Processing NLP techniques to analyze and understand the sentiments expressed in the user generated content. One such state of the art NLP model is Bidirectional Encoder Representations from Transformers BERT which has been widely adapted in sentiment analysis. BERT is susceptible to adversarial attacks. This paper aims to scrutinize the inherent vulnerabilities of such models in Twitter sentiment analysis. It aims to formulate a framework for constructing targeted adversarial texts capable of deceiving these models, while maintaining stealth. In contrast to conventional methodologies, such as Importance Reweighting, this framework core idea resides in its reliance on gradients to prioritize the importance of individual words within the text. It uses a whitebox approach to attain fine grained sensitivity, pinpointing words that exert maximal influence on the classification outcome. This paper is organized into three interdependent phases. It starts with fine-tuning a pre-trained BERT model on Twitter data. It then analyzes gradients of the model to rank words on their importance, and iteratively replaces those with feasible candidates until an acceptable solution is found. Finally, it evaluates the effectiveness of the adversarial text against the custom trained sentiment classification model. This assessment would help in gauging the capacity of the adversarial text to successfully subvert classification without raising any alarm.
Problem

Research questions and friction points this paper is trying to address.

Examines BERT's vulnerabilities in Twitter sentiment analysis
Develops gradient-based adversarial text framework for misclassification
Evaluates stealthy attack effectiveness on custom sentiment models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses gradient-based word importance ranking
Whitebox approach for fine-grained sensitivity
Iterative word replacement for adversarial text
🔎 Similar Papers
No similar papers found.
A
Akil Raj Subedi
School of Computer Science Engineering and Information Systems, Vellore Institute of Technology, Vellore, India
T
Taniya Shah
School of Computer Science Engineering and Information Systems, Vellore Institute of Technology, Vellore, India
Aswani Kumar Cherukuri
Aswani Kumar Cherukuri
Vellore Institute of Technology, Vellore
Quantum ComputingInformation SecurityMachine Learning
T
Thanos Vasilakos
Center for AI Research (CAIR), University of Agder(UiA), Grimstad, Norway, College of Computer Science and Information Technology, IAU, Saudi Arabia