Content Word Replacement for Training as Defense Against Word-level Adversarial Perturbations in Sentiment Analysis

Main Article Content

Sai Reethi Pydi, Sudha Pelluri

Abstract

Ever since their discovery, adversarial attacks have been plaguing machine learning models. These attacks happen when perturbations are added to benign text by replacing it with similar words. This often causes models to misclassify input. In this paper, we aim to use these attacks to improve our model’s accuracy and in turn, their reliability. For this purpose, we propose a novel attack called the WordReplace attack where we generate adversarial examples by combining the power of text rank algorithm, parts of speech consistency and word2vec synonym replacement. Adversarial training has been the de facto defense in these scenarios. It involves using attacked samples to train the model. In this paper, we explore the impact of such attacks and how they can be utilized to improve the performance of a model by introducing a modified version of adversarial training by finetuning our models on the adversarial inputs. Our results show that on the Internet Movie Database dataset, our attack is successful in reducing the performance of both the transformer based bidirectional encoder representations from transformers model, BERT and the recurrent neural network based long short-term memory model, LSTM. Our experiments showed an improvement in the accuracy of the attacked LSTM model by 39.15 percent and 21.75 percent for the attacked BERT model.

Article Details

How to Cite
Sai Reethi Pydi, Sudha Pelluri. (2026). Content Word Replacement for Training as Defense Against Word-level Adversarial Perturbations in Sentiment Analysis. International Journal of Special Education, 41(7s), 1104–1113. Retrieved from https://internationalsped.com/index.php/ijse/article/view/3321
Section
General