Please use this identifier to cite or link to this item: http://hdl.handle.net/1893/38346
Appears in Collections:Computing Science and Mathematics Journal Articles
Peer Review Status: Refereed
Title: A Comprehensive Analysis of Adversarial Attacks against Spam Filters
Author(s): Hotoğlu, Esra
Sen, Sevil
Can, Burcu
Contact Email: burcu.can@stir.ac.uk
Keywords: Email security
Spam detection
Adversarial learning
Natural language processing
Deep learning * Corresponding author
Issue Date: Dec-2026
Date Deposited: 29-Jul-2026
Citation: Hotoğlu E, Sen S & Can B (2026) A Comprehensive Analysis of Adversarial Attacks against Spam Filters. <i>Computers and Security</i>, 171, Art. No.: 105066. https://doi.org/10.1016/j.cose.2026.105066
Abstract: Deep learning has revolutionized email filtering, which is critical to protect users from cyber threats such as spam, malware, and phishing. However, the increasing sophistication of adversarial attacks poses a significant challenge to the effectiveness of these filters. This study investigates the impact of adversarial attacks on deep learning-based spam detection systems using real-world datasets. Six prominent deep learning models are evaluated on these datasets, analyzing attacks at the word, character sentence, and AI-generated paragraph-levels. Novel scoring functions, including spam weights and attention weights, are introduced to improve attack effectiveness. A key contribution of this study is the analysis of spam-weight-and attention-weight-based scoring functions, highlighting their role in improving the effectiveness and efficiency of adversarial attacks. This comprehensive analysis sheds light on the vulnerabilities of spam filters and contributes to efforts to improve their security against evolving adversarial threats. Experimental results show that word-and sentence-level attacks markedly increase false negatives, while character-level perturbations disrupt token representations with minimal semantic change. AI-generated paragraph-level attacks remain challenging even for transformer-based models. In addition, spam-weight-based scoring consistently enables more effective adversarial attacks than alternative scoring strategies with lower computational cost.
DOI Link: 10.1016/j.cose.2026.105066
Notes: (Sevil Sen)
Licence URL(s): http://creativecommons.org/licenses/by-nc-nd/4.0/

Files in This Item:
File Description SizeFormat 
A Comprehensive Analysis of Adversarial Attacks against Spam Filters.pdfFulltext - Accepted Version1.71 MBAdobe PDFUnder Embargo until 2028-07-24    Request a copy

Note: If any of the files in this item are currently embargoed, you can request a copy directly from the author by clicking the padlock icon above. However, this facility is dependent on the depositor still being contactable at their original email address.



This item is protected by original copyright



A file in this item is licensed under a Creative Commons License Creative Commons

Items in the Repository are protected by copyright, with all rights reserved, unless otherwise indicated.

The metadata of the records in the Repository are available under the CC0 public domain dedication: No Rights Reserved https://creativecommons.org/publicdomain/zero/1.0/

If you believe that any material held in STORRE infringes copyright, please contact library@stir.ac.uk providing details and we will remove the Work from public display in STORRE and investigate your claim.