Detection of AI-Generated Phishing Emails: Comparing The Efficiency Of SVM, Random Forest, CNN And BiLSTM In Detecting AI-Generated Phishing Emails
2025 (English)Independent thesis Basic level (degree of Bachelor), 10 credits / 15 HE credits
Student thesis
Abstract [en]
Background: Phishing attacks have become progressively more complex with the introduction of Large Language Models (LLMs) that are able to produce human-like emails. Consequently, automated phishing emails are difficult to differentiate from legitimate emails using conventional security measures. Accordingly, a new level of sophistication around phishing attacks is pushing for a next generation of detection mechanisms.
Objectives: This thesis intends to detect and classify Artificial Intelligence (AI) generated emails from human-written phishing emails. Traditional machine learning models, Support Vector Machine (SVM) and Random Forest, are evaluated as well as DL models, Convolutional Neural Network (CNN) and Bidirectional Long Short Term Memory (BiLSTM).
Methods: The study utilizes a multiclass email dataset from Kaggle and another spear-phishing dataset with AI-generated emails. Text preprocessing was done using natural language processing (NLP) processing steps, which included tokenizing and lemmatizing the text. The traditional machine learning models were executed using TF-IDF vectorized emails and the CNN and BiLSTM were executed using padded token sequences. In all cases the models were examined based on metrics of Accuracy, Precision, Recall, F1-score, and AUC-ROC.
Results: All four models performed quite well, reaching perfect (100%) accuracy, precision, recall, F1-score, and AUC-ROC with SVM and BiLSTM. CNN and Random Forest were also effective, albeit with a few false positives. The study found a stark contrast in linguistic variables between AI-generated texts vs human-created texts, allowing us to classify the results accurately.
Conclusions: Both traditional and DL models perform effectively in identifying AI generated phishing emails, with SVM and BiLSTM demonstrating the best performance. As this research illustrates, automated detection systems will be a practical tool in the arsenal of modern commonplace phishing defense. It is important however, to remain agile to evolutionary progression of AI.
Place, publisher, year, edition, pages
2025. , p. 46
Keywords [en]
Phishing Email Detection, AI-generated Emails, Machine Learning, Deep Learning, Natural Language Processing (NLP)
National Category
Computer and Information Sciences
Identifiers
URN: urn:nbn:se:bth-28339OAI: oai:DiVA.org:bth-28339DiVA, id: diva2:1982231
Subject / course
DV1478 Bachelor Thesis in Computer Science
Educational program
DVGDT Bachelor Qualification Plan in Computer Science 60.0 hp
Supervisors
Examiners
2025-08-052025-07-072025-09-30Bibliographically approved