Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Detection of AI-Generated Phishing Emails: Comparing The Efficiency Of SVM, Random Forest, CNN And BiLSTM In Detecting AI-Generated Phishing Emails
Blekinge Institute of Technology, Faculty of Computing, Department of Computer Science.
Blekinge Institute of Technology, Faculty of Computing, Department of Computer Science.
2025 (English)Independent thesis Basic level (degree of Bachelor), 10 credits / 15 HE creditsStudent thesis
Abstract [en]

Background: Phishing attacks have become progressively more complex with the introduction of Large Language Models (LLMs) that are able to produce human-like emails. Consequently, automated phishing emails are difficult to differentiate from legitimate emails using conventional security measures. Accordingly, a new level of sophistication around phishing attacks is pushing for a next generation of detection mechanisms.

Objectives: This thesis intends to detect and classify Artificial Intelligence (AI) generated emails from human-written phishing emails. Traditional machine learning models, Support Vector Machine (SVM) and Random Forest, are evaluated as well as DL models, Convolutional Neural Network (CNN) and Bidirectional Long Short Term Memory (BiLSTM).

Methods: The study utilizes a multiclass email dataset from Kaggle and another spear-phishing dataset with AI-generated emails. Text preprocessing was done using natural language processing (NLP) processing steps, which included tokenizing and lemmatizing the text. The traditional machine learning models were executed using TF-IDF vectorized emails and the CNN and BiLSTM were executed using padded token sequences. In all cases the models were examined based on metrics of Accuracy, Precision, Recall, F1-score, and AUC-ROC.

Results: All four models performed quite well, reaching perfect (100%) accuracy, precision, recall, F1-score, and AUC-ROC with SVM and BiLSTM. CNN and Random Forest were also effective, albeit with a few false positives. The study found a stark contrast in linguistic variables between AI-generated texts vs human-created texts, allowing us to classify the results accurately.

Conclusions: Both traditional and DL models perform effectively in identifying AI generated phishing emails, with SVM and BiLSTM demonstrating the best performance. As this research illustrates, automated detection systems will be a practical tool in the arsenal of modern commonplace phishing defense. It is important however, to remain agile to evolutionary progression of AI. 

Place, publisher, year, edition, pages
2025. , p. 46
Keywords [en]
Phishing Email Detection, AI-generated Emails, Machine Learning, Deep Learning, Natural Language Processing (NLP)
National Category
Computer and Information Sciences
Identifiers
URN: urn:nbn:se:bth-28339OAI: oai:DiVA.org:bth-28339DiVA, id: diva2:1982231
Subject / course
DV1478 Bachelor Thesis in Computer Science
Educational program
DVGDT Bachelor Qualification Plan in Computer Science 60.0 hp
Supervisors
Examiners
Available from: 2025-08-05 Created: 2025-07-07 Last updated: 2025-09-30Bibliographically approved

Open Access in DiVA

fulltext(1658 kB)1145 downloads
File information
File name FULLTEXT01.pdfFile size 1658 kBChecksum SHA-512
61c88e78399ff595101c4a337ae49c2f85404e45567da0bc60806b1b96b034917dca3bb2f1bf7d88a07543c1c63f5bd1afc12555923d42e68e69ac00a1977d44
Type fulltextMimetype application/pdf

By organisation
Department of Computer Science
Computer and Information Sciences

Search outside of DiVA

GoogleGoogle Scholar
Total: 1146 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 769 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf