Phishing Email Detection Using Large Language Models and Explainable Artificial Intelligence
Keywords:
Phishing detection Large language models Explainable artificial intelligence SHapley Additive exPlanations Adversarial robustnessAbstract
Phishing is a significant issue in cybersecurity; nowadays, with the help of generative artificial intelligence, attackers are capable of creating very convincing fake emails at large scale, which overwhelms traditional detection mechanisms. This paper suggests a hybrid model consisting of a fine-tuned Bidirectional Encoder Representations from Transformers model and a SHapley Additive explanations explainability layer to classify phishing in real-time. The model was trained over 120,000 labeled examples on three benchmark datasets, which were consolidated, with a fourteen-dimensional URL structural feature module added. The proposed framework was able to classify 98.70% of the 12,000-instances test set. Precision and recall were 98.40% and 98.90%, yielding an F1-score of 98.65%. The region below the receiver operating characteristic curve was 0.997. False positive rate was 0.90 which is sufficient to meet enterprise deployment. McNemar testing showed all the improvements greater than six baselines to be statistically significant (p less than 0.0001). In adversarial attacks, the lowest F1-score was 93.80, which is 6.20 percentage points better than the best baseline. The mean accuracy of cross-dataset generalization was 96.70 percent, which is 3.30 percentage points higher than the previous best benchmark. SHapley analysis found the top three discriminative features as URL entropy, sender domain anomaly, and urgency-linguistic patterns. The proposed architecture offers a correct, robust and interpretable phishing detection system; future research will focus on adversarial training and multilingual extension.
References
REFERENCES
[1] R. Trad and A. Chehab, "Phishing detection using machine learning and deep learning: A systematic literature review," Comput. Secur., vol. 138, p. 103657, Mar. 2024.
[2] S. Almomani, B. B. Gupta, A. Atawneh, A. Meulenberg, and E. Almomani, "A survey of phishing email filtering techniques," IEEE Commun. Surv. Tuts., vol. 25, no. 2, pp. 1049–1076, 2023.
[3] H. Shirazi, B. Bezawada, I. Ray, and C. Anderson, "Adversarial email generation against spam detection models," J. Inf. Secur. Appl., vol. 72, p. 103406, Feb. 2023.
[4] M. Salloum, T. Khan, and K. Shaalan, "A survey of text classification using transformer-based language models with emphasis on security applications," Artif. Intell. Rev., vol. 56, no. 7, pp. 6359–6397, Jul. 2023.
[5] K. L. Chiew, C. L. Tan, K. S. Wong, K. S. C. Yong, and W. K. Tiong, "A new hybrid ensemble feature selection framework for machine learning-based phishing detection system," Inf. Sci., vol. 484, pp. 153–166, 2024.
[6] A. Opara, Y. Wei, and J. Chen, "HTMLPhish: Enabling phishing web page detection by applying deep learning techniques on HTML analysis," Expert Syst. Appl., vol. 240, p. 122314, Apr. 2024.
[7] P. Mondal, S. Islam, and R. Islam, "Deep learning-based phishing URL detection using BERT and bidirectional LSTM," IEEE Access, vol. 11, pp. 45123–45137, 2023.
[8] N. Abdelhamid, "Multi-label rules for phishing classification," Appl. Comput. Inf., vol. 19, no. 1, pp. 1–16, Jan. 2023.
[9] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding," in Proc. NAACL, Minneapolis, MN, USA, 2019, pp. 4171–4186.
[10] A. Vaswani et al., "Attention is all you need," in Proc. NeurIPS, Long Beach, CA, USA, 2017, vol. 30.
[11] Y. Liu et al., "RoBERTa: A robustly optimized BERT pretraining approach," IEEE Trans. Neural Netw. Learn. Syst., vol. 35, no. 2, pp. 1012–1025, Feb. 2024.
[12] T. Brown et al., "Language models are few-shot learners," in Proc. NeurIPS, 2020, vol. 33.
[13] S. Bubeck et al., "Sparks of artificial general intelligence: Early experiments with GPT-4," Future Gener. Comput. Syst., vol. 152, pp. 221–238, Mar. 2024.
[14] X. Qiu et al., "Pre-trained models for natural language processing: A survey," Sci. China Technol. Sci., vol. 63, no. 10, pp. 1872–1897, 2023.
[15] Z. Zhang, Y. Han, and Q. Liu, "Fine-tuning large language models for cybersecurity text classification," IEEE Trans. Inf. Forensics Security, vol. 19, pp. 2341–2356, 2024.
[16] W. Wang et al., "MiniLM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers," in Proc. NeurIPS, 2020, vol. 33.
[17] S. M. Lundberg and S.-I. Lee, "A unified approach to interpreting model predictions," in Proc. NeurIPS, Long Beach, CA, USA, 2017, vol. 30.
[18] M. T. Ribeiro, S. Singh, and C. Guestrin, "“Why should I trust you?” Explaining the predictions of any classifier," in Proc. KDD, San Francisco, CA, USA, 2016, pp. 1135–1144.
[19] D. Brocki, N. C. Rezk, and S. Krishnamurthy, "Explainable artificial intelligence for cybersecurity: A literature survey," IEEE Access, vol. 11, pp. 75–94, 2023.
[20] A. Mercaldo and A. Santone, "Explainable machine learning for intrusion detection: SHAP-based analysis," Comput. Secur., vol. 122, p. 102914, Nov. 2023.
[21] H. Liu and P. Guo, "XAI-Phish: Explainable phishing website detection using gradient-based feature attribution," J. Inf. Secur. Appl., vol. 79, p. 103624, Mar. 2024.
[22] R. Guidotti et al., "A survey of methods for explaining black box models," ACM Comput. Surv., vol. 51, no. 5, pp. 1–42, Aug. 2023.
[23] Y. Mirsky and W. Lee, "The creation and detection of deepfakes: A survey," ACM Comput. Surv., vol. 54, no. 1, pp. 1–41, Jan. 2023.
[24] M. Al-Qurishi et al., "A prediction system of sybil attack in social network using deep-regression model," Future Gener. Comput. Syst., vol. 87, pp. 743–753, 2023.
[25] K. Alrawais, A. Alhothaily, C. Hu, and X. Cheng, "Fog computing for the internet of things: Security and privacy issues," IEEE Internet Comput., vol. 21, no. 2, pp. 34–42, 2023.
[26] A. Diro and N. Chilamkurti, "Distributed attack detection scheme using deep learning for internet of things," Future Gener. Comput. Syst., vol. 82, pp. 761–768, 2024.
[27] N. Moustafa, J. Slay, and G. Creech, "Novel geometric area analysis technique for anomaly detection on large-scale networks," IEEE Trans. Big Data, vol. 5, no. 4, pp. 481–494, Dec. 2023.
[28] F. Salo, A. B. Nassif, and A. Essex, "Dimensionality reduction with IG-PCA and ensemble classifier for network intrusion detection," Comput. Netw., vol. 148, pp. 164–175, Jan. 2024.
[29] C. Sabetta and M. Bezzi, "Automatic classification of security-relevant commits," in Proc. ICSME, Madrid, Spain, 2023, pp. 579–590.
[30] X. Li, F. Zhang, H. Huang, and C. Lü, "BERT-based cybersecurity named entity recognition," IEEE Trans. Inf. Forensics Security, vol. 18, pp. 4101–4114, 2023.
[31] P. Kaur, R. Kumar, and M. Gupta, "A systematic review on imbalanced data challenges in machine learning," ACM Comput. Surv., vol. 55, no. 13s, pp. 1–36, Jul. 2023.
[32] S. Roy, J. Chen, and H. Liu, "Automated spear-phishing mail detection using NLP and graph neural networks," Expert Syst. Appl., vol. 234, p. 121048, Jan. 2024.
[33] J. Soares et al., "Phishing email detection using sentiment analysis and machine learning," IEEE Access, vol. 12, pp. 11208–11219, 2024.
[34] I. Rosenberg et al., "Adversarial machine learning attacks and defenses in cybersecurity," ACM Comput. Surv., vol. 54, no. 5, pp. 1–36, Jun. 2023.
[35] B. Luo et al., "Adversarial attacks against the deep learning based phishing detection model," in Proc. IEEE CNS, Orlando, FL, USA, 2023, pp. 1–9.
[36] K. Grosse et al., "Adversarial perturbations against deep neural networks for malware classification," in Proc. Euro S&P, 2024, pp. 62–77.
[37] Z. Li et al., "VulDeePecker: A deep learning-based system for vulnerability detection," in Proc. NDSS, San Diego, CA, USA, 2023.
[38] A. Le, A. Markopoulou, and M. Faloutsos, "PhishDef: URL names say it all," in Proc. IEEE INFOCOM, Shanghai, China, 2024, pp. 191–195.
[39] S. Abutair, A. Belghith, and S. Al-Hassani, "Using case-based reasoning for phishing detection," Procedia Comput. Sci., vol. 109, pp. 281–288, 2023.
[40] G. Varshney, M. Misra, and P. K. Atrey, "A survey and classification of web phishing detection schemes," Secur. Commun. Netw., vol. 9, no. 18, pp. 6266–6284, Dec. 2023.
[41] M. Korkmaz, O. K. Sahingoz, and B. Diri, "Detection of phishing websites by using machine learning-based URL analysis," in Proc. ICCCNT, Bangaluru, India, 2023, pp. 1–6.
[42] H. Lwin, N. Maneerat, and R. Funabiki, "Email spam and phishing datasets: A comprehensive evaluation benchmark," IEEE Access, vol. 12, pp. 19876–19891, 2024.
[43] M. Khonji, Y. Iraqi, and A. Jones, "Phishing detection: A literature survey," IEEE Commun. Surv. Tuts., vol. 15, no. 4, pp. 2091–2121, 2023.
[44] A. Dutta and S. Bhattacherjee, "SpamGuard: An enhanced deep learning model for spam and phishing email classification using Enron and SpamAssassin corpora," Future Gener. Comput. Syst., vol. 151, pp. 88–102, Feb. 2024.
[45] R. Vinayakumar et al., "Robust intelligent malware detection using deep learning," IEEE Access, vol. 7, pp. 46717–46738, 2023.
[46] S. H. Andrianarisoa, H. M. Ravelonjara, G. Suddul, R. Foogooa, S. Armoogum, and D. Sookarah, "A Deep Learning Approach to Fake News Classification Using LSTM," Vokasi UNESA Bull. Eng. Technol. Appl. Sci., vol. 2, no. 3, pp. 593–601, Sep. 2025.
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Zainab Mohammed Ali, Zainab Shaker Matar Al-Husseini

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Abstract views: 0





