Machine Learning for Phishing URL Detection in Social Media: A Systematic Literature Review

Authors

  • Frankline Kashindi Luseno Co-operative University of Kenya
  • Simon Karume Co-operative University of Kenya
  • Philemon Kasyoka South Eastern University of Kenya

Keywords:

Phishing Detection; Social Media Security; Machine Learning; URL Analysis; Adversarial Robustness.

Abstract

Social media platforms have become the primary conduit for phishing, exploiting trust networks and viral propagation that static blacklists and platform-native controls only partially counter. Conducted per PRISMA 2020 with an OSF-registered protocol, this systematic literature review critically synthesizes 58 primary studies (96 model-level experiments) selected from 1,245 records across seven databases (2016–2026). Publication output peaked in 2023–2024 as the field transitioned from feature-engineered classifiers to representation-rich models: traditional machine learning (34.5%) and deep learning (29.3%) dominate, while graph-based and transformer/LLM approaches form the growth frontier. Lexical features are near-universal (84.5%), yet social-context, behavioural, and multimodal features the signals most diagnostic of social phishing appear in only 31.0%, 12.1%, and 8.6% of studies. Public phishing lists underpin 38% of datasets, no community benchmark exists, and accuracy (82.8%) is privileged over operationally critical latency (13.8%) and false-positive reporting. Reported F1-scores rise from 0.88–0.95 (traditional ML) to 0.93–0.99 (transformer/LLM), yet heterogeneity precludes ranking, and reproducibility is weak (55% rated low). Coverage lags operational importance for rapid propagation, concept drift, and adversarial evasion. Future gains will derive less from novel architectures than from socially contextualized benchmarks, temporally and adversarially rigorous evaluation, and explainable, deployment-aware design.

References

Abdelhamid, N., Ayesh, A., & Thabtah, F. (2014). Phishing detection based associative classification data mining. Expert Systems with Applications, 41(13).

Adeyemi-Onih, V. (2024). Phishing Detection Using Machine Learning: A Model Development and Integration. International Journal of Scientific and Management Research, 07(04), 27–63.

Alkhalil, Z., Hewage, C., Nawaf, L., & Khan, I. (2021). Phishing attacks: A recent comprehensive survey

and a new anatomy. Computers & Security, 105, 104221.

Almousa, A., et al. (2023). Machine Learning Approaches for Social Media Phishing Detection: A Comprehensive Review. IEEE Access, 11, 45120-45140.

Alshehri, S. M. (2025). A hybrid graph neural network framework for malicious URL detection. Electronics Anti-

Phishing Working Group. (2023). Phishing attack trends report 2023. APWG.

APWG (2024). Phishing Activity Trends Report. Anti-Phishing Working Group.

Arp, D., et al. (2022). Dos and don'ts of machine learning in computer security. USENIX Security.

Azam, M., et al. (2024). A Review on Large Language Models: Architectures, Applications, Taxonomies,

Open Issues and Challenges. IEEE Access, 12, 1–1.

Bilot, T., et al. (2022). PhishGNN: A phishing website detection framework using graph neural networks.

SECRYPT

Cao, K., Zhang, T., & Huang, J. (2024). Advanced hybrid LSTM-transformer architecture for real-time multi-task prediction in engineering systems. Scientific Reports, 14(1), 4890.

Çatal, C., et al. (2022). Applications of deep learning for phishing detection: A systematic literature review

Chiew, K. L., Chang, E. H., Sze, S. N., & Yong, K. S. M. (2018). Utilisation of website logo for phishing

detection. Computers & Security, 74, 1–20.

Dalton, T., et al. (2025). PhreshPhish: A Real-World, High-Quality, Large-Scale Phishing Website Dataset and Benchmark. ArXiv.

Davis, J., & Goadrich, M. (2006). The relationship between precision and recall for an imbalanced dataset.

ICML.

Dutta, A. K., et al. (2021). Detecting phishing websites using machine learning technique. PLOS ONE

Garera, S., Provos, N., Chew, M., & Rubin, A. D. (2007). A framework for detection and measurement of

phishing attacks. WORM.

Ghalechyan, H. (2024). Phishing URL detection with neural networks: An empirical study. Scientific Reports

Goldenits, G., et al. (2026). Small language models for phishing website detection

Gupta, N., et al. (2014). bit.ly/malicious: Deep dive into short URL based e-crime detection

Hassan, M., Jameel, M., & Bashir, M. (2023). An Exploratory Study of Malicious Link Posting on Social Media Applications. Symposium on Usable Security and Privacy (USEC).

Jagatic, T., Johnson, N., Jakobsson, M., & Menczer, F. (2007). Social phishing. Communications of the ACM, 50(10).

Jain, A., & Gupta, R. (2022). Adversarial Robustness in URL-based Phishing Detection. Journal of Cybersecurity, 14(1), 1-18.

Jarczewski, M., Białczak, P., & Mazurczyk, W. (2026). Phishing Website Impersonation: Comparative Analysis of Detection and Target Recognition Methods. Applied Sciences, 16(2), 640.

Ketmanto Wangsa, Karim, S., & Sandu, R. (2026). Leveraging AI to Fight Phishing: A Systematic Review.

Lecture Notes in Networks and Systems, 158–180.

Khonji, M., Iraqi, Y., & Jones, A. (2013). Phishing detection: A literature survey. IEEE Communications Surveys & Tutorials, 15(4).

Kitchenham, B., & Charters, S. (2007). Guidelines for performing systematic literature reviews in software engineering.

Korkmaz, E., et al. (2022). Ensemble Learning for Real-Time Phishing Detection in Social Media. IEEE Transactions on Information Forensics and Security, 17, 2100-2115.

Le, H., Pham, Q., Sahoo, D., & Hoi, S. C. H. (2018). URLNet: Learning a URL representation with deep learning for malicious URL detection. arXiv:1802.03162.

Lee, J., et al. (2024). Multimodal large language models for phishing webpage detection. arXiv:2408.05941

Lee, K., & Caverlee, J. (2011). Understanding and combating link farming in the Twitter network. WWW.

Li, S., & Dib, O. (2024). Enhancing Online Security: A Novel Machine Learning Framework for Robust Detection of Known and Unknown Malicious URLs. Journal of Theoretical and Applied Electronic Commerce Research, 19(4), 2919–2960.

Li, W., et al. (2024). Machine learning-enabled attacks on anti-phishing blacklists. IEEE Access

Lin, Y., Liu, R., Divakaran, D. M., Ng, J. Y., et al. (2021).

Phishpedia: A hybrid deep learning based approach to visually identify phishing webpages. USENIX Security

Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. NeurIPS.

Ma, J., Saul, L. K., Savage, S., & Voelger, G. M. (2009). Beyond blacklists: Learning to detect malicious web sites from suspicious URLs. KDD.

Maggi, F., et al. (2013). Two years of short URLs internet measurement: Security assessment and case studies

Marchal, S., François, J., State, R., & Engel, T. (2014). PhishStorm: Detecting phishing with streaming analytics. IEEE TNSM, 11(4), 458–471

McMahan, B., et al. (2017). Communication-efficient learning of deep networks from decentralized data.

AISTATS.

Oest, A., et al. (2018). Sunrise to sunset: Analyzing the end-to-end life cycle and effectiveness of phishing attacks at scale. USENIX Security.

Oest, A., et al. (2019). PhishFarm: A scalable framework for measuring the effectiveness of evasion techniques against browser phishing blacklists. IEEE S&P

Ouyang, L., et al. (2021). Phishing web page detection with HTML-level graph neural network. TrustCom Page, M. J., et al. (2021). The PRISMA 2020 statement. BMJ, 372, n71.

Qazi, E. ul H., Faheem, M. H., & Ahmad, I. (2024). Detecting Phishing URLs Based on a Deep Learning Approach to Prevent Cyber-Attacks. Applied Sciences, 14(22), 10086.

Rao, R. S., & Pais, A. R. (2019). Detection of phishing websites using an efficient feature-based machine learning framework. Neural Computing & Applications, 32.

Rashid, H., et al. (2025). Framework for detecting phishing crimes on Twitter using selective feature selection

Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). "Why should I trust you?" Explaining the predictions of

any classifier. KDD.

Safi, A., et al. (2023). A systematic literature review on phishing website detection techniques. Journal of King Saud University – Computer and Information Sciences

Sahingoz, O. K., Buber, E., Demir, Ö., & Diri, B. (2019). Machine learning based phishing detection from URLs. Expert Systems with Applications, 117, 345–357.

Sahingoz, O. K., et al. (2019). Machine learning-based phishing detection from URLs. Expert Systems with Applications, 117, 345-357.

Saxe, J., & Berlin, K. (2017). eXpose: A character-level convolutional neural network with embeddings for detecting malicious URLs, files, and e-mails.

Sheng, S., et al. (2010). Who falls for phish? A demographic analysis of phishing susceptibility. CHI. Sohrab, M. G., et al. (2024). URLTran: Improving Phishing URL Detection Using Transformers. IEEE

Access, 12, 23–33.

Wilk-Jakubowski, J. L. (2025). Machine learning and neural networks for phishing detection. Electronics Xiang, G., Hong, J., Rose, C. P., & Cranor, L. (2011). CANTINA+: A feature-rich machine learning framework for detecting

phishing web sites. ACM TISSEC, 14(2).

Zamir, A., et al. (2020). Phishing web site detection using diverse machine learning algorithms. The Computer Journal, 63(1), 65–80.

Zhao, Y., et al. (2022). Social Media Security: Threats, Defenses, and Future Directions. ACM Computing Surveys, 55(4), 1-38.

Downloads

Published

2026-10-05

How to Cite

Frankline Kashindi Luseno, Simon Karume, & Philemon Kasyoka. (2026). Machine Learning for Phishing URL Detection in Social Media: A Systematic Literature Review . African Journal of Education,Science and Technology (AJEST), 9(1), 23–33. Retrieved from https://ajest.org/index.php/ajest/article/view/1046

Issue

Section

Articles

Similar Articles

<< < 16 17 18 19 20 21 22 23 24 25 > >> 

You may also start an advanced similarity search for this article.