Phishing attacks constitute one of the most pervasive cyber threats, exploiting human behaviour and system vulnerabilities to steal sensitive information such as login credentials and financial data. Existing detection systems frequently rely on singular modelling approaches and thus fall short in addressing the multidimensional and continuously evolving nature of these attacks. To overcome this challenge, the present work proposes a hybrid phishing detection framework that integrates three complementary techniques: DistilBERT (Distilled Bidirectional Encoder Representations from Transformers) for semantic analysis of URL text, Graph Neural Networks (GNN) for modelling structural relationships among URL components, and LightGBM (Light Gradient Boosting Machine) for efficient metadata-based feature classification. The outputs of the three models are fused through a weighted ensemble voting mechanism, enabling the system to leverage textual, structural, and statistical perspectives simultaneously. By replacing the computationally heavy BERT model used in prior work with the lighter DistilBERT variant, the proposed system achieves reduced computational complexity, faster inference, and lower memory consumption while preserving detection quality. Experimental evaluation on a large-scale Kaggle phishing URL dataset demonstrates that the hybrid ensemble yields a weighted phishing probability of 86.9%, with individual model contributions of 0.628 (DistilBERT), 1.000 (GNN), and 0.999 (LightGBM). The system is deployed via a Flask-based web interface that enables real-time URL classification. This multi-level approach underscores the importance of strengthening online security through robust, efficient, and scalable detection mechanisms.
Introduction
This study proposes a hybrid phishing website detection system that combines DistilBERT, Graph Neural Networks (GNN), and LightGBM to improve the accuracy and efficiency of identifying malicious websites. Phishing attacks have become a major cybersecurity threat, as attackers create fake websites that imitate legitimate ones to steal sensitive information such as passwords and financial data. Traditional rule-based and conventional machine learning methods often struggle to detect sophisticated phishing techniques, resulting in high false-positive rates and poor adaptability to evolving attacks.
To overcome these limitations, the proposed framework integrates three complementary models. DistilBERT performs semantic analysis of URL text, GNN captures structural relationships and identifies clusters of related phishing websites, and LightGBM analyses metadata features such as domain age, SSL certificates, and WHOIS information. Their predictions are combined using a weighted ensemble voting mechanism, enabling more robust and accurate phishing detection. Replacing the computationally intensive BERT model with DistilBERT significantly reduces memory usage, training time, and inference time, making the system suitable for real-time deployment and resource-constrained devices.
The literature review highlights previous research demonstrating the effectiveness of ensemble learning, transformer models, graph neural networks, and gradient boosting for phishing detection. However, existing systems generally rely on a single learning paradigm or computationally expensive hybrid models. The proposed approach addresses these shortcomings by balancing detection accuracy with computational efficiency.
The methodology consists of data collection, preprocessing, feature extraction, hybrid model processing, ensemble classification, and deployment. The dataset, obtained from Kaggle, contains approximately 549,346 URLs, including both phishing and legitimate websites. During preprocessing, duplicate records and missing values are removed, URLs are normalised and tokenised, features are standardised, and SMOTE is applied to balance the training data.
A total of 31 URL features are extracted, including lexical characteristics (such as URL length, suspicious keywords, and special characters), metadata features (including HTTPS usage, entropy, domain age, and WHOIS information), and HTML/JavaScript attributes that indicate malicious behaviour. DistilBERT analyses URL semantics, GNN models relationships among URLs to identify coordinated phishing campaigns, and LightGBM performs the final classification using engineered features. The ensemble combines their outputs using optimised weights (DistilBERT: 0.35, GNN: 0.30, LightGBM: 0.35) through soft voting, producing the final phishing prediction.
The system is implemented using Python, Flask, PyTorch, PyTorch Geometric, LightGBM, and the Hugging Face Transformers library, with a Flask-based web interface that allows users to submit URLs and receive real-time predictions. Experimental results demonstrate that the hybrid ensemble effectively combines the strengths of all three models. For a sample phishing URL, the ensemble produced a phishing confidence score of 86.9%, with GNN and LightGBM contributing high-confidence predictions while DistilBERT provided complementary semantic analysis. Compared with the earlier BGL-PhishNet framework, the proposed system achieves comparable detection performance while significantly reducing computational cost, memory consumption, and inference time, making it a practical and scalable solution for real-time phishing website detection.
Conclusion
This paper presented a hybrid phishing detection system that integrates DistilBERT, Graph Neural Networks, and LightGBM within a weighted ensemble framework. Each model contributes a unique analytical perspective: DistilBERT captures semantic patterns in URL text, GNN identifies structural relationships and coordinated attack clusters, and LightGBM efficiently classifies based on engineered metadata features. The weighted voting mechanism combines these complementary outputs into a robust final prediction.
By replacing BERT with DistilBERT, the proposed system achieves a significant reduction in computational complexity, memory usage, and inference time, making it suitable for real-time deployment and resource-constrained environments. Experimental evaluation demonstrates that the hybrid ensemble produces a strong composite phishing probability of 86.9%, with GNN and LightGBM contributing near-perfect individual scores. The system is deployed through a user-friendly Flask-based web interface that supports both single-URL and batch-CSV classification.
The work underscores the importance of multi-dimensional feature analysis in modern phishing detection, combining textual, structural, and statistical perspectives. The proposed approach paves the way for further research in automated and scalable phishing detection systems, contributing to stronger cybersecurity measures and safer online experiences.
References
[1] D. M. Divakaran and A. Oest, “Phishing detection leveraging machine learning and deep learning: A review,” IEEE Secur. Privacy, vol. 20, no. 5, pp. 86–95, Sep. 2022.
[2] S. M. Wajid, T. Javed, and M. M. Su’ud, “Ensemble learning powered URL phishing detection: A performance driven approach,” J. Informat. Web Eng., vol. 3, no. 2, pp. 134–145, Jun. 2024.
[3] N. Q. Do, A. Selamat, H. Fujita, and O. Krejcar, “An integrated model based on deep learning classifiers and pre-trained transformer for phishing URL detection,” Future Gener. Comput. Syst., vol. 161, pp. 269–285, Dec. 2024.
[4] M. S. Saravanan and C. Lokesh, “Performance enhancement to improve accuracy for the novel CNN GoogleNet compared against ResNet in predicting cyber hacking breaches,” in Proc. Int. Conf. Self Sustain. Artif. Intell. Syst. (ICSSAS), Oct. 2023, pp. 158–163.
[5] S. Wang, R. K. L. Ko, G. Bai, N. Dong, T. Choi, and Y. Zhang, “Evasion attack and defense on machine learning models in cyber-physical systems: A survey,” IEEE Commun. Surveys Tuts., vol. 26, no. 2, pp. 930–966, 2nd Quart., 2024.
[6] H. Abusaimeh and Y. Alshareef, “Detecting the phishing website with the highest accuracy,” TEM J., vol. 10, no. 2, pp. 947–953, May 2021.
[7] L. R. Kalabarige, R. S. Rao, A. Abraham, and L. A. Gabralla, “Multilayer stacked ensemble learning model to detect phishing websites,” IEEE Access, vol. 10, pp. 79543–79552, 2022.
[8] M. M. Ali and N. F. M. Zaharon, “Phishing—A cyber fraud: The types, implications and governance,” Int. J. Educ. Reform, vol. 33, no. 1, pp. 101–121, Jan. 2024.
[9] S. P. Ripa, F. Islam, and M. Arifuzzaman, “The emergence threat of phishing attack and the detection techniques using machine learning models,” in Proc. Int. Conf. Autom., Control Mechatronics Ind. (ACMI), Jul. 2021, pp. 1–6.
[10] H. Huang, X. Zhang, J. Wang, C. Gao, X. Li, R. Zhu, and Q. Ma, “PEAE-GNN: Phishing detection on Ethereum via augmentation ego-graph based on graph neural network,” IEEE Trans. Computat. Social Syst., vol. 11, no. 3, pp. 4326–4339, Jun. 2024.
[11] R. Yang, K. Zheng, B. Wu, D. Li, Z. Wang, and X. Wang, “Predicting user susceptibility to phishing based on multidimensional features,” Comput. Intell. Neurosci., vol. 2022, Jan. 2022, Art. no. 7058972.
[12] L. Zhou, D. Zhang, and Z. Liu, “A stage model for understanding phishing victimization behavior in embedded training,” in Proc. IEEE Int. Conf. Intell. Secur. Informat. (ISI), Oct. 2023, pp. 1–6.
[13] C. L. Tan, K. L. Chiew, K. S. C. Yong, S. N. Sze, J. Abdullah, and Y. Sebastian, “A graph-theoretic approach for the detection of phishing webpages,” Comput. Secur., vol. 95, Aug. 2020, Art. no. 101793.
[14] L. Tang and Q. H. Mahmoud, “A deep learning-based framework for phishing website detection,” IEEE Access, vol. 10, pp. 1509–1521, 2022.
[15] H. Feng, F. Wang, N. Li, Q. Xu, G. Zheng, X. Sun, M. Hu, G. Xing, and G. Zhang, “A random forest model for peptide classification based on virtual docking data,” Int. J. Mol. Sci., vol. 24, no. 14, p. 11409, Jul. 2023.