Multi-Source Generalization-Aware Phishing URL Detection Using Calibrated Stacked Ensemble and False-Positive Control
view PDF
view PDF

How to Cite

S., Mohammed Elias Basha, and Mallikharjuna Rao N. 2026. “Multi-Source Generalization-Aware Phishing URL Detection Using Calibrated Stacked Ensemble and False-Positive Control”. Journal of Trends in Computer Science and Smart Technology 8 (3): 838-61. https://doi.org/10.36548/jtcsst.2026.3.020.

Keywords

Phishing URL Detection
GAFPNet
False-Positive Control
Generalisation-Aware Learning
Calibrated Stacked Ensemble
Real-Time Phishing Detection

Abstract

Phishing remains a persistent cybersecurity threat, with over 1.3 million attacks reported in a single quarter of 2023. Despite strong benchmark performance, many machineand deep-learning models exhibit limited deployment reliability because they are evaluated using balanced, single-source datasets with randomized splits. This paper addresses this gap in two phases. First, it presents a controlled multi-source empirical study of four baseline phishing URL detection models–Logistic Regression, Support Vector Machine, Random Forest, and XGBoost–under four increasingly realistic evaluation conditions, including cross-source generalization, temporal drift, and class imbalance. Second, it introduces GAFPNet (Generalization-Aware and False-Positive Controlled Framework Network), a five-module stacked-ensemble framework. GAFPNet uses dataset-neutral lexical and structural URL feature extraction, SMOTE-based imbalance correction, Platt scaling for ensemble calibration, and a tunable false-positive control scheme. Experiments using a consolidated 13,000-sample set from PhishTank, OpenPhish, and Tranco Top-Sites show that baseline accuracy decreases by 13.16 to 20.22 percentage points in cross-source testing and by 2.41 to 8.14 percentage points in standard testing. GAFPNet achieves 99.12% accuracy, a 98.97% F1-score, an AUC-ROC of 0.9943, an MCC of 0.988, and a false positive rate of 0.74%, outperforming the evaluated baselines in all four scenarios. An ablation study confirms the contribution of each module. These results position GAFPNet as a generalization-aware and deployment-oriented phishing URL classifier for real-time filtering applications.

References

  1. Anti-Phishing Working Group, “Phishing activity trends report: 3rd quarter 2023,” 2023. Available: https://docs.apwg.org/reports/apwg_trends_report_q3_2023.pdf
  2. Federal Bureau of Investigation, “Internet Crime Complaint Center releases 2022 statistics,” 2023. Available: https://www.fbi.gov/contact-us/field-offices/springfield/news/internetcrime-complaint-center-releases-2022-statistics
  3. S. Anupam and A. K. Kar, “Phishing website detection using support vector machines and nature-inspired optimization algorithms,” Telecommunication Systems, vol. 76, no. 1, 2021,17–32.
  4. Z. Alshingiti, R. Alaqel, J. Al-Muhtadi, Q. E. Ul Haq, K. Saleem, and M. H. Faheem, “A deep learning-based phishing detection system using CNN, LSTM, and LSTM-CNN,” Electronics, vol. 12, no. 1, 2023, 232.
  5. S. Aslam, H. Aslam, A. Manzoor, H. Chen, and A. Rasool, “AntiPhishStack: LSTM-based stacked generalization model for optimized phishing URL detection,” Symmetry, vol. 16, no. 2, 2024.,248.
  6. H. Ghalechyan, E. Israyelyan, A. Arakelyan, G. Hovhannisyan, and A. Davtyan, “Phishing URL detection with neural networks: an empirical study,” Scientific Reports, vol. 14, no. 1, 2024,25134.
  7. S. Asiri, Y. Xiao, S. Alzahrani, S. Li, and T. Li, “A survey of intelligent detection designs of HTML URL phishing attacks,” IEEE Access, vol. 11, 2023,6421–6443.
  8. M. Osmanoglu, D. Gupta, M. Ozkan, Y. Ar, and Ö. Aslan, “A comprehensive review of malicious URLs: Detection techniques, features and datasets,” Computers and Electrical Engineering, vol. 136, 2026,111186.
  9. O. K. Sahingoz, E. Buber, O. Demir, and B. Diri, “Machine learning based phishing detection from URLs,” Expert Systems with Applications, vol. 117, 2019,345–357.
  10. S. Hamadouche, O. Boudraa, and M. Gasmi, “Combining lexical, host, and content-based features for phishing websites detection using machine learning models,” EAI Endorsed Transactions on Scalable Information Systems, vol. 11, 2024,1–15.
  11. M. Sánchez-Paniagua, E. Fidalgo Fernández, E. Alegre, W. Al-Nabki, and V. González-Castro, “Phishing URL detection: A real-case scenario through login URLs,” IEEE Access, vol. 10, 2022,42949–42960.
  12. I. Kara, M. Ok, and A. Ozaday, “Characteristics of understanding URLs and domain names features: The detection of phishing websites with machine learning methods,” IEEE Access, vol. 10, 2022,124420–124428.
  13. A. Ozcan, C. Catal, E. Donmez, and B. Senturk, “A hybrid DNN-LSTM model for detecting phishing URLs,” Neural Computing and Applications, vol. 35, no. 7, 2023,4957–4973.
  14. S. Joshi and S. M. Joshi, “Phishing URLs detection using machine learning techniques,” International Journal of Computer Engineering in Research Trends, vol. 6, no. 6, 2019,326–333.
  15. L. Tang and Q. H. Mahmoud, “A deep learning-based framework for phishing website detection,” IEEE Access, vol. 10, 2021,1509–1521.
  16. S. Asiri, Y. Xiao, S. Alzahrani, and T. Li, “PhishingRTDS: A real-time detection system for phishing attacks using a deep learning model,” Computers & Security, vol. 141, Article 103843, 2024. doi: 10.1016/j.cose.2024.103843.
  17. M. Nanda and S. Goel, “URL based phishing attack detection using BiLSTM-gated highway attention block convolutional neural network,” Multimedia Tools and Applications, vol. 83, no. 27, 2024,69345–69375.
  18. M. Elsadig, A. O. Ibrahim, S. Basheer, M. A. Alohali, S. Alshunaifi, H. Alqahtani, N. Alharbi, and W. Nagmeldin, “Intelligent deep machine learning cyber phishing URL detection based on BERT features extraction,” Electronics, vol. 11, no. 22, 2022, 3647.
  19. K. S. Jishnu and B. Arthi, “Real-time phishing URL detection framework using knowledge distilled ELECTRA,” Automatika: časopis za automatiku, mjerenje, elektroniku, računarstvo i komunikacije, vol. 65, no. 4, 2024,1621–1639,.
  20. A. S. Bozkir, F. C. Dalgic, and M. Aydos, “GramBeddings: A new neural network for URL based identification of phishing web pages through n-gram embeddings,” Computers & Security, vol. 124, 2023,102964.
  21. M. Mia, D. Derakhshan, and M. M. A. Pritom, “Can features for phishing URL detection be trusted across diverse datasets? A case study with explainable AI,” in Proceedings of the 11th International Conference on Networking, Systems, and Security, 2024,37–145.
  22. R. Zaimi, M. Hafidi, and L. Mahnane, “A deep learning mechanism to detect phishing URLs using the permutation importance method and SMOTE-Tomek link,” The Journal of Supercomputing, vol. 80, no. 12, 2024,17159–17191.
  23. J. L. Wilk-Jakubowski, L. Pawlik, G. Wilk-Jakubowski, and A. Sikora, “Machine learning and neural networks for phishing detection: A systematic review (2017–2024),” Electronics, vol. 14, no. 18, p. 3744, 2025.
  24. A. Aljofey, Q. Jiang, A. Rasool, H. Chen, W. Liu, Q. Qu, and Y. Wang, “An effective detection approach for phishing websites using URL and HTML features,” Scientific Reports, vol. 12, no. 1, 2022,8842.
  25. S. Kavya and D. Sumathi, “Staying ahead of phishers: A review of recent advances and emerging methodologies in phishing detection,” Artificial Intelligence Review, vol. 58, no. 2, 2024,50.
  26. S. Remya, M. J. Pillai, K. K. Nair, S. R. Subbareddy, and Y. Y. Cho, “An effective detection approach for phishing URL using ResMLP,” IEEE Access, vol. 12, 2024,79367–79382.
  27. S. M. Alasmari, H. Sakly, N. Kraiem, and A. Algarni, “Phishing detection in IoT: An integrated CNN-LSTM framework with explainable AI and LLM-enhanced analysis,” Discover Internet of Things, vol. 5, no. 1, 2025,102.
  28. S. Mukherjee, S. Chatterjee, M. Bachhar, S. Maji, G. Paul, S. Sengupta, and R. Das, “Enhanced phishing URL detection using machine learning technique,” in International Conference on Computational Intelligence, Data Science and Cloud Computing, Singapore: Springer Nature Singapore, 2025,209–221.
  29. PhishTank, “Phish archive,” 2026. [Online]. Available: https://data.dev.phishtank.com/phi sh_archive.php [Accessed: Aug. 5, 2026].
  30. OpenPhish, “Phishing feeds,” 2026. [Online]. Available: https://openphish.com/phishing_f eeds.html [Accessed: Aug. 5, 2026].
  31. V. Le Pochat, T. Van Goethem, S. Tajalizadehkhoob, and W. Joosen, “Tranco: A research-oriented top sites ranking hardened against manipulation,” arXiv preprint arXiv:1806.01156, 2018.
  32. N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic minority over-sampling technique,” Journal of Artificial Intelligence Research, vol. 16, 2002,321–357.
  33. J. Platt, “Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods,” in Advances in Large Margin Classifiers, vol. 10, no. 3, 1999,61–74.