VeriSphere: An NLP and XGBoost-Based Framework for Early Detection of Cyberbullying in Social Media
view PDF
view PDF

How to Cite

A M., Aririthissh, Prasanth J., Yayathiraja N., and Preethi Monika K. 2026. “VeriSphere: An NLP and XGBoost-Based Framework for Early Detection of Cyberbullying in Social Media”. Journal of ISMAC 8 (3): 277-91. https://doi.org/10.36548/jismac.2026.3.006.

Keywords

Cyberbullying Detection
Natural Language Processing (NLP)
XGBoost Classification
TF–IDF Feature Representation
Automated Content Moderation
Real-Time Text Analytics

Abstract

The increasing adoption of social media platforms has increased the incidences of cyberbullying which poses serious problems in terms of user security, psychological safety and effective content moderation on the internet. This paper presents VeriSphere which is an intelligent cyberbullying detection framework that incorporates NLP with the use of the XGBoost classifier for identifying cyberbullying at an early stage. The proposed system uses a robust preprocessing pipeline that consists of text normalization, tokenization, stop-word removal, lemmatization and TF–IDF feature extraction to derive discriminating features for classification purposes. XGBoost classifier is used to classify abusive content from non-abusive texts in an efficient way and with good generalization through the use of regularization. In addition, automatic notification system sends emails to administrators as soon as cyberbullying is identified. Experiments have shown that the proposed system can classify cyberbullying cases with 92.84% accuracy, 91.76% precision, 92.31% recall, 92.03% F1-score and 93.90% ROC-AUC which indicate high performance and good discrimination ability.

References

  1. Subramanian, Malliga, Veerappampalayam Easwaramoorthy Sathiskumar, G. Deepalakshmi, Jaehyuk Cho, and G. Manikandan. "A Survey on Hate Speech Detection and Sentiment Analysis using Machine Learning and Deep Learning Models." Alexandria Engineering Journal 2023, vol. 80: 110-121.
  2. Alkomah, Fatimah, and Xiaogang Ma. "A Literature Review of Textual Hate Speech Detection Methods and Datasets." Information 2022, vol. 13, no. 6: 273.
  3. Khan, Shakir, Mohd Fazil, Vineet Kumar Sejwal, Mohammed Ali Alshara, Reemiah Muneer Alotaibi, Ashraf Kamal, and Abdul Rauf Baig. "BiCHAT: BiLSTM with Deep CNN and Hierarchical Attention for Hate Speech Detection." Journal of King Saud University-Computer and Information Sciences 2022, vol. 34, no. 7: 4335-4344.
  4. Mehta, Harshkumar, and Kalpdrum Passi. "Social Media Hate Speech Detection using Explainable Artificial Intelligence (XAI)." Algorithms 2022, vol.15, no. 8: 291.
  5. Saleh, Hind, Areej Alhothali, and Kawthar Moria. "Detection of Hate Speech Using BERT and Hate Speech Word Embedding with Deep Model." Applied Artificial Intelligence 2023, vol 37, no. 1: 2166719.
  6. Pérez, Juan Manuel, Franco M. Luque, Demian Zayat, Martín Kondratzky, Agustín Moro, Pablo Santiago Serrati, Joaquín Zajac et al. "Assessing the Impact of Contextual Information in Hate Speech Detection." IEEE Access 2023, vol. 11: 30575-30590.
  7. Bilal, Muhammad, Atif Khan, Salman Jan, and Shahrulniza Musa. "Context-Aware Deep Learning Model for Detection of Roman Urdu Hate Speech on Social Media Platform." IEEE Access 2022, vol.10: 121133-121151.
  8. Raza Ur Rehman, Hafiz Muhammad, Mahpara Saleem, Muhammad Zeeshan Jhandir, Eduardo Silva Alvarado, Helena Garay, and Imran Ashraf. "Detecting Hate in Diversity: A Survey of Multilingual Code-Mixed Image and Video Analysis." Journal of Big Data 2025, vol.12, no. 1: 109.
  9. Karim, Md Rezaul, Sumon Kanti Dey, Tanhim Islam, Md Shajalal, and Bharathi Raja Chakravarthi. "Multimodal Hate Speech Detection from Bengali Memes and Texts." in International Conference on Speech and Language Technologies for Low-resource Languages, Cham: Springer International Publishing, 2022: 293-308.
  10. Jahan, Md Saroar, and Mourad Oussalah. "A Systematic Review of Hate Speech Automatic Detection using Natural Language Processing." Neurocomputing 2023, vol. 546: 126232.
  11. F. Elsafoury, Cyberbullying Datasets, Mendeley Data, vol. 1, 2020. Available: https://www.kaggle.com/datasets/saurabhshahane/cyberbullying-dataset