⚠ Official Notice: www.ijisrt.com is the official website of the International Journal of Innovative Science and Research Technology (IJISRT) Journal for research paper submission and publication. Please beware of fake or duplicate websites using the IJISRT name.



Evaluation Artefacts in Deep Hate Speech Detection: A Reproduction Study and a Character-Aware Model Robust to Adversarial Obfuscation


Authors : Hafsat Mahe Omar; Abdulhakeem Ibrahim; Karatu Musa Tanimu; Sirajo Abdullahi Bakura

Volume/Issue : Volume 11 - 2026, Issue 9 - September


Google Scholar : https://tinyurl.com/y73vhjeu

DOI : https://doi.org/10.38124/ijisrt/26sep325

Note : A published paper may take 4-5 working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and ResearchGate.


Abstract : Automated hate speech detection is routinely reported at accuracies above 98%. We show that for two widely cited systems such figures are substantially artefacts of how their evaluation corpora were built, and we propose a model and protocol designed to avoid the same traps. First, we reproduce Ali et al. (2023) and Kibriya et al. (2024). Re-collecting the hashtag space of the former yields a corpus that is 91.3% duplicated; under the paper's stated protocol a lookup table that performs no learning scores 0.9804 accuracy, and deduplicating before splitting reduces the best model from 0.9915 to 0.5842. For the latter we rebuild the detection corpus from the paper's own recipe, matching its reported hate-class count exactly (22,862 tweets), and show that because the corpus merges two sources along the label axis, corpus identity nearly determines the label: the two sources are 98.8% separable from text alone, a provenance-only baseline reaches 0.8837 against a reported 0.933, and cross-corpus transfer falls below the target corpus's majority-class baseline in both directions. Second, we present a compact character-aware model a character CNN and corpus-trained WordPiece embedding feeding a BiLSTM with additive attention and a per-document graph convolutional branch evaluated over three seeds on a single human-annotated corpus with deduplication and an annotator-column guard. It attains 0.7354 macro-F1 (mean of three seeds) with 713k parameters, exceeding faithful re-implementations of both published architectures by 0.047 and 0.056 macro-F1 at a fraction of their size. Under five character-level evasion attacks it loses 0.252 macro-F1 against 0.359-0.379 for the neural baselines. We further report a quantitative evaluation of SHAP, LIME and attention against a randomdeletion control, finding all three faithful but attention clearly the weakest. We report two negative results: the graph branch's contribution lies inside the seed noise, and a TF-IDF linear baseline matches the proposed model on clean data and degrades comparably under attack. All code, splits and artefacts are released.

Keywords : Hate Speech Detection, Reproducibility, Dataset Bias, Adversarial Robustness, Explainable AI, Graph Convolutional Networks.

References :

  1. Kibriya H, Siddiqa A, Khan WZ, Khan MK. Towards safer online communities: Deep learning and explainable AI for hate speech detection and classification. Computers and Electrical Engineering. 2024;116:109153. doi:10.1016/j.compeleceng.2024.109153
  2. Ali M, Hassan M, Kifayat K, Kim JY, Hakak S, Khan MK. Social media content classification and community detection using deep learning and graph analytics. Technological Forecasting and Social Change. 2023;188:122252.
  3. Davidson T, Warmsley D, Macy M, Weber I. Automated hate speech detection and the problem of offensive language. In: Proceedings of the International AAAI Conference on Web and Social Media (ICWSM). 2017;11(1):512-515.
  4. Lundberg SM, Lee S-I. A unified approach to interpreting model predictions. In: Advances in Neural Information Processing Systems (NeurIPS). 2017:4768-4777.
  5. Ribeiro MT, Singh S, Guestrin C. “Why should I trust you?”: Explaining the predictions of any classifier. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016:1135-1144.
  6. Kipf TN, Welling M. Semi-supervised classification with graph convolutional networks. In: International Conference on Learning Representations (ICLR). 2017.
  7. Hochreiter S, Schmidhuber J. Long short-term memory. Neural Computation. 1997;9(8):1735-1780.
  8. Schuster M, Paliwal KK. Bidirectional recurrent neural networks. IEEE Transactions on Signal Processing. 1997;45(11):2673-2681.
  9. Bahdanau D, Cho K, Bengio Y. Neural machine translation by jointly learning to align and translate. In: International Conference on Learning Representations (ICLR). 2015.
  10. Wu Y, Schuster M, Chen Z, et al. Google's neural machine translation system: Bridging the gap between human and machine translation. arXiv:1609.08144. 2016.
  11. Devlin J, Chang M-W, Lee K, Toutanova K. BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of NAACL-HLT. 2019:4171-4186.
  12. Jain S, Wallace BC. Attention is not Explanation. In: Proceedings of NAACL-HLT. 2019:3543-3556.
  13. DeYoung J, Jain S, Rajani NF, Lehman E, Xiong C, Socher R, Wallace BC. ERASER: A benchmark to evaluate rationalized NLP models. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL). 2020:4443-4458.
  14. Wiegand M, Ruppenhofer J, Kleinbauer T. Detection of Abusive Language: the Problem of Biased Datasets. In: Proceedings of NAACL-HLT. 2019:602-608.
  15. Gorman K, Bedrick S. We need to talk about standard splits. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL). 2019:2786-2791.
  16. Reimers N, Gurevych I. Reporting score distributions makes a difference: Performance study of LSTM-networks for sequence tagging. In: Proceedings of EMNLP. 2017:338-348.
  17. Fazil M, Khan S, Albahlal BM, Alotaibi RM, Siddiqui T, Shah MA. Attentional multi-channel convolution with bidirectional LSTM cell toward hate speech prediction. IEEE Access. 2023;11:16801-16811.
  18. Mozafari M, Farahbakhsh R, Crespi N. A BERT-based transfer learning approach for hate speech detection in online social media. In: Complex Networks and Their Applications VIII. Springer; 2019:928-940.
  19. Saleh H, Alhothali A, Moria K. Detection of hate speech using BERT and hate speech word embedding with deep model. Applied Artificial Intelligence. 2023;37(1):2166719.
  20. Mazari AC, Boudoukhani N, Djeffal A. BERT-based ensemble learning for multi-aspect hate speech detection. Cluster Computing. 2024;27(1):325-339.
  21. Newman MEJ, Girvan M. Finding and evaluating community structure in networks. Physical Review E. 2004;69(2):026113.
  22. Kingma DP, Ba J. Adam: A method for stochastic optimization. In: International Conference on Learning Representations (ICLR). 2015.
  23. Basile V, Bosco C, Fersini E, et al. SemEval-2019 Task 5: Multilingual detection of hate speech against immigrants and women in Twitter. In: Proceedings of the 13th International Workshop on Semantic Evaluation. 2019:54-63.
  24. Awal MR, Lee RK-W, Tanwar E, Garg T, Chakraborty T. Model-agnostic meta-learning for multilingual hate speech detection. IEEE Transactions on Computational Social Systems. 2023.
  25. Zhou X, Yong Y, Fan X, et al. Hate speech detection based on sentiment knowledge sharing. In: Proceedings of ACL-IJCNLP. 2021:7158-7166.
  26. Arango A, Pérez J, Poblete B. Hate speech detection is not as easy as you may think: A closer look at model validation. In: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval. 2019:45-54.

Automated hate speech detection is routinely reported at accuracies above 98%. We show that for two widely cited systems such figures are substantially artefacts of how their evaluation corpora were built, and we propose a model and protocol designed to avoid the same traps. First, we reproduce Ali et al. (2023) and Kibriya et al. (2024). Re-collecting the hashtag space of the former yields a corpus that is 91.3% duplicated; under the paper's stated protocol a lookup table that performs no learning scores 0.9804 accuracy, and deduplicating before splitting reduces the best model from 0.9915 to 0.5842. For the latter we rebuild the detection corpus from the paper's own recipe, matching its reported hate-class count exactly (22,862 tweets), and show that because the corpus merges two sources along the label axis, corpus identity nearly determines the label: the two sources are 98.8% separable from text alone, a provenance-only baseline reaches 0.8837 against a reported 0.933, and cross-corpus transfer falls below the target corpus's majority-class baseline in both directions. Second, we present a compact character-aware model a character CNN and corpus-trained WordPiece embedding feeding a BiLSTM with additive attention and a per-document graph convolutional branch evaluated over three seeds on a single human-annotated corpus with deduplication and an annotator-column guard. It attains 0.7354 macro-F1 (mean of three seeds) with 713k parameters, exceeding faithful re-implementations of both published architectures by 0.047 and 0.056 macro-F1 at a fraction of their size. Under five character-level evasion attacks it loses 0.252 macro-F1 against 0.359-0.379 for the neural baselines. We further report a quantitative evaluation of SHAP, LIME and attention against a randomdeletion control, finding all three faithful but attention clearly the weakest. We report two negative results: the graph branch's contribution lies inside the seed noise, and a TF-IDF linear baseline matches the proposed model on clean data and degrades comparably under attack. All code, splits and artefacts are released.

Keywords : Hate Speech Detection, Reproducibility, Dataset Bias, Adversarial Robustness, Explainable AI, Graph Convolutional Networks.

Paper Submission Last Date
30 - September - 2026

SUBMIT YOUR PAPER CALL FOR PAPERS
Video Explanation for Published paper

Never miss an update from Papermashup

Get notified about the latest tutorials and downloads.

Subscribe by Email

Get alerts directly into your inbox after each post and stay updated.
Subscribe
OR

Subscribe by RSS

Add our RSS to your feedreader to get regular updates from us.
Subscribe