Authors :
Hafsat Mahe Omar; Abdulhakeem Ibrahim; Karatu Musa Tanimu; Sirajo Abdullahi Bakura
Volume/Issue :
Volume 11 - 2026, Issue 9 - September
Google Scholar :
https://tinyurl.com/y73vhjeu
DOI :
https://doi.org/10.38124/ijisrt/26sep325
Note : A published paper may take 4-5
working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and
ResearchGate.
Abstract :
Automated hate speech detection is routinely reported at accuracies above 98%. We show that for two widely
cited systems such figures are substantially artefacts of how their evaluation corpora were built, and we propose a model
and protocol designed to avoid the same traps. First, we reproduce Ali et al. (2023) and Kibriya et al. (2024). Re-collecting
the hashtag space of the former yields a corpus that is 91.3% duplicated; under the paper's stated protocol a lookup table
that performs no learning scores 0.9804 accuracy, and deduplicating before splitting reduces the best model from 0.9915 to
0.5842. For the latter we rebuild the detection corpus from the paper's own recipe, matching its reported hate-class count
exactly (22,862 tweets), and show that because the corpus merges two sources along the label axis, corpus identity nearly
determines the label: the two sources are 98.8% separable from text alone, a provenance-only baseline reaches 0.8837 against
a reported 0.933, and cross-corpus transfer falls below the target corpus's majority-class baseline in both directions. Second,
we present a compact character-aware model a character CNN and corpus-trained WordPiece embedding feeding a
BiLSTM with additive attention and a per-document graph convolutional branch evaluated over three seeds on a single
human-annotated corpus with deduplication and an annotator-column guard. It attains 0.7354 macro-F1 (mean of three
seeds) with 713k parameters, exceeding faithful re-implementations of both published architectures by 0.047 and 0.056
macro-F1 at a fraction of their size. Under five character-level evasion attacks it loses 0.252 macro-F1 against 0.359-0.379
for the neural baselines. We further report a quantitative evaluation of SHAP, LIME and attention against a randomdeletion control, finding all three faithful but attention clearly the weakest. We report two negative results: the graph
branch's contribution lies inside the seed noise, and a TF-IDF linear baseline matches the proposed model on clean data and
degrades comparably under attack. All code, splits and artefacts are released.
Keywords :
Hate Speech Detection, Reproducibility, Dataset Bias, Adversarial Robustness, Explainable AI, Graph Convolutional Networks.
References :
- Kibriya H, Siddiqa A, Khan WZ, Khan MK. Towards safer online communities: Deep learning and explainable AI for hate speech detection and classification. Computers and Electrical Engineering. 2024;116:109153. doi:10.1016/j.compeleceng.2024.109153
- Ali M, Hassan M, Kifayat K, Kim JY, Hakak S, Khan MK. Social media content classification and community detection using deep learning and graph analytics. Technological Forecasting and Social Change. 2023;188:122252.
- Davidson T, Warmsley D, Macy M, Weber I. Automated hate speech detection and the problem of offensive language. In: Proceedings of the International AAAI Conference on Web and Social Media (ICWSM). 2017;11(1):512-515.
- Lundberg SM, Lee S-I. A unified approach to interpreting model predictions. In: Advances in Neural Information Processing Systems (NeurIPS). 2017:4768-4777.
- Ribeiro MT, Singh S, Guestrin C. “Why should I trust you?”: Explaining the predictions of any classifier. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016:1135-1144.
- Kipf TN, Welling M. Semi-supervised classification with graph convolutional networks. In: International Conference on Learning Representations (ICLR). 2017.
- Hochreiter S, Schmidhuber J. Long short-term memory. Neural Computation. 1997;9(8):1735-1780.
- Schuster M, Paliwal KK. Bidirectional recurrent neural networks. IEEE Transactions on Signal Processing. 1997;45(11):2673-2681.
- Bahdanau D, Cho K, Bengio Y. Neural machine translation by jointly learning to align and translate. In: International Conference on Learning Representations (ICLR). 2015.
- Wu Y, Schuster M, Chen Z, et al. Google's neural machine translation system: Bridging the gap between human and machine translation. arXiv:1609.08144. 2016.
- Devlin J, Chang M-W, Lee K, Toutanova K. BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of NAACL-HLT. 2019:4171-4186.
- Jain S, Wallace BC. Attention is not Explanation. In: Proceedings of NAACL-HLT. 2019:3543-3556.
- DeYoung J, Jain S, Rajani NF, Lehman E, Xiong C, Socher R, Wallace BC. ERASER: A benchmark to evaluate rationalized NLP models. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL). 2020:4443-4458.
- Wiegand M, Ruppenhofer J, Kleinbauer T. Detection of Abusive Language: the Problem of Biased Datasets. In: Proceedings of NAACL-HLT. 2019:602-608.
- Gorman K, Bedrick S. We need to talk about standard splits. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL). 2019:2786-2791.
- Reimers N, Gurevych I. Reporting score distributions makes a difference: Performance study of LSTM-networks for sequence tagging. In: Proceedings of EMNLP. 2017:338-348.
- Fazil M, Khan S, Albahlal BM, Alotaibi RM, Siddiqui T, Shah MA. Attentional multi-channel convolution with bidirectional LSTM cell toward hate speech prediction. IEEE Access. 2023;11:16801-16811.
- Mozafari M, Farahbakhsh R, Crespi N. A BERT-based transfer learning approach for hate speech detection in online social media. In: Complex Networks and Their Applications VIII. Springer; 2019:928-940.
- Saleh H, Alhothali A, Moria K. Detection of hate speech using BERT and hate speech word embedding with deep model. Applied Artificial Intelligence. 2023;37(1):2166719.
- Mazari AC, Boudoukhani N, Djeffal A. BERT-based ensemble learning for multi-aspect hate speech detection. Cluster Computing. 2024;27(1):325-339.
- Newman MEJ, Girvan M. Finding and evaluating community structure in networks. Physical Review E. 2004;69(2):026113.
- Kingma DP, Ba J. Adam: A method for stochastic optimization. In: International Conference on Learning Representations (ICLR). 2015.
- Basile V, Bosco C, Fersini E, et al. SemEval-2019 Task 5: Multilingual detection of hate speech against immigrants and women in Twitter. In: Proceedings of the 13th International Workshop on Semantic Evaluation. 2019:54-63.
- Awal MR, Lee RK-W, Tanwar E, Garg T, Chakraborty T. Model-agnostic meta-learning for multilingual hate speech detection. IEEE Transactions on Computational Social Systems. 2023.
- Zhou X, Yong Y, Fan X, et al. Hate speech detection based on sentiment knowledge sharing. In: Proceedings of ACL-IJCNLP. 2021:7158-7166.
- Arango A, Pérez J, Poblete B. Hate speech detection is not as easy as you may think: A closer look at model validation. In: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval. 2019:45-54.
Automated hate speech detection is routinely reported at accuracies above 98%. We show that for two widely
cited systems such figures are substantially artefacts of how their evaluation corpora were built, and we propose a model
and protocol designed to avoid the same traps. First, we reproduce Ali et al. (2023) and Kibriya et al. (2024). Re-collecting
the hashtag space of the former yields a corpus that is 91.3% duplicated; under the paper's stated protocol a lookup table
that performs no learning scores 0.9804 accuracy, and deduplicating before splitting reduces the best model from 0.9915 to
0.5842. For the latter we rebuild the detection corpus from the paper's own recipe, matching its reported hate-class count
exactly (22,862 tweets), and show that because the corpus merges two sources along the label axis, corpus identity nearly
determines the label: the two sources are 98.8% separable from text alone, a provenance-only baseline reaches 0.8837 against
a reported 0.933, and cross-corpus transfer falls below the target corpus's majority-class baseline in both directions. Second,
we present a compact character-aware model a character CNN and corpus-trained WordPiece embedding feeding a
BiLSTM with additive attention and a per-document graph convolutional branch evaluated over three seeds on a single
human-annotated corpus with deduplication and an annotator-column guard. It attains 0.7354 macro-F1 (mean of three
seeds) with 713k parameters, exceeding faithful re-implementations of both published architectures by 0.047 and 0.056
macro-F1 at a fraction of their size. Under five character-level evasion attacks it loses 0.252 macro-F1 against 0.359-0.379
for the neural baselines. We further report a quantitative evaluation of SHAP, LIME and attention against a randomdeletion control, finding all three faithful but attention clearly the weakest. We report two negative results: the graph
branch's contribution lies inside the seed noise, and a TF-IDF linear baseline matches the proposed model on clean data and
degrades comparably under attack. All code, splits and artefacts are released.
Keywords :
Hate Speech Detection, Reproducibility, Dataset Bias, Adversarial Robustness, Explainable AI, Graph Convolutional Networks.