BiCHAT: A Hybrid CNN-BiLSTM Attention Transformer for Hate Speech Detection in Kenyan Multilingual Political Discourse

Authors

  • Kimani Julia Waceke Karatina University, Kenya
  • Kennedy Ndenga Malanga Karatina University, Kenya
  • Josphat Mwai Karani Karatina University, Kenya

Keywords:

Hate Speech Detection, Transformer, CNN, BiLSTM, Attention, Multilingual NLP, Kenyan Political Discourse, Code-Switching, Counterfactual Data Augmentation, Class Imbalance.

Abstract

Detecting hate speech on social media presents significant challenges in multilingual, low-resource contexts. In Kenya, political discourse frequently alternates between English and Swahili, emphasizes ethnic identities, and suffers from insufficient annotated training data for hate speech detection. This study introduces BiCHAT (Bidirectional CNN-LSTM with Hierarchical Attention Transformer), a hybrid deep learning model that integrates a pre-trained cross-lingual transformer (XLM-RoBERTa), a convolutional block, a bidirectional LSTM encoder, and a linear attention mechanism. A lexicon-based labeling system was applied to 121,149 political tweets, resulting in a highly imbalanced binary training set (HATE: 1.45%). Counterfactual Data Augmentation (CDA) was implemented to mitigate identity-group bias. Class-weighted cross-entropy loss helped tackle the class imbalance issue. When tested on a separate set of 11,994 samples, BiCHAT achieved an overall accuracy of 99.10%. Its hate-class F1 score was 0.6250, and it had a ROC AUC of 0.8358. By optimizing the threshold at 0.43 and applying temperature scaling (T = 1.053), the study increased the validation-set hate-class F1 to 0.6548. We have set up infrastructure for language debiasing and systematic ablation studies, ready for use. This study shows that hybrid transformer models, combined with thoughtful data preparation and fair training, can effectively detect hate speech in multilingual low-resource contexts.

References

Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzman, F., & Stoyanov, V. (2020). Unsupervised cross-lingual representation learning at scale. In Proceedings of ACL 2020 (pp. 8440-8451).

Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT 2019 (pp. 4171-4186).

Dixon, L., Li, J., Sorensen, J., Thain, N., & Vasserman, L. (2018). Measuring and mitigating unintended bias in text classification. In Proceedings AAAI/ACM AIES 2018 (pp. 67-73).

Fortuna, P., & Nunes, S. (2018). A survey on automatic detection of hate speech in text. ACM Computing Surveys, 51(4), 1-30.

Founta, A. M., Djouvas, C., Chatzakou, D., Leontiadis, I., Blackburn, J., Stringhini, G., & Kourtellis, N. (2018). Large scale crowdsourcing and characterization of Twitter abusive behavior. In Proceedings of ICWSM 2018 (pp. 491-500).

Kim, Y. (2014). Convolutional neural networks for sentence classification. In Proceedings of EMNLP 2014 (pp. 1746-1751).

Lu, K., Mardziel, P., Wu, F., Amancharla, P., & Datta, A. (2020). Gender bias in neural natural language processing. In Logic, Language, Information, and Computation (pp. 189-202). Springer.

Mozafari, M., Farahbakhsh, R., & Crespi, N. (2020). A BERT-based transfer learning approach for hate speech detection. In Complex Networks Conference (pp. 928-940).

Ngugi, J. M., Fleisch, E., & Afande, F. O. (2021). Swahili-English code-switching in Kenyan digital media. Journal of African Languages and Linguistics, 42(1), 89-117.

Nkemdirim, C., Obiora, O., & Adegoke, T. (2022). Hate speech detection in Nigerian Pidgin. African Language Technology, 3(2), 45-58.

Patwa, P., Aguilar, G., Kar, S., Pandya, S., Pykl, S., Gelbukh, A., & Solorio, T. (2020). SemEval-2020 task 9: Sentiment analysis of code-mixed tweets. In SemEval Workshop (pp. 774-790).

Ranasinghe, T., & Zampieri, M. (2021). MUDES: Multilingual detection of offensive spans. In NAACL-HLT 2021 (pp. 3659-3666).

Schmidt, A., & Wiegand, M. (2017). A survey on hate speech detection using natural language processing. In SocialNLP Workshop (pp. 1-10).

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.

Wamuyu, P. K. (2017). Bridging the digital divide among low income urban communities. Telematics and Informatics, 34(8), 1709-1720.

Downloads

Published

2026-09-21

How to Cite

Kimani Julia Waceke, Kennedy Ndenga Malanga, & Josphat Mwai Karani. (2026). BiCHAT: A Hybrid CNN-BiLSTM Attention Transformer for Hate Speech Detection in Kenyan Multilingual Political Discourse. African Journal of Education,Science and Technology (AJEST), 8(4), 207–214. Retrieved from http://ajest.org/index.php/ajest/article/view/1033

Issue

Section

Articles

Similar Articles

<< < 30 31 32 33 34 35 36 37 38 39 > >> 

You may also start an advanced similarity search for this article.