Submitted:
30 October 2025
Posted:
31 October 2025
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Materials and Methods
2.1. Research Questions
- RQ1: What are the most used NLP and machine learning techniques for detecting emotional tone and/or hate speech on social media?
- RQ2: What tools do authors use to implement NLP and machine learning techniques for detecting hate speech and/or emotional tone on social media?
- RQ3: What are the most detected emotions in hate speech?
- RQ4: What are the main challenges, limitations, and future research directions for using NLP and machine learning techniques for detecting emotional tone in hate speech?
- RQ5: Which NLP or machine learning models perform best in emotional tone classification based on metrics such as precision, recall, or F1 score?
2.2. Eligibility Criteria
2.3. Data Sources
2.4. Search Strategy
- P: “Hate Speech”, “Online Hate Speech”, “Hate Speech Against Women”, “Offensive Messages on Social Media”, “Hate Messages Against Women”, “Emotional tone”
- I: “Natural Language Processing”, “Machine Learning”, “Techniques”, “Classification”, “Supervised/Unsupervised Machine Learning”, “RNN”, “BERT”, “GPT”, “Emotion Detection”
- C: “Deep Learning Models and Pre-Trained Embeddings vs. Traditional NLP Classification Techniques”
- O: “Precision”, “Accuracy”, “Recall”, “Detection Rate”, “Precision”, “F1-Score”, “False Positive Rate”, “False Negative Rate”.
- S: “Empirical Studies”, “Comparative Analyses”, “Correlational Studies”, “Inferential Statistical Analysis”.
2.5. Inclusion Criteria
- Primary studies using NLP and ML techniques to detect emotional tone in hate speech.
- Only studies with quantitative results, i.e., with precise measurements.
- Written only in English.
- From the last 6 years (2019 - 2025).
2.6. Exclusion Criteria
- Studies that do not present empirical results related to the detection of emotional tone in hate speech.
- Research that uses datasets that are unrepresentative, irrelevant, or unrelated to hate speech.
- Studies that lack a precise, reproducible, and evaluable methodology for emotional tone classification.
- Studies focused on theoretical or conceptual aspects of natural language processing or machine learning without providing applicable or measurable results.
2.7. Data Classification
3. Results
4. Discussion
5. Conclusions and Future Work
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Ramos, G.; Batista, F.; Ribeiro, R.; Fialho, P.; Moro, S.; Fonseca, A.; Guerra, R.; Carvalho, P.; Marques, C.; Silva, C. A comprehensive review on automatic hate speech detection in the age of the transformer. 14. [CrossRef]
- Zhang, F.; Chen, J.; Tang, Q.; Tian, Y. Evaluation of emotion classification schemes in social media text: an annotation-based approach. 12. [CrossRef]
- Min, C.; Lin, H.; Li, X.; Zhao, H.; Lu, J.; Yang, L.; Xu, B. Finding hate speech with auxiliary emotion detection from self-training multi-label learning perspective. 96, E: Publisher. [CrossRef]
- Paul, J.; Mallick, S.; Mitra, A.; Roy, A.; Sil, J. Multi-modal Twitter Data Analysis for Identifying Offensive Posts Using a Deep Cross-Attention–based Transformer Framework. 19, A: Publisher. [CrossRef]
- Baruah, A.; Wahlang, L.; Jyrwa, F.; Shadap, F.; Barbhuiya, F.; Dey, K. Abusive Language Detection in Khasi Social Media Comments. A: Publisher. [CrossRef]
- Kastrati, M.; Imran, A.S.; Hashmi, E.; Kastrati, Z.; Daudpota, S.M.; Biba, M. Unlocking language barriers: Assessing pre-trained large language models across multilingual tasks and unveiling the black box with Explainable Artificial Intelligence. 149, E: Publisher. [CrossRef]
- Li, Y.; Chan, J.; Peko, G.; Sundaram, D. An explanation framework and method for AI-based text emotion analysis and visualisation. 178, E: Publisher. [CrossRef]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. Declaración PRISMA 2020: una guía actualizada para la publicación de revisiones sistemáticas. 74, E: Publisher. [CrossRef]
- Frandsen, T.F.; Bruun Nielsen, M.F.; Lindhardt, C.L.; Eriksen, M.B. Using the full PICO model as a search tool for systematic reviews resulted in lower recall for some PICO elements. 127, E: Publisher. [CrossRef]
- Zapata, J.I.; Garcés, E.; Fuertes, W. Ransomware Detection with Machine Learning: Techniques, Challenges, and Future Directions - A Systematic Review. 15, S: Publisher. [CrossRef]
- Macas, M.; Wu, C.; Fuertes, W. Adversarial examples: A survey of attacks and defenses in deep learning-enabled cybersecurity systems. Expert Systems with Applications 2024, 238, 122223. [Google Scholar] [CrossRef]
- Benavides-Astudillo, E.; Fuertes, W.; Sanchez-Gordon, S.; Nuñez-Agurto, D.; Rodríguez-Galán, G. A Phishing-Attack-Detection Model Using Natural Language Processing and Deep Learning. Applied Sciences 2023, 13, 5275. [Google Scholar] [CrossRef]
- Haynes, R.B. Forming research questions. 59, E: Publisher.
- Al-Hashedi, M.; Soon, L.K.; Goh, H.N.; Lim, A.H.L.; Siew, E.G. Cyberbullying Detection Based on Emotion. 11, I: Publisher, 5391. [Google Scholar] [CrossRef]
- Alvarez-Gonzalez, N.; Kaltenbrunner, A.; Gómez, V. Uncovering the Limits of Text-based Emotion Detection. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2021. Association for Computational Linguistics. [Google Scholar] [CrossRef]
- Bashynska, I.; Sarafanov, M.; Manikaeva, O. Research and Development of a Modern Deep Learning Model for Emotional Analysis Management of Text Data. 14, M: Publisher. [CrossRef]
- Chakraborty, P.; Nawar, F.; Chowdhury, H.A. Sentiment Analysis of Bengali Facebook Data Using Classical and Deep Learning Approaches. In Innovation in Electrical Power Engineering, Communication, and Computing Technology; Springer Singapore; pp. 209–218. 1: ISSN, 1876. [Google Scholar] [CrossRef]
- de León Languré, A.; Zareei, M. Improving Text Emotion Detection Through Comprehensive Dataset Quality Analysis. 12, I: Publisher, 1665. [Google Scholar] [CrossRef]
- Kastrati, M.; Kastrati, Z.; Shariq Imran, A.; Biba, M. Leveraging distant supervision and deep learning for twitter sentiment and emotion classification. 62, S: Publisher, 1070. [Google Scholar] [CrossRef]
- Keya, A.J.; Kabir, M.M.; Shammey, N.J.; Mridha, M.F.; Islam, M.R.; Watanobe, Y. G-BERT: An Efficient Method for Identifying Hate Speech in Bengali Texts on Social Media. 11, I: Publisher, 7970. [Google Scholar] [CrossRef]
- Koufakou, A.; Garciga, J.; Paul, A.; Morelli, J.; Frank, C. Automatically Classifying Emotions based on Text: A Comparative Exploration of Different Datasets. In Proceedings of the 2022 IEEE 34th International Conference on Tools with Artificial Intelligence (ICTAI). IEEE; pp. 342–346. [CrossRef]
- Liapis, C.M.; Karanikola, A.; Kotsiantis, S. Enhancing sentiment analysis with distributional emotion embeddings. 634, E: Publisher. [CrossRef]
- Pan, R.; García-Díaz, J.A.; Valencia-García, R. Spanish MTLHateCorpus 2023: Multi-task learning for hate speech detection to identify speech type, target, target group and intensity. 94, E: Publisher. [CrossRef]
- Priya, P.; Firdaus, M.; Ekbal, A. A multi-task learning framework for politeness and emotion detection in dialogues for mental health counselling and legal aid. 224, E: Publisher. [CrossRef]
- Rodriguez, A.; Chen, Y.L.; Argueta, C. FADOHS: Framework for Detection and Integration of Unstructured Data of Hate Speech on Facebook Using Sentiment and Emotion Analysis. 10, I: Publisher, 2241. [Google Scholar] [CrossRef]
- Sasidhar, T.T.; B, P.; P, S.K. Emotion Detection in Hinglish(Hindi+English) Code-Mixed Social Media Text. 171, E: Publisher, 1352. [Google Scholar] [CrossRef]
- Sohail, T.; Aiman, A.; Hashmi, E.; Imran, A.S.; Daudpota, S.M.; Yayilgan, S.Y. Hate Speech Detection in Code-Mixed Datasets Using Pretrained Embeddings and Transformers. In Proceedings of the 2024 International Conference on Frontiers of Information Technology (FIT). IEEE; pp. 1–6. [CrossRef]
- Takawane, G.; Phaltankar, A.; Patwardhan, V.; Patil, A.; Joshi, R.; Takalikar, M.S. Language augmentation approach for code-mixed text classification. 5, E: Publisher. [CrossRef]
- Vallecillo-Rodríguez, M.E.; Plaza-del Arco, F.M.; Montejo-Ráez, A. Combining profile features for offensiveness detection on Spanish social media. 272, E: Publisher. [CrossRef]
- Alaeddini, M. Emotion Detection in Reddit: Comparative Study of Machine Learning and Deep Learning Techniques.
- Ngo, A.; Kocoń, J. Integrating personalized and contextual information in fine-grained emotion recognition in text: A multi-source fusion approach with explainability. 118, E: Publisher. [CrossRef]
- Touahri, I.; Mazroui, A. Enhancement of a multi-dialectal sentiment analysis system by the detection of the implied sarcastic features. 227, E: Publisher. [CrossRef]
- Lecourt, F.; Croitoru, M.; Todorov, K. " Only ChatGPT gets me": An Empirical Analysis of GPT versus other Large Language Models for Emotion Detection in Text.
- Zhang, T.; Irsan, I.C.; Thung, F.; Lo, D. Revisiting Sentiment Analysis for Software Engineering in the Era of Large Language Models. 34, A: Publisher. [CrossRef]
- Fillies, J.; Paschke, A. Youth language and emerging slurs: tackling bias in BERT-based hate speech detection. S: Publisher. [CrossRef]
- García-Díaz, J.A.; Cánovas-García, M.; Colomo-Palacios, R.; Valencia-García, R. Detecting misogyny in Spanish tweets. An approach based on linguistics features and word embeddings. 114, E: Publisher. [CrossRef]
- Kaminska, O.; Cornelis, C.; Hoste, V. Fuzzy rough nearest neighbour methods for detecting emotions, hate speech and irony. 625, E: Publisher. [CrossRef]
- Ankita., *!!! REPLACE !!!*; Rani, S.; Bashir, A.K.; Alhudhaif, A.; Koundal, D.; Gunduz, E.S. Ankita.; Rani, S.; Bashir, A.K.; Alhudhaif, A.; Koundal, D.; Gunduz, E.S. An efficient CNN-LSTM model for sentiment detection in #BlackLivesMatter. 193, E: Publisher. [CrossRef]
- Mittal, U. Detecting Hate Speech Utilizing Deep Convolutional Network and Transformer Models. In Proceedings of the 2023 International Conference on Electrical, Electronics, Communication and Computers (ELEXCOM). IEEE; pp. 1–4. [CrossRef]
- Zhu, X.; Lou, Y.; Deng, H.; Ji, D. Leveraging bilingual-view parallel translation for code-switched emotion detection with adversarial dual-channel encoder. 235, E: Publisher. [CrossRef]
- Najafi, A.; Varol, O. TurkishBERTweet: Fast and reliable large language model for social media analysis. 255, E: Publisher. [CrossRef]
- Andrade, R.O.; Fuertes, W.; Cazares, M.; Ortiz-Garcés, I.; Navas, G. An Exploratory Study of Cognitive Sciences Applied to Cybersecurity. Electronics 2022, 11, 1692. [Google Scholar] [CrossRef]
- Yadollahi, A.; Shahraki, A.G.; Zaiane, O.R. Current State of Text Sentiment Analysis from Opinion to Emotion Mining. 50. [CrossRef]
- Wang, S.; Shibghatullah, A.S.; Iqbal, T.J.; Keoy, K.H. A review of multimodal-based emotion recognition techniques for cyberbullying detection in online social media platforms. 36, S: Publisher, 2195. [Google Scholar] [CrossRef]






| Component | Definition | Guiding Question |
|---|---|---|
| P | Population | Who am I referring to? |
| I | Intervention | What action or intervention do I want to study? |
| C | Comparison | What can I compare it to? Are there other options? |
| O | Outcomes | What effect or outcome do I hope to observe? |
| S | Study Design | What type of study will I use to answer the question? |
| Database | Search Strings | Number of Articles |
|---|---|---|
| IEEE Xplore | ("hate speech" OR "hate speech against women") AND "emotion detection" AND ("natural language processing" OR "machine learning" OR BERT OR GPT OR "deep learning") | 16 |
| Science Direct | ("hate speech" OR "hate speech against women") AND "emotion detection" AND ("natural language processing" OR "machine learning" OR BERT OR GPT OR "deep learning") | 82 |
| ACM | ("hate speech" OR "hate speech against women") AND "emotion detection" AND ("natural language processing" OR "machine learning" OR BERT OR GPT OR "deep learning") | 85 |
| Springer | ("hate speech" OR "hate speech against women") AND "emotion detection" AND ("natural language processing" OR "machine learning" OR BERT OR GPT OR "deep learning") | 112 |
| WILEY | ("hate speech" OR "hate speech against women") AND "emotion detection" AND ("natural language processing" OR "machine learning" OR BERT OR GPT OR "deep learning") | 7 |
| Total | 302 |
| Technique | Articles/Authors | Category | Frequency |
|---|---|---|---|
| Tokenization | [3,4,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29] | Text preprocessing | 17 |
| Bag of Words | [3,5,17,19,20,21,25,27,30,31,32] | Traditional text representation | 11 |
| Emotional Processing | [7,14,18,22,25,31,33,34] | Semantic/emotional analysis | 9 |
| Lexical Features | [7,14,18,31,34,35,36] | Feature extraction | 8 |
| BERT Embeddings | [3,4,5,14,15,16,20,37] | Modern contextual representation | 8 |
| Stopword Removal | [17,20,21,25,27,30] | Text preprocessing | 6 |
| Word2Vec | [5,22,30,31,38] | Vector representation | 5 |
| GloVe | [3,19,30,31,39] | Vector representation | 5 |
| FastText | [5,15,19,27,39] | Vector representation | 5 |
| Lemmatization | [14,27,30,39] | Linguistic normalization | 4 |
| Contextual Embeddings | [3,5,14,22] | Modern contextual representation | 4 |
| Stemming | [17,25,39] | Linguistic normalization | 3 |
| Feature Engineering | [25,26,40] | Feature extraction | 3 |
| Attention Mechanisms | [7,24,38] | Interpretation mechanisms | 3 |
| Total | 91 |
| Technique | Articles/Authors | Category | Frequency |
|---|---|---|---|
| BERT | [3,4,14,16,18,19,20,21,23,24,27,28,29,31,34,35,36,39,40,41] | Pre-trained Transformer | 20 |
| RoBERTA | [3,4,19,21,27,28,29,34,35,36] | Pre-trained Transformer | 11 |
| SVM | [3,5,17,21,22,25,27,30,31,32,36] | Discriminative Classifier | 11 |
| BiLSTM | [3,4,5,7,15,18,19,26,31,40] | Recurrent Neural Network (RNN) | 10 |
| Random Forest | [5,15,17,20,22,27,30,31,32,36] | Tree-Based Classifier | 10 |
| LSTM | [3,17,19,26,30,31,32,39,40] | Recurrent Neural Network (RNN) | 9 |
| Naive Bayes | [5,15,17,19,21,22,27,30,31] | Probabilistic Classifier | 9 |
| CNN | [5,17,19,26,31,39,40] | Convolutional Neural Network | 7 |
| GPT-3 | [6,33,41] | Large Language Model (LLM) | 3 |
| GPT-4 | [6,33,41] | Large Language Model (LLM) | 3 |
| Gemini | [6,33,41] | Large Language Model (LLM) | 3 |
| LLaMA 2 | [33,34,41] | Large Language Model (LLM) | 3 |
| Gemma | [23,33] | Large Language Model (LLM) | 2 |
| Total | 101 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).