Submitted:
05 October 2023
Posted:
05 October 2023
You are already at the latest version
Abstract
Keywords:
MSC: 68T01
1. Introduction
2. Related Work
3. Materials and Methods
3.1. Data Collection
3.2. Text Preprocessing
3.2.1. Removal Methods for Stop-Words
- (a)
- Iterate the list of genres against the movie reviews to create a vocabulary without repeated words and the counter of each term for each genre.
- (b)
- Iterate the vocabulary without repeated words against the list of genres, and if the word is in all the genres, then the word is added.
- (c)
- Finally, a threshold is set (minimum times a word must appear in each genre to be added). Iterate the vocabulary list in each genre against the list of genres. If the word is in all the genres and the word occurrences are equal or greater than the threshold, then the word is added.
| Algorithm 1 Threshold method. |
![]() |
3.2.2. TF-IDF method
- N represents the total number of documents in the corpus
- : is the number of documents in which the term t appears.
3.2.3. Word2Vec Method
3.3. Clustering Approaches
3.3.1. K-Means Clustering
3.3.2. Mini Batch K-means
3.4. Visualization Stage
3.5. Post-Processing Stage
3.5.1. Logistic Regression Approach
3.5.2. Support Vector Machine Approach
3.5.3. Evaluation Metrics
4. Experimental Results
5. Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Verma, P.; Gupta, P.; Singh, V. A Smart Movie Recommendation System Using Machine Learning Predictive Analysis. Proceedings of Data Analytics and Management: ICDAM 2022; 2023; pp. 45–56. [Google Scholar] [CrossRef]
- Lou, Y. Deep learning-based sentiment analysis of movie reviews. Third International Conference on Machine Learning and Computer Application (ICMLCA 2022). SPIE 2023, 12636, 177–184. [Google Scholar] [CrossRef]
- IMDb. IMDb datasets, 2022.
- Maas, A.L.; Daly, R.E.; Pham, P.T.; Huang, D.; Ng, A.Y.; Potts, C. Learning Word Vectors for Sentiment Analysis. Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies; Association for Computational Linguistics: Portland, Oregon, USA, 2011; pp. 142–150. [Google Scholar]
- Banik, R. The Movies Dataset, 2017.
- Lakshmi, P. IMDb Dataset of 50K Movie Reviews, 2019.
- Adinarayana, S.; Ilavarasan, E. A Hybrid Imbalanced Data Learning Framework to Tackle Opinion Imbalance in Movie Reviews. Communication Software and Networks; Satapathy, S.C., Bhateja, V., Ramakrishna Murty, M., Gia Nhu, N., Kotti, J., Eds.; Springer Singapore: Singapore, 2021; pp. 453–462. [Google Scholar] [CrossRef]
- Unal, F.Z.; Guzel, M.S.; Bostanci, E.; Acici, K.; Asuroglu, T. Multilabel Genre Prediction Using Deep-Learning Frameworks. Applied Sciences 2023, 13, 8665. [Google Scholar] [CrossRef]
- Battu, V.; Batchu, V.; Gangula, R.R.R.; Dakannagari, M.M.K.R.; Mamidi, R. Predicting the Genre and Rating of a Movie Based on its Synopsis. Proceedings of the 32nd Pacific Asia Conference on Language, Information and Computation; Association for Computational Linguistics: Hong Kong, 2018. [Google Scholar]
- wing Ho, K. Movies’ Genres Classification by Synopsis. Movies’ Genres Classification by Synopsis, 2011.
- Hoang, Q. Predicting Movie Genres Based on Plot Summaries, 2018. [CrossRef]
- Pal, A.; Barigidad, A.; Mustafi, A. Identifying movie genre compositions using neural networks and introducing GenRec-a recommender system based on audience genre perception. 2020 5th International Conference on Computing, Communication and Security (ICCCS), 2020; pp. 1–7. [Google Scholar] [CrossRef]
- Wissler, L.; Almashraee, M.; Monett, D.; Paschke, A. The Gold Standard in Corpus Annotation. 5th IEEE Germany Student Conference, 2014. [CrossRef]
- Zhang, X.; Zhao, J.; LeCun, Y. Character-level Convolutional Networks for Text Classification, 2015. [CrossRef]
- Lehmann, J.; Isele, R.; Jakob, M.; Jentzsch, A.; Kontokostas, D.; Mendes, P.; Hellmann, S.; Morsey, M.; Van Kleef, P.; Auer, S.; Bizer, C. DBpedia - A Large-scale, Multilingual Knowledge Base Extracted from Wikipedia. Semantic Web Journal 2014, 6. [Google Scholar] [CrossRef]
- Han, J.; Pei, J.; Tong, H. Data Mining: Concepts and Techniques; Elsevier Science, 2012. [Google Scholar]
- Arora, P.; Deepali; Varshney, S. Analysis of K-Means and K-Medoids Algorithm For Big Data. Procedia Computer Science 2016, 78, 507–512, 1st International Conference on Information Security & Privacy 2015. [Google Scholar] [CrossRef]
- Rose, R.L.; Puranik, T.G.; Mavris, D.N. Natural Language Processing Based Method for Clustering and Analysis of Aviation Safety Narratives. Aerospace 2020, 7. [Google Scholar] [CrossRef]
- Weißer, T.; Saßmannshausen, T.; Ohrndorf, D.; Burggräf, P.; Wagner, J. A clustering approach for topic filtering within systematic literature reviews. MethodsX 2020, 7, 100831. [Google Scholar] [CrossRef]
- Adinugroho, S.; Wihandika, R.C.; Adikara, P.P. Newsgroup topic extraction using term-cluster weighting and Pillar K-Means clustering. International Journal of Computers and Applications 2022, 44, 357–364. [Google Scholar] [CrossRef]
- Kumari, A.; Shashi, M. Vectorization of Text Documents for Identifying Unifiable News Articles. International Journal of Advanced Computer Science and Applications 2019, 10, 305. [Google Scholar] [CrossRef]
- Soni, R.; Mathai, K.J. Improved Twitter Sentiment Prediction through Cluster-then-Predict Model, 2015. [CrossRef]
- Zhao, Y. R and Data Mining: Examples and Case Studies; Elsevier, 2012. [Google Scholar]
- Kaur, N. A Combinatorial Tweet Clustering Methodology Utilizing Inter and Intra Cosine Similarity. PhD thesis, Faculty Of Graduate Studies And Research, University Of Regina, 2015. [Google Scholar]
- Miyamoto, S.; Suzuki, S.; Takumi, S. Clustering in tweets using a fuzzy neighborhood model. 2012 IEEE International Conference on Fuzzy Systems, 2012; pp. 1–6. [Google Scholar] [CrossRef]
- Kadhim, A.I.; Cheah, Y.N.; Ahamed, N.H. Text Document Preprocessing and Dimension Reduction Techniques for Text Document Clustering. 2014 4th International Conference on Artificial Intelligence with Applications in Engineering and Technology, 2014; pp. 69–73. [Google Scholar] [CrossRef]
- Fodeh, S.J.; Al-Garadi, M.; Elsankary, O.; Perrone, J.; Becker, W.; Sarker, A. Utilizing a multi-class classification approach to detect therapeutic and recreational misuse of opioids on Twitter. Computers in Biology and Medicine 2021, 129, 104132. [Google Scholar] [CrossRef]
- Bird, S.; Loper, E. NLTK: The Natural Language Toolkit. Proceedings of the ACL Interactive Poster and Demonstration Sessions; Association for Computational Linguistics: Barcelona, Spain, 2004; pp. 214–217. [Google Scholar] [CrossRef]
- Gene, D.; Suriyawongkul, A. stopwords-iso. Github, 2020.
- Piantadosi, S.T. Zipf’s word frequency law in natural language: A critical review and future directions. Psychonomic Bulletin & Review 2014, 21, 1112–1130. [Google Scholar] [CrossRef]
- Rajaraman, A.; Ullman, J.D. Data Mining. In Mining of Massive Datasets; Cambridge University Press, 2011; pp. 1–17. [Google Scholar] [CrossRef]
- Mikolov, T.; Chen, K.; Corrado, G.; Dean, J. Efficient Estimation of Word Representations in Vector Space, 2013. [CrossRef]
- Blashfield, R.K.; Aldenderfer, M.S. The Literature On Cluster Analysis. Multivariate Behavioral Research 1978, 13, 271–295. [Google Scholar] [CrossRef] [PubMed]
- Wei, D.; Jiang, Q.; Wei, Y.; Wang, S. A novel hierarchical clustering algorithm for gene sequences. BMC Bioinformatics 2012, 13, 174. [Google Scholar] [CrossRef] [PubMed]
- Filipovych, R.; Resnick, S.M.; Davatzikos, C. Semi-supervised cluster analysis of imaging data. NeuroImage 2011, 54, 2185–2197. [Google Scholar] [CrossRef] [PubMed]
- Punj, G.; Stewart, D.W. Cluster Analysis in Marketing Research: Review and Suggestions for Application. Journal of Marketing Research 1983, 20, 134–148. [Google Scholar] [CrossRef]
- Cooley, R.; Mobasher, B.; Srivastava, J. Data Preparation for Mining World Wide Web Browsing Patterns. Knowledge and Information Systems 1999, 1, 5–32. [Google Scholar] [CrossRef]
- Fonseca, J.R. Clustering in the field of social sciences: That is your choice. International Journal of Social Research Methodology 2013, 16, 403–428. [Google Scholar] [CrossRef]
- Dhanachandra, N.; Manglem, K.; Chanu, Y.J. Image Segmentation Using K -means Clustering Algorithm and Subtractive Clustering Algorithm. Procedia Computer Science 2015, 54, 764–771. [Google Scholar] [CrossRef]
- Fahad, A.; Alshatri, N.; Tari, Z.; Alamri, A.; Khalil, I.; Zomaya, A.Y.; Foufou, S.; Bouras, A. A Survey of Clustering Algorithms for Big Data: Taxonomy and Empirical Analysis. IEEE Transactions on Emerging Topics in Computing 2014, 2, 267–279. [Google Scholar] [CrossRef]
- Gallardo Garcia, R.; Beltran, B.; Vilariño, D.; Zepeda, C.; Martínez, R. Comparison of Clustering Algorithms in Text Clustering Tasks. Computacion y Sistemas 2020, 24. [Google Scholar] [CrossRef]
- Hadifar, A.; Sterckx, L.; Demeester, T.; Develder, C. A Self-Training Approach for Short Text Clustering". Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019); Association for Computational Linguistics: Florence, Italy, 2019; pp. 194–199. [Google Scholar] [CrossRef]
- Naeem, S.; Wumaier, A. Study and Implementing K-mean Clustering Algorithm on English Text and Techniques to Find the Optimal Value of K. International Journal of Computer Applications 2018, 182, 7–14. [Google Scholar] [CrossRef]
- Tou, J.; Gonzalez, R.C. Pattern recognition principles; Addison-Wesley Publishing Company, 1974. [Google Scholar]
- Sculley, D. Web-Scale k-Means Clustering. Proceedings of the 19th International Conference on World Wide Web; Association for Computing Machinery: New York, NY, USA, 2010. [Google Scholar] [CrossRef]
- Thorndike, R.L. Who belongs in the family? Psychometrika 1953, 18, 267–276. [Google Scholar] [CrossRef]
- Rousseeuw, P.J. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics 1987, 20, 53–65. [Google Scholar] [CrossRef]
- Caliński, T.; Harabasz, J. A dendrite method for cluster analysis. Communications in Statistics 1974, 3, 1–27. [Google Scholar] [CrossRef]
- Shahapure, K.R.; Nicholas, C. Cluster Quality Analysis Using Silhouette Score. 2020 IEEE 7th International Conference on Data Science and Advanced Analytics (DSAA), 2020; pp. 747–748. [Google Scholar] [CrossRef]
- Tixier, A.J.P. Notes on deep learning for nlp. arXiv 2018. [Google Scholar]
- Franc, V.; Zien, A.; Schölkopf, B. Support vector machines as probabilistic models. ICML, 2011.
- N., V.V. Adaptive and Learning Systems for Signal Processing Communications, and control. Statistical Learning Theory 1998.
- Wei, Q.; Dunbrack, R.L., Jr. The Role of Balanced Training and Testing Data Sets for Binary Classifiers in Bioinformatics. PLOS ONE 2013, 8, 1–12. [Google Scholar] [CrossRef]
- Rostom, E. Unsupervised Clustering and Multi-Label Classification of Ticket Data. Master’s thesis, Freie Universitat Berlin, 2018. [Google Scholar]
- Saif, H.; Fernandez, M.; He, Y.; Alani, H. On Stopwords, Filtering and Data Sparsity for Sentiment Analysis of Twitter. Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14); European Language Resources Association (ELRA): Reykjavik, Iceland, 2014; pp. 810–817. [Google Scholar]
- Ghag, K.V.; Shah, K. Comparative analysis of effect of stopwords removal on sentiment classification. 2015 International Conference on Computer, Communication and Control (IC4), 2015; pp. 1–6. [Google Scholar] [CrossRef]
- Blanchard, A. Understanding and customizing stopword lists for enhanced patent mapping. World Patent Information 2007, 29, 308–316. [Google Scholar] [CrossRef]












| Method | Description |
|---|---|
| Baseline | The baseline method of this study is the non-removal of stop-words. We preprocessed the original dataset as described before, but non-stopwords were removed. |
| Classic method: Pre-compiled stop-words lists | In this study, we employed three different stop-words lists: NLTK [28], ISO-en [29], and the list generated by [12] referred to as Genrec. |
| Method based on Zipf’s law: Z-Method | There are many different methods to compute custom stoplists for various languages. The most common ones are based on Zipf’s law [30]. In this method, the frequency of each distinct word in the text is calculated, and then the words are sorted in decreasing order of their frequencies. Then, the top k most frequent words are added to a stopword list and removed from the text. |
| Proposed method: Threshold | The proposed method verifies if a proposed stop-word is present in each review across all genres; if this condition is met, then checks if the frequency of the word is at least a designated threshold (see Algorithm 1 below). |
| Combined stop-words list | Datasets were cleaned with the list of stop-words mentioned above, and then either with the Z-Method or Threshold method or custom stop-words were detected and removed. |
| Clustering Algorithm Type | Description | Examples |
|---|---|---|
| Partitioning-Based | The number of groups is determined in the beginning. Then, the partitioning algorithms divide data objects into many partitions, each representing a cluster. | K-means, K-mediods, K-modes, PAM, CLARA, CLARANS and FCM. |
| Hierarchical-Based | Data are organized hierarchically depending on the medium of proximity. The intermediate nodes obtain proximities. A dendrogram represents the datasets, where leaf nodes present individual data. | BIRCH, CURE, ROCK and Chameleon |
| Density-Based | Data objects are separated based on density, connectivity, and boundary regions. They are closely related to point-nearest neighbors. A cluster, defined as a connected dense component, grows in any direction that density leads. | DBSCAN, OPTICS, DBCLASD and DENCLUE |
| Grid-Based | Grid-based clustering algorithms partition the data space into a finite number of cells to form a grid structure and then form clusters from the cells in the grid structure. | Wave-Cluster and STING |
| Model-Based | Such a method optimizes the fit between the given data and some (predefined) mathematical models. It is based on the assumption that a mixture of underlying probability distributions generates the data. Also, it automatically determines the number of clusters based on standard statistics, taking noise (outliers) into account and thus yielding a robust clustering method. | MCLUST and COBWEB |
| Genre | Reviews |
|---|---|
| Biography | 33,295 |
| History | 24,397 |
| Action | 28,723 |
| Drama | 36,565 |
| Horror | 20,477 |
| Comedy | 31,181 |
| Romance | 24,671 |
| Adventure | 25,048 |
| Crime | 24,557 |
| Mystery | 16,033 |
| Sci-Fi | 13,985 |
| Thriller | 23,646 |
| Fantasy | 13,982 |
| War | 19,111 |
| Music | 20,814 |
| Western | 19,293 |
| Sport | 18,129 |
| Animation | 24,530 |
| Total | 418,437 |
| Datasets | Number | Pre-Compiled List | Custom List | Total Stopwords |
|---|---|---|---|---|
| Original | 1 | NA | NA | 0 |
| Not_removed | 1 | NA | NA | 0 |
| NLTK | 1 | 127 | 0 | 127 |
| NTLK + Z-Method | 2 | 127 | 50-98 | 177-225 |
| ISO | 1 | 1,298 | 0 | 1,298 |
| ISO + Z-Method | 2 | 1,298 | 47-89 | 1,345-1,387 |
| ISO + Threshold | 3 | 1,298 | 167-412 | 1,465-1,710 |
| Genrec | 1 | 720 | 0 | 720 |
| Genrec + Z-Method | 2 | 720 | 48-91 | 768-811 |
| Z-Method | 3 | NA | 288-985 | 288-985 |
| Threshold | 8 | NA | 160-1,346 | 160-1,346 |
| Total | 25 |
| Cluster | Top 10 Features | Genre |
|---|---|---|
| 0 | horror, film, movi, horror film, horror movi, like, scare, scari, origin,good | Horror |
| 1 | comedi, funni, movi, laugh, film, joke, like, good, time, make, | Comedy |
| 2 | film, like, charact, stori, time, good, make, watch, realli, scene, | Undefined |
| 3 | music, danc, song, movi, film, sing, love, like, stori, great | Music |
| 4 | anim, disney, film, movi, voic, charact, stori, like, kid, good | Animation |
| 5 | film, life, movi, stori, play, man, time, make, love, famili | Undefined |
| 6 | western, film, eastwood, movi, west, good, town, charact, great, time, | Western |
| 7 | war, film, movi, soldier, german, battl, stori, scene, american, like, | War |
| 8 | wayn, john, western, ford, film, movi, charact, play, great, indian | Western |
| 9 | rocki, fight, movi, film, box, train, like, good, son, seri | Sport |
| 10 | movi, like, realli, film, good, charact, stori, time, thing, make | Undefined |
| 11 | movi, watch, like, good, realli, great, stori, time, charact, make | Undefined |
| 12 | film, perform, best, stori, role, movi, oscar, actor, great, play | Undefined |
| 13 | action, film, movi, good, scene, charact, like, time, plot, sequenc, | Action |
| Cluster | Top 10 Features | Genre |
|---|---|---|
| 0 | movi, realli, film, think, definit, great, good, sure, honestli, certainli | Undefined |
| 1 | go, happen, want, know, get, find, decid, peopl, thing, actual | Undefined |
| 2 | movi, realli, think, sure, actual, honestli, lot, good, enjoy, said | Undefined |
| 3 | film, movi, howev, realli, think, feel, simpli, actual, stori, mani | Undefined |
| 4 | peopl, movi, film, actual, think, howev, fact, one, understand, therefor | Undefined |
| 5 | film, movi, realli, think, one, sure, actual, certainli, howev, great | Undefined |
| 6 | think, realli, movi, actual, sure, know, anyway, honestli, guess, kind | Undefined |
| 7 | movi, think, realli, actual, sure, know, thing, one, said, anyway | Undefined |
| 8 | movi, realli, film, think, actual, sure, much, howev, lot, overal | Undefined |
| 9 | think, movi, realli, way, actual, understand, know, howev, peopl, one | Undefined |
| 10 | movi, realli, film, think, sure, still, actual, enjoy, one, definit | Undefined |
| 11 | movi, think, realli, film, actual, sure, one, said, howev, know | Undefined |
| 12 | movi, realli, actual, think, film, sure, howev, one, thing, much | Undefined |
| 13 | movi, think, realli, sure, actual, film, watch, honestli, said, definit | Undefined |
| Cluster | Top 10 Features | Genre |
|---|---|---|
| 0 | hi, br, thi, film, ha, life, movi, br br, wa, play | Undefined |
| 1 | br, movi, br br, thi, wa, thi movi, like, good, watch, just | Undefined |
| 2 | movi, thi, thi movi, wa, watch, like, good, just, great, realli | Undefined |
| 3 | war, br, film, movi, thi, soldier, wa, br br, german, hi | War |
| 4 | film, thi, thi film, br, wa, veri, like, br br, good, watch | Undefined |
| 5 | br, br br, film, thi, hi, wa, stori, ha, movi, like | Undefined |
| 6 | wa, movi, thi, did, br, film, like, good, just, veri | Undefined |
| 7 | anim, br, disney, film, thi, movi, voic, br br, wa, stori | Animation |
| 8 | br, br br, thi, wa, film, movi, hi, like, good, stori | Undefined |
| 9 | thi, film, movi, stori, good, ha, like, wa, time, charact | Undefined |
| Cluster | Top 10 features | Genre |
|---|---|---|
| Cluster 0 | horror, film, movi, horror film, horror movi, like, scare, scari | Horror |
| Cluster 1 | comedi, funni, movi, laugh, film, joke, like, good, time, make | Comedy |
| Cluster 2 | film, like, charact, stori, time, good, make, watch, realli, scene | Undefined |
| Cluster 3 | music, danc, song, movi, film, sing, love, like, stori, great | Musical |
| Cluster 4 | anim, disney, film, movi, voic, charact, stori, like, kid, good | Animation |
| Cluster 5 | film, life, movi, stori, play, man, time, make, love, famili | Undefined |
| Cluster 6 | western, film, eastwood, movi, west, good, town, charact, great, time | Western |
| Cluster 7 | war, film, movi, soldier, german, battl, stori, scene, american, like | War |
| Cluster 8 | wayn, john, western, ford, film, movi, charact, play, great, indian | Western |
| Cluster 9 | rocki, fight, movi, film, box, train, like, good, son, seri | Sport |
| Cluster 10 | movi, like, realli, film, good, charact, stori, time, thing, make, | Undefined |
| Cluster 11 | movi, watch, like, good, realli, great, stori, time, charact, make | Undefined |
| Cluster 12 | film, perform, best, stori, role, movi, oscar, actor, great, play | Undefined |
| Cluster 13 | action, film, movi, good, scene, charact, like, time, plot, sequenc | Action |
| Cluster | Top 10 Features | Genre |
|---|---|---|
| 0 | horror, scari, scare, gore, genr, dead, creepi, hous, zombi, remak | Horror |
| 1 | war, soldier, german, battl, american, men, fight, action, histori, privat | War |
| 2 | rocki, fight, box, train, seri, son, franchis, sequel, ring, match | Sport |
| 3 | western, wayn, eastwood, west, clint, town, ford, genr, stewart, indian | Western |
| 4 | kid, adult, child, famili, fun, anim, parent, funni, school, voic | Family |
| 5 | famili, book, base, power, girl, human, portray, drama, american, oscar | Drama |
| 6 | anim, disney, voic, song, child, famili, fun, music, featur, kid | Animation |
| 7 | funni, comedi, laugh, joke, hilari, humor, fun, sandler, romant, stupid | Comedy |
| 8 | music, danc, song, sing, rock, band, singer, soundtrack, number, girl | Music |
| 9 | action, sequenc, fight, fun, action sequenc, seri, sequel, special, kill, hero | Action |
| Precision | Recall | F1-score | Support | |
|---|---|---|---|---|
| Cluster 0 | 0.86 | 0.72 | 0.78 | 4,548 |
| Cluster 1 | 0.65 | 0.37 | 0.47 | 6,137 |
| Cluster 2 | 0.87 | 0.77 | 0.82 | 6,028 |
| Cluster 3 | 0.91 | 0.78 | 0.84 | 1,762 |
| Cluster 4 | 0.86 | 0.74 | 0.79 | 5,576 |
| Cluster 5 | 0.58 | 0.35 | 0.44 | 10,823 |
| Cluster 6 | 0.85 | 0.69 | 0.76 | 4,819 |
| Cluster 7 | 0.9 | 0.81 | 0.86 | 2,130 |
| Cluster 8 | 0.92 | 0.84 | 0.88 | 5,105 |
| Cluster 9 | 0.83 | 0.78 | 0.8 | 12,541 |
| micro avg | 0.81 | 0.65 | 0.72 | 59,469 |
| macro avg | 0.82 | 0.69 | 0.74 | 59,469 |
| weighted avg | 0.79 | 0.65 | 0.71 | 59,469 |
| samples avg | 0.59 | 0.65 | 0.61 | 59,469 |
| Precision | Recall | F1-score | Support | |
|---|---|---|---|---|
| Cluster 0 | 0.93 | 0.84 | 0.88 | 2,463 |
| Cluster 1 | 0.89 | 0.78 | 0.83 | 3,681 |
| Cluster 2 | 0.88 | 0.77 | 0.82 | 8,735 |
| Cluster 3 | 0.92 | 0.79 | 0.85 | 2,346 |
| Cluster 4 | 0.95 | 0.89 | 0.92 | 3,061 |
| Cluster 5 | 0.82 | 0.74 | 0.77 | 12,140 |
| Cluster 6 | 0.93 | 0.82 | 0.87 | 1,625 |
| Cluster 7 | 0.94 | 0.85 | 0.9 | 2,660 |
| Cluster 8 | 0.97 | 0.83 | 0.9 | 501 |
| Cluster 9 | 0.99 | 0.91 | 0.95 | 396 |
| Cluster 10 | 0.72 | 0.47 | 0.57 | 9,496 |
| Cluster 11 | 0.9 | 0.82 | 0.86 | 7,444 |
| Cluster 12 | 0.81 | 0.59 | 0.68 | 5,798 |
| Cluster 13 | 0.87 | 0.69 | 0.77 | 3,640 |
| micro avg | 0.86 | 0.72 | 0.78 | 63,986 |
| macro avg | 0.89 | 0.77 | 0.83 | 63,986 |
| weighted avg | 0.85 | 0.72 | 0.78 | 63,986 |
| samples avg | 0.67 | 0.72 | 0.69 | 63,986 |
| Precision | Recall | F1-score | Support | |
|---|---|---|---|---|
| Cluster 0 | 0.96 | 0.88 | 0.92 | 2,695 |
| Cluster 1 | 0.96 | 0.88 | 0.92 | 3,122 |
| Cluster 2 | 0.99 | 0.9 | 0.94 | 368 |
| Cluster 3 | 0.96 | 0.89 | 0.92 | 2,045 |
| Cluster 4 | 0.92 | 0.72 | 0.81 | 2,382 |
| Cluster 5 | 0.92 | 0.94 | 0.93 | 32,226 |
| Cluster 6 | 0.95 | 0.9 | 0.92 | 3,076 |
| Cluster 7 | 0.95 | 0.86 | 0.9 | 5,045 |
| Cluster 8 | 0.95 | 0.83 | 0.89 | 3,752 |
| Cluster 9 | 0.92 | 0.83 | 0.88 | 4,757 |
| micro avg | 0.93 | 0.9 | 0.92 | 59,468 |
| macro avg | 0.95 | 0.85 | 0.9 | 59,468 |
| weighted avg | 0.93 | 0.9 | 0.92 | 59,468 |
| samples avg | 0.88 | 0.9 | 0.89 | 59,468 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
